Data systems · E11 · Implementation audit

Build a foveal time-series state without borrowing the future

A multi-resolution feature builder combines fast detail and slower context. Its aggregation boundary reveals a subtle source of look-ahead.

Pandas groupbyNumPyCustom feature configuration
Coarse and fine views of time need different completion rules. This cover illustrates the causal distinction the archived builder does not fully respect.
Figure 1. One state. Several clocks.. Coarse and fine views of time need different completion rules. This cover illustrates the causal distinction the archived builder does not fully respect. Illustrative resolution pyramid. Original vector illustration.

Follow the information

From input to outcome

The archived builder uses completed coarse aggregates at finer timestamps. Those aggregates can include later bars in the same bucket. The warning marks this information leak; a causal prefix-only redesign is not silently substituted here.

The archived builder uses completed coarse aggregates at finer timestamps. Those aggregates can include later bars in the same bucket. The warning marks this information leak; a causal prefix-only redesign is not silently substituted here.
Figure 2. Information flow. Solid arrows carry observations, tensors or artifacts; other routes are explicitly labelled. Signal shapes, matrices and network icons are schematic, not measured samples or literal neuron counts. Open full-size SVG ↗ On narrow screens, scroll the diagram horizontally.

Read this alongside Figure 1: Coarse and fine views of time need different completion rules. This cover illustrates the causal distinction the archived builder does not fully respect. The module map and layer-level figures below expand the operations in this route.

Build a foveal time-series state without borrowing the future: architectureFive-second bars: Grouped by ticker / date → Resolution bank: 5, 10, 20 … 60 seconds → OHLCV aggregation: Open / extrema / close / sum → Relative features: Latest close + local context → Policy input: Fixed ordered vector. A high-level module map; comparison branches and training details are explained in the article.DATA SYSTEMS / E11 / MODULE MAP01 INPUTFive-second barsGrouped by ticker / date02 MODULEResolution bank5, 10, 20 … 60 seconds03 MODULEOHLCV aggregationOpen / extrema / close / sum04 MODULERelative featuresLatest close + local context05 OUTPUTPolicy inputFixed ordered vector
Source-grounded module map. Boxes summarize operations, not individual neurons; comparison arms and training paths are detailed below. On a small screen, scroll the diagram horizontally.
Five-second bars — Grouped by ticker / date

The architecture in context

The system we are building

A foveal representation spends more coordinates on recent detail and fewer on older or slower context. The archive builds several resolutions from five-second bars, normalizes price geometry relative to the current close and adds volume and other contextual features. A compact policy can then see multiple time scales without receiving the entire raw history.

Who does what in the stack

Pandas groupby
Separates ticker/session histories.
NumPy
Aggregates multi-resolution OHLCV arrays.
Custom feature configuration
Defines ordered policy inputs across training and inference.

The feature library owns the resolution specification and the assembly of a consistent vector shared by training and inference. Pandas groups sessions; NumPy computes aggregated OHLCV values. That is a substantial part of the ML system even though it contains no trainable layer.

Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph.
Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph. Open full-size SVG ↗

Open up the implementation

Why a foveal feature can accidentally see the future

A concrete operation-level view of this implementation; no unobserved neural architecture is implied.
A concrete operation-level view of this implementation; no unobserved neural architecture is implied. Open full-size SVG ↗

The code aggregates complete coarse buckets, then uses the current bucket index while iterating finer bars. High, low and volume can therefore contain values from later in the same coarse bucket. An efficient tensor lookup does not make those values observable. The mathematical distinction is between a prefix statistic and a completed-interval statistic.

The mathematical contract

Htcausal=max⁡s∈[b(t),t]ps≠max⁡s∈[b(t),b(t)+Δ)psH_t^{\rm causal}=\max_{s\in[b(t),t]}p_s\neq\max_{s\in[b(t),b(t)+\Delta)}p_s

Mixed resolution can allocate detail to the recent past while retaining long-range context cheaply. The cost is semantic complexity: each feature needs a cut-off and a completion rule. Closed-bucket history and a separate evolving current bucket are a possible redesign, not the behavior silently attributed to the source.

Implementation and resource card

Capacity / budget
Feature-vector width depends on resolution and lookback configuration. No independent neural layer or training time belongs to this feature builder.
Execution evidence
This revision inspects and explains the archived implementation. It does not rerun the original workload. No unrecorded convergence time, throughput or accelerator result is supplied.
Current reproduction context
Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.

From explanation to a reproducible check

Perturb only the last fine bar of a coarse bucket. All earlier feature vectors should remain unchanged for a causal implementation. Run this at each supported resolution; testing only the native five-second stream will miss cross-resolution leakage.

Preserve input identities, configuration and failure records with the result. A successful numerical check only establishes the operation it exercises: it does not certify an entire dataset, model or deployed system. Reproduce the interface on a small deterministic input before optimizing throughput or increasing workload size.

A closer look at the implementation

The code that carries the idea

The excerpt selects current_agg_idx = i // n_bars and reads the already aggregated high, low and volume for that bucket. In the preceding code, each bucket was computed from its full chunk. For a bar inside an unfinished chunk, those values can therefore include later source bars. Using the current close alone does not make every other coordinate causal.

Python · file · lines 2349–2366
        # For each bar in the group
        for i in range(len(indices)):
            # CLOSE0 is always the most recent 5-second bar close for all resolutions
            most_recent_close = close_values[i]
            
            # Set resolution-specific VOL0 values and reference values for each time resolution
            for res_name, n_bars in time_resolutions.items():
                current_agg_idx = i // n_bars
                
                if current_agg_idx < len(agg_candles[res_name]['volumes']):
                    # VOL0 is the volume of the most recent candle at this resolution
                    group_df.loc[indices[i], f'{res_name}_VOL_0'] = agg_candles[res_name]['volumes'][current_agg_idx]
                    
                    # Store other reference values for convenience in calculations
                    group_df.loc[indices[i], f'{res_name}_CLOSE_0'] = most_recent_close  # Same as close_values[i]
                    group_df.loc[indices[i], f'{res_name}_OPEN_0'] = agg_candles[res_name]['opens'][current_agg_idx]
                    group_df.loc[indices[i], f'{res_name}_HIGH_0'] = agg_candles[res_name]['highs'][current_agg_idx]
                    group_df.loc[indices[i], f'{res_name}_LOW_0'] = agg_candles[res_name]['lows'][current_agg_idx]

Verbatim archive excerpt from neat_trader_lib_i9.py. Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.

The boundary that matters

This article documents the existing implementation and its boundary; it does not claim the historical features were live-safe. A corrected version must either use only completed coarse buckets or build a partial current bucket from the observed prefix.

Keep building

Other posts of interest