The architecture in context
The system we are building
A foveal representation spends more coordinates on recent detail and fewer on older or slower context. The archive builds several resolutions from five-second bars, normalizes price geometry relative to the current close and adds volume and other contextual features. A compact policy can then see multiple time scales without receiving the entire raw history.
Who does what in the stack
- Pandas groupby
- Separates ticker/session histories.
- NumPy
- Aggregates multi-resolution OHLCV arrays.
- Custom feature configuration
- Defines ordered policy inputs across training and inference.
The feature library owns the resolution specification and the assembly of a consistent vector shared by training and inference. Pandas groups sessions; NumPy computes aggregated OHLCV values. That is a substantial part of the ML system even though it contains no trainable layer.
Open up the implementation
Why a foveal feature can accidentally see the future
The code aggregates complete coarse buckets, then uses the current bucket index while iterating finer bars. High, low and volume can therefore contain values from later in the same coarse bucket. An efficient tensor lookup does not make those values observable. The mathematical distinction is between a prefix statistic and a completed-interval statistic.
The mathematical contract
Mixed resolution can allocate detail to the recent past while retaining long-range context cheaply. The cost is semantic complexity: each feature needs a cut-off and a completion rule. Closed-bucket history and a separate evolving current bucket are a possible redesign, not the behavior silently attributed to the source.
Implementation and resource card
- Capacity / budget
- Feature-vector width depends on resolution and lookback configuration. No independent neural layer or training time belongs to this feature builder.
- Execution evidence
- This revision inspects and explains the archived implementation. It does not rerun the original workload. No unrecorded convergence time, throughput or accelerator result is supplied.
- Current reproduction context
- Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.
From explanation to a reproducible check
Perturb only the last fine bar of a coarse bucket. All earlier feature vectors should remain unchanged for a causal implementation. Run this at each supported resolution; testing only the native five-second stream will miss cross-resolution leakage.
Preserve input identities, configuration and failure records with the result. A successful numerical check only establishes the operation it exercises: it does not certify an entire dataset, model or deployed system. Reproduce the interface on a small deterministic input before optimizing throughput or increasing workload size.
A closer look at the implementation
The code that carries the idea
The excerpt selects current_agg_idx = i // n_bars and reads the already aggregated high, low and volume for that bucket. In the preceding code, each bucket was computed from its full chunk. For a bar inside an unfinished chunk, those values can therefore include later source bars. Using the current close alone does not make every other coordinate causal.
# For each bar in the group
for i in range(len(indices)):
# CLOSE0 is always the most recent 5-second bar close for all resolutions
most_recent_close = close_values[i]
# Set resolution-specific VOL0 values and reference values for each time resolution
for res_name, n_bars in time_resolutions.items():
current_agg_idx = i // n_bars
if current_agg_idx < len(agg_candles[res_name]['volumes']):
# VOL0 is the volume of the most recent candle at this resolution
group_df.loc[indices[i], f'{res_name}_VOL_0'] = agg_candles[res_name]['volumes'][current_agg_idx]
# Store other reference values for convenience in calculations
group_df.loc[indices[i], f'{res_name}_CLOSE_0'] = most_recent_close # Same as close_values[i]
group_df.loc[indices[i], f'{res_name}_OPEN_0'] = agg_candles[res_name]['opens'][current_agg_idx]
group_df.loc[indices[i], f'{res_name}_HIGH_0'] = agg_candles[res_name]['highs'][current_agg_idx]
group_df.loc[indices[i], f'{res_name}_LOW_0'] = agg_candles[res_name]['lows'][current_agg_idx]Verbatim archive excerpt from neat_trader_lib_i9.py. Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.
The boundary that matters
This article documents the existing implementation and its boundary; it does not claim the historical features were live-safe. A corrected version must either use only completed coarse buckets or build a partial current bucket from the observed prefix.