Data systems · E19 · Implementation

Keep the policy snapshot separate from the future tape

An episode dataset has two time directions: observations describe what was known; the future tape measures what happened next.

PandasNumPyCustom event extraction
The policy receives information available at the decision time. Future observations belong on the outcome side of the boundary.
Figure 1. Two time directions, one sample. The policy receives information available at the decision time. Future observations belong on the outcome side of the boundary. Causality schematic. Original vector illustration.

Follow the information

From input to outcome

Future bars enter only the simulator/scorer. There is no arrow from the future tape back to the observation or policy; that boundary is the essential contract.

Future bars enter only the simulator/scorer. There is no arrow from the future tape back to the observation or policy; that boundary is the essential contract.
Figure 2. Information flow. Solid arrows carry observations, tensors or artifacts; other routes are explicitly labelled. Signal shapes, matrices and network icons are schematic, not measured samples or literal neuron counts. Open full-size SVG ↗ On narrow screens, scroll the diagram horizontally.

Read this alongside Figure 1: The policy receives information available at the decision time. Future observations belong on the outcome side of the boundary. The module map and layer-level figures below expand the operations in this route.

Keep the policy snapshot separate from the future tape: architectureObserved bar prefix: Setup becomes available → Geometry snapshot: Eight normalized features → Policy observation: Frozen at entry → Future bar tape: Twelve-bar horizon → Simulation outcome: Scoring only. A high-level module map; comparison branches and training details are explained in the article.DATA SYSTEMS / E19 / MODULE MAP01 INPUTObserved bar prefixSetup becomes available02 MODULEGeometry snapshotEight normalized features03 MODULEPolicy observationFrozen at entry04 MODULEFuture bar tapeTwelve-bar horizon05 OUTPUTSimulation outcomeScoring only
Source-grounded module map. Boxes summarize operations, not individual neurons; comparison arms and training paths are detailed below. On a small screen, scroll the diagram horizontally.
Observed bar prefix — Setup becomes available

The architecture in context

The system we are building

The preparation script extracts setup-centered episodes. A compact vector describes volatility, trend, geometry, direction and structural stop distances. A separate future segment supports simulation. This division is broadly useful in offline learning: observation construction and outcome construction may share a source table, but they must not share an information cutoff.

Who does what in the stack

Pandas
Loads and aligns historical bars.
NumPy
Builds fixed-shape feature and future arrays.
Custom event extraction
Defines setup observability and geometry.

The project-specific work is the event detector and the eight-feature representation, including ATR-relative scales and bounded transforms. NumPy produces dense arrays suitable for a batched evaluator; Pandas handles file and time-series preparation. The active future-window constant is twelve bars in this file.

Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph.
Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph. Open full-size SVG ↗

Open up the implementation

Keep a label tape out of the input tensor

A concrete operation-level view of this implementation; no unobserved neural architecture is implied.
A concrete operation-level view of this implementation; no unobserved neural architecture is implied. Open full-size SVG ↗

The preparation code combines geometry, context and subsequent bars into a stored sample. The resulting file can contain both causal input and future outcome data; file membership does not make every column a permitted feature. Column roles, cut-off times and row identity are the critical interfaces to the learner.

The mathematical contract

xi=f(Fti),yi=g(pti+1:ti+12)x_i=f(\mathcal F_{t_i}),\qquad y_i=g(p_{t_i+1:t_i+12})

A longer target horizon changes overlap and label maturity. It also increases opportunities for leakage if upstream aggregations use complete future-containing buckets. The short policy’s eight-input interface should be enforced structurally rather than relying on a programmer to remember which columns to exclude.

Implementation and resource card

Capacity / budget
Eight feature columns; future tape is target/evaluation data, not policy input. No training occurs in preparation.
Execution evidence
This revision inspects and explains the archived implementation. It does not rerun the original workload. No unrecorded convergence time, throughput or accelerator result is supplied.
Current reproduction context
Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.

From explanation to a reproducible check

Randomize the future tape and assert that all eight inputs remain unchanged. Shift one row’s label tape and make the identity check fail. At chronological split boundaries, exclude labels whose maturity crosses into the next partition.

Preserve input identities, configuration and failure records with the result. A successful numerical check only establishes the operation it exercises: it does not certify an entire dataset, model or deployed system. Reproduce the interface on a small deterministic input before optimizing throughput or increasing workload size.

A closer look at the implementation

The code that carries the idea

The excerpt builds direction and structural-distance coordinates, then assembles the feature vector in a fixed order. A tanh transform bounds magnitude, but its denominator still needs a meaningful unit and a nonzero scale. Feature order is an API shared with the policy adapter.

Python · file · lines 255–276
        f_ratio = ratio

        # [4] Child Bar Size (relative to ATR)
        f_size = np.tanh((h[i] - l[i]) / curr_atr)

        # [5] Direction
        f_dir = direction

        # [6] Structural Stop Distance — Child (ATR-normalized)
        f_struct_child = np.tanh(child_sl_dist / curr_atr)

        # [7] Structural Stop Distance — Mother (ATR-normalized)
        f_struct_mother = np.tanh(mother_sl_dist / curr_atr)

        features = np.array([
            f_vol_regime,
            f_vol_opp,
            f_trend,
            f_ratio,
            f_size,
            f_dir,
            f_struct_child,

Verbatim archive excerpt from prep_data_pgpe.py. Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.

The boundary that matters

Separating arrays is necessary, not sufficient, for causality. Every feature’s upstream rolling statistic and the setup’s observability time still need checking. A future tape is an evaluation object, never an extra policy input. The current article does not reopen the completed trading studies.

Keep building

Other posts of interest