Training systems · E20 · Implementation

Put the state machine inside Numba, not the experiment design

A compiled simulation kernel can accelerate repeated evaluations. It cannot turn an in-sample grid search into walk-forward validation.

NumbaNumPyPython orchestration
A compiled transition kernel is still a state machine whose event order and accounting must be preserved.
Figure 1. Compile the transition. A compiled transition kernel is still a state machine whose event order and accounting must be preserved. Illustrative state machine. Original vector illustration.

Follow the information

From input to outcome

A parameter tuple and aligned observations drive one sequential state machine. Parallel candidate search is outside this transition kernel; compilation must preserve event order and accounting.

A parameter tuple and aligned observations drive one sequential state machine. Parallel candidate search is outside this transition kernel; compilation must preserve event order and accounting.
Figure 2. Information flow. Solid arrows carry observations, tensors or artifacts; other routes are explicitly labelled. Signal shapes, matrices and network icons are schematic, not measured samples or literal neuron counts. Open full-size SVG ↗ On narrow screens, scroll the diagram horizontally.

Read this alongside Figure 1: A compiled transition kernel is still a state machine whose event order and accounting must be preserved. The module map and layer-level figures below expand the operations in this route.

Put the state machine inside Numba, not the experiment design: architectureAligned bar arrays: Signal and reference clocks → Parameter tuple: Stops / targets / filters → Numba state machine: Sequential event semantics → Per-case results: Outcomes + counts → Search orchestrator: Candidate comparison. A high-level module map; comparison branches and training details are explained in the article.TRAINING SYSTEMS / E20 / MODULE MAP01 INPUTAligned bar arraysSignal and reference clocks02 MODULEParameter tupleStops / targets / filters03 MODULENumba state machineSequential event semantics04 MODULEPer-case resultsOutcomes + counts05 OUTPUTSearch orchestratorCandidate comparison
Source-grounded module map. Boxes summarize operations, not individual neurons; comparison arms and training paths are detailed below. On a small screen, scroll the diagram horizontally.
Aligned bar arrays — Signal and reference clocks

The architecture in context

The system we are building

The archive separates numeric simulation from Python orchestration. Its core receives contiguous price, indicator and timestamp arrays, plus explicit policy parameters. This is a good compilation boundary: the inner loop follows state transitions, while file discovery, candidate enumeration and reporting remain outside.

Who does what in the stack

Numba
Compiles the numeric state-transition loop.
NumPy
Provides typed arrays and index mappings.
Python orchestration
Enumerates candidates and records results.

The custom kernel encodes the setup logic, ATR band, stop/target choices and signal-to-reference index map. Numba supplies machine-code compilation; it does not decide when a signal is observable or which transition wins when several conditions occur together. Those are properties of the algorithm being compiled.

Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph.
Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph. Open full-size SVG ↗

Open up the implementation

Compile a transition, not a DataFrame workflow

A concrete operation-level view of this implementation; no unobserved neural architecture is implied.
A concrete operation-level view of this implementation; no unobserved neural architecture is implied. Open full-size SVG ↗

A Numba kernel is most useful when its inputs are contiguous numeric arrays and its outputs have fixed meaning. The transition order is part of the algorithm: a stop check before a target check is not interchangeable when both boundaries occur in one observation. The signal-to-reference index map also determines when a decision becomes observable.

The mathematical contract

st+1=T(st,xt;α)s_{t+1}=T(s_t,x_t;\alpha)

Moving Python object work outside the loop can accelerate repeated evaluation, but a faster candidate search does not provide a better validation design. The archived runner’s name does not establish walk-forward evaluation. Keep the reference implementation available as a semantic oracle when optimizing.

Implementation and resource card

Capacity / budget
State-machine implementation; no neural parameter count. Separate compilation time, warmed kernel time and orchestration time.
Execution evidence
This revision inspects and explains the archived implementation. It does not rerun the original workload. No unrecorded convergence time, throughput or accelerator result is supplied.
Current reproduction context
Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.

From explanation to a reproducible check

Run an identical tiny tape through a readable scalar transition function and the compiled version. Include an empty tape, a gap, simultaneous boundaries and the last valid index. Require identical state traces before measuring speed.

Preserve input identities, configuration and failure records with the result. A successful numerical check only establishes the operation it exercises: it does not certify an entire dataset, model or deployed system. Reproduce the interface on a small deterministic input before optimizing throughput or increasing workload size.

A closer look at the implementation

The code that carries the idea

The excerpt is the kernel’s argument contract. Separate signal and reference arrays make the two clocks visible, and s_to_r_idx connects them. Passing precomputed arrays avoids repeated DataFrame operations in the hot loop, but requires an independently checked alignment.

Python · file · lines 337–352
def core_logic_numba_atr_band(
    s_open, s_high, s_low, s_close, s_ema, s_atr, s_time,
    r_open, r_high, r_low, r_close, r_time,
    s_to_r_idx,
    min_body, tick_size,
    trade_longs, trade_shorts,
    atr_min, atr_max,  # ATR BAND
    use_fixed,         # True = fixed points, False = ATR multipliers
    sl_val, tp_val,    # Either points or multipliers depending on use_fixed
    trail_on,
    ema_filter_mode_close, color_order,
    valid_bars
):
    """
    Core backtest logic with ATR band filter (min AND max).
    Supports both fixed point and ATR multiplier SL/TP.

Verbatim archive excerpt from walk_forward_optimization_base_numba.py. Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.

The boundary that matters

Despite the filename, the inspected runner performs parameter search rather than establishing a complete chronological walk-forward procedure. Faster evaluation can make selection bias easier to accumulate. No runtime speedup or strategy profitability is newly measured here.

Keep building

Other posts of interest