Evaluation practice · E13 · Evaluation

An OOD notebook needs an explicit distribution boundary

Loading a trained genome is straightforward. Establishing that its evaluation is genuinely out of distribution takes a separate evidence trail.

Genome loaderShared feature builderCustom inference runner
A stable schema and a changed data distribution are separate questions. The two illustrative clouds make that distinction visible.
Figure 1. What actually shifted?. A stable schema and a changed data distribution are separate questions. The two illustrative clouds make that distinction visible. Illustrative feature-space shift. Original vector illustration.

Follow the information

From input to outcome

The cohort supplies data; the saved artifact supplies the decoder and weights. Loading the artifact is not a test of distribution shift or independence.

The cohort supplies data; the saved artifact supplies the decoder and weights. Loading the artifact is not a test of distribution shift or independence.
Figure 2. Information flow. Solid arrows carry observations, tensors or artifacts; other routes are explicitly labelled. Signal shapes, matrices and network icons are schematic, not measured samples or literal neuron counts. Open full-size SVG ↗ On narrow screens, scroll the diagram horizontally.

Read this alongside Figure 1: A stable schema and a changed data distribution are separate questions. The two illustrative clouds make that distinction visible. The module map and layer-level figures below expand the operations in this route.

An OOD notebook needs an explicit distribution boundary: system and evaluation mapFrozen genome: Checkpoint + config → Evaluation cohort: Identity / date boundary → Feature adapter: Same ordered schema → Inference runner: No parameter updates → Grouped outcomes: Compared with controls. A high-level module map; comparison branches and training details are explained in the article.EVALUATION PRACTICE / E13 / MODULE MAP01 INPUTFrozen genomeCheckpoint + config02 MODULEEvaluation cohortIdentity / date boundary03 MODULEFeature adapterSame ordered schema04 MODULEInference runnerNo parameter updates05 OUTPUTGrouped outcomesCompared with controls
Source-grounded module map. Boxes summarize operations, not individual neurons; comparison arms and training paths are detailed below. On a small screen, scroll the diagram horizontally.
Frozen genome — Checkpoint + config

The architecture in context

What this comparison asks

The notebook loads a saved genome and routes a new feature table into the same inference interface used by the training project. Separating inference from evolution is a good architectural boundary: the evaluator should consume a frozen object rather than mutate the population while inspecting the new cohort.

Who does what in the stack

Genome loader
Restores the archived network and configuration.
Shared feature builder
Maintains inference-time input ordering.
Custom inference runner
Separates policy decisions from outcome accounting.

The project-specific adapter accepts raw observations, unified features, aligned outcome information and a feature configuration. Those inputs have different roles. Outcome-derived quantities can be used to score a completed decision, but they must not enter the policy observation at decision time.

Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph.
Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph. Open full-size SVG ↗

Open up the implementation

Reconstruct the policy before evaluating shift

A concrete operation-level view of this implementation; no unobserved neural architecture is implied.
A concrete operation-level view of this implementation; no unobserved neural architecture is implied. Open full-size SVG ↗

The deployed object is a triple: genome g, decoding configuration c and feature schema s. An OOD filename does not establish an independent population or a meaningful distribution shift. The evaluation needs cohort identity, timestamp coverage and prior exposure in addition to a successfully loaded network.

The mathematical contract

π(x)=Decode⁡(g,c)(Features⁡(x;s))\pi(x)=\operatorname{Decode}(g,c)(\operatorname{Features}(x;s))

A changed feature normalizer or action mapping can masquerade as distribution shift. Freeze those components before asking whether the model transfers. Conversely, a deliberately adapted normalizer is a new adaptation protocol and should not be described as unchanged inference.

Implementation and resource card

Capacity / budget
Inherited variable topology; this evaluation file does not introduce a new trainable architecture.
Execution evidence
This revision inspects and explains the archived implementation. It does not rerun the original workload. No unrecorded convergence time, throughput or accelerator result is supplied.
Current reproduction context
Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.

From explanation to a reproducible check

Evaluate a saved deterministic input before and after reload. Then compare schema hashes and source cohort overlap. Record invalid inputs and abstentions alongside outputs so loading success is not confused with predictive success.

Preserve input identities, configuration and failure records with the result. A successful numerical check only establishes the operation it exercises: it does not certify an entire dataset, model or deployed system. Reproduce the interface on a small deterministic input before optimizing throughput or increasing workload size.

A closer look at the implementation

The code that carries the idea

The snippet makes the distinction inspectable by naming each argument to run_inference_with_trained_genome. It also shows that configuration accompanies the genome. The correct unit of deployment is that pair plus the feature schema—not an isolated pickled network.

Python · cell 6 · lines 7–25
genome, config = load_trained_genome(entry_genome)

# feature_config
features_ood, feature_helpers = get_feature_config_unified(df_unified_features, enable_all=True) # 10 inputs for 5sec features


# Run inference on OOD data
results = run_inference_with_trained_genome(
    genome=genome,
    config=config,
    df_l2_raw_filtered=df_l2_raw,
    df_unified_features_filtered=df_unified_features,
    potential_df_aligned_filtered=potential_df_aligned,
    feature_config=features_ood,  # Choose to use all or specific features
    potential_threshold=1.0,        # 1% minimum profit threshold
    earliest_trading_time="09:30",  # No trading before 7 AM
    max_entry_rate=None,           # Max 2% of bars can be entries
    stop_loss_pct=0.2,            # 0.05 = 5% stop loss
    take_profit_pct=0.03,          # 3% take profit

Verbatim archive excerpt from ood_test_neat_dual_algo_9.ipynb. Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.

The boundary that matters

The filename “ood_test” does not certify a held-out date range, unseen instrument, unseen regime or absence of development exposure. This article treats it as an evaluation harness, not a verified OOD success. Untrusted pickle files must never be loaded merely to inspect a checkpoint.

Keep building

Other posts of interest