The architecture in context
The system we are building
Topology evolution can change which connections exist, but the meaning of the external input coordinates still needs to remain fixed. The notebook constructs a feature configuration, selects a training cohort and passes both into a NEAT entry-detection runner. The search objective combines a potential-based reward with a sparsity penalty for issuing decisions.
Who does what in the stack
- NEAT runner
- Evolves topology and connection parameters.
- Custom feature helpers
- Assemble the policy’s ordered input contract.
- Notebook orchestration
- Selects cohorts, fitness settings and resume points.
The project adds a domain-specific fitness function, feature helpers, time constraints and checkpoint/resume orchestration around the evolutionary library. Those adapters are the engineering contribution. NEAT itself is an upstream method, and this notebook is not evidence that evolving topology is universally better than a fixed network.
Open up the implementation
A NEAT checkpoint describes a graph, not a layer list
NEAT searches both weights and connectivity. A model card must identify input/output nodes, enabled edges, activation functions and the exact configuration used to decode a genome. Evolutionary generations measure repeated fitness evaluations, not gradient steps or passes through a labeled dataset.
The mathematical contract
Topology growth can add capacity only if fitness rewards useful behavior rather than data leakage or a simulator loophole. Sparsity penalties change that tradeoff. Feature order is as important as edge weights: loading the same genome with a different input list changes the policy without changing the checkpoint bytes.
Implementation and resource card
- Capacity / budget
- Configured maximum 2,000 generations in the selected training cell; topology and active parameter count vary by genome. No fixed “three-layer” architecture is implied.
- Execution evidence
- This revision inspects and explains the archived implementation. It does not rerun the original workload. No unrecorded convergence time, throughput or accelerator result is supplied.
- Current reproduction context
- Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.
From explanation to a reproducible check
Export a tiny genome as a node-edge table and compute one forward pass by hand in topological order. Compare the decoded network. Count enabled and disabled genes separately, preserve the configuration hash, and do not claim a full-run speed from the maximum generation setting.
Preserve input identities, configuration and failure records with the result. A successful numerical check only establishes the operation it exercises: it does not certify an entire dataset, model or deployed system. Reproduce the interface on a small deterministic input before optimizing throughput or increasing workload size.
A closer look at the implementation
The code that carries the idea
The excerpt places get_feature_config_unified immediately before the training call. Its comment warns that enabling another resolution also requires changing num_inputs in the configuration. That is a concrete schema coupling: a checkpoint can deserialize successfully while reading the wrong feature order or count.
#features, feature_helpers = get_feature_config_unified(df_unified_features, enable_all=True)
features, feature_helpers = get_feature_config_unified(df_unified_features_filtered, enable_all=True) # 10 inputs for 5sec features
# Apply helper functions if needed --> set enable_all=False above
#feature_helpers['enable_10sec_features']() # NEED TO SET num_inputs IN CONFIG ACCORDINGLY, OTHERWISE BUG !!
#feature_helpers['enable_20sec_features']() # can enable multiple features
winner = run_entry_point_detection_training(
df_l2_raw_filtered, #df_l2_raw,
df_unified_features_filtered, #df_unified_features,
potential_df_aligned_filtered, #potential_df_aligned,
config_path="config_entry_model_l2",
max_generations=MAX_GENERATIONS,
potential_threshold=POTENTIAL_THRESHOLD, # Minimum % profit to consider good
sparsity_factor=SPARSITY_FACTOR, # Penalty for each trade
potential_power=POTENTIAL_POWER, # Power for rewarding higher potentials
earliest_trading_time=EARLIEST_TRADING_TIME, # No trading before e.g. 9:30 AM Eastern Time
max_entry_rate=MAX_ENTRY_RATE, # Maximum entry rate (e.g. 0.02 = 2% of the portfolio)
feature_config=features,
checkpoint_dir=DIR_CHECKPOINT,Verbatim archive excerpt from neat_dual_algo_9.ipynb. Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.
The boundary that matters
The archived objective comments also flag a percent-versus-decimal convention. Unit mismatches can change selection pressure by orders of magnitude without causing a runtime error. Saved fitness does not establish a deployable trading policy.