The architecture in context
The system we are building
The archived optimizer defines a trial objective that samples thresholds and stop parameters, runs a simulation and returns a score. Multiple workers use persistent study storage so completed trials survive process boundaries. The architectural value is a common ledger rather than a collection of unrelated worker-local maxima.
Who does what in the stack
- Optuna
- Suggests parameters and records trial states.
- Persistent storage
- Coordinates trials across workers and restarts.
- Custom evaluator
- Defines the actual score and data semantics.
The custom objective adapts a domain simulator to Optuna’s suggest-and-return interface. It also records process identity and elapsed time. The archive’s worker count is a historical setting, not a safe default for this shared machine; no optimizer or worker pool was started for this article.
Open up the implementation
The study database is part of the optimizer
Optuna owns candidate suggestion and trial state; the custom objective owns simulator semantics and the scalar being optimized. A persistent study can be resumed, but only if the dataset and objective remain comparable. Reusing a study name after changing either can combine incomparable trials.
The mathematical contract
Concurrency raises throughput while also increasing peak memory and database contention. It does not multiply independent evidence. The final model needs an untouched evaluation after all trial selection, and failures should remain explicit trial outcomes rather than being replaced with favorable defaults.
Implementation and resource card
- Capacity / budget
- Historical worker setting 16; not a recommendation to occupy all 16 CPUs on the shared machine. Number of trials and data exposure matter.
- Execution evidence
- This revision inspects and explains the archived implementation. It does not rerun the original workload. No unrecorded convergence time, throughput or accelerator result is supplied.
- Current reproduction context
- Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.
From explanation to a reproducible check
Record objective version, cohort identity and parameter schema with each study. Test one trial serially, then a small worker count. Never publish private database connection details as part of a reproducibility example.
Preserve input identities, configuration and failure records with the result. A successful numerical check only establishes the operation it exercises: it does not certify an entire dataset, model or deployed system. Reproduce the interface on a small deterministic input before optimizing throughput or increasing workload size.
A closer look at the implementation
The code that carries the idea
The excerpt shows named trial parameters being passed directly into the evaluator. Names, ranges and units are part of the study definition. A resumed study with a changed objective or data cohort is no longer a continuation of the same statistical experiment merely because the database name is unchanged.
def objective(trial):
import os
import time
print(f"Starting trial {trial.number} on PID {os.getpid()}")
start = time.time()
long_th = trial.suggest_float("long_th", 0.5, 0.95)
short_th = -trial.suggest_float("short_th_abs", 0.5, 0.95)
adverse_long = trial.suggest_float("adverse_long", -0.95, -0.5)
adverse_short = trial.suggest_float("adverse_short", 0.5, 0.95)
trail_stop = trial.suggest_float("trail_stop", 0.5, 5.0)
hard_stop = -trial.suggest_float("hard_stop_abs", 0.5, 3.0)
df_trades = backtest_regime_strategy_trailing_stop_adverse_ibkr(
quotes, sizes, START_TIME, END_TIME,
long_threshold=long_th,
short_threshold=short_th,
adverse_long_threshold=adverse_long,
adverse_short_threshold=adverse_short,
fee_per_trade=total_fees,
trailing_stop_ticks=trail_stop,
hard_stop_ticks=hard_stop,
verbose=FalseVerbatim archive excerpt from hft_v3_analysis_ibkr_param_opt.py. Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.
The boundary that matters
Parallelism multiplies memory as well as throughput, particularly when every worker retains its own quote arrays. Selection consumes the evaluation feedback. A separately frozen final split is still required after choosing a candidate.