Training systems · E23 · Implementation

Parallel hyperparameter search needs one experiment ledger

Optuna workers can share a study, but they must also share a frozen objective, data identity and resource budget.

OptunaPersistent storageCustom evaluator
Parallel workers remain comparable only when they report to the same trial ledger under the same evaluator.
Figure 1. Parallel workers, one experiment. Parallel workers remain comparable only when they report to the same trial ledger under the same evaluator. Study orchestration schematic. Original vector illustration.

Follow the information

From input to outcome

Workers receive separate trial parameters and return results to one shared ledger. Data identity and objective semantics must be constant across workers.

Workers receive separate trial parameters and return results to one shared ledger. Data identity and objective semantics must be constant across workers.
Figure 2. Information flow. Solid arrows carry observations, tensors or artifacts; other routes are explicitly labelled. Signal shapes, matrices and network icons are schematic, not measured samples or literal neuron counts. Open full-size SVG ↗ On narrow screens, scroll the diagram horizontally.

Read this alongside Figure 1: Parallel workers remain comparable only when they report to the same trial ledger under the same evaluator. The module map and layer-level figures below expand the operations in this route.

Parallel hyperparameter search needs one experiment ledger: architecturePersistent study: Trial identities + states → Worker processes: Bounded concurrency → Suggested parameters: One candidate per trial → Fixed evaluator: Same data / objective → Trial result: Score + diagnostics. A high-level module map; comparison branches and training details are explained in the article.TRAINING SYSTEMS / E23 / MODULE MAP01 INPUTPersistent studyTrial identities + states02 MODULEWorker processesBounded concurrency03 MODULESuggested parametersOne candidate per trial04 MODULEFixed evaluatorSame data / objective05 OUTPUTTrial resultScore + diagnostics
Source-grounded module map. Boxes summarize operations, not individual neurons; comparison arms and training paths are detailed below. On a small screen, scroll the diagram horizontally.
Persistent study — Trial identities + states

The architecture in context

The system we are building

The archived optimizer defines a trial objective that samples thresholds and stop parameters, runs a simulation and returns a score. Multiple workers use persistent study storage so completed trials survive process boundaries. The architectural value is a common ledger rather than a collection of unrelated worker-local maxima.

Who does what in the stack

Optuna
Suggests parameters and records trial states.
Persistent storage
Coordinates trials across workers and restarts.
Custom evaluator
Defines the actual score and data semantics.

The custom objective adapts a domain simulator to Optuna’s suggest-and-return interface. It also records process identity and elapsed time. The archive’s worker count is a historical setting, not a safe default for this shared machine; no optimizer or worker pool was started for this article.

Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph.
Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph. Open full-size SVG ↗

Open up the implementation

The study database is part of the optimizer

A concrete operation-level view of this implementation; no unobserved neural architecture is implied.
A concrete operation-level view of this implementation; no unobserved neural architecture is implied. Open full-size SVG ↗

Optuna owns candidate suggestion and trial state; the custom objective owns simulator semantics and the scalar being optimized. A persistent study can be resumed, but only if the dataset and objective remain comparable. Reusing a study name after changing either can combine incomparable trials.

The mathematical contract

α∗=arg⁡max⁡α∈AS(α;Dselect)\alpha^*=\arg\max_{\alpha\in\mathcal A}S(\alpha;D_{\rm select})

Concurrency raises throughput while also increasing peak memory and database contention. It does not multiply independent evidence. The final model needs an untouched evaluation after all trial selection, and failures should remain explicit trial outcomes rather than being replaced with favorable defaults.

Implementation and resource card

Capacity / budget
Historical worker setting 16; not a recommendation to occupy all 16 CPUs on the shared machine. Number of trials and data exposure matter.
Execution evidence
This revision inspects and explains the archived implementation. It does not rerun the original workload. No unrecorded convergence time, throughput or accelerator result is supplied.
Current reproduction context
Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.

From explanation to a reproducible check

Record objective version, cohort identity and parameter schema with each study. Test one trial serially, then a small worker count. Never publish private database connection details as part of a reproducibility example.

Preserve input identities, configuration and failure records with the result. A successful numerical check only establishes the operation it exercises: it does not certify an entire dataset, model or deployed system. Reproduce the interface on a small deterministic input before optimizing throughput or increasing workload size.

A closer look at the implementation

The code that carries the idea

The excerpt shows named trial parameters being passed directly into the evaluator. Names, ranges and units are part of the study definition. A resumed study with a changed objective or data cohort is no longer a continuation of the same statistical experiment merely because the database name is unchanged.

Python · file · lines 900–922
def objective(trial):
    import os
    import time
    print(f"Starting trial {trial.number} on PID {os.getpid()}")
    start = time.time()

    long_th  = trial.suggest_float("long_th", 0.5, 0.95)
    short_th = -trial.suggest_float("short_th_abs", 0.5, 0.95)
    adverse_long  = trial.suggest_float("adverse_long", -0.95, -0.5)
    adverse_short = trial.suggest_float("adverse_short", 0.5, 0.95)
    trail_stop = trial.suggest_float("trail_stop", 0.5, 5.0)
    hard_stop  = -trial.suggest_float("hard_stop_abs", 0.5, 3.0)

    df_trades = backtest_regime_strategy_trailing_stop_adverse_ibkr(
        quotes, sizes, START_TIME, END_TIME,
        long_threshold=long_th,
        short_threshold=short_th,
        adverse_long_threshold=adverse_long,
        adverse_short_threshold=adverse_short,
        fee_per_trade=total_fees,
        trailing_stop_ticks=trail_stop,
        hard_stop_ticks=hard_stop,
        verbose=False

Verbatim archive excerpt from hft_v3_analysis_ibkr_param_opt.py. Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.

The boundary that matters

Parallelism multiplies memory as well as throughput, particularly when every worker retains its own quote arrays. Selection consumes the evaluation feedback. A separately frozen final split is still required after choosing a candidate.

Keep building

Other posts of interest