Evaluation practice · E21 · Evaluation

Aggregate a strategy evaluation without losing the quiet days

Session-level accounting, zero-action days and regime labels belong in the result—not just the winning events.

NumPyCustom sequential simulatorReporting layer
Observed inactivity is a valid outcome. An unavailable session is a different kind of missing information.
Figure 1. Quiet days still count. Observed inactivity is a valid outcome. An unavailable session is a different kind of missing information. Illustrative session ledger. Original vector illustration.

Follow the information

From input to outcome

A quiet observed day contributes a zero-activity record. Missing observation coverage is a different state and must not be converted into either a zero or an omitted inconvenient day.

A quiet observed day contributes a zero-activity record. Missing observation coverage is a different state and must not be converted into either a zero or an omitted inconvenient day.
Figure 2. Information flow. Solid arrows carry observations, tensors or artifacts; other routes are explicitly labelled. Signal shapes, matrices and network icons are schematic, not measured samples or literal neuron counts. Open full-size SVG ↗ On narrow screens, scroll the diagram horizontally.

Read this alongside Figure 1: Observed inactivity is a valid outcome. An unavailable session is a different kind of missing information. The module map and layer-level figures below expand the operations in this route.

Aggregate a strategy evaluation without losing the quiet days: system and evaluation mapDaily event records: All eligible sessions → Sequential state: Daily rules / lockouts → Session summaries: Outcome + activity counts → Monthly aggregation: Regime context retained → Diagnostics: All days and active days. A high-level module map; comparison branches and training details are explained in the article.EVALUATION PRACTICE / E21 / MODULE MAP01 INPUTDaily event recordsAll eligible sessions02 MODULESequential stateDaily rules / lockouts03 MODULESession summariesOutcome + activity counts04 MODULEMonthly aggregationRegime context retained05 OUTPUTDiagnosticsAll days and active days
Source-grounded module map. Boxes summarize operations, not individual neurons; comparison arms and training paths are detailed below. On a small screen, scroll the diagram horizontally.
Daily event records — All eligible sessions

The architecture in context

What this comparison asks

This validator collects daily outcomes and event counts before producing monthly summaries. It distinguishes total days, active days and zero-trade days. That distinction is useful beyond trading: a detector that rarely fires can look excellent on its accepted events while serving almost none of the population it was meant to handle.

Who does what in the stack

NumPy
Computes summary statistics over recorded sessions.
Custom sequential simulator
Maintains event and daily state.
Reporting layer
Separates coverage, activity and conditional outcomes.

The project adds daily stopping rules and context-dependent reporting around a simulator. The filename includes “parallel”, but the inspected configuration selects sequential processing. Execution order matters whenever one session or event updates a state consumed by the next.

Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph.
Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph. Open full-size SVG ↗

Open up the implementation

Choose the denominator before aggregation

A concrete operation-level view of this implementation; no unobserved neural architecture is implied.
A concrete operation-level view of this implementation; no unobserved neural architecture is implied. Open full-size SVG ↗

A no-action day with complete observation contributes a valid zero under an all-session outcome definition. An unobserved day is not a zero. The validator’s sequential setting must also be retained when one day’s state affects another; a filename mentioning parallelism is not proof that sessions were evaluated independently.

The mathematical contract

r‾all=∑d∈Drd∣D∣,r‾active=∑d∈Ard∣A∣\overline r_{\rm all}=\frac{\sum_{d\in D}r_d}{|D|},\qquad \overline r_{\rm active}=\frac{\sum_{d\in A}r_d}{|A|}

Session weighting treats days equally; event weighting gives more influence to busy days. Both can be useful, but switching denominators after looking at results creates an apparent improvement without a model change. Regime labels must be defined before selecting a favorable subgroup.

Implementation and resource card

Capacity / budget
No trainable model. Session count, activity count and coverage replace parameter/epoch fields.
Execution evidence
This revision inspects and explains the archived implementation. It does not rerun the original workload. No unrecorded convergence time, throughput or accelerator result is supplied.
Current reproduction context
Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.

From explanation to a reproducible check

Reconcile eligible=observed+missing and observed=active+inactive. Include one missing and one inactive day in a toy ledger. Confirm that the report preserves both instead of silently merging them.

Preserve input identities, configuration and failure records with the result. A successful numerical check only establishes the operation it exercises: it does not certify an entire dataset, model or deployed system. Reproduce the interface on a small deterministic input before optimizing throughput or increasing workload size.

A closer look at the implementation

The code that carries the idea

The excerpt separately computes total outcome, event count, wins, daily standard deviation and zero-action days. These are different denominators. A mean over active days and a mean over all eligible days answer different questions, so both need explicit labels rather than being swapped to improve a headline.

Python · file · lines 451–472
def calc_summary_metrics(daily_pnls, trade_counts, trade_objs, multiplier, vol_metrics=None):
    if not daily_pnls:
        keys =["TotalPnL", "Total$", "Exp", "Sharpe", "TotalTrds", "WR", "AvgDay",
                "MinDay", "NegDays", "ZeroTrd", "MCL", "GrnDays", "GrnDaysEx",
                "ATR", "ATR_RTH", "ATR_Opn", "ATR_PM",
                "Rel", "Rel_RTH", "Rel_Opn", "Rel_PM", "Corr_PM", "Corr_Opn"]
        return {k: 0 for k in keys}
    tot_pnl = sum(daily_pnls)
    n_trd = sum(trade_counts)
    wins = len([t for t in trade_objs if t.result == TradeResult.WIN])
    avg_day = np.mean(daily_pnls)
    std_day = np.std(daily_pnls)
    target_hits = sum(1 for p in daily_pnls if p >= DAILY_TARGET)
    total_days = len(daily_pnls)
    zero_trd_days = sum(1 for c in trade_counts if c == 0)
    active_days = total_days - zero_trd_days
    if vol_metrics:
        avg_atr = np.mean([v[0] for v in vol_metrics])
        avg_rth = np.mean([v[1] for v in vol_metrics])
        avg_opn = np.mean([v[2] for v in vol_metrics])
        avg_pm  = np.mean([v[3] for v in vol_metrics])
        avg_rel = np.mean([v[4] for v in vol_metrics])

Verbatim archive excerpt from validate_monthly_umbrella_v1b_v4_regime_aware_parallel_v2.py. Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.

The boundary that matters

Zero is not a substitute for an unknown outcome. Missing recordings, censored horizons and a genuine no-action day need different states upstream. Regime breakdowns are descriptive unless their definitions and comparisons were fixed before inspecting outcomes.

Keep building

Other posts of interest