Ask a narrow predictive question
Given otherwise similar native setups, does knowing how they developed improve prediction beyond their current geometry? The study reconstructs 2,923 NQ opportunities from 484 qualified five-second files. It includes setups a simulated account would skip while busy. The labels describe five-minute bar paths, not fills, achievable returns, or an executable strategy.
The causal prefix is part of the architecture
The cohort builder, feature extractor, and labeler are as consequential as the predictor. A setup’s chart location can precede the bar close at which it becomes observable. Every feature must be computed from the latter’s available prefix. Native opportunities are retained even when a simulated account would be busy, so the population is not silently restricted to historically selected trades.
Five-second bars also limit the target. If both boundaries occur in one bar, their order is ambiguous. Missing coverage is censoring, not a neither-boundary outcome. These distinctions remain in the accounting before the same low-capacity logistic model is fitted to each representation.
The timestamp is part of the model
A pattern may be drawn at an earlier turning point but become observable only when a later bar closes. Features use that observable trigger, not the retrospective chart position. A bar crossing both boundaries remains ordering-ambiguous; missing future coverage remains censored rather than being called “neither.” These choices prevent a convenient dataset from inventing information.
Freeze a small comparison
Development uses 2024, validation the first half of 2025, and study-final evaluation the second half. Frequency, eight geometry features, fourteen geometry-plus-recent-statistic features, and eighteen features including setup development form the comparison. The fitted arms share low-capacity logistic ridge. Only three fits are run: no validation refit, nonlinear spline rescue, target change, or optimizer sweep.
The result is smaller than the story we wanted
Final recording-day-weighted Brier scores are 0.538502 for frequency, 0.537364 for geometry, 0.536172 with recent statistics, and 0.539079 with development descriptors. The recent-statistic gain is below the frozen 0.01 requirement; paired block intervals include zero. All fitted arms lose to frequency on validation. The richer development description worsens final prediction under both reported weightings.
A reusable causal pipeline and a clean stopping decision
This is a negative information experiment, not a failed execution strategy. The small recent-statistic improvement does not meet the frozen requirement, and richer development descriptors worsen the final result. The reusable contribution is a qualified causal pipeline and a clean stopping decision—not a reason to reopen the same split with more spline capacity.
A negative result that still improves the infrastructure
The cohort builder, causal prefix tests, censoring ledger, and independent score checks remain useful. The final set is not globally untouched: prior research exposure and overlapping earlier dates are recorded. No winning quarter or direction is promoted after the fact. Closing this test preserves what it learned: this particular information proposal did not earn continuation, and adding splines now would begin a new search rather than confirm the old hypothesis.
Evidence & further reading
The links below distinguish the project record from foundational literature. This revised story does not add a new application-validation experiment.
- Larger NQ bar-history test: continuation gate failed. Spline research archive (2026). Local archive snapshot.