# Larger NQ bar-history test: continuation gate failed The bounded experiment is complete. **Ordinary recent statistics did not establish transferable incremental value, and the compact setup-development description worsened final-period prediction.** Do not start a spline rescue search or transfer experiment from this result. This is a finding about this fixed setup, observation contract, representation and low-capacity model—not a proof that all market history is useless. ## What was actually tested 2,923 native ABC two-minute C-failure opportunities across 482 event-bearing recording days, reconstructed from 484 qualified 2024–2025 NQ files. The within-file cohort contains 6,602,149 five-second bars and 275,068 complete two-minute bars in 516 contiguous segments. All native opportunities in those segments were retained, irrespective of simulated account occupancy. Missing segments, cold starts and file-date cropping mean this is not every opportunity of a fully observed, warm-started exchange-session stream. Observation time is the close of the native trigger bar, not its earlier chart timestamp. Reference is the latest completed close; boundaries are symmetric at max(half pre-trigger ATR, one NQ point), over five minutes. First-bar outcomes are favorable, adverse, both/ordering ambiguous, or neither. Unknown future gaps and recording ends are not labeled neither. These are **bar-path classes, not achievable P&L or execution outcomes**. | Split | Files | Event days | Opportunities | Resolved | F / A / B / N | |---|---:|---:|---:|---:|---| | Development, 2024 | 248 | 248 | 1,512 | 1,505 | 713 / 732 / 2 / 58 | | Validation, first half 2025 | 115 | 114 | 685 | 682 | 318 / 334 / 3 / 27 | | Study-final, second half 2025 | 121 | 120 | 726 | 724 | 350 / 346 / 0 / 28 | Twelve observations remain unresolved: nine recording-end censorings, two gaps before resolution and one fixed day-boundary purge. All are retained in the ledger and paired-loss bounds. Five genuinely ambiguous outcomes remain explicit in development/validation; no final observations happen to be ambiguous. The fixed comparison uses frequency; G geometry (8 features); GT geometry plus ordinary recent bar statistics (14); and GTD with four additional setup-path summaries (18). The three fitted models share logistic ridge, preprocessing and training weights. Only three fits were performed; no validation selection, validation refit, optimizer sweep, new target, spline basis or ticker search. ## Results Unnormalized four-class Brier, lower is better. “Session” below means equal **recording-day proxy** weight, not independently verified exchange-session weight. Opportunity weighting is reported alongside it, not selected afterward. | Model | Validation session | Validation opportunity | Final session | Final opportunity | |---|---:|---:|---:|---:| | Frequency | 0.539650 | 0.541251 | 0.538502 | 0.536633 | | Geometry | 0.541843 | 0.542351 | 0.537364 | 0.536659 | | Geometry + recent statistics | 0.547866 | 0.547821 | 0.536172 | 0.535013 | | Above + setup development | 0.547662 | 0.545372 | 0.539079 | 0.537052 | The more detailed representations improved training fit, not consistently later prediction. All three lost to frequency on validation. On study-final GT's absolute gain over frequency is only 0.002329 / 0.001621 (session/opportunity), well below the frozen 0.01 continuation threshold. GTD is worse than GT under both final weightings. | Final paired difference | Session delta [95% block interval] | Opportunity delta [95% block interval] | |---|---|---| | GT − frequency | −0.002329 [−0.009979, +0.005175] | −0.001621 [−0.008151, +0.004872] | | GT − G | −0.001191 [−0.006564, +0.004070] | −0.001647 [−0.006289, +0.002907] | | GTD − GT | +0.002906 [−0.000689, +0.006359] | +0.002039 [−0.001593, +0.005469] | These are 10,000 paired resamples of five consecutive event-bearing days. Final coverage is 120 day proxies, or 24 five-day blocks—not 724 independent setups. All five registered paired comparisons, both weights, the wider 99.5% family intervals and iid-day sensitivity appear in [ALL_RESULTS.md](ALL_RESULTS.md). No registered simultaneous interval establishes the required improvement. At this resolution the data do more than reproduce a tiny pilot: GT–geometry's 99.5% intervals are approximately [−0.00930,+0.00608] and [−0.00836,+0.00495], which do not include the predeclared −0.01 gain under this conditional-on-fit block model. They are not universal bounds on achievable history-model performance or a prospective power certificate. ## Where the apparent gain comes from The predeclared breakdowns do not show uniform transfer. GT–geometry is worse in Q3 (+0.004230 / +0.003191) and better in Q4 (−0.006435 / −0.006327). It is worse in the calendar roll-window proxy and better outside it. Long/short and source- hour comparisons also differ. **These are diagnostics, not filters to deploy.** All positive and negative groups are tabulated; no winning subset was selected. Removing each comparison's five best recording days reverses GT's aggregate gain: versus frequency the deltas become +0.001955 / +0.001999; versus geometry they become +0.001892 / +0.000915. Roughly 23–25% of positive daily gains sit in those five days. This is not literally a single-day result, but it fails the frozen concentration requirement. GTD is worse even before any removals. GT–geometry's session Brier difference decomposes into approximately −0.001149 for F, +0.000242 for A, effectively zero for B, and −0.000284 for N. Thus the small observed gain is not primarily an ambiguous-class trick; it also does not show a clear, stable directional edge. ## Calibration, coverage and robustness Final GT log loss is 0.832761 / 0.826735 versus frequency 0.837154 / 0.832250; GTD worsens it to 0.840330 / 0.830948. Those are secondary diagnostics, not a replacement winning objective. Final session-weighted GT probabilities average 47.28% favorable, 48.25% adverse, 0.10% ambiguous, 4.37% neither; observations are 48.80%, 47.14%, 0%, 4.07%. Mean agreement does not certify conditional calibration. All fixed-bin reliability tables for all models, splits and weights are in [CALIBRATION.md](CALIBRATION.md); rare B outcomes cannot support a strong calibration assessment. The full-60-bar sensitivity leaves the ranking unchanged: final session Brier is 0.538476 frequency, 0.537491 G, 0.536277 GT, 0.539207 GTD (722 observations, no refit). Assigning either of the two unknown final outcomes any possible class does not overturn the descriptive point ranking: GT−G session paired-loss bounds are [−0.001418,−0.001009], GTD−GT [+0.002741,+0.002978]. These are missing-label identification bounds, NOT sampling confidence intervals. Missing execution receipts were not used as exclusions and are not the blocker here. ## Provenance and verification The broader archive audit is in [DATA_AUDIT.md](DATA_AUDIT.md). Three date- contaminated files were quarantined; 56 reserved/adjacent file references remained unopened. Exact contracts, rolls and timezone/arrival semantics for naive files remain unverified. We reset within files, use no inferred exchange-clock feature, and make only the explicit bar-start/bar-completion observation assumption. “Final” means final within THIS study. The archive has prior strategy-research exposure; 2024 has known descriptor exposure, and 22 modeled final-period dates also appeared in the earlier tick-development inventory. That overlap is recorded in `exposure_ledger.json`. None is relabeled a globally untouched holdout. The native strategy itself was previously researched, so this chronology cannot remove earlier strategy-selection bias. Its reported uncertainty is conditional on this fitted model and archive, not on the full earlier research process. Verification: 15 synthetic tests; 6,964 native-prefix comparisons; feature-prefix checks on all 2,923 events; independent vectorized re-labeling of every event; 48 independently recomputed aggregate scores; 484 source hashes rechecked. Scalar-recursive EMA agreement is within 1.03e-11 normalized units. Old-study hashes and frozen inputs remain unchanged. The initial exact-float assertion stop and its verification-only amendment are preserved in [IMPLEMENTATION_LOG.md](IMPLEMENTATION_LOG.md), with both freezes retained. One CPU process, two numerical threads, no GPU. The final resumed invocation took 16.2 seconds with 341 MiB peak resident memory; that excludes earlier audit, cached reconstruction and independent verification, so is NOT an end-to-end benchmark. Research outputs remain about 10 MB, without raw archive copies. ## Decision Both frozen continuation gates fail. **Stop this bounded experiment without transfer or a spline-specific addition.** The question was executable as a bar-defined research test, and substantially more data did not validate the small pilot's suggestion. Do not blame missing fills for this predictive result. Retain the causal cohort builder, explicit ambiguity/censoring ledger and scoring contract as reusable infrastructure. Any future idea must name a genuinely different mechanism and prospectively frozen test; changing horizons, selecting Q4/short trades or adding nonlinear splines now would be a new search, not a confirmation of this result. Nothing here establishes profitable admission, realistic order execution, or a spline-specific advantage.