Historical source. Some claims in older records were subsequently corrected. The associated article states the adopted interpretation. This record preserves the original source alongside its rendered reading view.
Rendered archival Markdown
This reading view preserves headings, tables, lists, code fragments and mathematical notation from the local research record.
Larger NQ bar-history test: continuation gate failed
The bounded experiment is complete. Ordinary recent statistics did not establish transferable incremental value, and the compact setup-development description worsened final-period prediction. Do not start a spline rescue search or transfer experiment from this result. This is a finding about this fixed setup, observation contract, representation and low-capacity model—not a proof that all market history is useless.
What was actually tested
2,923 native ABC two-minute C-failure opportunities across 482 event-bearing recording days, reconstructed from 484 qualified 2024–2025 NQ files. The within-file cohort contains 6,602,149 five-second bars and 275,068 complete two-minute bars in 516 contiguous segments. All native opportunities in those segments were retained, irrespective of simulated account occupancy. Missing segments, cold starts and file-date cropping mean this is not every opportunity of a fully observed, warm-started exchange-session stream.
Observation time is the close of the native trigger bar, not its earlier chart timestamp. Reference is the latest completed close; boundaries are symmetric at max(half pre-trigger ATR, one NQ point), over five minutes. First-bar outcomes are favorable, adverse, both/ordering ambiguous, or neither. Unknown future gaps and recording ends are not labeled neither. These are bar-path classes, not achievable P&L or execution outcomes.
| Split | Files | Event days | Opportunities | Resolved | F / A / B / N |
|---|
| Development, 2024 | 248 | 248 | 1,512 | 1,505 | 713 / 732 / 2 / 58 |
| Validation, first half 2025 | 115 | 114 | 685 | 682 | 318 / 334 / 3 / 27 |
| Study-final, second half 2025 | 121 | 120 | 726 | 724 | 350 / 346 / 0 / 28 |
Twelve observations remain unresolved: nine recording-end censorings, two gaps before resolution and one fixed day-boundary purge. All are retained in the ledger and paired-loss bounds. Five genuinely ambiguous outcomes remain explicit in development/validation; no final observations happen to be ambiguous.
The fixed comparison uses frequency; G geometry (8 features); GT geometry plus ordinary recent bar statistics (14); and GTD with four additional setup-path summaries (18). The three fitted models share logistic ridge, preprocessing and training weights. Only three fits were performed; no validation selection, validation refit, optimizer sweep, new target, spline basis or ticker search.
Results
Unnormalized four-class Brier, lower is better. “Session” below means equal recording-day proxy weight, not independently verified exchange-session weight. Opportunity weighting is reported alongside it, not selected afterward.
| Model | Validation session | Validation opportunity | Final session | Final opportunity |
|---|
| Frequency | 0.539650 | 0.541251 | 0.538502 | 0.536633 |
| Geometry | 0.541843 | 0.542351 | 0.537364 | 0.536659 |
| Geometry + recent statistics | 0.547866 | 0.547821 | 0.536172 | 0.535013 |
| Above + setup development | 0.547662 | 0.545372 | 0.539079 | 0.537052 |
The more detailed representations improved training fit, not consistently later prediction. All three lost to frequency on validation. On study-final GT's absolute gain over frequency is only 0.002329 / 0.001621 (session/opportunity), well below the frozen 0.01 continuation threshold. GTD is worse than GT under both final weightings.
| Final paired difference | Session delta [95% block interval] | Opportunity delta [95% block interval] |
|---|
| GT − frequency | −0.002329 [−0.009979, +0.005175] | −0.001621 [−0.008151, +0.004872] |
| GT − G | −0.001191 [−0.006564, +0.004070] | −0.001647 [−0.006289, +0.002907] |
| GTD − GT | +0.002906 [−0.000689, +0.006359] | +0.002039 [−0.001593, +0.005469] |
These are 10,000 paired resamples of five consecutive event-bearing days. Final coverage is 120 day proxies, or 24 five-day blocks—not 724 independent setups. All five registered paired comparisons, both weights, the wider 99.5% family intervals and iid-day sensitivity appear in ALL_RESULTS.md. No registered simultaneous interval establishes the required improvement. At this resolution the data do more than reproduce a tiny pilot: GT–geometry's 99.5% intervals are approximately [−0.00930,+0.00608] and [−0.00836,+0.00495], which do not include the predeclared −0.01 gain under this conditional-on-fit block model. They are not universal bounds on achievable history-model performance or a prospective power certificate.
Where the apparent gain comes from
The predeclared breakdowns do not show uniform transfer. GT–geometry is worse in Q3 (+0.004230 / +0.003191) and better in Q4 (−0.006435 / −0.006327). It is worse in the calendar roll-window proxy and better outside it. Long/short and source- hour comparisons also differ. These are diagnostics, not filters to deploy. All positive and negative groups are tabulated; no winning subset was selected.
Removing each comparison's five best recording days reverses GT's aggregate gain: versus frequency the deltas become +0.001955 / +0.001999; versus geometry they become +0.001892 / +0.000915. Roughly 23–25% of positive daily gains sit in those five days. This is not literally a single-day result, but it fails the frozen concentration requirement. GTD is worse even before any removals.
GT–geometry's session Brier difference decomposes into approximately −0.001149 for F, +0.000242 for A, effectively zero for B, and −0.000284 for N. Thus the small observed gain is not primarily an ambiguous-class trick; it also does not show a clear, stable directional edge.
Calibration, coverage and robustness
Final GT log loss is 0.832761 / 0.826735 versus frequency 0.837154 / 0.832250; GTD worsens it to 0.840330 / 0.830948. Those are secondary diagnostics, not a replacement winning objective. Final session-weighted GT probabilities average 47.28% favorable, 48.25% adverse, 0.10% ambiguous, 4.37% neither; observations are 48.80%, 47.14%, 0%, 4.07%. Mean agreement does not certify conditional calibration. All fixed-bin reliability tables for all models, splits and weights are in CALIBRATION.md; rare B outcomes cannot support a strong calibration assessment.
The full-60-bar sensitivity leaves the ranking unchanged: final session Brier is 0.538476 frequency, 0.537491 G, 0.536277 GT, 0.539207 GTD (722 observations, no refit). Assigning either of the two unknown final outcomes any possible class does not overturn the descriptive point ranking: GT−G session paired-loss bounds are [−0.001418,−0.001009], GTD−GT [+0.002741,+0.002978]. These are missing-label identification bounds, NOT sampling confidence intervals. Missing execution receipts were not used as exclusions and are not the blocker here.
Provenance and verification
The broader archive audit is in DATA_AUDIT.md. Three date- contaminated files were quarantined; 56 reserved/adjacent file references remained unopened. Exact contracts, rolls and timezone/arrival semantics for naive files remain unverified. We reset within files, use no inferred exchange-clock feature, and make only the explicit bar-start/bar-completion observation assumption.
“Final” means final within THIS study. The archive has prior strategy-research exposure; 2024 has known descriptor exposure, and 22 modeled final-period dates also appeared in the earlier tick-development inventory. That overlap is recorded in exposure_ledger.json. None is relabeled a globally untouched holdout. The native strategy itself was previously researched, so this chronology cannot remove earlier strategy-selection bias. Its reported uncertainty is conditional on this fitted model and archive, not on the full earlier research process.
Verification: 15 synthetic tests; 6,964 native-prefix comparisons; feature-prefix checks on all 2,923 events; independent vectorized re-labeling of every event; 48 independently recomputed aggregate scores; 484 source hashes rechecked. Scalar-recursive EMA agreement is within 1.03e-11 normalized units. Old-study hashes and frozen inputs remain unchanged. The initial exact-float assertion stop and its verification-only amendment are preserved in IMPLEMENTATION_LOG.md, with both freezes retained.
One CPU process, two numerical threads, no GPU. The final resumed invocation took 16.2 seconds with 341 MiB peak resident memory; that excludes earlier audit, cached reconstruction and independent verification, so is NOT an end-to-end benchmark. Research outputs remain about 10 MB, without raw archive copies.
Decision
Both frozen continuation gates fail. Stop this bounded experiment without transfer or a spline-specific addition. The question was executable as a bar-defined research test, and substantially more data did not validate the small pilot's suggestion. Do not blame missing fills for this predictive result.
Retain the causal cohort builder, explicit ambiguity/censoring ledger and scoring contract as reusable infrastructure. Any future idea must name a genuinely different mechanism and prospectively frozen test; changing horizons, selecting Q4/short trades or adding nonlinear splines now would be a new search, not a confirmation of this result. Nothing here establishes profitable admission, realistic order execution, or a spline-specific advantage.
View raw MD source
# Larger NQ bar-history test: continuation gate failed
The bounded experiment is complete. **Ordinary recent statistics did not
establish transferable incremental value, and the compact setup-development
description worsened final-period prediction.** Do not start a spline rescue
search or transfer experiment from this result. This is a finding about this
fixed setup, observation contract, representation and low-capacity model—not a
proof that all market history is useless.
## What was actually tested
2,923 native ABC two-minute C-failure opportunities across 482 event-bearing
recording days, reconstructed from 484 qualified 2024–2025 NQ files. The
within-file cohort contains 6,602,149 five-second bars and 275,068 complete
two-minute bars in 516 contiguous segments. All native opportunities in those
segments were retained, irrespective of simulated account occupancy. Missing
segments, cold starts and file-date cropping mean this is not every opportunity
of a fully observed, warm-started exchange-session stream.
Observation time is the close of the native trigger bar, not its earlier chart
timestamp. Reference is the latest completed close; boundaries are symmetric
at max(half pre-trigger ATR, one NQ point), over five minutes. First-bar outcomes
are favorable, adverse, both/ordering ambiguous, or neither. Unknown future
gaps and recording ends are not labeled neither. These are **bar-path classes,
not achievable P&L or execution outcomes**.
| Split | Files | Event days | Opportunities | Resolved | F / A / B / N |
|---|---:|---:|---:|---:|---|
| Development, 2024 | 248 | 248 | 1,512 | 1,505 | 713 / 732 / 2 / 58 |
| Validation, first half 2025 | 115 | 114 | 685 | 682 | 318 / 334 / 3 / 27 |
| Study-final, second half 2025 | 121 | 120 | 726 | 724 | 350 / 346 / 0 / 28 |
Twelve observations remain unresolved: nine recording-end censorings, two gaps
before resolution and one fixed day-boundary purge. All are retained in the
ledger and paired-loss bounds. Five genuinely ambiguous outcomes remain explicit
in development/validation; no final observations happen to be ambiguous.
The fixed comparison uses frequency; G geometry (8 features); GT geometry plus
ordinary recent bar statistics (14); and GTD with four additional setup-path
summaries (18). The three fitted models share logistic ridge, preprocessing and
training weights. Only three fits were performed; no validation selection,
validation refit, optimizer sweep, new target, spline basis or ticker search.
## Results
Unnormalized four-class Brier, lower is better. “Session” below means equal
**recording-day proxy** weight, not independently verified exchange-session
weight. Opportunity weighting is reported alongside it, not selected afterward.
| Model | Validation session | Validation opportunity | Final session | Final opportunity |
|---|---:|---:|---:|---:|
| Frequency | 0.539650 | 0.541251 | 0.538502 | 0.536633 |
| Geometry | 0.541843 | 0.542351 | 0.537364 | 0.536659 |
| Geometry + recent statistics | 0.547866 | 0.547821 | 0.536172 | 0.535013 |
| Above + setup development | 0.547662 | 0.545372 | 0.539079 | 0.537052 |
The more detailed representations improved training fit, not consistently later
prediction. All three lost to frequency on validation. On study-final GT's
absolute gain over frequency is only 0.002329 / 0.001621 (session/opportunity),
well below the frozen 0.01 continuation threshold. GTD is worse than GT under
both final weightings.
| Final paired difference | Session delta [95% block interval] | Opportunity delta [95% block interval] |
|---|---|---|
| GT − frequency | −0.002329 [−0.009979, +0.005175] | −0.001621 [−0.008151, +0.004872] |
| GT − G | −0.001191 [−0.006564, +0.004070] | −0.001647 [−0.006289, +0.002907] |
| GTD − GT | +0.002906 [−0.000689, +0.006359] | +0.002039 [−0.001593, +0.005469] |
These are 10,000 paired resamples of five consecutive event-bearing days. Final
coverage is 120 day proxies, or 24 five-day blocks—not 724 independent setups.
All five registered paired comparisons, both weights, the wider 99.5% family
intervals and iid-day sensitivity appear in [ALL_RESULTS.md](ALL_RESULTS.md).
No registered simultaneous interval establishes the required improvement.
At this resolution the data do more than reproduce a tiny pilot: GT–geometry's
99.5% intervals are approximately [−0.00930,+0.00608] and
[−0.00836,+0.00495], which do not include the predeclared −0.01 gain under this
conditional-on-fit block model. They are not universal bounds on achievable
history-model performance or a prospective power certificate.
## Where the apparent gain comes from
The predeclared breakdowns do not show uniform transfer. GT–geometry is worse in
Q3 (+0.004230 / +0.003191) and better in Q4 (−0.006435 / −0.006327). It is worse
in the calendar roll-window proxy and better outside it. Long/short and source-
hour comparisons also differ. **These are diagnostics, not filters to deploy.**
All positive and negative groups are tabulated; no winning subset was selected.
Removing each comparison's five best recording days reverses GT's aggregate
gain: versus frequency the deltas become +0.001955 / +0.001999; versus geometry
they become +0.001892 / +0.000915. Roughly 23–25% of positive daily gains sit in
those five days. This is not literally a single-day result, but it fails the
frozen concentration requirement. GTD is worse even before any removals.
GT–geometry's session Brier difference decomposes into approximately −0.001149
for F, +0.000242 for A, effectively zero for B, and −0.000284 for N. Thus the
small observed gain is not primarily an ambiguous-class trick; it also does not
show a clear, stable directional edge.
## Calibration, coverage and robustness
Final GT log loss is 0.832761 / 0.826735 versus frequency 0.837154 / 0.832250;
GTD worsens it to 0.840330 / 0.830948. Those are secondary diagnostics, not a
replacement winning objective. Final session-weighted GT probabilities average
47.28% favorable, 48.25% adverse, 0.10% ambiguous, 4.37% neither; observations
are 48.80%, 47.14%, 0%, 4.07%. Mean agreement does not certify conditional
calibration. All fixed-bin reliability tables for all models, splits and weights
are in [CALIBRATION.md](CALIBRATION.md); rare B outcomes cannot support a strong
calibration assessment.
The full-60-bar sensitivity leaves the ranking unchanged: final session Brier
is 0.538476 frequency, 0.537491 G, 0.536277 GT, 0.539207 GTD (722 observations,
no refit). Assigning either of the two unknown final outcomes any possible
class does not overturn the descriptive point ranking: GT−G session paired-loss
bounds are [−0.001418,−0.001009], GTD−GT [+0.002741,+0.002978]. These are
missing-label identification bounds, NOT sampling confidence intervals. Missing
execution receipts were not used as exclusions and are not the blocker here.
## Provenance and verification
The broader archive audit is in [DATA_AUDIT.md](DATA_AUDIT.md). Three date-
contaminated files were quarantined; 56 reserved/adjacent file references remained
unopened. Exact contracts, rolls and timezone/arrival semantics for naive files
remain unverified. We reset within files, use no inferred exchange-clock feature,
and make only the explicit bar-start/bar-completion observation assumption.
“Final” means final within THIS study. The archive has prior strategy-research
exposure; 2024 has known descriptor exposure, and 22 modeled final-period dates
also appeared in the earlier tick-development inventory. That overlap is recorded
in `exposure_ledger.json`. None is relabeled a globally untouched holdout. The
native strategy itself was previously researched, so this chronology cannot
remove earlier strategy-selection bias. Its reported uncertainty is conditional
on this fitted model and archive, not on the full earlier research process.
Verification: 15 synthetic tests; 6,964 native-prefix comparisons; feature-prefix
checks on all 2,923 events; independent vectorized re-labeling of every event;
48 independently recomputed aggregate scores; 484 source hashes rechecked.
Scalar-recursive EMA agreement is within 1.03e-11 normalized units. Old-study
hashes and frozen inputs remain unchanged. The initial exact-float assertion
stop and its verification-only amendment are preserved in
[IMPLEMENTATION_LOG.md](IMPLEMENTATION_LOG.md), with both freezes retained.
One CPU process, two numerical threads, no GPU. The final resumed invocation
took 16.2 seconds with 341 MiB peak resident memory; that excludes earlier audit,
cached reconstruction and independent verification, so is NOT an end-to-end
benchmark. Research outputs remain about 10 MB, without raw archive copies.
## Decision
Both frozen continuation gates fail. **Stop this bounded experiment without
transfer or a spline-specific addition.** The question was executable as a
bar-defined research test, and substantially more data did not validate the
small pilot's suggestion. Do not blame missing fills for this predictive result.
Retain the causal cohort builder, explicit ambiguity/censoring ledger and scoring
contract as reusable infrastructure. Any future idea must name a genuinely
different mechanism and prospectively frozen test; changing horizons, selecting
Q4/short trades or adding nonlinear splines now would be a new search, not a
confirmation of this result. Nothing here establishes profitable admission,
realistic order execution, or a spline-specific advantage.