Evaluation practice · E22 · Evaluation audit

Ten thousand bootstrap draws are still twenty observed days

A small resampling notebook illustrates the difference between Monte Carlo precision and the amount of independent evidence.

NumPy random indexingSciPy statisticsNotebook summaries
A large synthetic resampling array still comes from a small observed sample; the additional draws are not new historical years.
Figure 1. 10,000 draws. Twenty observed days.. A large synthetic resampling array still comes from a small observed sample; the additional draws are not new historical years. Source sizes; illustrative resampling. Original vector illustration.

Follow the information

From input to outcome

The sampler selects indices into the same small empirical sample. More synthetic draws reduce Monte Carlo noise; they do not create independent historical sessions.

The sampler selects indices into the same small empirical sample. More synthetic draws reduce Monte Carlo noise; they do not create independent historical sessions.
Figure 2. Information flow. Solid arrows carry observations, tensors or artifacts; other routes are explicitly labelled. Signal shapes, matrices and network icons are schematic, not measured samples or literal neuron counts. Open full-size SVG ↗ On narrow screens, scroll the diagram horizontally.

Read this alongside Figure 1: A large synthetic resampling array still comes from a small observed sample; the additional draws are not new historical years. The module map and layer-level figures below expand the operations in this route.

Ten thousand bootstrap draws are still twenty observed days: system and evaluation mapObserved daily values: Small historical sample → IID index sampler: 10,000 × 252 draws → Synthetic annual sums: Repeated empirical values → Tail summaries: Conditional on IID model → Interpretation: Not new observed years. A high-level module map; comparison branches and training details are explained in the article.EVALUATION PRACTICE / E22 / MODULE MAP01 INPUTObserved daily valuesSmall historical sample02 MODULEIID index sampler10,000 × 252 draws03 MODULESynthetic annual sumsRepeated empirical values04 MODULETail summariesConditional on IID model05 OUTPUTInterpretationNot new observed years
Source-grounded module map. Boxes summarize operations, not individual neurons; comparison arms and training paths are detailed below. On a small screen, scroll the diagram horizontally.
Observed daily values — Small historical sample

The architecture in context

What this comparison asks

The notebook repeatedly samples daily values with replacement, adds 252 of them and summarizes the distribution of synthetic annual totals. This is a compact way to explore a specified resampling model. It is not a way to manufacture more independent market history or establish that future days follow the same distribution.

Who does what in the stack

NumPy random indexing
Implements the empirical IID resampling model.
SciPy statistics
Evaluates distribution tails.
Notebook summaries
Report conditional simulation statistics, not forecasts.

NumPy makes the full simulation matrix explicit; SciPy supplies a t-distribution calculation elsewhere in the notebook. The engineering lesson is to label exactly what randomness was introduced. Repeated bootstrap draws reduce simulation noise conditional on the empirical sample, not uncertainty about whether that sample represents a new regime.

Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph.
Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph. Open full-size SVG ↗

Open up the implementation

Open the bootstrap tensor

A concrete operation-level view of this implementation; no unobserved neural architecture is implied.
A concrete operation-level view of this implementation; no unobserved neural architecture is implied. Open full-size SVG ↗

Resampling conditions on the empirical distribution of the observed days. More draws reduce Monte Carlo error in the synthetic quantiles but do not reduce uncertainty caused by an unrepresentative sample or serial dependence. The IID assumption is doing real work in the annual aggregation.

The mathematical contract

Rb∗=∑j=1252rIbj,Ibj∼Uniform⁡1,…,nR_b^*=\sum_{j=1}^{252}r_{I_{bj}},\qquad I_{bj}\sim\operatorname{Uniform}{1,\ldots,n}

Blocked resampling may address some dependence, but choosing block length needs a defensible protocol and does not conjure missing regimes. The notebook’s absolute-value t-statistic in a directional tail calculation also requires correction before interpreting significance.

Implementation and resource card

Capacity / budget
10,000 simulation rows are not 10,000 independent observed years. Memory is proportional to simulations × sampled days.
Execution evidence
This revision inspects and explains the archived implementation. It does not rerun the original workload. No unrecorded convergence time, throughput or accelerator result is supplied.
Current reproduction context
Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.

From explanation to a reproducible check

Use an all-negative sample and check that a directional profitability test does not return a misleading small positive-direction p-value. Distinguish a numerical simulation check from a test of the empirical assumptions.

Preserve input identities, configuration and failure records with the result. A successful numerical check only establishes the operation it exercises: it does not certify an entire dataset, model or deployed system. Reproduce the interface on a small deterministic input before optimizing throughput or increasing workload size.

A closer look at the implementation

The code that carries the idea

The index matrix has one row per simulation and one column per resampled day. Advanced indexing gathers observed values, then summing across columns produces a hypothetical year. The saved example uses only a small daily sample, so dependence and unobserved regimes dominate any interpretation of its tails.

Python · cell 1 · lines 30–46
    simulations = 10000
    sim_days = 252 # Trading days in a year
    
    # Generate random indices
    indices = np.random.randint(0, n, size=(simulations, sim_days))
    # Map indices to values
    bootstrapped_years = data[indices]
    # Sum each year
    yearly_pnls = np.sum(bootstrapped_years, axis=1)
    
    loss_prob = np.sum(yearly_pnls < 0) / simulations * 100
    worst_5_pct = np.percentile(yearly_pnls, 5)
    avg_yearly = np.mean(yearly_pnls)
    
    print(f"\n--- 2. BOOTSTRAP (1 year projection) ---")
    print(f"Avg Yearly PnL: {avg_yearly:.2f}")
    print(f"Prob of Loss:   {loss_prob:.2f}%")

Verbatim archive excerpt from analyze_significance.ipynb. Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.

The boundary that matters

The separate t-test code uses the absolute t statistic with a one-tail survival function. That is not a valid directional positive-mean p-value for negative t. More generally, an IID bootstrap omits serial dependence; block resampling would be a different declared model, not a retroactive guarantee.

Keep building

Other posts of interest