Research record

Consolidated limitations and research boundaries

Historical source. Some claims in older records were subsequently corrected. The associated article states the adopted interpretation. This record preserves the original source alongside its rendered reading view.

Rendered archival TeX

This is an HTML reading rendition of the local TeX record. Mathematical notation is rendered with KaTeX; archived figures are included when their source assets are part of this collection.

Limitations and honest boundaries

A central commitment of this work is to state where the approach does not win, where a claim is mechanism-only and not yet demonstrated at scale, and where a baseline we do not beat is the honest reference point. The SSP interpretation (Section [sec:theory-ssp]) is a conditional modeling hypothesis, not a theorem predicting the success or failure of every feature map. The earlier universal risk dichotomy is withdrawn: nontrivial operators alone do not guarantee improvement, and ordinary ridge is not automatically sparsity-adaptive. Reported experimental numbers are retained under their stated protocols.

Correction of the historical no-forgetting and privacy interpretation.

Exact fixed-feature pooled ridge preserves objective contributions, not each old task's optimal prediction. Conflicting labels give an elementary counterexample even with a perfectly accumulated Gram. Older descriptions of ``no forgetting by construction'' must therefore be read as either empirical results on their stated protocols or the separate, scoped preservation of an immutable program/protected support; they are not a general consequence of Eq. eqref(eq:joint-opt). Floating-point sums are not bitwise associative. Likewise, no raw replay does not establish privacy: a one-record scalar Gram and cross-statistic can reconstruct that record. Exact subtraction requires the deleted contribution, fixed preprocessing/features and consistent regularization, and does not unlearn a pretrained backbone. New nonlinear features require historical cross-statistics that the previous Gram need not contain. Five executable counterexamples/regressions accompany this correction. The enlarged research program also includes ordinary gradient-trained neural and SAC experiments; a gradient-free inner solver does not make the complete pipeline gradient-free. Historical numerical comparisons are retained, not promoted to universal or current-SOTA guarantees.

The singularity telescope is a reduced analogue, not a Navier–Stokes result.

Similarity-coordinate compilation and guarded map fitting are tested on an analytic two-dimensional divergence-free field with the published anisotropic scaling, not on the full three-dimensional construction. The experiments do not establish blowup, validate the presented proof, or show that ordinary turbulence follows the same coordinates. Within the reduced analogue, the theory-constrained two-variable model does support a calibrated short horizon: two independent 40-stream panels at 3% noise have no error above 10% at τ=0.012\tau=0.012, including continuously de-aliased event times, schedules, and exponents. This does not transfer automatically across sensing regimes. With 33233^2 sensors, 19/20 streams fail even at 0.0005 lead under 5% noise and all 20 fail under 8%; scalar ridge does not repair the floor. Raising the sensor grid to 49249^2 restores 40/40 safety at 5% noise, but increases median CPU time to 15.59 seconds and has occasional 40–54 second refinement tails. Transfer to an executable three-dimensional leading profile, real-time batched execution, and validation on external fluid data remain open. A rank-two profile prior plus multiresolution map search does recover and exceed the dense arm's median accuracy at 33233^2, but one of 40 fresh streams reaches 10.17% error (Wilson upper unsafe risk 12.88%) and median CPU time is 53.25 seconds. Validation-selected snapshot centering subsequently removes that observed tail on 40 further streams: median/p90/worst error becomes 5.71/6.57/7.45% at 33233^2, with 40/40 safe and 15.79-second median CPU time. This closes the sparse-sensing gate for the analytic generator, but is not yet a safety certificate beyond it: the centroid candidate set, rank-two profile, and similarity manifold are matched to this construction. External three-dimensional JHTDB transfer is now tested for reconstruction, but not the full leading-profile construction: exact divergence residuals improve instantaneous reconstruction while worsening particle trajectories on all six dynamic streams. Thus operator satisfaction cannot be substituted for a downstream rollout metric.

Darcy energy assimilation is not strong-form closure.

On 134 prospective PDEBench Darcy fields, conductivity-weighted cardinal assimilation with 32 sensors lowers blind and weighted-gradient error relative to a selected DCT residual. It nevertheless has larger mismatch under our finite-volume strong operator (1.356 versus 0.936), because the dataset's exact interface/discretization convention remains unidentified. We therefore claim sparse physical-energy evidence propagation, not recovery of the hidden generator, exact flux conservation, or a blind coefficient-to-solution solver.

The nano-drone result is a systems frontier, not SOTA accuracy.

The real-robot transaction is prospective and independently admitted, but it uses Melon runs 1–2 for target adaptation and confirms on run 3. The original Physics+Residual and newer ASIA results train only on Square/Random/Chirp and report the complete held-out Melon trajectory. Our confirmation errors are numerically below ASIA in all four cumulative groups, yet these protocols are not interchangeable. Gate 199 closes this gap with an adequate pinned ASIA reproduction and matched sparse target adaptation. ASIA then beats the operator in all four run-3 cumulative groups by 20.3–47.1%. The operator's supported advantage is systems-level: 1,020×\times less program storage, 65.9×\times faster compilation, fewer target labels, and exact transactional state. A subsequent typed multi-execution consolidation narrows the mean physical ratio gap to 12.2% and removes its recurrent teacher from runtime, but still loses all four absolute groups to adapted ASIA. This evidence is post-exposure on the same three Melon flights and is not a prospective SOTA result. Gate 198 further rejects universal direct-flow dominance: translation favors the recursive short-step model even while direct prediction sharply improves long-horizon rotation. No MCU energy or closed-loop control result is claimed.

The corrected KUKA transaction is post-exposure.

An input-only audit found that the first KUKA analysis paired adjacent target blocks even though repeated input programs are (0,3),(1,4),(2,5)(0,3),(1,4),(2,5). The earlier claim that every sparse update harmed its paired repeat is retracted. The source-only ridge selector had made the same pairing error; correcting it before the target phase changes the ridge from 100 to 0.01. Sparse then improves every repeat and retains 97.45% of pooled dense gain with 3.189×\times fewer labels; all pairwise retentions exceed 89% and the frozen mechanism gate passes. Because target outputs were already inspected, this is a scientific correction, not fresh prospective confirmation.

Continuous zero-forget growth is conditional, not universal.

The coupled lifelong operator network gives exact noninterference because typed experiments isolate scalar edges, the missing laws are sparse cardinal atoms, and every accepted task declares an operating interval whose complete basis support is frozen. This is a strong continuous-function guarantee, but it does not apply to arbitrary entangled deep features. Grid-aligned one-atom defects also make the present localization problem favorable. An actual law change inside a protected interval is a conflict and requires explicit versioning or unlearning; silently adapting it would contradict the guarantee. Off-grid defects, inferred region bounds, multi-atom growth, and external measurements remain necessary before claiming a practical lifelong digital twin.

Hermite-score coordinate discovery currently assumes a known score.

The multi-index result uses Gaussian whitening and the exact third Hermite/Stein identity. It requires smooth ridge components with nonzero third Hermite coefficients, a specified rank, several passes over evidence, and currently a tensor or contracted-tensor operator. It does not yet apply unchanged to unknown non-Gaussian laws, even-symmetric components invisible at third order, or arbitrary entangled dynamics. Distribution-specific score operators, multiple Hermite orders, or a validated transport to Gaussian coordinates are needed for those regimes.

The compact ParWH compiler is not accuracy state of the art.

Its prospective stationary multisine result tracks internal development, but the frozen growing-amplitude result fails badly. The later Hermite-tail rule is developed only on estimation-data amplitude holdouts and is not rescored on the opened official target. Its replicated 16.5–24.0% advantage over clipping is therefore an extrapolation-mechanism result, not a repaired official benchmark number. The common-pole SVD also recovers branch count but not the physical input/output pole allocation. A fresh drift benchmark and shift-coupled branch factorization are required before claiming general nonlinear-system identification or field-level superiority.

Hermite continuation is an extrapolation hypothesis, not a safety certificate.

On input-only F-16 coordinates, affine cardinal tails are robust across three guarded boundaries and improve every prospectively evaluated channel, but the official mean gain is 7.64%, below the frozen 15% target. On recurrent Cascaded Tanks dynamics, the same flexibility can compound error: an exact cardinal residual lowers quiet-regime error yet worsens held-overflow RMSE from 0.5400.540 V for its latent physical core to 1.3561.356 V. Clipping the recurrent edge while retaining Hermite feedforward edges is worse still, so graph topology alone cannot select a true exterior law. A practical lifelong model must verify proposals on every affected event class and preserve an immutable core for exact rollback. The older reported 0.550.55 V two-tank latent result is retrospective: its code accessed the test maximum and initial output; correcting the initial value leaves 0.5510.551 V but does not restore prospective status.

Zero forgetting is conditional on a frozen support map.

Compact support alone is insufficient when normalization or learned coordinates move historical samples across basis regions. Our immutable coordinate envelope certifies exactly the finite archived Levels-1/3 samples, not their full generating distribution. The affected-support credit rule and quadratic step guarantee empirical non-degradation on the four retained Level-5 verifier blocks; they are not distribution-free safety guarantees. Moreover, the 7.29% later-regime gain is retrospective because Level 7 was already inspected. A fresh streaming benchmark must freeze representation, envelope, event partition, energy/credit thresholds, and future regimes before the claimed safe-plasticity composition is prospectively tested.

The fixed global JHTDB flow-map law does not transfer across velocity fields, and a one-prefix law has a finite temporal trust horizon. Receding compilation with a fixed one-eighth trust step subsequently passes on six new nine-frame streams (41/42 decisions, 3.90% median gain), and transfers from observed to disjoint tracers on exposed fields. A second fully prospective fields-plus- particles panel gives 39/42 strict wins and 42/42 non-regressions, but formally misses its per-stream strict-win gate because one stream invokes exact identity three times. These remain small tests from one DNS dataset, one sensor layout, and one tracker-noise model. Coordinate-noise transfer is now positive: on six further sealed streams, a 256-track compiler obtains 40/42 wins and 3.60% median gain; over eight fresh particle/noise populations the operator-smoothed version obtains 320/336 wins, 100% safety, and 3.59% median gain. However, the single-population prospective smoothing advantage misses its frozen margin, and the multi-population confirmation necessarily reuses exposed fields. Two-view physical-space fusion plus circulant smoothing subsequently passes on six new streams with 40/42 wins and 3.985% median gain using 128 readings per time, versus 3.620% from a 256-reading single view. This factor-of-two measurement result assumes zero-mean Gaussian coordinate noise; a symmetric cross-Gram errors-in-variables proposal does not beat ordinary averaging. Multi-population stress tests retain the 3% target through correlation 0.5 and 25% secondary-view dropout, but fail it at correlation 0.75 and 50% dropout. Inverse-variance weighted Grams cross the exposed 50%-availability threshold but fail on fresh populations, showing that known sensor variance alone does not determine forecast utility. Evidence-triggered nested acquisition is more robust: a frozen 5% confidence policy nearly halves mean sensing and matches or improves dense coverage/safety on fresh populations. On a different field panel it still halves sensing and ties dense wins, but has one unsafe dense escalation and formally misses its gain margin. The current validation split therefore decides acquisition, not safe deployment; a common untouched event block and transactional rollback are necessary. A subsequent retrospective mechanism test and fresh-population confirmation implement this split: the frozen transactional compiler uses 2.21-times fewer measurements and exceeds its dense-240-plus-16-holdout control in wins, safety, and median gain. This is still one DNS dataset with exposed fields; the admission block is only 16 particles and certifies empirical local events, not future distributions. Across three panels the policy improves pooled wins and safety while halving sensing, but the difficult Gate-157 panel misses its per-population and dense- median clauses. The aggregate must not erase that negative result. A subsequently acquired one-shot field panel strengthens both sides of the boundary: transactional sensing is perfectly safe, exceeds 4% median gain, and uses 2.56-times fewer readings, yet trails dense by 0.529 points and two wins. Thus neither the pooled result nor low-sensor confidence establishes universal dense-quality equivalence. Retrospective value-of-information modeling does not yet close this gap. Nonlinear candidate-score signatures carry a repeatable but weak ranking signal (Gate-181 Spearman about 0.30); changing the target to admitted-transaction value and adding 16 passive sentinel observations both fail the frozen policy criteria. A same-budget oracle shows large effect-size headroom, so these failures diagnose missing decision information rather than an adequate learned router. Branch-local reuse of the sentinel improves coverage to 249/252 at 2.57-times fewer readings, but still closes only 8.5% of the dense-median gap. It is a retrospective mechanism result, not a prospective adaptive-sensing claim. The next test must use a pilot-induced program change, a disjoint admission block, and a newly sealed panel; thresholds or classifiers tuned on the current receipt would not answer the scientific question. The explicit 64/8/8/176 pilot/verifier test also fails: despite positive training-panel ranking, it gives 247 wins and 4.021% at 99.75 readings. We therefore close deterministic receipt routing on this exposed panel. A valid continuation needs randomized exploration that observes action value online, or a second physical system; another deterministic model is not independent evidence. Pure random acquisition also fails as a complete policy. Replacing only eight of 28 branch-safe acquisitions by random transactions is much less costly: across 1,000 trials every decision is safe and 99.9% retain at least 248 wins. This calibrates a retrospective exploration allowance; it does not demonstrate that an online contextual ledger uses those labels effectively or transfers to new fields. Indeed, the frozen 8/20 mixture fails on a new sealed panel: it is perfectly safe at 2.57-times fewer readings but reaches only 2.959% median and 2/6 population medians above 3%. Its deterministic control reaches 3.107%; the old confidence policy has a −12.208%-12.208\% unsafe event. Exploration-cost transfer is therefore rejected; only transactional safety transfers. The acquired labels cannot be used retrospectively to relabel this gate as online learning. These are empirical boundaries on exposed fields, not guarantees or a prospective independent-system result. Delayed or biased tracking; heavier or structured missingness; partial state matching across moving sensors; stronger online data-assimilation baselines; much longer operation; and an independent physical dataset are required before claiming a general fast-adapting digital twin. The 32-tracer and two-transition-pooling failures also show that exact sufficient-statistic memory cannot replace current spatial coverage or justify stale evidence under drift.

Sub-cubic score sketches need a weak-mode certificate.

The matrix-free contraction experiment shows that large predictive gains can coexist with a missed physical direction. Its fixed 32-by-24 sketch lowers storage but does not certify that all low-energy tensor components survive. The next method must compare independently accumulated score subspaces and grow evidence or range only when their weakest principal direction is reproducible; until then, the full-tensor confirmation is a mechanism result rather than a scalable general discovery engine.

Circuit tangent gains and transfer are post-exposure.

The 0.364-mV held-out transfer-function-jet frontier is obtained only after the official circuit target had been opened for Gate 215; it is source-held-out mechanism evidence, not a new prospective benchmark score. A complete-source frozen refit reaches 0.367 mV when replayed on that target, close to its 0.338-mV source score but still above the reported 0.241-mV deep encoder. This supports transfer of the compact mechanism, not prospective SOTA. The next extension must be selected on disjoint source evidence and ultimately tested on a fresh physical record; target-driven residual fitting would invalidate the claim.

Broadband natural-signal fitting: closed form loses to deep SGD.

Our operator-matched, closed-form principle is decisive on structured signals but not on broadband ones, and implicit neural representations make the boundary sharp. Fitting an image as a coordinate-to-color map, a deep SGD-trained SIREN reaches roughly 6060 dB PSNR, whereas a closed-form random-Fourier-feature ridge reaches about 2828 dB and a closed-form representation matched to the dominant modes does worse still (∼ ⁣10\sim\!10 dB, underfitting). The closed-form solver is far faster (over two orders of magnitude) but a shallow one-layer feature ridge cannot match a deep network's learned-frequency fidelity on broadband natural content. This is exactly the SSP prediction: when the generating operator is trivial and the innovation is dense, broad frequency coverage beats energy-concentrated matched modes, and deep learned features beat both. Matched closed-form wins on structured signals (dynamical systems, PDEs); it does not win on broadband natural images, and we do not claim otherwise. The sharper statement is that the boundary is per factor, not global: a signal can be broadband in one factor and structured in another, and the right design matches structure factor-by-factor. Dynamic 3D Gaussian splatting is a clean example—the appearance is broadband (kept as an expressive learned Gaussian field) while the deformation is a smooth function of continuous time; replacing the time-deformation MLP with a closed-form continuous-time (liquid) cell that bakes the time structure into the architecture matches or beats the MLP at lower compute, with gains concentrated exactly on the high-frequency, structured motion [li2026cfc,hasani2022cfc]. The lesson we carry is to structure the factors that are structured and leave the broadband factors to learned features, rather than choosing closed-form versus deep wholesale.

Atari mastery via behavior cloning fails: imitation is not reinforcement learning.

We balance and play simple games from pixels with no backpropagation, but scaling to Atari-style mastery by cloning a planning teacher does not work. Even when gradient-free perception clones the teacher with reasonable per-frame accuracy (∼ ⁣0.79\sim\!0.79), the resulting policy scores at the floor (−21-21, losing every point), a textbook behavior-cloning failure: average action accuracy does not imply competent closed-loop play, because the small fraction of mispredicted actions compounds over a trajectory and the agent never sees the states its own errors lead to. Closing this requires interactive reinforcement learning at scale, not imitation, and we have not demonstrated that.

From-scratch LLMs at full scale: mechanism proven, scale is a compute question.

We trained a real autoregressive character-level GPT (token and positional embeddings, causal transformer blocks, and an LM head) by the per-block local sweep with no global backward pass; it reaches a validation bits-per-character within about 0.0130.013 of global backpropagation and generates coherent corpus-style text. This settles the mechanism: predictive-coding-style local learning matches backpropagation on a working language model, not merely on a probe task. It does not settle scale. Training a GPT-2/3-scale model gradient-free is a compute and cluster-engineering question that we have not run; we claim the mechanism, not a scaled result.

Evolution strategies are a niche tool, not a universal booster.

We examined whether evolution strategies / PGPE could improve the framework broadly and concluded that they cannot. On a low-dimensional control toy, a simple genetic algorithm trivially solved the task while our quick ES stalled on the flat, deceptive near-random reward landscape — the same deceptive-landscape failure that motivated novelty search [conti2018novelty]. A properly tuned ES is known to solve such toys but offers no advantage over the genetic algorithm there. The honest verdict is that ES/PGPE fits a narrow slot — high-dimensional, smooth-landscape, model-free, non-differentiable optimization — and is not a general accelerator for local or closed-form learning.

Self-improvement is bounded: no free lunch from self-reference.

We map self-improvement precisely (Section [sec:res-selfimprove]) and the map is as much a boundary as a result. Self-labelling and self-consistency-style loops plateau—they add no information the model does not already contain, and a self-derived verifier fails outright because its errors are correlated with the model's own. Reinforcement learning does not rescue this within our framework: the closed-form memory's stability requires stationary targets, and RL's bootstrapped, policy-dependent targets are not, so it is no more stable than gradient TD. Real compounding needs an external verifier, and even then it lifts only the knowledge/readout layer to its ceiling—improving the representation itself, or bootstrapping from a near-chance start (below a viability threshold, where there is nothing correct to verify), still requires external teaching. Our real-LLM demonstration is on arithmetic, the canonical cheap-verifier domain, and lifts to a plateau via retrieved verified exemplars rather than by changing weights; scaling to code (unit-test verifiers) or proof at LLM scale is scoped but not run.

Predictive coding does not, by itself, remove weight transport.

We are careful not to overclaim biological plausibility. Classical predictive coding uses the transposed forward weights (W⊤W^\top) for its feedback, i.e.\ it assumes weight transport; its actual claim is the absence of a global backward pass with local error signals, not the absence of weight transport. Removing weight transport is the separate concern of random-feedback methods [lillicrap2016randomfeedback,nokland2016directfeedbackalignment], which we evaluate separately and which carry their own accuracy cost. When we report no-global-backward results and no-weight-transport results, we keep them distinct and do not conflate the two guarantees.

Single-task accuracy versus tuned backprop.

In early experiments, single-task accuracy of the gradient-free stack sat below tuned backpropagation, and on some architectures (e.g.\ convolutional stacks) local learning matches backprop only to within one to two points rather than cleanly beating it. The gap has since been closed on vision through deep local perception, but we flag the boundary honestly: our headline claim is a unified, closed-form, no-forgetting principle that is decisive where structure is known and ties strong backprop baselines at one-to-three orders of magnitude less compute where structure is absent — not that it beats every accuracy number in every single-task setting.

Deliberately deferred to cloud scale: the open benchmark backlog.

The experiments in this paper were run on a single workstation, which caps demonstrations at CIFAR/smallNORB scale and at continual learning over frozen pretrained backbones. We separate mechanism, which we demonstrate, from frontier-scale absolute numbers, which require multi-GPU compute we have not yet provisioned. To make the boundary explicit — and to mark these as planned, scoped work rather than oversights — we list the deferred benchmarks; each runs the same gradient-free substrate (per-block local sweep, closed-form Gram memory, topology evolution), with scripts that auto-detect CUDA and need only big-batch/epoch configs and dataset download on the cloud box.

We report these so that the reader can distinguish ``not shown because it fails'' from ``not shown because it needs a cluster we name and have scoped.''

Nominal VOI is not yet a safe experiment policy.

The hysteresis acquisition experiments are prospective but synthetic and local. A directional symmetric-Gram probe improves the intended damping coefficient without uniformly improving displacement, while a constrained waveform mixture preserves broadband performance but fails robustly across four nearby laws. We have therefore not demonstrated closed-loop safe self-directed experimentation or a measured-hardware advantage. The next claim requires risk-sensitive covariance over archived operator versions, explicit actuation cost and safety constraints, and a physical prospective panel. Whole-transaction simulation gives the correct conservative direction on a subsequent panel, but its 19.92% worst-risk gain misses the frozen 20% margin and redundant endpoint normalization breaks bitwise waveform identity. This is promising mechanism evidence, not a certified or hardware-validated acquisition policy. An eight-system follow-up passes the sensing-policy clauses with a 92.61% worst-risk advantage, but the complete controller still misses one independent model-adequacy ceiling; its successful repair uses exposed validation. Thus the two transactional layers are first demonstrated separately. A subsequent frozen eight-system run composes them prospectively: four certificate failures route to bounded monotone relinearization, all close, and all unseen errors remain below 0.001273 mm while the acquisition planner safely retains generic. This is still synthetic evidence under one Bouc–Wen family. No real actuator, sensor noise model, energy comparison, safety envelope, or open equation vocabulary has yet validated the complete loop. We now have a deterministic interactive MuJoCo receipt with exact plant and wrapper-state rollback, but it contains no learned correction or control-policy comparison. It should not be interpreted as closing the external-actuation gap. The gauge-fixed MuJoCo cardinal compiler is likewise only an adaptation-set algebra result until disjoint trajectories are opened.

Original: paper/v2_sections/06_limitations.tex · Raw source file

View raw TEX source
\section{Limitations and honest boundaries}\label{sec:limits}

A central commitment of this work is to state where the approach does not win, where a claim is mechanism-only and not yet demonstrated at scale, and where a baseline we do not beat is the honest reference point. The SSP interpretation (Section~\ref{sec:theory-ssp}) is a conditional modeling hypothesis, not a theorem predicting the success or failure of every feature map. The earlier universal risk dichotomy is withdrawn: nontrivial operators alone do not guarantee improvement, and ordinary ridge is not automatically sparsity-adaptive. Reported experimental numbers are retained under their stated protocols.

\paragraph{Correction of the historical no-forgetting and privacy interpretation.}
Exact fixed-feature pooled ridge preserves objective contributions, not each
old task's optimal prediction. Conflicting labels give an elementary
counterexample even with a perfectly accumulated Gram. Older descriptions of
``no forgetting by construction'' must therefore be read as either empirical
results on their stated protocols or the separate, scoped preservation of an
immutable program/protected support; they are not a general consequence of
Eq.~\eqref{eq:joint-opt}. Floating-point sums are not bitwise associative.
Likewise, no raw replay does not establish privacy: a one-record scalar Gram
and cross-statistic can reconstruct that record. Exact subtraction requires
the deleted contribution, fixed preprocessing/features and consistent
regularization, and does not unlearn a pretrained backbone. New nonlinear
features require historical cross-statistics that the previous Gram need not
contain. Five executable counterexamples/regressions accompany this correction.
The enlarged research program also includes ordinary gradient-trained neural
and SAC experiments; a gradient-free inner solver does not make the complete
pipeline gradient-free. Historical numerical comparisons are retained, not
promoted to universal or current-SOTA guarantees.

\paragraph{The singularity telescope is a reduced analogue, not a Navier--Stokes result.}
Similarity-coordinate compilation and guarded map fitting are tested on an
analytic two-dimensional divergence-free field with the published anisotropic
scaling, not on the full three-dimensional construction.  The experiments do
not establish blowup, validate the presented proof, or show that ordinary
turbulence follows the same coordinates.  Within the reduced analogue, the
theory-constrained two-variable model does support a calibrated short horizon:
two independent 40-stream panels at 3\% noise have no error above 10\% at
$\tau=0.012$, including continuously de-aliased event times, schedules, and
exponents.  This does not transfer automatically across sensing regimes.  With
$33^2$ sensors, 19/20 streams fail even at 0.0005 lead under 5\% noise and all
20 fail under 8\%; scalar ridge does not repair the floor.  Raising the sensor
grid to $49^2$ restores 40/40 safety at 5\% noise, but increases median CPU time
to 15.59 seconds and has occasional 40--54 second refinement tails.  Transfer
to an executable three-dimensional leading profile, real-time batched
execution, and validation on external fluid data remain open.  A rank-two
profile prior plus multiresolution map search does recover and exceed the dense
arm's median accuracy at $33^2$, but one of 40 fresh streams reaches 10.17\%
error (Wilson upper unsafe risk 12.88\%) and median CPU time is 53.25 seconds.
Validation-selected snapshot centering subsequently removes that observed tail
on 40 further streams: median/p90/worst error becomes 5.71/6.57/7.45\% at
$33^2$, with 40/40 safe and 15.79-second median CPU time.  This closes the
sparse-sensing gate for the analytic generator, but is not yet a safety
certificate beyond it: the centroid candidate set, rank-two profile, and
similarity manifold are matched to this construction.  External three-dimensional
JHTDB transfer is now tested for reconstruction, but not the full leading-profile
construction: exact divergence residuals improve instantaneous reconstruction
while worsening particle trajectories on all six dynamic streams.  Thus
operator satisfaction cannot be substituted for a downstream rollout metric.

\paragraph{Darcy energy assimilation is not strong-form closure.}
On 134 prospective PDEBench Darcy fields, conductivity-weighted cardinal
assimilation with 32 sensors lowers blind and weighted-gradient error relative
to a selected DCT residual.  It nevertheless has larger mismatch under our
finite-volume strong operator (1.356 versus 0.936), because the dataset's exact
interface/discretization convention remains unidentified.  We therefore claim
sparse physical-energy evidence propagation, not recovery of the hidden
generator, exact flux conservation, or a blind coefficient-to-solution solver.

\paragraph{The nano-drone result is a systems frontier, not SOTA accuracy.}
The real-robot transaction is prospective and independently admitted, but it
uses Melon runs 1--2 for target adaptation and confirms on run 3. The original
Physics+Residual and newer ASIA results train only on Square/Random/Chirp and
report the complete held-out Melon trajectory. Our confirmation errors are
numerically below ASIA in all four cumulative groups, yet these protocols are
not interchangeable. Gate 199 closes this gap with an adequate pinned ASIA
reproduction and matched sparse target adaptation. ASIA then beats the operator
in all four run-3 cumulative groups by 20.3--47.1\%. The operator's supported
advantage is systems-level: 1,020$\times$ less program storage, 65.9$\times$
faster compilation, fewer target labels, and exact transactional state. A
subsequent typed multi-execution consolidation narrows the mean physical ratio
gap to 12.2\% and removes its recurrent teacher from runtime, but still loses
all four absolute groups to adapted ASIA. This evidence is post-exposure on the
same three Melon flights and is not a prospective SOTA result. Gate 198 further rejects
universal direct-flow dominance: translation favors the recursive short-step
model even while direct prediction sharply improves long-horizon rotation.
No MCU energy or closed-loop control result is claimed.

\paragraph{The corrected KUKA transaction is post-exposure.}
An input-only audit found that the first KUKA analysis paired adjacent target
blocks even though repeated input programs are $(0,3),(1,4),(2,5)$. The earlier
claim that every sparse update harmed its paired repeat is retracted. The
source-only ridge selector had made the same pairing error; correcting it before
the target phase changes the ridge from 100 to 0.01. Sparse then improves every
repeat and retains 97.45\% of pooled dense gain with 3.189$\times$ fewer labels;
all pairwise retentions exceed 89\% and the frozen mechanism gate passes.
Because target outputs were already inspected, this is a scientific correction,
not fresh prospective confirmation.

\paragraph{Continuous zero-forget growth is conditional, not universal.}
The coupled lifelong operator network gives exact noninterference because typed
experiments isolate scalar edges, the missing laws are sparse cardinal atoms,
and every accepted task declares an operating interval whose complete basis
support is frozen. This is a strong continuous-function guarantee, but it does
not apply to arbitrary entangled deep features. Grid-aligned one-atom defects
also make the present localization problem favorable. An actual law change
inside a protected interval is a conflict and requires explicit versioning or
unlearning; silently adapting it would contradict the guarantee. Off-grid
defects, inferred region bounds, multi-atom growth, and external measurements
remain necessary before claiming a practical lifelong digital twin.

\paragraph{Hermite-score coordinate discovery currently assumes a known score.}
The multi-index result uses Gaussian whitening and the exact third Hermite/Stein
identity. It requires smooth ridge components with nonzero third Hermite
coefficients, a specified rank, several passes over evidence, and currently a
tensor or contracted-tensor operator. It does not yet apply unchanged to
unknown non-Gaussian laws, even-symmetric components invisible at third order,
or arbitrary entangled dynamics. Distribution-specific score operators,
multiple Hermite orders, or a validated transport to Gaussian coordinates are
needed for those regimes.

\paragraph{The compact ParWH compiler is not accuracy state of the art.}
Its prospective stationary multisine result tracks internal development, but
the frozen growing-amplitude result fails badly. The later Hermite-tail rule is
developed only on estimation-data amplitude holdouts and is not rescored on the
opened official target. Its replicated 16.5--24.0\% advantage over clipping is
therefore an extrapolation-mechanism result, not a repaired official benchmark
number. The common-pole SVD also recovers branch count but not the physical
input/output pole allocation. A fresh drift benchmark and shift-coupled branch
factorization are required before claiming general nonlinear-system
identification or field-level superiority.

\paragraph{Hermite continuation is an extrapolation hypothesis, not a safety certificate.}
On input-only F-16 coordinates, affine cardinal tails are robust across three
guarded boundaries and improve every prospectively evaluated channel, but the
official mean gain is 7.64\%, below the frozen 15\% target.  On recurrent
Cascaded Tanks dynamics, the same flexibility can compound error: an exact
cardinal residual lowers quiet-regime error yet worsens held-overflow RMSE from
$0.540$ V for its latent physical core to $1.356$ V.  Clipping the recurrent
edge while retaining Hermite feedforward edges is worse still, so graph
topology alone cannot select a true exterior law.  A practical lifelong model
must verify proposals on every affected event class and preserve an immutable
core for exact rollback.  The older reported $0.55$ V two-tank latent result is
retrospective: its code accessed the test maximum and initial output; correcting
the initial value leaves $0.551$ V but does not restore prospective status.

\paragraph{Zero forgetting is conditional on a frozen support map.}
Compact support alone is insufficient when normalization or learned
coordinates move historical samples across basis regions.  Our immutable
coordinate envelope certifies exactly the finite archived Levels-1/3 samples,
not their full generating distribution.  The affected-support credit rule and
quadratic step guarantee empirical non-degradation on the four retained
Level-5 verifier blocks; they are not distribution-free safety guarantees.
Moreover, the 7.29\% later-regime gain is retrospective because Level 7 was
already inspected.  A fresh streaming benchmark must freeze representation,
envelope, event partition, energy/credit thresholds, and future regimes before
the claimed safe-plasticity composition is prospectively tested.

The fixed global JHTDB flow-map law does not transfer across velocity fields,
and a one-prefix law has a finite temporal trust horizon. Receding compilation
with a fixed one-eighth trust step subsequently passes on six new nine-frame
streams (41/42 decisions, 3.90\% median gain), and transfers from observed to
disjoint tracers on exposed fields. A second fully prospective fields-plus-
particles panel gives 39/42 strict wins and 42/42 non-regressions, but formally
misses its per-stream strict-win gate because one stream invokes exact identity
three times. These remain small tests from one DNS dataset, one sensor layout,
and one tracker-noise model. Coordinate-noise transfer is now positive: on six
further sealed streams, a 256-track compiler obtains 40/42 wins and 3.60\%
median gain; over eight fresh particle/noise populations the operator-smoothed
version obtains 320/336 wins, 100\% safety, and 3.59\% median gain. However, the
single-population prospective smoothing advantage misses its frozen margin,
and the multi-population confirmation necessarily reuses exposed fields.
Two-view physical-space fusion plus circulant smoothing subsequently passes on
six new streams with 40/42 wins and 3.985\% median gain using 128 readings per
time, versus 3.620\% from a 256-reading single view. This factor-of-two
measurement result assumes zero-mean Gaussian coordinate noise; a symmetric
cross-Gram errors-in-variables proposal does not beat ordinary averaging.
Multi-population stress tests retain the 3\% target through correlation 0.5 and
25\% secondary-view dropout, but fail it at correlation 0.75 and 50\% dropout.
Inverse-variance weighted Grams cross the exposed 50\%-availability threshold
but fail on fresh populations, showing that known sensor variance alone does
not determine forecast utility. Evidence-triggered nested acquisition is more
robust: a frozen 5\% confidence policy nearly halves mean sensing and matches
or improves dense coverage/safety on fresh populations. On a different field
panel it still halves sensing and ties dense wins, but has one unsafe dense
escalation and formally misses its gain margin. The current validation split
therefore decides acquisition, not safe deployment; a common untouched event
block and transactional rollback are necessary. A subsequent retrospective
mechanism test and fresh-population confirmation implement this split: the
frozen transactional compiler uses 2.21-times fewer measurements and exceeds
its dense-240-plus-16-holdout control in wins, safety, and median gain. This is
still one DNS dataset with exposed fields; the admission block is only 16
particles and certifies empirical local events, not future distributions.
Across three panels the policy improves pooled wins and safety while halving
sensing, but the difficult Gate-157 panel misses its per-population and dense-
median clauses. The aggregate must not erase that negative result.
A subsequently acquired one-shot field panel strengthens both sides of the
boundary: transactional sensing is perfectly safe, exceeds 4\% median gain,
and uses 2.56-times fewer readings, yet trails dense by 0.529 points and two
wins. Thus neither the pooled result nor low-sensor confidence establishes
universal dense-quality equivalence.
Retrospective value-of-information modeling does not yet close this gap.
Nonlinear candidate-score signatures carry a repeatable but weak ranking signal
(Gate-181 Spearman about 0.30); changing the target to admitted-transaction
value and adding 16 passive sentinel observations both fail the frozen policy
criteria. A same-budget oracle shows large effect-size headroom, so these
failures diagnose missing decision information rather than an adequate learned
router. Branch-local reuse of the sentinel improves coverage to 249/252 at
2.57-times fewer readings, but still closes only 8.5\% of the dense-median gap.
It is a retrospective mechanism result, not a prospective adaptive-sensing
claim. The next test must use a pilot-induced program change, a disjoint
admission block, and a newly sealed panel; thresholds or classifiers tuned on
the current receipt would not answer the scientific question.
The explicit 64/8/8/176 pilot/verifier test also fails: despite positive
training-panel ranking, it gives 247 wins and 4.021\% at 99.75 readings. We
therefore close deterministic receipt routing on this exposed panel. A valid
continuation needs randomized exploration that observes action value online,
or a second physical system; another deterministic model is not independent
evidence.
Pure random acquisition also fails as a complete policy. Replacing only eight
of 28 branch-safe acquisitions by random transactions is much less costly:
across 1,000 trials every decision is safe and 99.9\% retain at least 248 wins.
This calibrates a retrospective exploration allowance; it does not demonstrate
that an online contextual ledger uses those labels effectively or transfers to
new fields.
Indeed, the frozen 8/20 mixture fails on a new sealed panel: it is perfectly
safe at 2.57-times fewer readings but reaches only 2.959\% median and 2/6
population medians above 3\%. Its deterministic control reaches 3.107\%; the
old confidence policy has a $-12.208\%$ unsafe event. Exploration-cost transfer
is therefore rejected; only transactional safety transfers. The acquired
labels cannot be used retrospectively to relabel this gate as online learning.
These are empirical boundaries on exposed fields, not guarantees or a
prospective independent-system result. Delayed or biased tracking; heavier or
structured missingness; partial state matching across
moving sensors; stronger online data-assimilation baselines; much longer
operation; and an independent physical dataset are required before claiming a
general fast-adapting digital twin. The 32-tracer and two-transition-pooling
failures also show that exact sufficient-statistic memory cannot replace current
spatial coverage or justify stale evidence under drift.

\paragraph{Sub-cubic score sketches need a weak-mode certificate.}
The matrix-free contraction experiment shows that large predictive gains can
coexist with a missed physical direction. Its fixed 32-by-24 sketch lowers
storage but does not certify that all low-energy tensor components survive.
The next method must compare independently accumulated score subspaces and grow
evidence or range only when their weakest principal direction is reproducible;
until then, the full-tensor confirmation is a mechanism result rather than a
scalable general discovery engine.

\paragraph{Circuit tangent gains and transfer are post-exposure.}
The 0.364-mV held-out transfer-function-jet frontier is obtained only after the official
circuit target had been opened for Gate 215; it is source-held-out mechanism
evidence, not a new prospective benchmark score. A complete-source frozen
refit reaches 0.367 mV when replayed on that target, close to its 0.338-mV
source score but still above the reported 0.241-mV deep encoder. This supports
transfer of the compact mechanism, not prospective SOTA. The next extension
must be selected on disjoint source evidence and ultimately tested on a fresh
physical record; target-driven residual fitting would invalidate the claim.

\paragraph{Broadband natural-signal fitting: closed form loses to deep SGD.}
Our operator-matched, closed-form principle is decisive on \emph{structured} signals but not on broadband ones, and implicit neural representations make the boundary sharp. Fitting an image as a coordinate-to-color map, a deep SGD-trained SIREN reaches roughly $60$\,dB PSNR, whereas a closed-form random-Fourier-feature ridge reaches about $28$\,dB and a closed-form representation matched to the dominant modes does worse still ($\sim\!10$\,dB, underfitting). The closed-form solver is far faster (over two orders of magnitude) but a shallow one-layer feature ridge cannot match a deep network's learned-frequency fidelity on broadband natural content. This is exactly the SSP prediction: when the generating operator is trivial and the innovation is dense, broad frequency coverage beats energy-concentrated matched modes, and deep learned features beat both. Matched closed-form wins on structured signals (dynamical systems, PDEs); it does not win on broadband natural images, and we do not claim otherwise. The sharper statement is that the boundary is \emph{per factor}, not global: a signal can be broadband in one factor and structured in another, and the right design matches structure factor-by-factor. Dynamic 3D Gaussian splatting is a clean example---the appearance is broadband (kept as an expressive learned Gaussian field) while the deformation is a smooth function of \emph{continuous time}; replacing the time-deformation MLP with a closed-form continuous-time (liquid) cell that bakes the time structure into the architecture matches or beats the MLP at lower compute, with gains concentrated exactly on the high-frequency, structured motion \citep{li2026cfc,hasani2022cfc}. The lesson we carry is to structure the factors that are structured and leave the broadband factors to learned features, rather than choosing closed-form versus deep wholesale.

\paragraph{Atari mastery via behavior cloning fails: imitation is not reinforcement learning.}
We balance and play simple games from pixels with no backpropagation, but scaling to Atari-style mastery by cloning a planning teacher does not work. Even when gradient-free perception clones the teacher with reasonable per-frame accuracy ($\sim\!0.79$), the resulting policy scores at the floor ($-21$, losing every point), a textbook behavior-cloning failure: average action accuracy does not imply competent closed-loop play, because the small fraction of mispredicted actions compounds over a trajectory and the agent never sees the states its own errors lead to. Closing this requires interactive reinforcement learning at scale, not imitation, and we have not demonstrated that.

\paragraph{From-scratch LLMs at full scale: mechanism proven, scale is a compute question.}
We trained a real autoregressive character-level GPT (token and positional embeddings, causal transformer blocks, and an LM head) by the per-block local sweep with no global backward pass; it reaches a validation bits-per-character within about $0.013$ of global backpropagation and generates coherent corpus-style text. This settles the \emph{mechanism}: predictive-coding-style local learning matches backpropagation on a working language model, not merely on a probe task. It does not settle \emph{scale}. Training a GPT-2/3-scale model gradient-free is a compute and cluster-engineering question that we have not run; we claim the mechanism, not a scaled result.

\paragraph{Evolution strategies are a niche tool, not a universal booster.}
We examined whether evolution strategies / PGPE could improve the framework broadly and concluded that they cannot. On a low-dimensional control toy, a simple genetic algorithm trivially solved the task while our quick ES stalled on the flat, deceptive near-random reward landscape --- the same deceptive-landscape failure that motivated novelty search \cite{conti2018novelty}. A properly tuned ES is known to solve such toys but offers no advantage over the genetic algorithm there. The honest verdict is that ES/PGPE fits a narrow slot --- high-dimensional, smooth-landscape, model-free, non-differentiable optimization --- and is not a general accelerator for local or closed-form learning.

\paragraph{Self-improvement is bounded: no free lunch from self-reference.}
We map self-improvement precisely (Section~\ref{sec:res-selfimprove}) and the map is as much a boundary as a result. Self-labelling and self-consistency-style loops \emph{plateau}---they add no information the model does not already contain, and a self-derived verifier fails outright because its errors are correlated with the model's own. Reinforcement learning does not rescue this within our framework: the closed-form memory's stability requires \emph{stationary} targets, and RL's bootstrapped, policy-dependent targets are not, so it is no more stable than gradient TD. Real compounding needs an \emph{external} verifier, and even then it lifts only the \emph{knowledge/readout} layer to its ceiling---improving the representation itself, or bootstrapping from a near-chance start (below a viability threshold, where there is nothing correct to verify), still requires external teaching. Our real-LLM demonstration is on arithmetic, the canonical cheap-verifier domain, and lifts to a plateau via retrieved verified exemplars rather than by changing weights; scaling to code (unit-test verifiers) or proof at LLM scale is scoped but not run.

\paragraph{Predictive coding does not, by itself, remove weight transport.}
We are careful not to overclaim biological plausibility. Classical predictive coding uses the transposed forward weights ($W^\top$) for its feedback, i.e.\ it assumes weight transport; its actual claim is the absence of a \emph{global backward pass} with local error signals, not the absence of weight transport. Removing weight transport is the separate concern of random-feedback methods \cite{lillicrap2016randomfeedback,nokland2016directfeedbackalignment}, which we evaluate separately and which carry their own accuracy cost. When we report no-global-backward results and no-weight-transport results, we keep them distinct and do not conflate the two guarantees.

\paragraph{Single-task accuracy versus tuned backprop.}
In early experiments, single-task accuracy of the gradient-free stack sat below tuned backpropagation, and on some architectures (e.g.\ convolutional stacks) local learning matches backprop only to within one to two points rather than cleanly beating it. The gap has since been closed on vision through deep local perception, but we flag the boundary honestly: our headline claim is a unified, closed-form, no-forgetting principle that is decisive where structure is known and \emph{ties} strong backprop baselines at one-to-three orders of magnitude less compute where structure is absent --- not that it beats every accuracy number in every single-task setting.

\paragraph{Deliberately deferred to cloud scale: the open benchmark backlog.}
The experiments in this paper were run on a single workstation, which caps demonstrations at CIFAR/smallNORB scale and at continual learning over frozen pretrained backbones. We separate \emph{mechanism}, which we demonstrate, from \emph{frontier-scale absolute numbers}, which require multi-GPU compute we have not yet provisioned. To make the boundary explicit --- and to mark these as planned, scoped work rather than oversights --- we list the deferred benchmarks; each runs the \emph{same} gradient-free substrate (per-block local sweep, closed-form Gram memory, topology evolution), with scripts that auto-detect CUDA and need only big-batch/epoch configs and dataset download on the cloud box.
\begin{itemize}
  \item \textbf{Atari with no global backward pass} (the marquee RL first): (i) a local-sweep DQN --- the exact DQN net, replay, target network, and $\epsilon$-greedy, but TD-error credit assigned block-locally; (ii) a deep-GA / ES neuroevolution baseline (embarrassingly parallel); (iii) latent-MPC model-based control at pixel scale. We already play real Pong gradient-free via DAgger; the open item is the $49$-game, $50$M-step protocol.
  \item \textbf{ImageNet-1k top-1 with no backprop} (the landmark vision first): scale the deep local-sweep CNN/ViT that already matches or beats backprop on the same net at CIFAR scale; decoupled blocks across GPUs also turn the backward-unlock parallelism from a microbenchmark into a real speedup.
  \item \textbf{GPT-2 (124M) from scratch on OpenWebText, gradient-free}: the scale step beyond our character-level and \texttt{enwik8} parity-with-backprop results.
  \item \textbf{Proper self-supervised invariance at scale}: our from-scratch contrastive pretraining collapses at CIFAR-100 under a quick recipe; the deferred test is a full big-batch/LARS/long-schedule SSL stack on CIFAR-100/ImageNet, to scale the small-scale invariance-beats-memorization result.
  \item \textbf{ViT-backbone class-incremental learning} (now done): swapping the frozen ResNet-18 for a frozen DINOv2 ViT-L brings the zero-forgetting Gram memory to $0.906$ final accuracy on Split-CIFAR-100 (from $0.685$ on ResNet-18), confirming the backbone-change route into the near-SOTA regime; it is also granularity-invariant ($0.868$ at $5/10/20$ tasks) and beats a fair replay by $\sim\!9$ points. Remaining at cloud scale: Split-ImageNet-R.
  \item \textbf{Split-ImageNet-R whitening}: closing the residual gap to RanPAC with its exact second-moment whitening and $\lambda$-selection (Mac-feasible; queued).
  \item \textbf{PDEBench 2D/3D and MuJoCo/DM-Control suites}: Darcy sparse assimilation and external JHTDB reconstruction are now demonstrated; the remaining scale step is strong-form multi-operator forecasting across the full PDEBench suite, and extension of latent MPC from HalfCheetah to the full locomotion suite.
\end{itemize}
We report these so that the reader can distinguish ``not shown because it fails'' from ``not shown because it needs a cluster we name and have scoped.''

\paragraph{Nominal VOI is not yet a safe experiment policy.}
The hysteresis acquisition experiments are prospective but synthetic and local.
A directional symmetric-Gram probe improves the intended damping coefficient
without uniformly improving displacement, while a constrained waveform mixture
preserves broadband performance but fails robustly across four nearby laws. We
have therefore not demonstrated closed-loop safe self-directed experimentation
or a measured-hardware advantage. The next claim requires risk-sensitive
covariance over archived operator versions, explicit actuation cost and safety
constraints, and a physical prospective panel. Whole-transaction simulation
gives the correct conservative direction on a subsequent panel, but its 19.92\%
worst-risk gain misses the frozen 20\% margin and redundant endpoint
normalization breaks bitwise waveform identity. This is promising mechanism
evidence, not a certified or hardware-validated acquisition policy.
An eight-system follow-up passes the sensing-policy clauses with a 92.61\%
worst-risk advantage, but the complete controller still misses one independent
model-adequacy ceiling; its successful repair uses exposed validation. Thus the
two transactional layers are first demonstrated separately. A subsequent frozen
eight-system run composes them prospectively: four certificate failures route to
bounded monotone relinearization, all close, and all unseen errors remain below
0.001273 mm while the acquisition planner safely retains generic. This is still
synthetic evidence under one Bouc--Wen family. No real actuator, sensor noise
model, energy comparison, safety envelope, or open equation vocabulary has yet
validated the complete loop.
We now have a deterministic interactive MuJoCo receipt with exact plant and
wrapper-state rollback, but it contains no learned correction or control-policy
comparison. It should not be interpreted as closing the external-actuation gap.
The gauge-fixed MuJoCo cardinal compiler is likewise only an adaptation-set
algebra result until disjoint trajectories are opened.