Research record

Negative-result appendix

Historical source. Some claims in older records were subsequently corrected. The associated article states the adopted interpretation. This record preserves the original source alongside its rendered reading view.

Rendered archival TeX

This is an HTML reading rendition of the local TeX record. Mathematical notation is rendered with KaTeX; archived figures are included when their source assets are part of this collection.

Negative results and what did not work

Off-grid transactions and overlapping atoms.

Continuous variable projection on the coupled lifelong edge system removed fixed-grid quantization—on twenty independent streams mean edge error fell from 0.656640.65664 to 0.0012500.001250 and median rollout error from 0.120880.12088 to 1.46×10−41.46\times10^{-4}—while all committed-state drift audits were zero. The frozen gate nevertheless failed because its implementation counted a rejected boundary proposal as a support violation. We retain that negative and define support certificates on committed state, with separate bitwise rollback for rejected transactions. With two overlapping opposite-sign atoms per region, model order and rollout were recovered but individual centers were not: worst center error was 0.03830.0383 and mean edge error 0.04650.0465. Function recovery and latent atom identification are therefore distinct claims.

External raw-coordinate growth.

On measured coupled-electric-drive data, an eighth-order delay carrier plus a finite-difference-velocity cardinal residual produced unstable validation rollout and was rejected before external test access. The old route remained bitwise identical, and the banded Gram and four-tap evaluation audits were 6.09×10−166.09\times10^{-16} and 6.59×10−176.59\times10^{-17}. This locates the bottleneck in observable stable state construction and deployment-aligned evidence, rather than in spline evaluation speed.

Characteristic-coherence replication.

An independent 512-Halton-probe implementation reproduced the earlier JHTDB material-coherence negative. Its matrix-free 15,00015{,}000-variable SPD normal reduced characteristic inconsistency by 63.9%63.9\% and improved over independent divergence corrections on five of six streams, but still worsened trajectories relative to raw trilinear interpolation on all six. Exact minimization of a misaligned physical penalty is not downstream correctness.

A research program is only as honest as the failures it records. This appendix documents, deliberately and in detail, every approach we tried that did not work or only partly worked, together with the diagnosis of why it failed. The intent is twofold: to keep the program from wasting effort re-trying dead ends, and to make the boundary of the approach explicit. Each entry states what was tried, the result, and the lesson. We group the entries into clusters; the parenthetical numbers refer to the master scoreboard.

Predictive coding: the large-output-error regime

Hard-clamp predictive coding loses to backpropagation (53–55).

Our first cut at predictive coding (PC) on a deep MLP teacher–student task, using a hard-clamped one-hot target with Tinfer=25T_{\mathrm{infer}}=25 inference steps, reached only 0.5950.595 accuracy against 0.7250.725 for cross-entropy backpropagation — a clear loss. The cause was not a loss confound (backprop-MSE 0.713≈0.713 \approx backprop-CE 0.7250.725, while PC stayed at ∼0.55\sim 0.55) and not non-convergence. A per-layer gradient-alignment unit test (cosine of the PC free-energy gradient against the autograd backprop-MSE gradient) isolated the mechanism: hard-clamping a one-hot target on an untrained network is the worst case, in which the top-layer gradient deviates (cosine 0.920.92 at the output) and corrupts the descent direction. The fix is to stay in the small-perturbation regime: a small target nudge (b=0.1b=0.1) recovers uniform per-layer cosine ≈0.99\approx 0.99, and Z-IL's layerwise-timed updates recover the backprop identity exactly (cosine 1.0001.000). In the corrected regime PC equals backpropagation (backprop-MSE 0.6780.678 vs.\ PC-Z-IL 0.6760.676 vs.\ PC-nudged 0.6740.674; the hard-clamp artefact 0.5630.563). Lesson: PC does not lose to backpropagation per se; hard-clamping drives a large output error that degrades the top-layer gradient. Nudge or Z-IL is mandatory.

Deep relaxation PC attenuates the error toward early layers (56, 56b, 60).

Carrying nudged relaxation PC to convolutional networks lost again (backprop-CE 0.5410.541, backprop-MSE 0.5450.545, PC-nudged at T=15T=15 only 0.3990.399). The same gradient-alignment test diagnosed an intrinsic pathology of iterative relaxation: the output error propagates downward one layer at a time and attenuates with depth. Even at 300300 settling steps the per-parameter cosine is 1.01.0 at the output but decays to ≈0\approx 0 at conv1 — the ePC/EVPE depth-attenuation pathology. A precision-gain schedule helps monotonically (conv1 cosine 0.026→0.163→0.4560.026 \to 0.163 \to 0.456 as the gain goes 1→2→31 \to 2 \to 3 at T=150T{=}150) but does not remove the wall, because the true bottleneck is convergence: deep relaxation simply needs many settling steps for the error wave to reach the early layers undiminished. Fix and lesson: replace slow iterative relaxation with a closed-form, single-sweep equilibrium solve (HGF/ePC-style), which is exact at all depths by construction because there is no wave to attenuate. ``Replacing iterative relaxation with a closed-form equilibrium solve'' is precisely the OSNR/Gram principle, so PC folds into our closed-form-solve family.

Per-sample RMS error normalization breaks PC (56b).

A Meta-PCN-style per-layer RMS normalization of the prediction errors, intended to counter depth-attenuation, broke training instead: the output gradient cosine flipped to −0.89-0.89. Lesson: per-sample RMS normalization destroys the relative scale of the error signal that the gradient identity depends on; it is the wrong precision mechanism. Closed-form precision is the right one.

Categorical full-CE-force PC is again the large-error regime (59).

A cross-entropy readout for PC is implementable, but applying the full CE force (no nudge) reproduced the hard-clamp failure: on the deep MLP teacher–student task backprop-CE reached 0.7160.716 and backprop-MSE 0.7120.712, while full-force PC-CE managed only 0.6530.653. This is the same large-error effect that produced the hard-clamp MSE result. Lesson: the readout form is not the issue; applying full target force is. A nudged or precision-scaled categorical readout is the clean formulation if pursued.

Prospective configuration does not beat backprop on i.i.d.\ accuracy (61–62).

Song et al.'s prospective-configuration motivation predicts that relaxation finds a better activity configuration before plasticity, which could beat backpropagation on data efficiency or online learning. Two controlled attempts (matched objective, initialization, seed, and optimizer; three seeds each) said otherwise. On the data-efficiency sweep (train sizes 250250–20002000), PC did not beat backprop. Online/few-pass full-clamp PC — the genuine prospective regime — also failed to win: e.g.\ epoch 1 batch 16 gave backprop 0.41970.4197, PC-nudged 0.42860.4286, PC-hard 0.38350.3835. Lesson: nudged PC is provably ≈\approx backprop, so it cannot beat backprop by much; full-clamp prospective PC deviates from backprop but underperforms on accuracy. There is no robust PC-beats-backprop regime in our i.i.d.\ settings. The remaining documented regime where prospective configuration might still win is continual/interference, which is handled instead by our closed-form Gram memory.

PC-NEAT shows no evolution gain when the seed already matches the teacher or the task is linearly decodable (58, 66).

When PC is used as the inner learning rule for evolved NEAT topologies, the expected evolution win did not materialize on tasks where the seed network already matched the teacher or where the target was linearly decodable: the MSE/capacity ceiling is reached without any topology search, so evolution has nothing to add. The bridge mechanism itself is sound — PC trains an arbitrary evolved skip-DAG locally, matching backprop (e.g.\ PC-nudged 0.58740.5874 on the evolved graph) — but a clean PC-NEAT win needs a task where topology is load-bearing. Lesson: re-frame the probe on parity, where the topology is provably load-bearing, before claiming an evolution gain.

Bolt-on credit-assignment shortcuts on convolution

DFA / direct-feedback bolt-ons do not beat backprop on conv, and scale-up widens the gap (48–50).

We made three attempts to beat backpropagation on convolutional vision by adding a Direct Feedback Alignment (DFA) global-error broadcast to InfoPro-style local learning (the brain's putative two-signal arrangement). All three failed. A DFA broadcast at γ=1.0\gamma{=}1.0 gave 0.7410.741 and at γ=0.05\gamma{=}0.05 gave 0.7620.762; a scale-up attempt reached 0.8540.854 versus backprop's 0.8750.875 — i.e.\ widening width and epochs widened the gap rather than closing it. The diagnosis is that local block training operates on a moving, detached output target and is inherently less stable, while the DFA term adds noise; both hurt. Verdict and lesson: bio-local credit assignment matches backprop on conv within ∼1 σ\sim 1\,\sigma but does not beat it, and DFA is not the lever that turns a match into a win. We claim ``matches,'' not ``beats.''

Gradient-free deep credit assignment on adversarial tasks

Crude gradient-free deep credit assignment lags backprop on kernel-hard compositional tasks (29).

On the adversarially-hardest probe — parity of the signs of three hidden random hyperplanes in 3030 dimensions, where a representation must be learned — backprop reached 0.880.88, while CEM-evolved first-layer features plus a closed-form readout reached 0.610.61 and closed-form target propagation only 0.570.57 (random/greedy features sit at chance, 0.530.53). Gradient-free credit assignment is therefore real (it decisively beats random) but substantially weaker than gradient descent at learning deep representations on a genuinely compositional task. Lesson: on hard compositional deep credit assignment, backprop's gradient-through-depth is a real moat; the practical verdict is a hybrid program (backprop or a pretrained backbone for the deep representation, our closed-form/evolution machinery for the structured, continual, control, and self-scaling parts). Caveat: parity is the worst case, and sophisticated gradient-free methods (difference target propagation with learned inverses, large-budget CMA-ES, NEAT block-growth weight evolution) might narrow the gap — an open question, not a closed one. Note that for decoupled-greedy local learning specifically the wall was later overturned (it matches backprop, with the DFA beat attempt above being the part that failed); the residual hard case is precisely this crude-gradient-free compositional regime.

Label-free whole-image contrastive vision fails; language-from-scratch gradient-free is ruled out at scale (V1; 29 cluster).

An earlier label-free contrastive vision objective learned augmentation invariance (the loss dropped) but produced near-random downstream accuracy; the label-free win came instead from sparse coding, not whole-image contrastive learning. Separately, the broader ``learn language from scratch, fully gradient-free, at scale'' ambition was ruled out as a near-term claim: the mechanism is proven (a character-level GPT trains by the per-block local sweep to within ∼0.013\sim 0.013 bits-per-character of backprop), but a scaled result is a hybrid verdict, not a gradient-free-from-scratch one. Lesson: contrastive is the wrong label-free objective for our setting (sparse coding is right), and ``from-scratch gradient-free language at scale'' is a compute/cluster question we have not run.

Evolution strategies and exploration

ES/PGPE fails on cartpole; a genetic algorithm trivially solves it (76).

Asking whether evolution strategies could broadly boost the framework, we compared a genetic algorithm against a quick antithetic, rank-normalized ES on cartpole-balance (reactive policy, no backprop in any method). The GA solved it trivially (best 471471 in 320320 episodes), while the ES did not solve it across four configurations (best ∼8\sim 8–1010). The diagnosis is the deceptive landscape: near random initialization almost every policy dies in ∼8\sim 8 steps, so single-point ES perturbations do not differentiate and the mean-update stalls, whereas the GA's population effectively random-searches and finds a surviving policy by luck. This is exactly the deceptive flat-reward failure that motivated novelty search. (A properly tuned ES is known to solve this toy but offers no advantage over the GA on such a low-dimensional problem.) Lesson: ES/PGPE is a niche tool for high-dimensional, smooth-landscape, model-free, non-differentiable optimization — not a universal accelerator. Our reliable edge for control is model-based MPC (∼28,800×\sim 28{,}800\times sample efficiency over model-free) plus local/closed-form learning. The on-brand idea kept for later is closed-form ES (the PGPE mean-update as a return-weighted regression of returns on perturbations), to revisit only when we tackle high-dimensional model-free RL.

Closed-form representations on broadband signals

Closed-form SIREN with matched dominant modes underfits broadband images (80).

Fitting a coordinate-to-color image map, a deep SGD-trained SIREN reached 59.659.6 dB PSNR; the closed-form random-Fourier-feature ridge reached 28.128.1 dB (226×226\times faster, no backprop), and a closed-form representation matched to the energy-concentrated dominant modes did worse still at 10.410.4 dB (underfitting). Lesson and boundary: this is the SSP boundary made sharp. For broadband natural images the generating operator is trivial and the innovation is dense, so broad frequency coverage beats matched dominant-mode concentration, and a shallow one-layer feature ridge cannot match a deep network's learned-frequency fidelity. Operator-matched closed form is decisive on structured signals (PDEs, dynamical systems — two to three orders of magnitude over FNO/Neural-ODE) and explicitly not on broadband natural content; we do not claim otherwise.

Reinforcement learning at scale via imitation

Atari Pong behavior cloning fails to play (86).

On real ALE Pong with a no-backprop local-sweep convolutional policy behavior-cloning a heuristic teacher, the best attempt used a predictive teacher (extrapolating ball trajectory with wall bounces) and frame-diff input: the teacher was good (−10.6-10.6 versus random −20.8-20.8) and gradient-free perception cloned it decently (0.7940.794 BC accuracy), yet the cloned policy scored at the floor (−21.0-21.0, losing every point). Diagnosis: a textbook behavior-cloning failure — 79%79\% average action accuracy misses the rare critical intercept frames, and BC distribution shift compounds as the policy's own rollout drifts off the teacher's distribution, worsened by the two-pixel ball being hard to localize. Lesson: the gradient-free perception mechanism works, but BC ≠\neq RL; Atari mastery needs interactive RL at scale (DAgger, model-based MPC on learned latent dynamics, or DQN-scale frames). The no-backprop model-based recipe is validated where state is recoverable (Catch from pixels, 96.5%96.5\%).

Self-improvement without representation learning

Gradient-free closed-form self-training on a fixed feature map degrades; it does not compound.

The stable recursive self-improvement we demonstrate (App. [app:record], stable-RSI) uses backprop to self-improve a trainable representation. We asked whether the two stabilizers that made it work—the never-forget Gram memory and the confidence curriculum—suffice to make a fully gradient-free loop compound on a fixed feature map (no backprop anywhere): a tiny labelled seed, the closed-form Gram readout self-labels an unlabelled pool, confident labels are additively accumulated, re-solve. They do not. Sweeping seed strength (2525–500500/class) and two gradient-free bases (random-conv, supervised-ridge ceiling 0.5560.556; k-means Coates–Ng, ceiling 0.7040.704) on CIFAR-10, the seed-only static solve is always the best and every self-training run degrades monotonically in how much pool it admits (canonical: Coates–Ng, 100100/class, 0.503 ⁣→ ⁣0.4060.503\!\to\!0.406, peaking at round 1 with a marginal +0.007+0.007 before eroding; on the weak base it collapses 0.31 ⁣→ ⁣0.150.31\!\to\!0.15). The never-forget memory is consistently slightly more robust than naive re-fitting (its order-invariant accumulation does not oscillate) but cannot turn a net-negative signal positive. Diagnosis: on a fixed feature map the pseudo-labels are produced by the same features the readout consumes, so they carry no information the seed cannot already extract (they cannot exceed it) while they do carry label noise (they degrade it). Lesson, and it sharpens the positive result: self-improvement requires representation learning. Backprop-EMA RSI compounds precisely because backprop reshapes the features using the unlabelled data; closed-form on fixed features has no mechanism to change the representation, so it can only add noise. The closed-form never-forget memory is the right tool for continual learning with new labelled data/classes (it beats buffered replay, App. [app:record]) but not for bootstrapping from unlabelled data on a frozen representation.

A plastic encoder cannot beat a frozen one under a closed-form continual memory: capture vs.\ retain is irreducible.

The EMA slow-anchor recovers retention for a plastic representation when the encoder is already good (App. [app:record]), but does it ever win? We tested the genuine case—continually fine-tuning a pretrained ResNet18 on an out-of-domain medical stream (PathMNIST →\to BloodMNIST →\to DermaMNIST →\to RetinaMNIST, 2929 classes), where adaptation should pay. It does pay per task (live fine-tuning standalone 0.7030.703 vs frozen 0.6690.669), yet EMA-anchored plasticity never beats frozen across the whole momentum range: final accuracy 0.142/0.466/0.527/0.6520.142/0.466/0.527/0.652 at momentum 0.80/0.90/0.95/0.990.80/0.90/0.95/0.99 versus frozen 0.6620.662, the best (0.990.99) merely matching it. Diagnosis: an irreducible capture-vs-retain tradeoff. Realising the per-task gain requires the memory to read the current fast-moving encoder, but then old-task statistics are stale and forgetting is catastrophic (live 0.0590.059); reading a slow EMA retains (forgetting 0.0190.019 at momentum 0.990.99) but the EMA lags so far that it captures essentially none of the gain (standalone pinned at ∼\sim0.666≈0.666\approx frozen at every momentum). No setting captures and retains at once. Lesson: this delimits precisely where the slow anchor works. It stabilises self-improvement when the target is fixed (the same task, a representation improving toward a stable optimum—the EMA tracks it and the memory stays valid), but it cannot serve continual learning across different tasks, where the representation must move to new domains and a single memory cannot both follow it (to capture) and stay anchored (to retain). Freezing a strong pretrained encoder is therefore not merely pragmatic but essentially optimal for closed-form continual memory; making the shared encoder plastic creates a tradeoff EMA cannot break. This closes the ``anchor a drifting representation'' frontier with data.

Open problems

Instantaneous divergence repair does not improve Lagrangian trajectories.

On six preregistered five-frame JHTDB streams, the confirmed static cardinal divergence residual improves global blind velocity on all 30 frames but worsens particle endpoint and all-time errors on 6/6 streams. Static held-out sensors select the harmful full correction, and a matrix-free 15,000-coefficient material-coherence extension also fails 6/6. The result rules out the tempting inference that stronger incompressibility alone implies a better flow map; a compiled operator must govern the downstream functional being evaluated.

Pooled and empirical Gram memory do not protect a continuous function.

In the lifelong constitutive-edge gate, additive banded statistics recover the order-invariant pooled ridge optimum and improve physical rollout, yet an accepted noisy revisit produces up to 68.5% relative regression from the best prior regional error. Constraining updates to coefficient directions with zero old empirical-Gram diagonal makes every old sampled prediction exactly invariant, but unseen points within the same interval can regress by 138%. Closing the entire continuous basis support removes that drift exactly, but a dense one-shot local fit then misses its plasticity threshold. These three negatives motivate the successful sparse support-closed rule: sufficient statistics remember evidence, continuous support protects functions, and parsimonious growth is needed to avoid freezing noise.

A compact cardinal residual is not the right block for globally smooth EMPS mismatch.

On measured positioning dynamics, compiling the exact kinematic equation and fitting only acceleration makes every held-out 500-step rollout stable. The cardinal velocity edge is useful relative to the linear carrier (0.280770.28077 versus 0.340140.34014 mean rollout nRMSE), and its seven-diagonal Gram and four-tap evaluation agree with dense references to 7.40×10−167.40\times10^{-16} and 8.33×10−178.33\times10^{-17}. It nevertheless loses decisively to a three-term global polynomial residual (0.183460.18346); the external test was therefore not opened. This closes fixed-smoother/cardinal tuning on EMPS and yields a block routing rule: global smooth residuals belong in an operator reproduction space, whereas compact cardinal innovations require independent evidence of local structure.

Static and fixed-memory F-16 interface shortcuts do not transfer.

On full-scale F-16 ground-vibration estimation data, an operator-derived payload-interface coordinate is useful but does not by itself justify local tensor structure: its 2D cardinal map lowers carrier error 21.9% yet is 0.90% worse than a global polynomial on a held-out amplitude. Closing a map learned from observed interface state by fixed-point iteration either diverges (polynomial) or remains 32–33% away from a fixed point (compact cardinal). A cardinally regularized play bank improves internal error only 1.60%, below its frozen 2% commit threshold. Finally, an external regime router cannot transfer unchanged FullMSine coefficients monotonically to sine sweeps: its route complexity is 0,3,3,10,3,3,1, and the stable exponential block fails catastrophically at the highest estimation amplitude. Sine-sweep validation remains unopened. Coefficient-separable and combined impulse-response gain bounds are valid but force correction scales below 7×10−47\times10^{-4} and therefore collapse every dynamic route to the polynomial; a 50-sample passive router chooses the carrier at every amplitude because its shared startup transient is uninformative. These negatives rule out more knots, relaxation tuning, global worst-case projection, and force-RMS/startup routing; the open requirement is a reachable-state gain or passivity certificate.

Input features, low-rank weights, and unconstrained cardinal compression do not preserve F-16 free-run geometry.

On the published SpecialOddMSine split, the FullMSine cardinal–exponential residual does not transfer across random phases (0.828350.82835 versus 0.827940.82794 carrier), while a causal stable modal carrier is a 17.0% positive compression result (0.686960.68696). Adding all quadratic/cubic products of eight dominant modes changes this to only 0.687090.68709. A stable 64-lag ARX observer plus a four-coordinate cardinal innovation reaches 0.698100.69810; autoregression alone is not the missing mechanism.

A locally reproduced 256-unit sigmoid observer reaches 0.500610.50061, but directly compiling its 16 most response-informative ridge directions into uniform cardinal edges gives teacher-forced RMSE 0.165960.16596 and free-run RMSE 1.735931.73593. The divergence begins at the second autonomous sample. Exact derivative-Gram regularization repairs stability and improves the modal anchor to 0.646540.64654, but misses the prospective 0.600.60 gate and yields a 3.32×10−93.32\times10^{-9} independent-solve discrepancy in its strongest ill-conditioned system. Ordinary low-rank arguments also fail: approximations retaining 98–98.6% of teacher weight energy roll out above 3.43.4. These results rule out widening, rank-energy compression, and unconstrained edge pruning as explanations. The successful repair changes the recurrent state itself (Sec. [sec:f16-neural-compiler]), substituting stable modal history before cardinal compilation.

Static amplitude coordinates do not supply missing hysteretic state.

On 108,477-sample F-16 SineSweep estimation records, the teacher-free Gram–Krylov compiler learned only at low amplitudes transfers positively to Level 7: all channels improve and mean nRMSE falls 18.59%. This narrowly misses a frozen 20% gate, so even validation levels remain sealed. Three theory-motivated repairs then fail without accessing them. A single pooled cardinal map gives only 0.99% mean leave-one-level-out gain and regresses Level 5; exact zero-forget regional maps with force-RMS coefficient interpolation regress Level 5 by 24.4%; and a 1,361-feature amplitude–state tensor field, regularized by exact Kronecker value/derivative Grams, improves two interior folds only 1.3–3.1%. The tensor Gram action itself agrees to 2.22×10−162.22\times10^{-16}. These negatives distinguish correct spline calculus from sufficient physical state: record-level amplitude does not encode the evolving constitutive memory of a sweep.

A bounded residual cannot replace a missing operator nullspace.

For scalar-position EMPS, a strictly stable input-only exponential bank has confirmation nRMSE 3.6713.671; cardinal compilation halves it to 1.7691.769 but remains far behind a projection polynomial. Adding exactly the repeated-root coordinates of the double-integrator operator—single and double force integrals—reduces carrier error by 85.3% to 0.5400.540, and cardinal compilation improves it another 13.3% to 0.4680.468. A cubic projection residual still reaches 0.3260.326, so the historical test remains unopened. Exact cumulative- sum and chunked integrator implementations agree to 2.07×10−152.07\times10^{-15}. This is a representation boundary: repeated zero roots and polynomial reproduction are load-bearing, but input integration cannot replace observable velocity/friction state.

Covariance-dual coordinates do not solve coupled F-16 dynamics.

After independent synthetic confirmation, a three-edge Riesz map and a 15-edge multi-shift Riesz bank were frozen for the existing F-16 chronology. Validation selected the latter, but untouched realization 8 reaches 0.453540.45354, 0.20% worse than the raw-Krylov compiler, with only 1.06×1.06\times cardinal- component compression. All projected solves and cardinal audits pass below 2.14×10−152.14\times10^{-15}. This rejects universal transfer of the single-index elliptical argument: several lagged dynamic coordinates can contribute to one output, so first-order covariance geometry collapses necessary structure.

Fixed Hermite-score sketches can omit weak physical modes.

Matrix-free contracted score actions reduce operator workspace 3.29×3.29\times and predictive error 72–73% in two 64-dimensional confirmations. Yet one weak generating direction is recovered at cosine 0.88110.8811, below the frozen 0.90 threshold, even while its cardinal predictor beats linear regression by more than threefold. Predictive error is therefore not a completeness certificate. Adaptive sketches need independent evidence-ledger stability or another mode certificate; increasing sketch width after observing this miss would be post-hoc tuning.

Confidence is not value of information.

On a newly acquired JHTDB panel, confidence-triggered transactional sensing is safe and effective at 2.56-times fewer readings but trails dense sensing by 0.529 points. A cross-panel ridge predictor, a nonlinear 31-feature operator- receipt model, an action-aligned transaction-value target, and a 16-reading sentinel router all fail their frozen policy gates. The rich receipt does rank held-out uplift positively (Spearman 0.304), and a same-budget hindsight transaction exceeds the dense median, so the negative is not lack of headroom. It is an observability boundary: endogenous validation statistics are too weak to decide the counterfactual value of unseen evidence. Reusing paid sentinels in the non-acquisition branch recovers six wins without extra cost, but still misses the effect margin; applying that reuse to the old router adds sensing without fixing selection. Future acquisition must learn from controlled pilot updates and disjoint verification, not another threshold on the same receipt. That progressive 64/8/8/176 test is also negative: pilot-induced receipt changes rank admitted value across the training panels but transfer at only 0.189 Spearman, yielding 247 wins and 4.021% versus fixed confidence's 248 and 4.042% at equal cost. The corresponding hindsight policy reaches 4.581%, so the remaining gap requires observed exploration outcomes, not more deterministic receipt engineering on this panel. Pure random acquisition is likewise not a replacement policy: only 1.3% of 1,000 trials jointly match fixed confidence's coverage and median. The retained result is narrower—transactional admission lets eight of 28 actions be randomized inside the stronger branch-safe policy without a material distributional penalty. Learning from those labels remains untested. Prospective transfer overturns even that calibrated allowance: on new fields the 8/20 mixture remains perfectly safe but falls to 2.959% median and 2/6 populations above 3%, missing both operational and relative-effect clauses. The eight labels may support a future online update, but they do not rescue this frozen gate.

Fixed nonlinear coordinates do not identify a Wiener–Hammerstein factorization.

On a prospectively frozen source split of the measured SYSID 2009 electronic circuit, twelve independent uniform-cardinal functions of fixed DCT history coordinates reach 40.220 mV RMS. This improves a matched linear FIR by only 5.80% and is 0.656% worse than a smaller additive cubic control. Exact Gram compilation takes only 0.387 s, so increasing evaluation speed or grid density does not address the failure. The static nonlinearity acts on an unknown filtered mixture and therefore induces cross-coordinate terms absent from the additive program. The source gate fails and the official target remains sealed; learned low-rank directions or an explicit causal factorization are required next.

Hidden recurrent topology is not a one-step feature-ranking problem.

On Bouc–Wen source evidence, derivative-space pursuit selects a wrong five-atom law. Output-metric singleton growth selects only y˙∣z∣\dot y|z| and stops, while a local 55-pair tangent ranks the true pair ninth and its full step misses the frozen improvement gate. Only nested calibration of complete recurrent candidate subspaces exposes the complementary value of ∣y˙∣z|\dot y|z. These failures prohibit describing ordinary greedy spline/KAN edge growth as equation discovery in latent-memory systems.

Decision-equivalent sensitivities need not match a coarse nonsmooth finite difference.

The one-pass analytic state Jacobian is 2.86 times faster and chooses the same update, but Gate 266 fails frozen raw-Jacobian and scaled-update tolerances against a 10−310^{-3} relative central difference. Refinement through 10−510^{-5} reduces both discrepancies monotonically and aligns final scores within 0.0614%, yet Gate 267 still misses its 2×10−42\times10^{-4} raw-Jacobian bound. Both failures remain; the supported result is end-to-end time-to-model, not bit-level equality at absolute-value switching events.

Verified backtracking still needs a stopping contract.

Trying the largest candidate step and halving until two independent programs improve reaches lower error in 55.90 s. It nevertheless fails a 16-simulation validation budget because a final cycle, launched after both errors were already below target, rejects four step sizes and consumes eight unnecessary simulations. Safe rollback is not the same as efficient termination.

Probe quality is decision-relative, and relative repair is unstable near zero.

Gate 275's E-optimal probe fails as a scalar wear detector, even though Gates 279–280 show that it conditions arbitrary six-parameter updates extremely well. A directional RMS probe also narrowly fails; only the matched operator statistic detects 0.25% wear reliably. Conversely, Gate 279 formally fails an 80% relative-improvement rule when an already-small 0.00176-mm error falls to 0.00052 mm. Gate 280 then misses a stringent absolute 0.003-mm ceiling at two corners of a ±1%\pm1\% box despite 92.4–98.3% reductions. These are retained metric and tangent-range boundaries, not evidence against the compiler.

Unconditionally committing a Newton correction can regress an already adequate twin.

On a fresh ±2%\pm2\% panel, exact same-observation relinearization reduces the median first-step error by 86.1% and puts every device below 0.000666 mm, yet two tiny first-step errors rise slightly. This falsifies unconditional local monotonicity. Gate 282 resolves the mechanism prospectively with a frozen active-residual acceptance margin; it does not retroactively convert Gate 281 into a pass.

Equation adequacy is not a coefficient-transfer certificate.

Three complementary probes recover the missing y˙z\dot y z topology, but one six-device case still reaches 0.004971 mm on broadband evidence. A fourth longer probe produces a passing fresh panel, yet usually acts only as a below-threshold check. In the subsequent known/unknown-law panel, the adequacy ceiling perfectly commits legal structure and refuses both absent force laws, while one legal commit reaches 0.003357 mm and formally fails transfer. The residual threshold can certify vocabulary insufficiency without certifying every coefficient direction needed by an unseen excitation.

A longer experiment should not be forced merely because it exists.

Gate 292 always forms two full four-capsule precision proposals. Two of six cases reject both under the frozen nonlinear safeguard and total compilation exceeds its time budget. All six retained artifacts nevertheless transfer below 0.000241 mm. Thus more evidence is sometimes useful, sometimes verification- only, and sometimes a harmful local step. A value-of-information rule based on the task-metric Gram inverse is required; unconditional acquisition or updating is not supported.

Learning-time calculus should not leak into deployment.

Gate 293 discovers the newly legal force atom and passes every scientific clause, but its evaluator carries a seven-parameter sensitivity state and takes 77.29 s, formally failing a 15-s deployment budget. Erasing sensitivities gives exact predictions and a 3.04×3.04\times speedup but still takes 23.59 s. Vector fusion crosses the absolute budget at 12.56 s yet misses its 2×2\times relative clause. Only the separately frozen inlined kernel passes. These retained failures distinguish mathematical compileability from an adequately lowered executable.

Nominal value of information is not robust value of information.

Inside the original random-band candidate language, a deployment trace metric, damping-coordinate covariance, and generic minimum-eigenvalue design repeatedly choose the same probe; changing only the scalar objective creates no new information. Typed decay waveforms finally separate the designs and prospectively reduce median damping-coefficient error by 74.39%, but do not improve the complete displacement task. A nominally constrained hybrid then preserves 93.91% of generic conditioning and improves median broadband transfer on fresh systems, yet fails median damping and transient clauses. These failures rule out a single-parent Fisher/Gram certificate as a sufficient nonlinear measurement policy. The required continuation is ensemble-robust design over typed archived programs, not post-hoc adjustment of the mixture or threshold.

A safe acquisition planner must be transactional too.

Minimax aggregation over five local Grams reproduces the nominal hybrid rather than its realized nonlinear risk. Full propose–update–deploy simulation first fails from a non-neutral new-law initializer; the corrected replay safely retains generic but cannot satisfy a clause demanding improvement over generic itself. On new evidence, that refusal reduces worst joint risk by 19.92%, just short of the frozen 20%, and endpoint re-normalization violates bitwise waveform identity. These are not rounded into success. They motivate literal experiment artifacts and a rollback margin, mirroring immutable model commits.

The subsequent eight-system confirmation passes every acquisition-policy clause but still is not an end-to-end pass: one stacked residual exceeds the separate adequacy limit by 0.97%. A targeted extra update closes it and sharply improves known validation, but that repair is retrospective. This distinction prevents a successful sensing decision from laundering a failed model certificate.

The first fully prospective composition also fails and sharpens the stopping contract. A frozen single escalation lowers two large residuals but leaves them above the absolute ceiling, so downstream transfer fails despite a correct acquisition rollback. A retrospective continuation closes them, motivating—but not proving—a bounded residual-monotone loop. Only a further fresh Gate-313 panel validates that loop end to end. Even there the physical hybrid loses four of eight cases and has 2.921 times generic worst risk. The supported claim is safe proposal/rollback and certified recompilation, not a universal task-shaped probe.

Physical state alone is not complete transaction state.

The first interactive MuJoCo probe is force-bounded yet drops an unstable pendulum after ten steps, so action clipping is not an experiment-safety certificate. A stabilizing-policy continuation survives nominally but its first receipt truncates because repeated state restores omit Gymnasium's wrapper step counter. We retain both failures. Versioning that orchestration coordinate gives exact replay, but does not itself establish dynamics learning or control gain.

A symmetric cardinal Gram can still have an additive gauge.

Five partition-of-unity edge bases plus an intercept create five redundant constant coordinates. The first MuJoCo residual compiler is therefore rank deficient and fails independent-solve agreement even though local evaluation is exact and fitting error is small. Contrast projection resolves the algebra on the same sealed data, but external transfer remains untested until validation is generated after artifact freeze.

Prediction repair can have negligible control value.

Gate 319 prospectively reduces one-step error by 99.96% versus a nominal MuJoCo twin and 41.83% versus ridge, while tracking cost is slightly worse than both. All methods also lose the unstable 128-step open-loop forecast. The precommitted control and recursive clauses fail. A mild physical perturbation and a robust feedback controller can make even very large prediction gains operationally irrelevant. Subsequent benchmarks must measure recovery from changes that actually impair operation.

Command symmetry does not guarantee safe acquisition.

On the two-link actuator task, paired opposite commands drift under an asymmetric force law, violating the joint limit after 106 of 128 planned transitions. Feedback-free command symmetry is not a physical safety certificate. The panel stops before fitting and requires a fresh feedback-stabilized successor.

A good forward fit may be a poor inverse near a flat law.

The operational successor retains a model-class negative: fixed-grid smooth cardinal inversion repairs tracking strongly, yet its deadzone RMS is 38.74 times that of a correctly specified parametric law. Near a flat response, small forward approximation errors can have large inverse consequences. The parametric model in turn performs poorly on asymmetric actuation. These results motivate evidence-based law selection, not a universal basis claim.

Useful retrieval need not improve total operation.

The continuous actuator bank saves 79.02% of returning-condition probe actions against recalibration but has 4.446 times the whole-trace error of probe-free online linear adaptation. Most logged reuse events reselect the currently active law after avoidable alarms. A separate saturated-observer exception is caused by trying to fit sensor noise outside the actuator's feasible range, not physical instability. Both negatives require explicit accounting in the next supervisor.

Several boundaries remain genuinely open rather than diagnosed-and-closed. We have not found a cheap method that beats e-prop on dense recurrent networks: the natural candidate is a sparse approximation (SnAp-style) that keeps e-prop's locality while restoring the long-range credit it drops, but a clean cheap-beats-e-prop result on dense RNNs is not yet in hand. We have not demonstrated full Atari RL at scale (only state-recoverable model-based control and a behavior-cloning negative). We have not trained a from-scratch LLM at scale gradient-free (only the mechanism, at small scale). We have not exhibited a clean PC-beats-backprop regime on i.i.d.\ accuracy (only PC == backprop in the correct regime, with prospective configuration negative). And the closed-form ES idea — casting the PGPE mean-update as a return-weighted regression — remains a promising sketch we have not validated on a high-dimensional model-free problem. We list these so that the program treats them as targets, not as settled claims.

Original: paper/v2_sections/A3_negatives.tex · Raw source file

View raw TEX source
\section{Negative results and what did not work}\label{app:negatives}

\paragraph{Off-grid transactions and overlapping atoms.}
Continuous variable projection on the coupled lifelong edge system removed
fixed-grid quantization---on twenty independent streams mean edge error fell
from $0.65664$ to $0.001250$ and median rollout error from $0.12088$ to
$1.46\times10^{-4}$---while all committed-state drift audits were zero.  The
frozen gate nevertheless failed because its implementation counted a rejected
boundary proposal as a support violation.  We retain that negative and define
support certificates on committed state, with separate bitwise rollback for
rejected transactions.  With two overlapping opposite-sign atoms per region,
model order and rollout were recovered but individual centers were not: worst
center error was $0.0383$ and mean edge error $0.0465$.  Function recovery and
latent atom identification are therefore distinct claims.

\paragraph{External raw-coordinate growth.}
On measured coupled-electric-drive data, an eighth-order delay carrier plus a
finite-difference-velocity cardinal residual produced unstable validation
rollout and was rejected before external test access.  The old route remained
bitwise identical, and the banded Gram and four-tap evaluation audits were
$6.09\times10^{-16}$ and $6.59\times10^{-17}$.  This locates the bottleneck in
observable stable state construction and deployment-aligned evidence, rather
than in spline evaluation speed.

\paragraph{Characteristic-coherence replication.}
An independent 512-Halton-probe implementation reproduced the earlier JHTDB
material-coherence negative.  Its matrix-free $15{,}000$-variable SPD normal
reduced characteristic inconsistency by $63.9\%$ and improved over independent
divergence corrections on five of six streams, but still worsened trajectories
relative to raw trilinear interpolation on all six.  Exact minimization of a
misaligned physical penalty is not downstream correctness.

A research program is only as honest as the failures it records. This appendix
documents, deliberately and in detail, every approach we tried that did
\emph{not} work or only partly worked, together with the diagnosis of
\emph{why} it failed. The intent is twofold: to keep the program from wasting
effort re-trying dead ends, and to make the boundary of the approach explicit.
Each entry states what was tried, the result, and the lesson. We group the
entries into clusters; the parenthetical numbers refer to the master
scoreboard.

\subsection{Predictive coding: the large-output-error regime}
\label{app:neg-pc}

\paragraph{Hard-clamp predictive coding loses to backpropagation (53--55).}
Our first cut at predictive coding (PC) on a deep MLP teacher--student task,
using a hard-clamped one-hot target with $T_{\mathrm{infer}}=25$ inference
steps, reached only $0.595$ accuracy against $0.725$ for cross-entropy
backpropagation --- a clear loss. The cause was \emph{not} a loss confound
(backprop-MSE $0.713 \approx$ backprop-CE $0.725$, while PC stayed at $\sim
0.55$) and \emph{not} non-convergence. A per-layer gradient-alignment unit test
(cosine of the PC free-energy gradient against the autograd backprop-MSE
gradient) isolated the mechanism: hard-clamping a one-hot target on an
\emph{untrained} network is the worst case, in which the top-layer gradient
deviates (cosine $0.92$ at the output) and corrupts the descent direction. The
fix is to stay in the small-perturbation regime: a small target nudge
($b=0.1$) recovers uniform per-layer cosine $\approx 0.99$, and Z-IL's
layerwise-timed updates recover the backprop identity exactly (cosine
$1.000$). In the corrected regime PC equals backpropagation
(backprop-MSE $0.678$ vs.\ PC-Z-IL $0.676$ vs.\ PC-nudged $0.674$; the
hard-clamp artefact $0.563$). \textbf{Lesson:} PC does not lose to
backpropagation \emph{per se}; hard-clamping drives a large output error that
degrades the top-layer gradient. Nudge or Z-IL is mandatory.

\paragraph{Deep relaxation PC attenuates the error toward early layers (56,
56b, 60).}
Carrying nudged relaxation PC to convolutional networks lost again
(backprop-CE $0.541$, backprop-MSE $0.545$, PC-nudged at $T=15$ only $0.399$).
The same gradient-alignment test diagnosed an intrinsic pathology of iterative
relaxation: the output error propagates \emph{downward} one layer at a time and
\emph{attenuates with depth}. Even at $300$ settling steps the per-parameter
cosine is $1.0$ at the output but decays to $\approx 0$ at \texttt{conv1} ---
the ePC/EVPE depth-attenuation pathology. A precision-gain schedule helps
monotonically (\texttt{conv1} cosine $0.026 \to 0.163 \to 0.456$ as the gain
goes $1 \to 2 \to 3$ at $T{=}150$) but does not remove the wall, because the
true bottleneck is convergence: deep relaxation simply needs many settling
steps for the error wave to reach the early layers undiminished. \textbf{Fix
and lesson:} replace slow iterative relaxation with a closed-form,
single-sweep equilibrium solve (HGF/ePC-style), which is exact at all depths by
construction because there is no wave to attenuate. ``Replacing iterative
relaxation with a closed-form equilibrium solve'' is precisely the OSNR/Gram
principle, so PC folds into our closed-form-solve family.

\paragraph{Per-sample RMS error normalization breaks PC (56b).}
A Meta-PCN-style per-layer RMS normalization of the prediction errors, intended
to counter depth-attenuation, broke training instead: the output gradient
cosine flipped to $-0.89$. \textbf{Lesson:} per-sample RMS normalization
destroys the relative scale of the error signal that the gradient identity
depends on; it is the wrong precision mechanism. Closed-form precision is the
right one.

\paragraph{Categorical full-CE-force PC is again the large-error regime (59).}
A cross-entropy readout for PC is implementable, but applying the \emph{full}
CE force (no nudge) reproduced the hard-clamp failure: on the deep MLP
teacher--student task backprop-CE reached $0.716$ and backprop-MSE $0.712$,
while full-force PC-CE managed only $0.653$. This is the same large-error
effect that produced the hard-clamp MSE result. \textbf{Lesson:} the readout
form is not the issue; applying full target force is. A nudged or
precision-scaled categorical readout is the clean formulation if pursued.

\paragraph{Prospective configuration does not beat backprop on i.i.d.\
accuracy (61--62).}
Song et al.'s prospective-configuration motivation predicts that relaxation
finds a better activity configuration before plasticity, which could beat
backpropagation on data efficiency or online learning. Two controlled attempts
(matched objective, initialization, seed, and optimizer; three seeds each) said
otherwise. On the data-efficiency sweep (train sizes $250$--$2000$), PC did not
beat backprop. Online/few-pass full-clamp PC --- the genuine prospective
regime --- also failed to win: e.g.\ epoch~1 batch~16 gave backprop $0.4197$,
PC-nudged $0.4286$, PC-hard $0.3835$. \textbf{Lesson:} nudged PC is provably
$\approx$ backprop, so it cannot beat backprop by much; full-clamp prospective
PC deviates from backprop but underperforms on accuracy. There is no robust
PC-beats-backprop regime in our i.i.d.\ settings. The remaining documented
regime where prospective configuration might still win is
continual/interference, which is handled instead by our closed-form Gram
memory.

\paragraph{PC-NEAT shows no evolution gain when the seed already matches the
teacher or the task is linearly decodable (58, 66).}
When PC is used as the inner learning rule for evolved NEAT topologies, the
expected evolution win did not materialize on tasks where the seed network
already matched the teacher or where the target was linearly decodable: the
MSE/capacity ceiling is reached without any topology search, so evolution has
nothing to add. The bridge mechanism itself is sound --- PC trains an arbitrary
evolved skip-DAG locally, matching backprop (e.g.\ PC-nudged $0.5874$ on the
evolved graph) --- but a \emph{clean} PC-NEAT win needs a task where topology is
load-bearing. \textbf{Lesson:} re-frame the probe on parity, where the topology
is provably load-bearing, before claiming an evolution gain.

\subsection{Bolt-on credit-assignment shortcuts on convolution}
\label{app:neg-dfa}

\paragraph{DFA / direct-feedback bolt-ons do not beat backprop on conv, and
scale-up widens the gap (48--50).}
We made three attempts to \emph{beat} backpropagation on convolutional vision by
adding a Direct Feedback Alignment (DFA) global-error broadcast to InfoPro-style
local learning (the brain's putative two-signal arrangement). All three failed.
A DFA broadcast at $\gamma{=}1.0$ gave $0.741$ and at $\gamma{=}0.05$ gave
$0.762$; a scale-up attempt reached $0.854$ versus backprop's $0.875$ ---
i.e.\ widening width and epochs \emph{widened} the gap rather than closing it.
The diagnosis is that local block training operates on a moving, detached output
target and is inherently less stable, while the DFA term adds noise; both hurt.
\textbf{Verdict and lesson:} bio-local credit assignment \emph{matches}
backprop on conv within $\sim 1\,\sigma$ but does not beat it, and DFA is not
the lever that turns a match into a win. We claim ``matches,'' not ``beats.''

\subsection{Gradient-free deep credit assignment on adversarial tasks}
\label{app:neg-gfree}

\paragraph{Crude gradient-free deep credit assignment lags backprop on
kernel-hard compositional tasks (29).}
On the adversarially-hardest probe --- parity of the signs of three hidden
random hyperplanes in $30$ dimensions, where a representation \emph{must} be
learned --- backprop reached $0.88$, while CEM-evolved first-layer features plus
a closed-form readout reached $0.61$ and closed-form target propagation only
$0.57$ (random/greedy features sit at chance, $0.53$). Gradient-free credit
assignment is therefore \emph{real} (it decisively beats random) but
substantially weaker than gradient descent at learning deep representations on a
genuinely compositional task. \textbf{Lesson:} on hard compositional deep credit
assignment, backprop's gradient-through-depth is a real moat; the practical
verdict is a hybrid program (backprop or a pretrained backbone for the deep
representation, our closed-form/evolution machinery for the structured,
continual, control, and self-scaling parts). Caveat: parity is the worst case,
and sophisticated gradient-free methods (difference target propagation with
learned inverses, large-budget CMA-ES, NEAT block-growth weight evolution) might
narrow the gap --- an open question, not a closed one. Note that for
\emph{decoupled-greedy} local learning specifically the wall was later overturned
(it matches backprop, with the DFA \emph{beat} attempt above being the part that
failed); the residual hard case is precisely this crude-gradient-free
compositional regime.

\paragraph{Label-free whole-image contrastive vision fails;
language-from-scratch gradient-free is ruled out at scale (V1; 29 cluster).}
An earlier label-free contrastive vision objective learned augmentation
invariance (the loss dropped) but produced near-random downstream accuracy; the
label-free win came instead from sparse coding, not whole-image contrastive
learning. Separately, the broader ``learn language from scratch, fully
gradient-free, at scale'' ambition was ruled out as a near-term claim: the
mechanism is proven (a character-level GPT trains by the per-block local sweep
to within $\sim 0.013$ bits-per-character of backprop), but a scaled result is a
hybrid verdict, not a gradient-free-from-scratch one. \textbf{Lesson:} contrastive
is the wrong label-free objective for our setting (sparse coding is right), and
``from-scratch gradient-free language at scale'' is a compute/cluster question
we have not run.

\subsection{Evolution strategies and exploration}
\label{app:neg-es}

\paragraph{ES/PGPE fails on cartpole; a genetic algorithm trivially solves it
(76).}
Asking whether evolution strategies could broadly boost the framework, we
compared a genetic algorithm against a quick antithetic, rank-normalized ES on
cartpole-balance (reactive policy, no backprop in any method). The GA solved it
trivially (best $471$ in $320$ episodes), while the ES did not solve it across
four configurations (best $\sim 8$--$10$). The diagnosis is the deceptive
landscape: near random initialization almost every policy dies in $\sim 8$
steps, so single-point ES perturbations do not differentiate and the
mean-update stalls, whereas the GA's \emph{population} effectively random-searches
and finds a surviving policy by luck. This is exactly the deceptive flat-reward
failure that motivated novelty search. (A properly tuned ES is known to solve
this toy but offers no advantage over the GA on such a low-dimensional problem.)
\textbf{Lesson:} ES/PGPE is a niche tool for high-dimensional, smooth-landscape,
model-free, non-differentiable optimization --- not a universal accelerator. Our
reliable edge for control is model-based MPC ($\sim 28{,}800\times$ sample
efficiency over model-free) plus local/closed-form learning. The on-brand idea
kept for later is closed-form ES (the PGPE mean-update as a return-weighted
regression of returns on perturbations), to revisit only when we tackle
high-dimensional model-free RL.

\subsection{Closed-form representations on broadband signals}
\label{app:neg-broadband}

\paragraph{Closed-form SIREN with matched dominant modes underfits broadband
images (80).}
Fitting a coordinate-to-color image map, a deep SGD-trained SIREN reached
$59.6$\,dB PSNR; the closed-form random-Fourier-feature ridge reached $28.1$\,dB
($226\times$ faster, no backprop), and a closed-form representation matched to
the energy-concentrated dominant modes did worse still at $10.4$\,dB
(underfitting). \textbf{Lesson and boundary:} this is the SSP boundary made
sharp. For broadband natural images the generating operator is trivial and the
innovation is dense, so broad frequency \emph{coverage} beats matched
dominant-mode concentration, and a shallow one-layer feature ridge cannot match
a deep network's learned-frequency fidelity. Operator-matched closed form is
decisive on \emph{structured} signals (PDEs, dynamical systems --- two to three
orders of magnitude over FNO/Neural-ODE) and explicitly not on broadband natural
content; we do not claim otherwise.

\subsection{Reinforcement learning at scale via imitation}
\label{app:neg-bc}

\paragraph{Atari Pong behavior cloning fails to play (86).}
On real ALE Pong with a no-backprop local-sweep convolutional policy
behavior-cloning a heuristic teacher, the best attempt used a predictive teacher
(extrapolating ball trajectory with wall bounces) and frame-diff input: the
teacher was good ($-10.6$ versus random $-20.8$) and gradient-free perception
cloned it decently ($0.794$ BC accuracy), yet the cloned policy scored at the
floor ($-21.0$, losing every point). \textbf{Diagnosis:} a textbook
behavior-cloning failure --- $79\%$ average action accuracy misses the rare
\emph{critical} intercept frames, and BC distribution shift compounds as the
policy's own rollout drifts off the teacher's distribution, worsened by the
two-pixel ball being hard to localize. \textbf{Lesson:} the gradient-free
perception mechanism works, but BC $\neq$ RL; Atari mastery needs interactive RL
at scale (DAgger, model-based MPC on learned latent dynamics, or DQN-scale
frames). The no-backprop model-based recipe \emph{is} validated where state is
recoverable (Catch from pixels, $96.5\%$).

\subsection{Self-improvement without representation learning}
\label{app:neg-gfrsi}

\paragraph{Gradient-free closed-form self-training on a fixed feature map degrades; it does not compound.}
The stable recursive self-improvement we demonstrate (App.~\ref{app:record}, stable-RSI) uses backprop to
self-improve a \emph{trainable} representation. We asked whether the two stabilizers that made it work---the
never-forget Gram memory and the confidence curriculum---suffice to make a \emph{fully gradient-free} loop
compound on a \emph{fixed} feature map (no backprop anywhere): a tiny labelled seed, the closed-form Gram
readout self-labels an unlabelled pool, confident labels are additively accumulated, re-solve. They do not.
Sweeping seed strength ($25$--$500$/class) and two gradient-free bases (random-conv, supervised-ridge ceiling
$0.556$; k-means Coates--Ng, ceiling $0.704$) on CIFAR-10, the seed-only static solve is \emph{always} the best
and every self-training run degrades monotonically in how much pool it admits (canonical: Coates--Ng, $100$/class,
$0.503\!\to\!0.406$, peaking at round~1 with a marginal $+0.007$ before eroding; on the weak base it collapses
$0.31\!\to\!0.15$). The never-forget memory is consistently slightly more robust than naive re-fitting (its
order-invariant accumulation does not oscillate) but cannot turn a net-negative signal positive. \textbf{Diagnosis:}
on a fixed feature map the pseudo-labels are produced by the same features the readout consumes, so they carry no
information the seed cannot already extract (they cannot exceed it) while they do carry label noise (they degrade
it). \textbf{Lesson, and it sharpens the positive result:} self-improvement \emph{requires representation learning}.
Backprop-EMA RSI compounds precisely because backprop reshapes the features using the unlabelled data; closed-form
on fixed features has no mechanism to change the representation, so it can only add noise. The closed-form
never-forget memory is the right tool for continual learning with \emph{new labelled} data/classes (it beats
buffered replay, App.~\ref{app:record}) but not for bootstrapping from \emph{unlabelled} data on a frozen
representation.

\paragraph{A plastic encoder cannot beat a frozen one under a closed-form continual memory: capture vs.\ retain is irreducible.}
The EMA slow-anchor recovers retention for a plastic representation when the encoder is already good (App.~\ref{app:record}), but does it ever
\emph{win}? We tested the genuine case---continually fine-tuning a pretrained ResNet18 on an out-of-domain medical stream (PathMNIST $\to$
BloodMNIST $\to$ DermaMNIST $\to$ RetinaMNIST, $29$ classes), where adaptation should pay. It does pay per task (live fine-tuning standalone
$0.703$ vs frozen $0.669$), yet EMA-anchored plasticity never beats frozen across the whole momentum range: final accuracy $0.142/0.466/0.527/0.652$
at momentum $0.80/0.90/0.95/0.99$ versus frozen $0.662$, the best ($0.99$) merely matching it. \textbf{Diagnosis:} an irreducible capture-vs-retain
tradeoff. Realising the per-task gain requires the memory to read the \emph{current} fast-moving encoder, but then old-task statistics are stale and
forgetting is catastrophic (live $0.059$); reading a slow EMA retains (forgetting $0.019$ at momentum $0.99$) but the EMA lags so far that it captures
essentially none of the gain (standalone pinned at $\sim$$0.666\approx$ frozen at \emph{every} momentum). No setting captures and retains at once.
\textbf{Lesson:} this delimits precisely where the slow anchor works. It stabilises self-improvement when the target is \emph{fixed} (the same task,
a representation improving toward a stable optimum---the EMA tracks it and the memory stays valid), but it cannot serve continual learning across
\emph{different} tasks, where the representation must move to new domains and a single memory cannot both follow it (to capture) and stay anchored
(to retain). Freezing a strong pretrained encoder is therefore not merely pragmatic but essentially optimal for closed-form continual memory; making
the shared encoder plastic creates a tradeoff EMA cannot break. This closes the ``anchor a drifting representation'' frontier with data.

\subsection{Open problems}
\label{app:neg-open}

\paragraph{Instantaneous divergence repair does not improve Lagrangian
trajectories.}
On six preregistered five-frame JHTDB streams, the confirmed static cardinal
divergence residual improves global blind velocity on all 30 frames but worsens
particle endpoint and all-time errors on 6/6 streams.  Static held-out sensors
select the harmful full correction, and a matrix-free 15,000-coefficient
material-coherence extension also fails 6/6.  The result rules out the tempting
inference that stronger incompressibility alone implies a better flow map; a
compiled operator must govern the downstream functional being evaluated.

\paragraph{Pooled and empirical Gram memory do not protect a continuous function.}
In the lifelong constitutive-edge gate, additive banded statistics recover the
order-invariant pooled ridge optimum and improve physical rollout, yet an
accepted noisy revisit produces up to 68.5\% relative regression from the best
prior regional error. Constraining updates to coefficient directions with zero
old empirical-Gram diagonal makes every old sampled prediction exactly
invariant, but unseen points within the same interval can regress by 138\%.
Closing the entire continuous basis support removes that drift exactly, but a
dense one-shot local fit then misses its plasticity threshold. These three
negatives motivate the successful sparse support-closed rule: sufficient
statistics remember evidence, continuous support protects functions, and
parsimonious growth is needed to avoid freezing noise.

\paragraph{A compact cardinal residual is not the right block for globally smooth EMPS mismatch.}
On measured positioning dynamics, compiling the exact kinematic equation and
fitting only acceleration makes every held-out 500-step rollout stable. The
cardinal velocity edge is useful relative to the linear carrier
($0.28077$ versus $0.34014$ mean rollout nRMSE), and its seven-diagonal Gram
and four-tap evaluation agree with dense references to $7.40\times10^{-16}$
and $8.33\times10^{-17}$. It nevertheless loses decisively to a three-term
global polynomial residual ($0.18346$); the external test was therefore not
opened. This closes fixed-smoother/cardinal tuning on EMPS and yields a block
routing rule: global smooth residuals belong in an operator reproduction
space, whereas compact cardinal innovations require independent evidence of
local structure.

\paragraph{Static and fixed-memory F-16 interface shortcuts do not transfer.}
On full-scale F-16 ground-vibration estimation data, an operator-derived
payload-interface coordinate is useful but does not by itself justify local
tensor structure: its 2D cardinal map lowers carrier error 21.9\% yet is 0.90\%
worse than a global polynomial on a held-out amplitude. Closing a map learned
from observed interface state by fixed-point iteration either diverges
(polynomial) or remains 32--33\% away from a fixed point (compact cardinal).
A cardinally regularized play bank improves internal error only 1.60\%, below
its frozen 2\% commit threshold. Finally, an external regime router cannot
transfer unchanged FullMSine coefficients monotonically to sine sweeps: its
route complexity is $0,3,3,1$, and the stable exponential block fails
catastrophically at the highest estimation amplitude. Sine-sweep validation
remains unopened. Coefficient-separable and combined impulse-response gain
bounds are valid but force correction scales below $7\times10^{-4}$ and
therefore collapse every dynamic route to the polynomial; a 50-sample passive
router chooses the carrier at every amplitude because its shared startup
transient is uninformative. These negatives rule out more knots, relaxation
tuning, global worst-case projection, and force-RMS/startup routing; the open
requirement is a reachable-state gain or passivity certificate.

\paragraph{Input features, low-rank weights, and unconstrained cardinal
compression do not preserve F-16 free-run geometry.}
On the published SpecialOddMSine split, the FullMSine
cardinal--exponential residual does not transfer across random phases
($0.82835$ versus $0.82794$ carrier), while a causal stable modal carrier is a
17.0\% positive compression result ($0.68696$). Adding all quadratic/cubic
products of eight dominant modes changes this to only $0.68709$. A stable
64-lag ARX observer plus a four-coordinate cardinal innovation reaches
$0.69810$; autoregression alone is not the missing mechanism.

A locally reproduced 256-unit sigmoid observer reaches $0.50061$, but directly
compiling its 16 most response-informative ridge directions into uniform
cardinal edges gives teacher-forced RMSE $0.16596$ and free-run RMSE $1.73593$.
The divergence begins at the second autonomous sample. Exact derivative-Gram
regularization repairs stability and improves the modal anchor to $0.64654$,
but misses the prospective $0.60$ gate and yields a $3.32\times10^{-9}$
independent-solve discrepancy in its strongest ill-conditioned system.
Ordinary low-rank arguments also fail: approximations retaining 98--98.6\% of
teacher weight energy roll out above $3.4$. These results rule out widening,
rank-energy compression, and unconstrained edge pruning as explanations. The
successful repair changes the recurrent state itself
(Sec.~\ref{sec:f16-neural-compiler}), substituting stable modal history before
cardinal compilation.

\paragraph{Static amplitude coordinates do not supply missing hysteretic state.}
On 108,477-sample F-16 SineSweep estimation records, the teacher-free
Gram--Krylov compiler learned only at low amplitudes transfers positively to
Level 7: all channels improve and mean nRMSE falls 18.59\%.  This narrowly
misses a frozen 20\% gate, so even validation levels remain sealed.  Three
theory-motivated repairs then fail without accessing them.  A single pooled
cardinal map gives only 0.99\% mean leave-one-level-out gain and regresses
Level 5; exact zero-forget regional maps with force-RMS coefficient
interpolation regress Level 5 by 24.4\%; and a 1,361-feature amplitude--state
tensor field, regularized by exact Kronecker value/derivative Grams, improves
two interior folds only 1.3--3.1\%.  The tensor Gram action itself agrees to
$2.22\times10^{-16}$.  These negatives distinguish correct spline calculus
from sufficient physical state: record-level amplitude does not encode the
evolving constitutive memory of a sweep.

\paragraph{A bounded residual cannot replace a missing operator nullspace.}
For scalar-position EMPS, a strictly stable input-only exponential bank has
confirmation nRMSE $3.671$; cardinal compilation halves it to $1.769$ but
remains far behind a projection polynomial.  Adding exactly the repeated-root
coordinates of the double-integrator operator---single and double force
integrals---reduces carrier error by 85.3\% to $0.540$, and cardinal compilation
improves it another 13.3\% to $0.468$.  A cubic projection residual still
reaches $0.326$, so the historical test remains unopened.  Exact cumulative-
sum and chunked integrator implementations agree to $2.07\times10^{-15}$.
This is a representation boundary: repeated zero roots and polynomial
reproduction are load-bearing, but input integration cannot replace observable
velocity/friction state.

\paragraph{Covariance-dual coordinates do not solve coupled F-16 dynamics.}
After independent synthetic confirmation, a three-edge Riesz map and a
15-edge multi-shift Riesz bank were frozen for the existing F-16 chronology.
Validation selected the latter, but untouched realization 8 reaches $0.45354$,
0.20\% worse than the raw-Krylov compiler, with only $1.06\times$ cardinal-
component compression. All projected solves and cardinal audits pass below
$2.14\times10^{-15}$. This rejects universal transfer of the single-index
elliptical argument: several lagged dynamic coordinates can contribute to one
output, so first-order covariance geometry collapses necessary structure.

\paragraph{Fixed Hermite-score sketches can omit weak physical modes.}
Matrix-free contracted score actions reduce operator workspace $3.29\times$
and predictive error 72--73\% in two 64-dimensional confirmations. Yet one
weak generating direction is recovered at cosine $0.8811$, below the frozen
0.90 threshold, even while its cardinal predictor beats linear regression by
more than threefold. Predictive error is therefore not a completeness
certificate. Adaptive sketches need independent evidence-ledger stability or
another mode certificate; increasing sketch width after observing this miss
would be post-hoc tuning.

\paragraph{Confidence is not value of information.}
On a newly acquired JHTDB panel, confidence-triggered transactional sensing is
safe and effective at 2.56-times fewer readings but trails dense sensing by
0.529 points. A cross-panel ridge predictor, a nonlinear 31-feature operator-
receipt model, an action-aligned transaction-value target, and a 16-reading
sentinel router all fail their frozen policy gates. The rich receipt does rank
held-out uplift positively (Spearman 0.304), and a same-budget hindsight
transaction exceeds the dense median, so the negative is not lack of headroom.
It is an observability boundary: endogenous validation statistics are too weak
to decide the counterfactual value of unseen evidence. Reusing paid sentinels
in the non-acquisition branch recovers six wins without extra cost, but still
misses the effect margin; applying that reuse to the old router adds sensing
without fixing selection. Future acquisition must learn from controlled pilot
updates and disjoint verification, not another threshold on the same receipt.
That progressive 64/8/8/176 test is also negative: pilot-induced receipt
changes rank admitted value across the training panels but transfer at only
0.189 Spearman, yielding 247 wins and 4.021\% versus fixed confidence's 248 and
4.042\% at equal cost. The corresponding hindsight policy reaches 4.581\%, so
the remaining gap requires observed exploration outcomes, not more deterministic
receipt engineering on this panel.
Pure random acquisition is likewise not a replacement policy: only 1.3\% of
1,000 trials jointly match fixed confidence's coverage and median. The retained
result is narrower---transactional admission lets eight of 28 actions be
randomized inside the stronger branch-safe policy without a material
distributional penalty. Learning from those labels remains untested.
Prospective transfer overturns even that calibrated allowance: on new fields
the 8/20 mixture remains perfectly safe but falls to 2.959\% median and 2/6
populations above 3\%, missing both operational and relative-effect clauses.
The eight labels may support a future online update, but they do not rescue this
frozen gate.

\paragraph{Fixed nonlinear coordinates do not identify a Wiener--Hammerstein
factorization.}
On a prospectively frozen source split of the measured SYSID 2009 electronic
circuit, twelve independent uniform-cardinal functions of fixed DCT history
coordinates reach 40.220 mV RMS. This improves a matched linear FIR by only
5.80\% and is 0.656\% worse than a smaller additive cubic control. Exact Gram
compilation takes only 0.387 s, so increasing evaluation speed or grid density
does not address the failure. The static nonlinearity acts on an unknown
filtered mixture and therefore induces cross-coordinate terms absent from the
additive program. The source gate fails and the official target remains
sealed; learned low-rank directions or an explicit causal factorization are
required next.

\paragraph{Hidden recurrent topology is not a one-step feature-ranking
problem.}
On Bouc--Wen source evidence, derivative-space pursuit selects a wrong
five-atom law. Output-metric singleton growth selects only
$\dot y|z|$ and stops, while a local 55-pair tangent ranks the true pair ninth
and its full step misses the frozen improvement gate. Only nested calibration
of complete recurrent candidate subspaces exposes the complementary value of
$|\dot y|z$. These failures prohibit describing ordinary greedy spline/KAN
edge growth as equation discovery in latent-memory systems.

\paragraph{Decision-equivalent sensitivities need not match a coarse
nonsmooth finite difference.}
The one-pass analytic state Jacobian is 2.86 times faster and chooses the same
update, but Gate 266 fails frozen raw-Jacobian and scaled-update tolerances
against a $10^{-3}$ relative central difference. Refinement through $10^{-5}$
reduces both discrepancies monotonically and aligns final scores within
0.0614\%, yet Gate 267 still misses its $2\times10^{-4}$ raw-Jacobian bound.
Both failures remain; the supported result is end-to-end time-to-model, not
bit-level equality at absolute-value switching events.

\paragraph{Verified backtracking still needs a stopping contract.}
Trying the largest candidate step and halving until two independent programs
improve reaches lower error in 55.90 s. It nevertheless fails a 16-simulation
validation budget because a final cycle, launched after both errors were
already below target, rejects four step sizes and consumes eight unnecessary
simulations. Safe rollback is not the same as efficient termination.

\paragraph{Probe quality is decision-relative, and relative repair is unstable
near zero.}
Gate 275's E-optimal probe fails as a scalar wear detector, even though Gates
279--280 show that it conditions arbitrary six-parameter updates extremely
well. A directional RMS probe also narrowly fails; only the matched operator
statistic detects 0.25\% wear reliably. Conversely, Gate 279 formally fails an
80\% relative-improvement rule when an already-small 0.00176-mm error falls to
0.00052 mm. Gate 280 then misses a stringent absolute 0.003-mm ceiling at two
corners of a $\pm1\%$ box despite 92.4--98.3\% reductions. These are retained
metric and tangent-range boundaries, not evidence against the compiler.

\paragraph{Unconditionally committing a Newton correction can regress an
already adequate twin.}
On a fresh $\pm2\%$ panel, exact same-observation relinearization reduces the
median first-step error by 86.1\% and puts every device below 0.000666 mm, yet
two tiny first-step errors rise slightly. This falsifies unconditional local
monotonicity. Gate 282 resolves the mechanism prospectively with a frozen
active-residual acceptance margin; it does not retroactively convert Gate 281
into a pass.

\paragraph{Equation adequacy is not a coefficient-transfer certificate.}
Three complementary probes recover the missing $\dot y z$ topology, but one
six-device case still reaches 0.004971 mm on broadband evidence. A fourth longer
probe produces a passing fresh panel, yet usually acts only as a below-threshold
check. In the subsequent known/unknown-law panel, the adequacy ceiling perfectly
commits legal structure and refuses both absent force laws, while one legal
commit reaches 0.003357 mm and formally fails transfer. The residual threshold
can certify vocabulary insufficiency without certifying every coefficient
direction needed by an unseen excitation.

\paragraph{A longer experiment should not be forced merely because it exists.}
Gate 292 always forms two full four-capsule precision proposals. Two of six cases
reject both under the frozen nonlinear safeguard and total compilation exceeds
its time budget. All six retained artifacts nevertheless transfer below
0.000241 mm. Thus more evidence is sometimes useful, sometimes verification-
only, and sometimes a harmful local step. A value-of-information rule based on
the task-metric Gram inverse is required; unconditional acquisition or updating
is not supported.

\paragraph{Learning-time calculus should not leak into deployment.}
Gate 293 discovers the newly legal force atom and passes every scientific clause,
but its evaluator carries a seven-parameter sensitivity state and takes 77.29 s,
formally failing a 15-s deployment budget. Erasing sensitivities gives exact
predictions and a $3.04\times$ speedup but still takes 23.59 s. Vector fusion
crosses the absolute budget at 12.56 s yet misses its $2\times$ relative clause.
Only the separately frozen inlined kernel passes. These retained failures
distinguish mathematical compileability from an adequately lowered executable.

\paragraph{Nominal value of information is not robust value of information.}
Inside the original random-band candidate language, a deployment trace metric,
damping-coordinate covariance, and generic minimum-eigenvalue design repeatedly
choose the same probe; changing only the scalar objective creates no new
information. Typed decay waveforms finally separate the designs and
prospectively reduce median damping-coefficient error by 74.39\%, but do not
improve the complete displacement task. A nominally constrained hybrid then
preserves 93.91\% of generic conditioning and improves median broadband transfer
on fresh systems, yet fails median damping and transient clauses. These failures
rule out a single-parent Fisher/Gram certificate as a sufficient nonlinear
measurement policy. The required continuation is ensemble-robust design over
typed archived programs, not post-hoc adjustment of the mixture or threshold.

\paragraph{A safe acquisition planner must be transactional too.}
Minimax aggregation over five local Grams reproduces the nominal hybrid rather
than its realized nonlinear risk. Full propose--update--deploy simulation first
fails from a non-neutral new-law initializer; the corrected replay safely
retains generic but cannot satisfy a clause demanding improvement over generic
itself. On new evidence, that refusal reduces worst joint risk by 19.92\%, just
short of the frozen 20\%, and endpoint re-normalization violates bitwise
waveform identity. These are not rounded into success. They motivate literal
experiment artifacts and a rollback margin, mirroring immutable model commits.

The subsequent eight-system confirmation passes every acquisition-policy clause
but still is not an end-to-end pass: one stacked residual exceeds the separate
adequacy limit by 0.97\%. A targeted extra update closes it and sharply improves
known validation, but that repair is retrospective. This distinction prevents a
successful sensing decision from laundering a failed model certificate.

The first fully prospective composition also fails and sharpens the stopping
contract. A frozen single escalation lowers two large residuals but leaves them
above the absolute ceiling, so downstream transfer fails despite a correct
acquisition rollback. A retrospective continuation closes them, motivating---but
not proving---a bounded residual-monotone loop. Only a further fresh Gate-313
panel validates that loop end to end. Even there the physical hybrid loses four
of eight cases and has 2.921 times generic worst risk. The supported claim is
safe proposal/rollback and certified recompilation, not a universal task-shaped
probe.

\paragraph{Physical state alone is not complete transaction state.}
The first interactive MuJoCo probe is force-bounded yet drops an unstable
pendulum after ten steps, so action clipping is not an experiment-safety
certificate. A stabilizing-policy continuation survives nominally but its first
receipt truncates because repeated state restores omit Gymnasium's wrapper step
counter. We retain both failures. Versioning that orchestration coordinate gives
exact replay, but does not itself establish dynamics learning or control gain.

\paragraph{A symmetric cardinal Gram can still have an additive gauge.}
Five partition-of-unity edge bases plus an intercept create five redundant
constant coordinates. The first MuJoCo residual compiler is therefore rank
deficient and fails independent-solve agreement even though local evaluation is
exact and fitting error is small. Contrast projection resolves the algebra on
the same sealed data, but external transfer remains untested until validation is
generated after artifact freeze.

\paragraph{Prediction repair can have negligible control value.}
Gate 319 prospectively reduces one-step error by 99.96\% versus a nominal
MuJoCo twin and 41.83\% versus ridge, while tracking cost is slightly worse than
both. All methods also lose the unstable 128-step open-loop forecast. The
precommitted control and recursive clauses fail. A mild physical perturbation
and a robust feedback controller can make even very large prediction gains
operationally irrelevant. Subsequent benchmarks must measure recovery from
changes that actually impair operation.

\paragraph{Command symmetry does not guarantee safe acquisition.}
On the two-link actuator task,
paired opposite commands drift under an asymmetric force law, violating the
joint limit after 106 of 128 planned transitions. Feedback-free command
symmetry is not a physical safety certificate. The panel stops before fitting
and requires a fresh feedback-stabilized successor.

\paragraph{A good forward fit may be a poor inverse near a flat law.}
The operational successor retains a model-class negative: fixed-grid
smooth cardinal inversion repairs tracking strongly, yet its deadzone RMS is
38.74 times that of a correctly specified parametric law. Near a flat response,
small forward approximation errors can have large inverse consequences.
The parametric model in turn performs poorly on asymmetric actuation. These
results motivate evidence-based law selection, not a universal basis claim.

\paragraph{Useful retrieval need not improve total operation.}
The continuous
actuator bank saves 79.02\% of returning-condition probe actions against
recalibration but has 4.446 times the whole-trace error of probe-free online
linear adaptation. Most logged reuse events reselect the currently active law
after avoidable alarms. A separate saturated-observer exception is caused by
trying to fit sensor noise outside the actuator's feasible range, not physical
instability. Both negatives require explicit accounting in the next supervisor.

Several boundaries remain genuinely open rather than diagnosed-and-closed. We
have not found a \emph{cheap} method that beats e-prop on dense recurrent
networks: the natural candidate is a sparse approximation (SnAp-style) that
keeps e-prop's locality while restoring the long-range credit it drops, but a
clean cheap-beats-e-prop result on dense RNNs is not yet in hand. We have not
demonstrated \emph{full Atari RL at scale} (only state-recoverable model-based
control and a behavior-cloning negative). We have not trained a
\emph{from-scratch LLM at scale} gradient-free (only the mechanism, at small
scale). We have not exhibited a \emph{clean PC-beats-backprop regime} on i.i.d.\
accuracy (only PC $=$ backprop in the correct regime, with prospective
configuration negative). And the \emph{closed-form ES} idea --- casting the
PGPE mean-update as a return-weighted regression --- remains a promising sketch
we have not validated on a high-dimensional model-free problem. We list these so
that the program treats them as targets, not as settled claims.