Research record

Full experimental record

Historical source. Some claims in older records were subsequently corrected. The associated article states the adopted interpretation. This record preserves the original source alongside its rendered reading view.

Rendered archival TeX

This is an HTML reading rendition of the local TeX record. Mathematical notation is rendered with KaTeX; archived figures are included when their source assets are part of this collection.

Complete experimental record

This appendix is a faithful, archival record of every experiment in the program, so the paper is self-contained and nothing is lost. The single organizing principle throughout: encode known structure (an operator LL) in the representation and solve in closed form; where structure is absent, a generic random feature map with a quadratic/ridge readout is already optimal (the SSP law). Matched beats generic iff LL is non-trivial; when LL is trivial and the innovation is Gaussian (e.g.\ frozen vision features) matched == random, provably. The advantage is accuracy where LL is known and categorical efficiency (no backprop, no replay buffer) everywhere. Honest negatives are kept, not hidden. Each row cites a master-scoreboard entry (#); accuracies are test accuracy unless noted, and ``CF'' abbreviates our closed-form learner.

Operator-matched closed-form: system ID, PDEs, FRI, neural operators

The decisive regime: where LL is known, the matched closed-form solve beats deep neural operators by orders of magnitude in both error and wall-clock, and recovers exact structure from real measured data.

p3.0cmp4.2cmp4.6cmp2.2cm@ Experiment (#)SetupKey numbersVerdict
ODE system ID (S1)matched dictionary ++ ridge; autograd-free spline derivativesnon-poly extrap ∼\sim400–2000×\times lower error, ∼\sim1000×\times faster; N=20N{=}20 stable while poly library diverges by N=6N{=}6; conductance neuron 30–70×\timesdecisive (known LL)
PDE vs neural operators (S2)closed-form weak/strong-form ID vs FNO / DeepONet1D Burgers: 2-traj beats 32-traj FNO; 2D/3D react-diff Pareto-dominant; chaotic KS 0.19% (weak form); cross-regime re-IDs ν\nu, <<0.6%decisive; gap grows 1D→\to3D
3D PDE at scale (S2)M=32M{=}32, hard FNO 2000 ep on MPSOSNR ∼\sim0.001 nRMSE in <<0.7 s vs FNO 0.13–0.72 in 120 s ⇒\Rightarrow ∼\sim100–600×\times error, ∼\sim150–1700×\times speed; exact at n=2n{=}2decisive
Real measured data (S4)Silverbox, EMPS, Cascaded Tanks, Allen neuronSilverbox 1.8 mV free-run; EMPS 14 mm (matched friction →\to stability); Cascaded Tanks 0.55 V latent-state mechanism (retrospective protocol audit); Allen τ=28.7\tau{=}28.7 ms recoveredcompetitive / mixed
Closed-form SIREN (#80)Fourier-feature ridge vs deep sin-MLP, 64×\times64 imageRFF F=1024: 28.1 dB in 30 ms (226×\times faster); matched dominant-mode 10.4 dB (underfits); SGD SIREN 59.6 dBhonest: broadband natural →\to deep wins; SSP boundary
Boundaries (S6)matched feature maps on unstructured frozen features; blind eq.\ discoverymatched (det.\ \stochastic) ++ sparse/MAP all tie/lose to random++ridge (trivial LL); SINDy wins blind discovery
Operator-matched closed-form learning. The matched advantage requires LL non-trivial and the dynamics observable (scope law); on unstructured / broadband signals matched == random, as SSP predicts.

Continual learning, RSI, and multimodal generality

The generality test: one closed-form Gram memory matches deep continual-learning SOTA on identical frozen features while using 100100–1000×1000\times less compute and zero buffer, and is the stability primitive for recursive self-improvement and multimodal composition.

p2.8cmp4.4cmp4.8cmp2.0cm@ Experiment (#)SetupKey numbersVerdict
Class-IL benchmarks (S3)CF (RanPAC-style) on frozen backbones, no task labelsSplit-CIFAR-100/DINOv2-L 0.923 (DER++ 0.904, EWC 0.24); /ViT-B-21k 0.898; Split-ImageNet-R/DINOv2-L 0.896; /ViT-B 0.672; Split-MNIST 0.911match/beat SOTA, sec, 0 buffer
ODE continual stream (S3)matched dictionary, 20 regimes∼\sim290×\times lower error than EWC; flat to 20 regimes; instant revisit recognitiondecisive
RSI stability (S7)telephone-game self-train, frozen backbone, only update rule varies1 lab/cls: CF 0.49→\to0.62 (stable) vs SGD 0.40→\to0.34 (drifts); 2 lab/cls: CF 0.66→\to0.73 vs SGD 0.51→\to0.44CF = stability primitive
Grow ++ no-forget capstone (S8)100 skills, 10/gen, grow capacity + gated self-improveall-skills 0.885, first-skill retention 0.94 vs gradient-grow 0.012 / 0.0 (∼\sim70×\times gap)decisive
LLM continual (#26)frozen GPT-2, 20NG 10-task class-ILCF 0.673 == joint upper bound (0 forgetting) vs seq SGD 0.057; retention gap 0.616joint optimum
LLM few-label RSI (#26b)5 lab/cls ++ unlabeled, NCM-gated pseudo-labelsseed-only 0.452 →\to 0.641 (∼\sim95% of full-sup 0.667); SGD 0.055RSI on real LLM
Multimodal generalist (#28)one Gram substrate, ViT vision ++ GPT-2 language, 14 skills/120 clsfull-label 0.794 == joint optimum; few-label++RSI 0.752; seq SGD 0.056one mechanism, many modalities
Three-modality agent (#30)++ control (pendulum world-model++MPC) skillclasses all-seen 0.794 UNCHANGED after control; shared backprop net →\to 0.056interference-free
Gradient-free deep credit (#29/29b)parity-of-3-hidden-hyperplanes (kernel-hard)backprop 0.88; CEM-evolved 0.61; target-prop 0.57; random 0.53wall real but not absolute →\to hybrid (later overturned, #31–35)
Continual learning, RSI, multimodal. CF = closed-form Gram memory; retention is structural (order-invariant joint optimum), a property backprop continual provably cannot have.

Neuroevolution and model-based control

A GPU-native NEAT keeping NEAT's design but replacing its optimizer with a closed-form readout: compact circuits found ∼\sim100×\times faster than CPU NEAT, at scales NEAT cannot reach; model-based MPC on exact closed-form world models cracks hard-exploration control with hundreds of real interactions.

p2.9cmp4.3cmp4.8cmp2.0cm@ Experiment (#)SetupKey numbersVerdict
Parity ladder (#9)evolve topology, closed-form readout/candidate (∼\sim1–14 ms)45–130×\times faster than SGD scoring; curriculum solves parity 2–7 at 1.0, parity-8 0.984 (1→\to263 hidden); greedy collapses at 3topology load-bearing
Evolved recurrence (#9b)sequential parity, readout closed-form over time (no BPTT)recurrent 1.00/0.99/0.91/0.83/0.72 (T=4−20T{=}4{-}20) vs feedforward pinned ∼\sim0.5–0.6recurrence breaks memory ceiling
GPU-NEAT vs vanilla (#9c)N-bit parity from scratch vs neat-pythonours++speciation solves 2–7 (1/3/2/22/37 hidden); vanilla solves only XOR, bloats; ∼\sim1.6–4.8k evals/s (∼\sim100×\times); 5-seed 100% at 3–7crushes vanilla NEAT
Real-ish tasks (#9d)two-spirals, digitsspirals 0.994 (ties tuned 2-layer MLP); digits 0.934 (linear 0.920, tuned MLP 0.966)topology discovery, not raw-acc supremacy
Evolved-circuit lifelong (#9e)evolve 23-neuron ϕ\phi, freeze, class-ILCF all-seen 0.924 / retention 0.931 vs SGD head 0.428 / 0.069; 4-seed 0.940 / 0.926no-backprop WINS
RL neuroevolution (#10)GPU-batched rollouts, evolved recurrent policyno-velocity CartPole: recurrent solves 500 (3/3); feedforward fails 57–60 (0/3)recurrence load-bearing when non-Markovian
Scaling to 105^5 (#11)dense vs sparse edge-list forwardsparse 100k-neuron circuit: 0.21 s / 114 MB on one M4 GPUforward feasibility (search open)
Model-based RL (#12)closed-form dynamics from few transitions, evolve controller in model250 real transitions →\to 500 (solved), model R2=^2{=}1.0; model-free needs 209k ⇒\Rightarrow ∼\sim800×\timesdecisive
Pendulum / swing-up (#16,#18)exact world model ++ closed-loop CEM-MPCstabilize −74.3-74.3 (1k transitions) vs model-free −308.6-308.6 (28.8M) ⇒\Rightarrow ∼\sim28,800×\times; swing-up CEM-MPC −347-347 (planning cracks it)win (planning)
MountainCar / quadrotor (#19,#20)exact model (R2=^2{=}1.0) ++ CEM-MPCMountainCar success 1.0 (∼\sim111 steps, 5 seeds); PVTOL hover error 0.006 (4 seeds)robotics-relevant wins
Language pillar (#13–15)evolved recurrent circuit, char-level next-charevolved-from-reservoir 0.541 (vs unevolved 0.433, bigram 0.350, trigram 0.806); scales to 0.634 at 241 neurons; K-delay echo beats n-grams beyond windowsubstrate works, below trigram (honest)
Block-growth self-scaling (#14)grow in modules, minimal start, no seedingblock=8 0.599$\pm$0.004 vs block=1 0.442 (lang); digits 0.957; flagship 518 neurons →\to 0.70; curve 0.57→\to0.70autonomous-scaling unlock
Lifelong-evolving capstone (#17)block-grown ϕ\phi →\to class-IL Gram245 neurons: all-seen 0.961$\pm$0.004, retention 0.953±\pm0.029 (4 seeds)grow ++ no-forget, robust
ES/PGPE (#76)cartpole, ES vs GAGA solves (471, 320 ep); quick ES stalls (deceptive flat reward)niche tool, not universal
Neuroevolution and model-based control. Closed-form pays off in RL by removing the model bottleneck (exact model from few transitions →\to plan), not the rollout.

Bio-plausible learning and predictive coding

The thesis that backprop's apparent supremacy is rote memorization, not learning quality: local (no-global-backward) credit assignment matches backprop, generalizes better, is depth-robust, and predictive coding == backprop at the gradient level once run in its correct regime.

p2.9cmp4.3cmp4.8cmp2.0cm@ Experiment (#)SetupKey numbersVerdict
SSP on frozen vision (#21)evolved circuit vs random projection, ViT CIFAR-100linear 0.881, RanPAC 0.892, evolved 0.882 (ties)matched == random, as SSP predicts
Known-operator vision (#22–25)DoG++Gabor++scattering / steerable / Coates-Ng, CIFAR-10, no backpropraw 0.383 →\to operators 0.671 →\to learned dict 0.718; steerable 0.66; deep steerable plateau 0.682no-backprop ceiling ∼\sim0.72 (honest)
Steerable θ\theta-free (#27)paper's Question B, learned-θ\theta vs steer-to-allep1 0.289 vs 0.226 (++6.3pp); final parity; 35×\times fewer conv paramsθ\theta-bootstrap avoidable
DFA / decoupled-greedy (#31,#32,#32b)local credit, teacher-studentDFA 0.711 vs backprop 0.780; decoupled-greedy 0.796$\pm$0.013 ≥\ge backprop 0.782±\pm0.012 (5 seeds, full); small-data tiewall overturned; local matches backprop
Memorization vs learning (#33)random-label fit; 40% label noisebackprop memorizes random labels 1.000; under noise clean-test backprop 0.464 vs local 0.507local generalizes better, memorizes less
Depth scaling (#36)deep teacher, student depth 2–8backprop degrades 0.685→\to0.640; local flat 0.683 (gap widens to −-0.043)local depth-robust
Closed-form-local / OOD (#34,#35,#37)CF head into features; input-scale shift; label-free localCF-local 0.655 << greedy 0.788 (honest neg); OOD non-differentiating; contrastive fails 0.491, sparse-coding 0.718honest negatives bracketed
Reasoning / e-prop (#40–46)multiplicative register; working memory; RTRLWM: e-prop random-feedback 0.999/0.992/0.974 == BPTT (no BPTT, no weight transport); RTRL 0.475 == BPTT; algorithmic e-prop diagonal lags (off-diagonal credit)recurrent bio-learning matches BPTT
Conv / InfoPro scale (#47–50)local sweep on real conv CIFAR-10decoupled-greedy 0.839, InfoPro 0.853 vs backprop 0.864 (within 1.1pt); DFA bolt-on fails 0.74–0.76local matches, beat-A at scale only
Online recurrent (#51,#52)UORO rank-kk vs e-propUORO rank-1 0.285 →\to rank-16 0.394 << e-prop 0.568 << BPTT 0.659e-prop = practical sweet spot
Predictive coding (#53–62)PCN vs backprop, gradient-alignment unit testhard-clamp loses (0.563); Z-IL cos ==1.000, nudged ==0.99; PC-nudged 0.674 == backprop-MSE 0.678; no robust PC>>backpropPC == backprop in correct regime
PC on arbitrary topology (#57,#58,#66)skip-DAG; PC-NEAT growPC 0.587 == backprop-CE 0.581 on skip-DAG; PC-NEAT parity-8 chance→\to1.000 (grew 2→\to3 nodes)PC trains evolved wiring locally
Closed-form precision PC (#60)one local downward sweep (ePC/HGF)CF-PC 0.662 == backprop-CE 0.658; precision-whitening hurts (0.632)deep-capable closed-form local learner
Bio-plausible learning and predictive coding. Across MLP, conv, recurrent (e-prop) and skip-DAG topologies, local / no-global-backward credit assignment matches backprop; PC equals backprop at the gradient and accuracy level. The PC equilibrium solved in one sweep is the OSNR closed-form principle applied to the cortical learning rule.

The gradient-free cortex, scale-up, and application fronts

The synthesis: a fully gradient-free agent that builds its own architecture (NEAT), learns deep features from pixels by a per-block local error sweep (no global backward), and accumulates tasks with zero order-invariant forgetting (closed-form Gram), at competitive-to-SOTA accuracy and the 10510^5-neuron scale target, across vision, language, control, and games.

p2.9cmp4.3cmp4.8cmp2.0cm@ Experiment (#)SetupKey numbersVerdict
Federated cortex (#63–65)closed-form PC area ++ Gram memory, digits/CIFAR continual2-task digits CF 0.967, A-after-B 0.978 vs backprop 0.009; CIFAR 5-task front-end++Gram 0.649, gen-PC++ 0.655; backprop 0.189no-backprop ++ no-forget holds
Shallow ceiling diagnosed (#67,#70)richer front-end; head depthricher 0.683 (order-independent, exact); depth0–3 PC 0.681–0.685, even backprop plateaus ∼\sim0.690.68 is fixed-front-end ceiling, not no-backprop
Ceiling broken (#71,#72)deep conv from pixels by per-block LOCAL sweep ++ Gramlocal-from-pixels 0.842 (backprop 0.864); deep-local++Gram 0.889 5-task continual (reverse-order identical), backprop 0.189competitive no-backprop, no-forget
Self-constructing agent (#73,#74)NEAT evolves conv arch, local weights, Gramseed [32] 0.429 →\to [32,64,192] 0.693; full agent 0.864 5-task continual (order-independent) vs backprop 0.193builds own arch, no backprop, no forget
SOTA-class vision (#75,#84)deeper CNN ++ aug, per-block local sweep, CIFAR-10local 0.9024 ≥\ge backprop 0.8855 (same net); multi-seed local 0.8994$\pm$0.0004 >> backprop 0.8824±\pm0.0024 (3/3)local matches-or-beats backprop, locked
Multimodal deep / scale (#77,#78)deep no-bp vision ++ GPT-2; neuron accountingvision lifted 0.653→\to0.888, language 0.652 retained; base substrate 246k neurons/img, 0.855 in 12 ep, <<1% of 128 GB, 7k img/s10510^5-neuron target met
Games from pixels (#79,#82)local-sweep perception ++ (known/learned) dynamics ++ MPCCartPole from pixels 236/300 (∼\sim3k frames); Catch learned-dynamics R2^2 0.997, catch-rate 0.965 (∼\sim2.7k transitions)DQN recipe, no backprop, no millions of frames
Transformers / real GPT (#83,#85)per-block local sweep on attentioninduction-head 1.000 == backprop; real char-GPT (4 blocks) val bpc 1.946 vs backprop 1.933, coherent generationPC == backprop on the full architecture family
Atari Pong (#86,#90)local-sweep conv policy on real ALE; BC then DAgger (relabel policy-visited states with predictive teacher)BC −-21.0 (distribution shift); DAgger lifts −-21→\to$-$8.8 over 4 rounds (teacher −-5.4, random −-20.8), all no global backwardgradient-free policy plays real Pong from pixels
Online video (#81)deep local perception ++ online Gram, day→\tonight shiftonline Gram DAY 0.922 retained vs online backprop head 0.880 (forgets)autopilot-style no-forget perception
The gradient-free cortex and application fronts. One agent self-constructs its architecture, learns deep features from pixels with no global backward pass, and accumulates skills with structural (order-invariant) zero forgetting; backprop continual collapses to ∼\sim0.19 on the same streams. The learning rule was never the limit—architecture, augmentation, and scale were.

Standardized benchmark campaign

To test the substrate against the field's canonical benchmarks rather than bespoke setups, we ran a standardized campaign across continual learning, language, scientific computing, and single-task vision. Each result is gradient-free (no global backward pass) and reproduced by a single script. The pattern is consistent with the theory: on specifiable-structure tasks (PDEs) the operator-matched closed-form solve wins by orders of magnitude; on generic perception the local sweep matches or beats backprop on the same architecture; on continual streams the closed-form Gram memory is competitive with the no-gradient SOTA with exactly zero forgetting.

p3.0cmp4.4cmp5.0cmp1.6cm@ Benchmark (#)SetupKey numbersVerdict
Split-CIFAR-100 CIL (#87)frozen ViT-B/16-in21k, random-proj ++ closed-form Gram, 10×\times10final 0.900, avg-inc 0.936, order-invariant (zero forgetting), no gradient; RanPAC ∼\sim0.92competitive, zero-forget, gradient-free
Split-ImageNet-R CIL (#88,#89)same recipe, 200 classes, 10×\times20final 0.669 (avg-inc 0.690), order-invariant; below RanPAC ∼\sim0.78 (gap == second-moment whitening; reimpl did not close it)honest partial
enwik8 byte-GPT (#91)4-block causal transformer, byte-level, per-block local sweepTEST 2.24 bpc vs backprop 2.238 (identical), local FASTER (122 s vs 416 s); learned Wikipedia markuplocal == backprop on standard LLM corpus
PDEBench-formal 1D (#92)Advection/Burgers/Diff-React, OSNR closed-form vs trained FNO (74k p), nRMSEextrap nRMSE: adv 0.004 vs 0.314, Burgers 0.042 vs 0.243, DR 0.000 vs 0.023; ∼\sim400×\times faster, FLAT extrapolationorders-of-magnitude win (matched regime)
CIFAR-100 single-task (#93)deep 4-block CNN ++ aug, per-block local sweeplocal 0.660 >> backprop 0.596 (same net), faster, monotonic; from-scratch (not SOTA ∼\sim0.85)gradient-free beats backprop, 100 classes
MuJoCo HalfCheetah (#94–98)deep dynamics model by local sweep ++ short-horizon CEM-MPC, all no global backwardridge model underfits (R2^2 0.64, return −-60); deep local-sweep model R2^2 0.99; short-horizon CEM-MPC 345$\pm$39 (8 ep) vs random −-415, from ∼\sim6–8k transitionsgradient-free continuous control: runs forward stably
Standardized gradient-free campaign: continual learning, language, scientific computing, vision, and continuous control.
p3.0cmp4.4cmp5.0cmp1.6cm@ Benchmark (#)SetupKey numbersVerdict
Atari Freeway — ES baseline (#99)OpenAI-ES (Salimans 2017) evolving a conv policy from pixels; standard method, not ours21.0 vs random 0, from ∼\sim176k env steps (∼\sim2.2M total); converges by gen 1gradient-free baseline FLOOR (reproduces ES-plays-Atari)
Atari Freeway — ours (#100)local-sweep autoencoder ++ latent dynamics ++ discrete CEM-MPC, all no global backward21.3$\pm$0.5 vs random 0, from ∼\sim14k env steps (∼\sim12×\times fewer than ES; ∼\sim150×\times vs ES budget)our-framework win, sample-efficient; Freeway easy (caveat S[app:mbrl])
Atari Pong — model-based (#101)reward-MPC; ball-tracking stress test−-21.0, reward-R2≈0R^2\approx0 (reconstruction latent drops the ∼\sim2×\times3-px ball; sparse/delayed reward)wrong tool for sparse reward (S[app:mbrl])
Atari Pong — local-sweep DQN (#102–103)value-based (Double, n-step) TD, credit assignment by local sweep, no global backward, Q-net on GPU (arch: Fig. [fig:arch-dqn])wins from pixels, reproducibly: final greedy $+$18.5$\pm$0.8 over 3 seeds (4/4 incl.\ seed0 positive; ceiling ++21, random −-20.7)reliable, not luck; matches a DQN baseline (not SOTA)
Atari Breakout — same config (#105)identical stabilized DQN config, no game-specific tuningfinal greedy 58.9 (curve 0→\to31→\to66 peak; random ∼\sim1) over 1M stepsgeneralizes: the fix is principled, not Pong-overfit
Standardized gradient-free campaign: pixel-control benchmarks; remaining gaps are explicit.

Gradient-free model-based control: one recipe, state and pixels

The control results (MuJoCo HalfCheetah #94–98, Atari Freeway/Pong #99–100) share a single gradient-free recipe: learn a world model with the per-block local sweep, then plan with the cross-entropy method (CEM-MPC). No global backward pass appears anywhere—not in perception, not in the dynamics model, not in the planner. The only difference between the low-dimensional-state case (MuJoCo) and the pixel case (Atari) is a front-end autoencoder that supplies a latent state.

Components (all trained by the local sweep, Alg. [alg:mbrl]).

(1) Encoder/latent (pixels only). A convolutional autoencoder Eϕ,DψE_\phi,D_\psi is trained on collected frames by the per-block local error sweep on the reconstruction loss ∥Dψ(Eϕ(o))−o∥2\lVert D_\psi(E_\phi(o))-o\rVert^2; activations are detached between blocks and each block updates from its own local target (no backprop through the stack). The encoder EϕE_\phi then maps a frame to a dz=64d_z{=}64 latent zz. For low-dimensional state (MuJoCo) this step is skipped and zz is the raw state. (2) Latent dynamics + reward. A deep MLP fθ([z,a])↦(Δz,r^)f_\theta([z,a])\mapsto(\Delta z, \hat r) is trained by the local sweep to predict the next-latent increment and the immediate reward (one-hot actions for discrete control). (3) Planner. CEM-MPC rolls candidate action sequences through fθf_\theta in latent space and keeps the elite by predicted return: a per-step Gaussian refined to its elite mean/variance (continuous, MuJoCo) or a per-step categorical refined to its elite action frequencies (discrete, Atari); the first action is executed, receding-horizon. (4) Aggregation. Roll out the planner, append the visited transitions, refit—the dynamics analogue of DAgger (#90).

caption(Gradient-free model-based control (local-sweep world model + CEM-MPC)) begin(algorithmic)[1] State Collect random transitions D={(o,a,r,o′)}\mathcal{D}=\{(o,a,r,o')\} For(round =1,…,R=1,\dots,R) If(pixels) train autoencoder Eϕ,DψE_\phi,D_\psi on {o}\{o\} by the local sweep (recon loss); set z=Eϕ(o)z=E_\phi(o) Else z=oz=o EndIf State Train fθ([z,a]) ⁣→ ⁣(Δz,r^)f_\theta([z,a])\!\to\!(\Delta z,\hat r) on D\mathcal{D} by the local sweep (MSE) Comment(no global backward) State CEM-MPC: at each step, refine per-step action distribution by elite predicted return in fθf_\theta; execute first action State Append planner-visited transitions to D\mathcal{D} EndFor end(algorithmic)

Results and honest scope.

On MuJoCo HalfCheetah a deep local-sweep dynamics model (reward R2 0.99R^2\,0.99, vs 0.640.64 for a closed-form ridge+random-feature model that underfits contact dynamics) with short-horizon CEM-MPC runs the robot forward at 345±39345\pm39 (random −415-415); the horizon had to be shortened (25 ⁣→ ⁣1225\!\to\!12) to suppress model-exploitation over long plans (#94–98). On Atari Freeway the same pixel pipeline reaches 21.3±0.521.3\pm0.5 (random 00) from ∼\sim14k environment steps, ∼\sim12×\times fewer than the gradient-free Evolution-Strategies baseline (#99, 21.021.0 from ∼\sim176k steps) and ∼\sim150×\times fewer than the ES total budget—the sample-efficiency edge of model-based planning. We are explicit about the limits: Freeway's optimal policy is nearly trivial (press textsc(up)), and the reward head's R2R^2 degraded under aggregation while the score held, so Freeway demonstrates the pipeline but not a high-quality pixel world model. The stronger test, Pong (#101), is a clean negative: the same pipeline scores −21.0-21.0 (random −20.7-20.7) with reward-R2≈0R^2\approx0 from the first round. The cause is diagnosed, not hand-waved—the encoder is trained by reconstruction, whose pixel-MSE is dominated by large static structures (paddles, background, score), so Pong's ∼\sim2×\times3-pixel fast ball is dropped from the latent; with no ball in zz, reward (a point) is unpredictable and the planner is blind. This is the textbook failure of reconstruction-latent world models on small reward-relevant features (why Dreamer-v2/v3 and TD-MPC2 use reward/value-predictive or self-supervised forward-predictive latents). The fix is a task-aware latent—action-conditioned next-frame prediction, reward/value-predictive encoding, or motion (frame-difference) input, all local-sweep-compatible—and the scaling of all of this to the full 49-game suite is a cluster task (Track A3 in RUNPOD\_TODO.md).

Value-based control with local credit assignment (Pong).

Pong's reward is sparse and delayed—the wrong regime for reward-seeking MPC but the home turf of value bootstrapping, where TD learning propagates the rare ±1\pm1 backward into a dense learned value. We therefore ran a DQN (standard 84 ⁣× ⁣84 ⁣× ⁣4→84\!\times\!84\!\times\!4\toconv32/64/64→_{32/64/64}\toFC512→Q_{512}\to Q, replay buffer, target network, ϵ\epsilon-greedy) but with credit assignment by the per-block local sweep instead of backprop—the TD error on the taken action is propagated block-locally, no global backward pass, Q-network trained on the GPU. It learns Pong from pixels: greedy score −21 ⁣→ ⁣−9.3-21\!\to\!-9.3 over ∼\sim1.5M frames, action distribution spreading across all six actions. This establishes that the gradient-free local rule performs value-based RL—bootstrapped, non-stationary targets—on a real Atari game, not only supervised and model-based learning. Reproducible, not a lucky run: the initial setup was high-variance (1 of 5 single-machine runs reached −9.3-9.3), which we root-caused—not to the local rule (it computes the exact backprop gradient) but to a replay buffer 10×10\times too small (correlated samples →\to Q-collapse to a degenerate action) and 1-step TD too slow for the sparse reward. With the standard fixes (300k buffer, n-step returns n=3n{=}3, Double-DQN), three fresh seeds all converge to a winning score—final greedy +19.6/+17.8/+18.1+19.6/+17.8/+18.1, mean +18.5±0.8\mathbf{+18.5\pm0.8} (ceiling +21+21, random −20.7-20.7), every seed a smooth monotone-trend climb with no collapse. An earlier flat −21-21 traced to a preprocessing bug of ours (a missing frame-max de-flicker that erased Pong's flickering ball), not the learning rule. Honest scope: this is algorithmically a (Double, n-step) DQN—the contribution is the no-global-backward local rule doing value-based RL to a reproducible near-ceiling score, not a new RL algorithm or a SOTA result; the full 49-game suite is deferred to RunPod (Track A1).

Stable recursive self-improvement: two knobs make a self-training loop compound instead of collapse.

We frame recursive self-improvement (RSI) as a dynamical system in capability-space: the binding constraint is the stability of the closed self-training loop (a tiny labelled seed pseudo-labels an unlabelled pool, the model learns from its own labels, repeat), not the per-step improvement. Backprop is allowed—it is the baseline that collapses. On frozen DINOv2 features, naive backprop self-training collapses (0.61 ⁣→ ⁣0.320.61\!\to\!0.32, self-label error 0.39 ⁣→ ⁣0.680.39\!\to\!0.68) while the same loop anchored to a never-forget Gram memory is collapse-proof and self-improves (0.318 ⁣→ ⁣0.6610.318\!\to\!0.661): the additive, order-invariant memory is the Lyapunov anchor. For the harder case—a trainable CNN representation self-improving from scratch on CIFAR-10—a fixed reference anchor either caps or drags the net, but a slow EMA mean-teacher that tracks the improving student (the dissipative drift-following anchor) yields stable self-improvement. Two stability knobs are each necessary. (i) The EMA anchor must EMA the batch-norm statistics consistently with the weights; copying the fast student's BN stats onto slow-EMA weights produces a mismatched teacher that looks like collapse (worse the slower the EMA: at momentum 0.950.95 the teacher fell 0.43 ⁣→ ⁣0.220.43\!\to\!0.22)—a BN-consistency bug, not a property of slow anchoring. With it fixed, the EMA teacher climbs monotonically where naive is noisy or degrades: at 2525/class it compounds 0.183 ⁣→ ⁣0.3260.183\!\to\!0.326 over 3030 rounds (still rising, the per-round gain growing as the self-paced pool admitted grows 31 ⁣→ ⁣362 ⁣→ ⁣78831\!\to\!362\!\to\!788), while naive peaks early and settles back at chance; at 100100/class it is smooth-monotone 0.430 ⁣→ ⁣0.4780.430\!\to\!0.478. (ii) The confidence threshold is the curriculum: at conf=0.95\text{conf}{=}0.95 the loop admits only reliable labels and compounds, but relaxing to 0.900.90 floods the loop with the teacher's lower-confidence errors (2121–4949k of the ∼\sim5050k pool admitted) and reproduces the classic confirmation-bias collapse (EMA peaks∼ ⁣0.20 ⁣→ ⁣0.13\text{EMA peaks}\sim\!0.20\!\to\!0.13). Honest scope: absolute accuracies are modest (small CNN; 2525/class reaches 0.3260.326, 39%39\% of the 0.8270.827 supervised ceiling) and EMA mean-teaching is established semi-supervised learning—the contribution is the framing (RSI as a stable dynamical system, slow anchor ++ confidence curriculum as the stabilizers) and the precise characterization that stable self-improvement is a BN-consistency-plus-curriculum property of the anchor.

One frozen encoder (or several) + one closed-form memory = a cross-domain, cross-modal lifelong brain.

The closed-form never-forget memory composes into a concrete instantiation of ``attach our framework to any frozen model and get a lifelong-learning brain.'' On a single frozen DINOv2 ViT-S encoder we stream three very different datasets as a task-free sequence—CIFAR-100 (objects) →\to Flowers-102 →\to FGVC-Aircraft (planes), unified into 302302 classes, evaluated by predicting over all seen classes with no task-id. A finetuned head catastrophically forgets (final average 0.0350.035, forgetting +0.42+0.42); the closed-form memory retains every domain with near-zero forgetting (+0.011+0.011) at average 0.7540.754, beating nearest-class-mean (0.7200.720, the edge on fine-grained aircraft, 0.4980.498 vs 0.3680.368, from feature decorrelation). One honest pitfall surfaced and was fixed: plain accumulation is swamped by dataset-size imbalance (CIFAR's 5050k examples vs Flowers' 22k drive the ridge to fit the large set and ignore the small ones, 0.0000.000 on Flowers/Aircraft—imbalance, not forgetting, since CIFAR stays at 0.7640.764); class-balanced accumulation (G=∑iwizizi⊤G=\sum_i w_i z_i z_i^\top, wi=1/nciw_i=1/n_{c_i}) is the correct default for imbalanced multi-domain streams. The strong form crosses modalities: two different encoders (DINOv2 for images, BGE-base for text) feed one shared memory through per-modality random feature maps into a common space, on an interleaved vision/language stream (CIFAR-100, banking77 intents, Flowers, dbpedia topics, Aircraft; 393393 classes). One memory spans both modalities with near-zero forgetting (+0.006+0.006, average 0.8030.803) where the finetuned head collapses (0.0690.069, +0.27+0.27); the per-modality random maps occupy near-orthogonal subspaces, so vision and language do not interfere. The contribution is the architecture—any encoder, any modality, one gradient-free buffer-free closed-form lifelong memory at ∼\simzero forgetting—not a new accuracy record (it tracks the frozen encoders' ceilings, per the teacher-scaling law).

Bulletproofing the closed-form memory: it beats replay, learns new classes instantly, and is private against exemplar attack.

Three receipts harden the continual-learning claim, all on frozen DINOv2 features. (i) Beats buffered replay. On Split-CIFAR-100 class-incremental learning, the gradient-free never-forget memory reaches 0.8680.868 at average forgetting 0.0500.050 versus a 20002000-exemplar replay head at 0.8010.801/0.1830.183 and naive online SGD at 0.3350.335/0.7030.703; on the harder cross-modal stream a fair replay (balanced half-current/half-buffer batches) still collapses to 0.1880.188 against the memory's 0.8030.803, because 20002000 exemplars over 393393 classes (∼\sim5/class) cannot rehearse a high-dimensional head while the memory accumulates every example exactly as sufficient statistics. (ii) Instant few-shot class-incremental. A new class is one additive update (O(1)O(1), no epochs); cross-domain FSCIL (CIFAR-100 base, then Flowers/Aircraft KK-shot) keeps base retention exactly constant (0.7730.773–0.7750.775, zero forgetting) while new-class accuracy scales with shots (0.6630.663 at 11-shot to 0.7730.773 at 1010-shot). Finetuning faces an unwinnable stability–plasticity dilemma: few steps underfit the new classes (new 0.0000.000), more steps erase the base (base 0.0000.000). (iii) Privacy. A membership-inference attack cannot distinguish the parametric memory from a gradient head (attack AUC ∼ ⁣0.53\sim\!0.53 for both—privacy comes from frozen features plus regularisation, not the closed form per se), but both are near-private against the exemplar alternative: a replay/RAG store of raw features is trivially de-anonymised (AUC 1.01.0). The memory's privacy edge is therefore over the exemplar methods it replaces, exactly the buffer it does without.

The deep frontier: anchoring a never-forget memory to a plastic representation.

Every lifelong-brain result above freezes the encoder; the open problem is retaining old tasks while the representation keeps learning, where the memory's stored sufficient statistics go stale as features drift. Isolating it (a trainable MLP on frozen DINOv2 features, Split-CIFAR-100, 5×205\times20), a live plastic representation forgets catastrophically (0.2910.291, forgetting 0.8440.844) because the stored statistics no longer match the drifted features. The fix is the same slow anchor that stabilised recursive self-improvement: an EMA-slow copy of the representation feeds the memory, and as the EMA momentum increases (0.90→0.95→0.990.90\to0.95\to0.99) retention climbs monotonically (0.645→0.765→0.8370.645\to0.765\to0.837, forgetting 0.373→0.0600.373\to0.060), the 0.990.99 case nearly matching the frozen upper bound (0.8850.885) while the representation still trains on every task—the momentum is a clean stability–plasticity knob. Two honest bounds. First, freezing still edges EMA when the encoder is adequate (confirmed on CIFAR-100 and fine-grained Aircraft, 0.7050.705 vs 0.5410.541): plastic-plus-EMA is the controllable fallback for when a frozen encoder is insufficient, not a universal improvement, and demonstrating EMA>{>}frozen needs a genuinely out-of-domain stream. Second, the EMA anchor needs a good base to anchor toward: from a random initialisation a slow EMA simply preserves the random start (a from-scratch CNN gives EMA ≈\approx frozen-random ≈\approx chance), the same lesson as the gradient-free self-improvement negative—it resolves the staleness catch-22 given a usable representation but cannot manufacture one. The validated recipe therefore remains: freeze a strong pretrained encoder and let the closed-form memory carry the lifelong, cross-domain, cross-modal continual learning on top.

A unified agent brain: one closed-form memory that both perceives and acts.

The lifelong-brain wins are all perception; a real agent must also act. We give a single never-forget memory three faculties in one interleaved stream—vision (DINOv2 CIFAR-100), language (BGE banking77), and control (behavior-cloning a gradient-free CartPole expert found by random linear-policy search, return 500/500500/500)—through per-faculty random feature maps, and after the stream evaluate all three, the control faculty by deploying the memory as a policy in the live environment. One memory perceives (vision 0.6890.689, text 0.5500.550) and acts at expert level (CartPole 500/500500/500) with zero forgetting, while a sequentially-finetuned head catastrophically forgets every earlier faculty (vision/text at chance 0.03/0.040.03/0.04) and even underperforms on control (418418). The design that makes this work—diagnosed by first finding closed-form control underperforming in a naive shared memory (234234), then isolating that closed-form ridge clones the policy nearly as well as cross-entropy SGD when given its own features (478478 vs 492/500492/500, so the deficit was interference, not least-squares)—is fully independent per-faculty blocks: each faculty gets its own bias, random-feature coordinates, and regularization (perception λ=100\lambda{=}100 class-balanced, control λ=1\lambda{=}1 full-mass) inside one block-structured solve. Shared coordinates corrupt the joint Gram matrix (an early version gave vision 0.0000.000); a shared bias is swamped by the high-mass control faculty; class-balancing crushes the two-action control faculty. With independent blocks all three reach their standalone ceilings at once. The contribution is architectural: a single gradient-free, buffer-free memory spanning perception and action with structural zero-forgetting—the closed-form approach extends to decision-making, not only classification.

Add to remember, subtract to forget: exact machine unlearning falls out of the additive structure.

The same additivity that makes the memory order-invariant and zero-forgetting also makes the right-to-be-forgotten problem trivial. Because G=∑tZt⊤ZtG=\sum_t Z_t^\top Z_t and B=∑tZt⊤YtB=\sum_t Z_t^\top Y_t are sums over tasks, a task is removed exactly by subtracting its contribution (G ⁣− ⁣= ⁣Zk⊤ZkG\!-\!=\!Z_k^\top Z_k, B ⁣− ⁣= ⁣Zk⊤YkB\!-\!=\!Z_k^\top Y_k) followed by one re-solve—O(1)O(1) in the forgotten task's size, gradient-free, no access to the retained data. On Split-CIFAR-100 (frozen DINOv2 ViT-B, 10 tasks), unlearning a task this way yields a model bit-identical to one retrained from scratch without it (max⁡∣Wsubtract−Wscratch∣≈1.5×10−5\max|W_{\text{subtract}}-W_{\text{scratch}}|\approx1.5\times10^{-5}, retained accuracy 0.8720.872 in both), whereas an SGD head finetuned on the remaining data carries no such guarantee (its weights still encode the removed task; exact removal would require full retraining). Add to remember, subtract to forget—both exact, both O(1)O(1) per task—a privacy/compliance capability gradient training cannot match.

The limits of self-improvement: self-reference plateaus, verification compounds the readout, only teaching improves the representation.

Asking directly what makes an online model self-improve, we mapped which information source lets the loop compound. (i) Self-labels (no new information) degrade or plateau everywhere—a model training on its own predictions can only sharpen what its features already separate. (ii) We first hoped our stability machinery would tame online RL, but it is an honest negative: a closed-form gated memory is no more stable than gradient online-TD on CartPole (gradient final 256256/late-std 8484 vs memory 220220/std 144144), because RL targets are non-stationary (bootstrapped and policy-dependent) whereas our memory's advantage requires stationary targets—the deadly triad is a target problem the memory just faithfully fits. (iii) This predicted the right source: verification. A verifier of reliability vv (checking correctness is cheaper than generating the answer) supplies new and stationary information. Added to the very fixed-feature never-forget loop that degrades on self-labels (v=0.5v{=}0.5: 0.32→0.290.32{\to}0.29), it compounds monotonically to the supervised ceiling as vv rises (v=0.9→88%v{=}0.9\to88\%, v=1.0→0.51=102%v{=}1.0\to 0.51{=}102\%)—a threshold effect (vv must be ≳0.9\gtrsim0.9). But verification breaks only the readout plateau: with a trainable CNN it does not break the representation plateau (v=1.0v{=}1.0 max 0.4560.456 == v=0.5v{=}0.5 max 0.4550.455, both ∼55%\sim55\% of ceiling), because verified-correct labels concentrate on already-correctly-classified examples and carry no new representational signal (and below a viability threshold—a near-chance start—even a perfect verifier cannot bootstrap, as there is nothing correct to certify). (iv) What breaks the representation plateau is teaching: external true labels on pool examples lift the CNN from 0.460.46 to 0.630.63–0.660.66 (∼77%\sim77\% of ceiling) where verification stalls at 0.460.46, and hard-vs-random selection is a minor secondary effect—the decisive factor is new information about examples the model cannot already handle. Conclusion: self-improvement compounds only up to the information already latent in model-plus-data; verification (new correctness information) compounds the knowledge/readout layer to its ceiling, but improving the representation itself requires external teaching—it cannot be reached by self-reference or by verifying the model's own outputs. A precise, honest boundary on ``recursive self-improvement.'' (v) Deployability caveat. The verifier must be external: a self-derived cheap verifier—ensemble agreement / query-by-committee over random feature-subsets of the same model—does not substitute for it (max 0.3720.372 vs.\ self-training 0.3750.375 vs.\ oracle 0.4770.477), because members sharing the same features have errors correlated with the model's own (they agree-but-wrong on feature-confusable classes), so agreement ≈\approx confidence, not correctness. Verification-based self-improvement therefore works in domains with a genuine external checker—code (unit tests), math (execution/proof), simulation—not from a model's own self-consistency; there, never-forget memory ++ external verifier gives gradient-free compounding to the ceiling, with zero forgetting and exact unlearning. (vi) Instantiated with a real LLM. A frozen qwen2.5vl:7b on 3×33\times3-digit multiplication (a genuine weak spot, 00-shot 0.310.31), with EXECUTION as the external verifier and a never-forget memory of execution-verified worked solutions retrieved as few-shot context, self-improves to 0.490.49 (+18+18 points) over 120120 problems—no weight updates, no labels, only the execution signal. A control isolates verification as the cause: showing the model's own unverified solutions (including wrong ones) as exemplars hurts to 0.130.13, well below both verified (0.490.49) and 00-shot (0.310.31), because it imitates the wrong procedure. Execution-filtered exemplars help decisively; unfiltered ones harm. This is the constructive realisation of the map: a genuine external verifier plus a never-forget memory yields real, label-free, gradient-free self-improvement of a frozen model.

Full per-experiment detail, exact configurations, multi-seed statistics, and the code for every row above live in RESULTS\_SCOREBOARD.md and the corresponding bio\_growth/closed\_form\_neat\_*.py and bio\_growth/pde\_*.py scripts.

Independent robotic-system campaign (Gates 192–211).

The KUKA transfer first produces a representation-positive result: the operator map beats the published linear baseline but a compact MLP is more accurate. Its initial transaction analysis paired adjacent blocks incorrectly. Gate 209's input-only audit identifies repeated programs (0,3),(1,4),(2,5)(0,3),(1,4),(2,5); on these pairs sparse improves every repeat and retains 99.23% of pooled dense gain with 3.189×\times fewer labels, though one pair initially misses a frozen retention clause. Gate 210 then corrects the same pairing error in source-only ridge selection, changing the ridge from 100 to 0.01 and raising sparse retention to 97.45% of dense gain with every pair above 89%. All mechanism clauses pass. The earlier cross-program harm statement is retracted and archived as a protocol error; both corrections are post-exposure. The subsequent real Nano-Drone protocol commits source and independent-repeat receipts before confirmation and passes all twelve clauses with 2.500×2.500\times fewer labels. Frozen ablations attribute the gain to exponential actuator memory and reject the broad cardinal expansion. A matched MPS MLP loses accuracy but exposes a slow literal compiler; exact packed validation, sufficient-statistic composition, and Gram-only admission then reproduce the program in 0.935 s, 13.35×13.35\times faster than neural optimization. A final direct/recursive control formally fails three-group dominance while isolating large long-horizon attitude/rate gains. Finally, an adequate pinned ASIA reproduction with matched sparse personalization beats the compact operator in all four error groups, while the operator is 1,020×\times smaller and compiles 65.9×\times faster. Gates 200–208 then implement neural proposal followed by exact operator consolidation. One-execution projection compresses but does not transfer; typed local-cardinal recurrence repairs translation; and additive pseudo-transition Grams from two executions produce the best compact cumulative errors 1.646/5.368/2.624/15.9571.646/5.368/2.624/15.957. Generic/discovered interaction expansion and recency forgetting fail their gates. Protocols, phase seals, outputs, and negative results are retained in the Gate 192–211 Markdown records and corresponding industrial-breakthrough scripts. Gate 211 packages the resulting provenance rule into a typed evidence capsule: all six true KUKA Gram compositions are exact, all twelve wrong-program compositions are rejected, and verified symmetric-circulant normals invert 55×55\times faster at 1,024 coordinates with 4,096×4,096\times storage reduction at 4,096 coordinates. The complete repository suite passes 46/46 tests. Gate 212 then transfers the frozen API and an additive cardinal history model to a previously unused measured Wiener–Hammerstein circuit. Its source-only gate fails (40.220 mV versus 39.958 mV for the polynomial control), and the official target therefore remains sealed. The committed failure motivates operator-coordinate discovery rather than a denser cardinal grid. Gate 213's generic rank-4 Jennrich decomposition then has no real well-conditioned solution and produces no predictive score. Gate 214 instead ties the shifted nonlinear coordinate through the disclosed physical factorization; source RMS falls to 1.967 mV and every clause passes. Gate 215 commits the full-source states before opening the official target once. Its input-only cardinal program reaches 1.569 mV versus 43.338 mV linear and 8.486 mV factorized cubic, with a 3,920-byte executable. The result is prospective but remains above the published 0.241-mV deep-encoder frontier.

Gates 216–225 are post-target-exposure source-only diagnostics. Grid density, longer FIR memory, nominal inverse-Chebyshev transfer, lifted stable ARX, output-error refitting, and exact front/output pole variable projection fail to cross the 0.90-mV band. Gates 226–228 construct the complete quadratic and cubic stable-pole Hermite jets. Quadratic reaches 0.844451 mV; cubic regresses. The overlap-streamed branch compiler reproduces quadratic RMS to 1.11×10−131.11\times10^{-13} relative error at 584.8 MiB rather than 881.5 MiB peak. Gate 229 admits one cubic atom only after gains repeat on both validation halves, reaching 0.800018 mV. Gate 230 exposes exactly three independent numerator tangent directions after the scaling mode is removed and admits one at 0.596536 mV. Gate 231 admits a second at 0.454572 mV. Gate 232 rejects the last: it improves aggregate RMS to 0.420820 mV and the later half by 14.1% but harms the earlier half by 0.81%. All these refinements are source-only and do not constitute a second prospective official-target claim.

Gates 235–237 admit three repeatable transfer-function curvature atoms and reach 0.363553 mV; Gate 238 rejects all six remaining exposed atoms. Gate 239 then refits the frozen 16-branch topology on all source samples and commits the complete causal artifact before target execution. It reaches 0.338035 mV on source and 0.367337 mV on the official target, 76.59% below Gate 215, with 15,200 static bytes and 24,136 bytes including carried state. The target had already been opened, so this is retrospective transfer evidence rather than a new prospective benchmark claim.

Gates 240–243 extend output memory to 96 taps and replace coefficient-space compression with exact functional-Gram factorization. Rank 10 reaches 0.331004 mV, reproduces dense execution to 2.77×10−162.77\times10^{-16}, and transfers after full-source refit at 0.332097 mV. Gates 244–249 diagnose a colored residual but reject recent-input tails, autonomous stable ARX, exponential-reproduction edges, a second nonlinear chart, and recurrent exponential-channel memory. Gate 250 replays the six previously rejected finite atoms under the changed temporal metric; exactly one numerator-zero atom is now admitted on both segments, reaching 0.288277 mV. Gate 251 transfers its full-source refit at 0.304471 mV, narrowly missing the frozen 0.300-mV clause. Gates 252–254 reject cutoff retuning, higher-rank accuracy growth, and alternating edge–temporal refits; float32 functional rank 17 is prediction-neutral and occupies 12,948 static bytes. Gates 255–258 test pure and hybrid cardinal cubic Hermite edges. Their aggregate RMS reaches 0.272181/0.273758 mV, but both harm one of four contiguous blocks, so zero-forget rejects them and closes circuit engineering.

Gate 259 seals the Bouc–Wen benchmark, three independently phased source programs, and all code before fitting. A 1,624-byte physical Gram law improves official input-only multisine/sweep RMS from 1.559830/1.197468 to 0.158032/0.124041 mm; all prospective clauses pass. A tensor-cardinal residual is rejected at every nonzero admission step. Gate 260 differentiates the full recurrent simulator into three output-trajectory tangents. Four exact Gram cycles, admitted on two separate source programs, recover the disclosed hysteresis coefficients within 0.006% and reach source RMS 4.04×10−5/4.42×10−54.04\times10^{-5}/4.42\times10^{-5} mm. Its source artifact and official input-only trajectories are committed before retrospective scoring at 0.002212/0.005213 mm. The first Gate-260 run completed computation but failed JSON serialization; the type-only repair and identical rerun are both recorded.

Gates 261–265 remove disclosed coefficients and then the equation pair. Gate 261 jointly fits six physical parameters. Gates 262–264 retain their negative greedy, singleton, and local-pair receipts. Gate 265 enumerates all 55 pair topologies, calibrates the best ten on program 0, selects on program 1, and opens program 2 only once; it recovers the exact signed pair. Gates 266–267 retain two strict sensitivity-equivalence failures together with the complete finite-difference convergence tables. Gate 268 performs a matched end-to-end race from the same initializer and passes. Gate 269 records a faster backtracking endpoint but fails its validation-count clause. Gate 270 emits analytic atom gradients for the full dictionary and repeats the Gate-265 selection without loading its learned artifact. All charters, source hashes, cycle receipts, nonfinite rejections, timings, and frozen programs are stored with the corresponding scripts and JSON outputs.

Gate 271 seals four device files and a truth manifest in commit 71b6ad31 before its fitting executable. It freezes each selected program before moving to the next device, then opens third-program prefixes for routing and suffixes for scoring. Gate 272 seals the fifth device in b7634aa4, fixes the adequacy threshold from the earlier bank, executes abstention before fitting, and records all pre/post routes and hashes. Gate 273 seals a continuous switched trajectory in edcb9c09 before its rolling detector. Its complete per-window score/route trace, latencies, tail errors, model hashes, stream hash, runtime, and resource checks are retained. Generator modules are never imported by their corresponding fitting/detection executables.

Gates 274–276 freeze a device-0 wear ladder and two active-probe designs. The passive detector crosses its adequacy threshold between 0.25 and 0.5% wear; generic E-optimal and wear-directional RMS probes formally fail, at 1.153x and 1.467x passive separation. Gate 277 freezes a nominal-only matched detector before generating 100 nominal and 100 worn trials; all 200 decisions are correct. Gate 278 commits one noisy adaptation response, updater, repaired artifact, validation record, and evaluator in that order; untouched RMS falls 98.46%.

Gates 279–281 retain three strict near-failures. Unknown six-dimensional ±0.5%\pm0.5\% updates all finish below 0.001 mm but one small parent improves only 70.46%. A 24-device ±1%\pm1\% panel achieves at least 92.4% reduction but two final scores exceed 0.003 mm. Same-observation relinearization on 12 ±2%\pm2\% faults puts every final score below 0.000666 mm, but marginally worsens two already-excellent first steps. Each charter, generator, observation, updater, artifact set, validation set, evaluator, and outcome is a separate chronological commit.

Gate 282 freezes a 5% active-residual acceptance margin before its new ±3%\pm3\% panel. Sixteen observations generate step-one, step-two, and selected artifacts; all are committed before broadband validation exists. The evidence- only selector accepts 13 second steps. Every frozen clause passes, including no unseen regression versus step one, 95.71% minimum parent repair, 0.000830-mm maximum RMS, bitwise parent replay, stable hashes, runtime, and memory. The complete machine-readable chronology resides in the Gate 274–282 JSON outputs and Markdown charters/reports.

Gate 283 freezes 24 anonymous local/structural cases and routes them before labels open; the exact 12/12 split and all admitted validation checks pass. Gates 284–288 retain successive exhaustive, shared-dictionary, top-two, dual-probe, and triple-probe receipts, including each formal failure and runtime. Gate 289 freezes a fourth maximin capsule, four responses, compiler, artifacts, validation generator, outcomes, and evaluator in separate commits; every clause passes, although only one case actually accepts a four-capsule refinement.

Gate 290 freezes its charter and six anonymous devices before controller code. The controller's six decisions and executables are committed before the hidden three-drift/three-structural ordering and broadband records are generated. It routes 6/6 and passes every transfer, resource, hash, topology, and evidence- economy clause. The initial invocation stops before processing cases because of a sealed-array key mismatch; the one-line repair is a separate commit.

Gate 291 commits two programs and two rollback receipts before labels reveal that these are exactly the two in-dictionary and two out-of-dictionary cases. It formally fails one known-law transfer ceiling. Gate 292 separately freezes six cases and two mandatory full-capsule proposals per case. Four accept an update, two retain the preceding state, and the gate fails the per-case update and runtime clauses although all withheld accuracy clauses pass. The Markdown charters/reports and JSON records preserve these negative outcomes without threshold changes.

Gate 293 authorizes the exogenous-force operator before sealing four new cases. Its compiler, four artifacts, validation generator, outcomes, and evaluator are separate commits. Discovery and accuracy pass; the first evaluator formally fails its runtime clause. Gates 294–295 retain exact but formally negative value-only and vector-fused deployment attempts. Gate 296 freezes an inlined, precomputed vector kernel before execution and passes exactness, scientific score, hash, memory, absolute-time, and relative-speed clauses. None of these implementation repairs retroactively changes the earlier gate labels.

Gate 297 freezes its charter and Cholesky-risk designer before generating 20 candidate waveforms; its selected probe and complete score table are committed before any response exists. Gate 298 then freezes four new response sets, the two-branch compiler, eight artifacts, the validation generator, broadband data, and evaluator in chronological commits. Its selected branch wins all four independent comparisons and every preregistered clause passes. The report notes that this panel does not isolate VOI from minimum-eigenvalue design.

Gates 299–300 retain two outcome-free non-separation results: narrowed task risk and damping-coordinate variance both choose the generic 35–50-Hz probe. Gate 301 freezes an enlarged transient language before evaluation and records the first divergent design. Gate 302 then commits fresh evidence, two matched branches, and both branches' hashes before generating two validation families; its coefficient-identification result passes while the overall task gate fails. Gate 303 freezes and selects a constrained hybrid before any new responses. Gate 304 repeats the full evidence–compiler–artifact–validation chronology on new systems and retains its mixed broadband success and formal transient/ coefficient failure. No threshold is revised after outcomes.

Gate 305 freezes a five-version minimax Gram design and retains its failure to improve the nominal mixture. Gate 306 freezes full-transaction simulation and records a nonfinite structural-initialization failure without interpreting its NaN ranking. Gate 307 commits a one-change neutral-initialization wrapper before replay; it chooses the generic endpoint but fails the frozen novelty clause. Gate 308 freezes that acquisition rollback, a new panel, both matched branches, their hashes, and only then two validation families. Its 19.92% worst-risk advantage narrowly misses the 20% threshold, and the exact endpoint receipt also fails. Both clauses remain unchanged in the report and JSON.

Gate 309 freezes an eight-system acquisition-only panel, literal generic endpoint, 16 fixed-topology branch artifacts, and only then eight validation pairs. All acquisition-policy clauses pass, while one marginal evidence residual keeps the overall gate false. An initial validation invocation stops before sealing because a four-entry frequency schedule is exhausted; the eight-entry input-only repair is a separate commit. Gate 310 is explicitly retrospective: its evidence-only router and one-branch escalation are frozen before replay, and its passing receipt does not relabel Gate 309.

Gate 311 freezes an end-to-end two-level controller before eight new systems. Its acquisition decision passes after artifacts are committed and validation is generated, but the one-extra-step model contract fails on two branches. Gate 312 is explicitly retrospective: after validation exposure it demonstrates that a further certificate-monotone step closes both branches, without relabeling Gate 311. Gate 313 then freezes the bounded two-step rule, generator, compiler, validation generator, and evaluator. Adaptation evidence is sealed first; final compiler artifacts and hashes are committed second; only then is the eight-case validation panel generated, sealed, and opened. Four routed branches close in one step and every precommitted end-to-end clause passes.

The first Gate-313 compiler invocation ends after its parent phase when its execution session expires. Audit also identifies a dormant tuple/subtraction typo in the escalation branch that had not run. The interruption record and one-line repair are committed before a complete deterministic replay; evidence, programs, thresholds, routing, budgets, and validation code remain unchanged.

Gate 314 freezes an external MuJoCo environment receipt and retains its unsafe open-loop termination. Gate 315 freezes a feedback-filtered nominal selector, then formally fails because the snapshot omits Gymnasium's cumulative time-limit counter. Its output and corrected interpretation are committed. Gate 316 changes only that snapshot coordinate; identical controller, candidates, perturbation, seed, and thresholds then pass exact replay, observability, restoration, time, and memory clauses. No learning result is inferred from this environment receipt.

Gate 317 next seals 2,048 adaptation-only transition pairs and freezes linear and cardinal residual artifacts while validation does not exist. Its compiler fails an independent-solve clause because the raw additive cardinal dictionary has five partition-of-unity gauges. Gate 318 freezes a contrast projection and replays the same evidence: the compiler passes every algebra, rank, size, time, and memory clause, and its artifact is committed before any validation generator is authorized.

Original: paper/v2_sections/A2_record.tex · Raw source file

View raw TEX source
\section{Complete experimental record}\label{app:record}

This appendix is a faithful, archival record of every experiment in the program, so the paper is
self-contained and nothing is lost. The single organizing principle throughout: encode known
structure (an operator $L$) in the representation and solve in closed form; where structure is
absent, a generic random feature map with a quadratic/ridge readout is already optimal (the SSP
law). Matched beats generic \emph{iff} $L$ is non-trivial; when $L$ is trivial and the innovation is
Gaussian (e.g.\ frozen vision features) matched $=$ random, provably. The advantage is accuracy
where $L$ is known and categorical efficiency (no backprop, no replay buffer) everywhere. Honest
negatives are kept, not hidden. Each row cites a master-scoreboard entry (\#); accuracies are test
accuracy unless noted, and ``CF'' abbreviates our closed-form learner.

\subsection{Operator-matched closed-form: system ID, PDEs, FRI, neural operators}

The decisive regime: where $L$ is known, the matched closed-form solve beats deep neural operators
by orders of magnitude in both error and wall-clock, and recovers exact structure from real
measured data.

\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{@{}p{3.0cm}p{4.2cm}p{4.6cm}p{2.2cm}@{}}
\toprule
Experiment (\#) & Setup & Key numbers & Verdict \\
\midrule
ODE system ID (\S1) & matched dictionary $+$ ridge; autograd-free spline derivatives & non-poly extrap $\sim$400--2000$\times$ lower error, $\sim$1000$\times$ faster; $N{=}20$ stable while poly library diverges by $N{=}6$; conductance neuron 30--70$\times$ & decisive (known $L$) \\
PDE vs neural operators (\S2) & closed-form weak/strong-form ID vs FNO / DeepONet & 1D Burgers: 2-traj beats 32-traj FNO; 2D/3D react-diff Pareto-dominant; chaotic KS 0.19\% (weak form); cross-regime re-IDs $\nu$, $<$0.6\% & decisive; gap grows 1D$\to$3D \\
3D PDE at scale (\S2) & $M{=}32$, hard FNO 2000 ep on MPS & OSNR $\sim$0.001 nRMSE in $<$0.7\,s vs FNO 0.13--0.72 in 120\,s $\Rightarrow$ $\sim$100--600$\times$ error, $\sim$150--1700$\times$ speed; exact at $n{=}2$ & decisive \\
Real measured data (\S4) & Silverbox, EMPS, Cascaded Tanks, Allen neuron & Silverbox 1.8\,mV free-run; EMPS 14\,mm (matched friction $\to$ stability); Cascaded Tanks 0.55\,V latent-state mechanism (retrospective protocol audit); Allen $\tau{=}28.7$\,ms recovered & competitive / mixed \\
Closed-form SIREN (\#80) & Fourier-feature ridge vs deep sin-MLP, 64$\times$64 image & RFF F=1024: 28.1\,dB in 30\,ms (226$\times$ faster); matched dominant-mode 10.4\,dB (underfits); SGD SIREN 59.6\,dB & honest: broadband natural $\to$ deep wins; SSP boundary \\
Boundaries (\S6) & matched feature maps on unstructured frozen features; blind eq.\ discovery & matched (det.\ \& stochastic) $+$ sparse/MAP all tie/lose to random$+$ridge (trivial $L$); SINDy wins blind discovery & honest negatives \\
\bottomrule
\end{tabular}
\caption{Operator-matched closed-form learning. The matched advantage requires $L$ non-trivial and
the dynamics observable (scope law); on unstructured / broadband signals matched $=$ random, as SSP
predicts.}
\end{table}

\subsection{Continual learning, RSI, and multimodal generality}

The generality test: one closed-form Gram memory matches deep continual-learning SOTA on identical
frozen features while using $100$--$1000\times$ less compute and zero buffer, and is the stability
primitive for recursive self-improvement and multimodal composition.

\begin{table}[h]
\centering
\small
\begin{tabular}{@{}p{2.8cm}p{4.4cm}p{4.8cm}p{2.0cm}@{}}
\toprule
Experiment (\#) & Setup & Key numbers & Verdict \\
\midrule
Class-IL benchmarks (\S3) & CF (RanPAC-style) on frozen backbones, no task labels & Split-CIFAR-100/DINOv2-L \textbf{0.923} (DER++ 0.904, EWC 0.24); /ViT-B-21k \textbf{0.898}; Split-ImageNet-R/DINOv2-L \textbf{0.896}; /ViT-B 0.672; Split-MNIST 0.911 & match/beat SOTA, sec, 0 buffer \\
ODE continual stream (\S3) & matched dictionary, 20 regimes & $\sim$290$\times$ lower error than EWC; flat to 20 regimes; instant revisit recognition & decisive \\
RSI stability (\S7) & telephone-game self-train, frozen backbone, only update rule varies & 1 lab/cls: CF 0.49$\to$\textbf{0.62} (stable) vs SGD 0.40$\to$0.34 (drifts); 2 lab/cls: CF 0.66$\to$\textbf{0.73} vs SGD 0.51$\to$0.44 & CF = stability primitive \\
Grow $+$ no-forget capstone (\S8) & 100 skills, 10/gen, grow capacity + gated self-improve & all-skills \textbf{0.885}, first-skill retention \textbf{0.94} vs gradient-grow 0.012 / 0.0 ($\sim$70$\times$ gap) & decisive \\
LLM continual (\#26) & frozen GPT-2, 20NG 10-task class-IL & CF \textbf{0.673} $=$ joint upper bound (0 forgetting) vs seq SGD 0.057; retention gap 0.616 & joint optimum \\
LLM few-label RSI (\#26b) & 5 lab/cls $+$ unlabeled, NCM-gated pseudo-labels & seed-only 0.452 $\to$ \textbf{0.641} ($\sim$95\% of full-sup 0.667); SGD 0.055 & RSI on real LLM \\
Multimodal generalist (\#28) & one Gram substrate, ViT vision $+$ GPT-2 language, 14 skills/120 cls & full-label \textbf{0.794} $=$ joint optimum; few-label$+$RSI \textbf{0.752}; seq SGD 0.056 & one mechanism, many modalities \\
Three-modality agent (\#30) & $+$ control (pendulum world-model$+$MPC) skill & classes all-seen 0.794 UNCHANGED after control; shared backprop net $\to$ 0.056 & interference-free \\
Gradient-free deep credit (\#29/29b) & parity-of-3-hidden-hyperplanes (kernel-hard) & backprop 0.88; CEM-evolved 0.61; target-prop 0.57; random 0.53 & wall real but not absolute $\to$ hybrid (later overturned, \#31--35) \\
\bottomrule
\end{tabular}
\caption{Continual learning, RSI, multimodal. CF = closed-form Gram memory; retention is structural
(order-invariant joint optimum), a property backprop continual provably cannot have.}
\end{table}

\subsection{Neuroevolution and model-based control}

A GPU-native NEAT keeping NEAT's design but replacing its optimizer with a closed-form readout:
compact circuits found $\sim$100$\times$ faster than CPU NEAT, at scales NEAT cannot reach;
model-based MPC on exact closed-form world models cracks hard-exploration control with hundreds of
real interactions.

\begin{table}[h]
\centering
\small
\begin{tabular}{@{}p{2.9cm}p{4.3cm}p{4.8cm}p{2.0cm}@{}}
\toprule
Experiment (\#) & Setup & Key numbers & Verdict \\
\midrule
Parity ladder (\#9) & evolve topology, closed-form readout/candidate ($\sim$1--14\,ms) & 45--130$\times$ faster than SGD scoring; curriculum solves parity 2--7 at 1.0, parity-8 0.984 (1$\to$263 hidden); greedy collapses at 3 & topology load-bearing \\
Evolved recurrence (\#9b) & sequential parity, readout closed-form over time (no BPTT) & recurrent 1.00/0.99/0.91/0.83/0.72 ($T{=}4{-}20$) vs feedforward pinned $\sim$0.5--0.6 & recurrence breaks memory ceiling \\
GPU-NEAT vs vanilla (\#9c) & N-bit parity from scratch vs \texttt{neat-python} & ours$+$speciation solves 2--7 (1/3/2/22/37 hidden); vanilla solves only XOR, bloats; $\sim$1.6--4.8k evals/s ($\sim$100$\times$); 5-seed 100\% at 3--7 & crushes vanilla NEAT \\
Real-ish tasks (\#9d) & two-spirals, digits & spirals \textbf{0.994} (ties tuned 2-layer MLP); digits 0.934 (linear 0.920, tuned MLP 0.966) & topology discovery, not raw-acc supremacy \\
Evolved-circuit lifelong (\#9e) & evolve 23-neuron $\phi$, freeze, class-IL & CF all-seen \textbf{0.924} / retention \textbf{0.931} vs SGD head 0.428 / 0.069; 4-seed 0.940 / 0.926 & no-backprop WINS \\
RL neuroevolution (\#10) & GPU-batched rollouts, evolved recurrent policy & no-velocity CartPole: recurrent solves 500 (3/3); feedforward fails 57--60 (0/3) & recurrence load-bearing when non-Markovian \\
Scaling to 10$^5$ (\#11) & dense vs sparse edge-list forward & sparse 100k-neuron circuit: 0.21\,s / 114\,MB on one M4 GPU & forward feasibility (search open) \\
Model-based RL (\#12) & closed-form dynamics from few transitions, evolve controller in model & 250 real transitions $\to$ \textbf{500} (solved), model R$^2{=}$1.0; model-free needs 209k $\Rightarrow$ $\sim$800$\times$ & decisive \\
Pendulum / swing-up (\#16,\#18) & exact world model $+$ closed-loop CEM-MPC & stabilize $-74.3$ (1k transitions) vs model-free $-308.6$ (28.8M) $\Rightarrow$ $\sim$28{,}800$\times$; swing-up CEM-MPC $-347$ (planning cracks it) & win (planning) \\
MountainCar / quadrotor (\#19,\#20) & exact model (R$^2{=}$1.0) $+$ CEM-MPC & MountainCar success 1.0 ($\sim$111 steps, 5 seeds); PVTOL hover error 0.006 (4 seeds) & robotics-relevant wins \\
Language pillar (\#13--15) & evolved recurrent circuit, char-level next-char & evolved-from-reservoir 0.541 (vs unevolved 0.433, bigram 0.350, trigram 0.806); scales to 0.634 at 241 neurons; K-delay echo beats n-grams beyond window & substrate works, below trigram (honest) \\
Block-growth self-scaling (\#14) & grow in modules, minimal start, no seeding & block=8 \textbf{0.599$\pm$0.004} vs block=1 0.442 (lang); digits 0.957; flagship 518 neurons $\to$ 0.70; curve 0.57$\to$0.70 & autonomous-scaling unlock \\
Lifelong-evolving capstone (\#17) & block-grown $\phi$ $\to$ class-IL Gram & 245 neurons: all-seen \textbf{0.961$\pm$0.004}, retention 0.953$\pm$0.029 (4 seeds) & grow $+$ no-forget, robust \\
ES/PGPE (\#76) & cartpole, ES vs GA & GA solves (471, 320 ep); quick ES stalls (deceptive flat reward) & niche tool, not universal \\
\bottomrule
\end{tabular}
\caption{Neuroevolution and model-based control. Closed-form pays off in RL by removing the
\emph{model} bottleneck (exact model from few transitions $\to$ plan), not the rollout.}
\end{table}

\subsection{Bio-plausible learning and predictive coding}

The thesis that backprop's apparent supremacy is rote memorization, not learning quality: local
(no-global-backward) credit assignment matches backprop, generalizes better, is depth-robust, and
predictive coding $=$ backprop at the gradient level once run in its correct regime.

\begin{table}[h]
\centering
\small
\begin{tabular}{@{}p{2.9cm}p{4.3cm}p{4.8cm}p{2.0cm}@{}}
\toprule
Experiment (\#) & Setup & Key numbers & Verdict \\
\midrule
SSP on frozen vision (\#21) & evolved circuit vs random projection, ViT CIFAR-100 & linear 0.881, RanPAC 0.892, evolved 0.882 (ties) & matched $=$ random, as SSP predicts \\
Known-operator vision (\#22--25) & DoG$+$Gabor$+$scattering / steerable / Coates-Ng, CIFAR-10, no backprop & raw 0.383 $\to$ operators 0.671 $\to$ learned dict \textbf{0.718}; steerable 0.66; deep steerable plateau 0.682 & no-backprop ceiling $\sim$0.72 (honest) \\
Steerable $\theta$-free (\#27) & paper's Question B, learned-$\theta$ vs steer-to-all & ep1 0.289 vs 0.226 ($+$6.3pp); final parity; 35$\times$ fewer conv params & $\theta$-bootstrap avoidable \\
DFA / decoupled-greedy (\#31,\#32,\#32b) & local credit, teacher-student & DFA 0.711 vs backprop 0.780; decoupled-greedy \textbf{0.796$\pm$0.013} $\ge$ backprop 0.782$\pm$0.012 (5 seeds, full); small-data tie & wall overturned; local matches backprop \\
Memorization vs learning (\#33) & random-label fit; 40\% label noise & backprop memorizes random labels 1.000; under noise clean-test backprop 0.464 vs local \textbf{0.507} & local generalizes better, memorizes less \\
Depth scaling (\#36) & deep teacher, student depth 2--8 & backprop degrades 0.685$\to$0.640; local flat \textbf{0.683} (gap widens to $-$0.043) & local depth-robust \\
Closed-form-local / OOD (\#34,\#35,\#37) & CF head into features; input-scale shift; label-free local & CF-local 0.655 $<$ greedy 0.788 (honest neg); OOD non-differentiating; contrastive fails 0.491, sparse-coding 0.718 & honest negatives bracketed \\
Reasoning / e-prop (\#40--46) & multiplicative register; working memory; RTRL & WM: e-prop random-feedback \textbf{0.999/0.992/0.974} $=$ BPTT (no BPTT, no weight transport); RTRL 0.475 $=$ BPTT; algorithmic e-prop diagonal lags (off-diagonal credit) & recurrent bio-learning matches BPTT \\
Conv / InfoPro scale (\#47--50) & local sweep on real conv CIFAR-10 & decoupled-greedy 0.839, InfoPro 0.853 vs backprop 0.864 (within 1.1pt); DFA bolt-on fails 0.74--0.76 & local matches, beat-A at scale only \\
Online recurrent (\#51,\#52) & UORO rank-$k$ vs e-prop & UORO rank-1 0.285 $\to$ rank-16 0.394 $<$ e-prop 0.568 $<$ BPTT 0.659 & e-prop = practical sweet spot \\
Predictive coding (\#53--62) & PCN vs backprop, gradient-alignment unit test & hard-clamp loses (0.563); Z-IL cos $=$1.000, nudged $=$0.99; PC-nudged 0.674 $=$ backprop-MSE 0.678; no robust PC$>$backprop & PC $=$ backprop in correct regime \\
PC on arbitrary topology (\#57,\#58,\#66) & skip-DAG; PC-NEAT grow & PC 0.587 $=$ backprop-CE 0.581 on skip-DAG; PC-NEAT parity-8 chance$\to$\textbf{1.000} (grew 2$\to$3 nodes) & PC trains evolved wiring locally \\
Closed-form precision PC (\#60) & one local downward sweep (ePC/HGF) & CF-PC 0.662 $=$ backprop-CE 0.658; precision-whitening hurts (0.632) & deep-capable closed-form local learner \\
\bottomrule
\end{tabular}
\caption{Bio-plausible learning and predictive coding. Across MLP, conv, recurrent (e-prop) and
skip-DAG topologies, local / no-global-backward credit assignment matches backprop; PC equals
backprop at the gradient and accuracy level. The PC equilibrium solved in one sweep \emph{is} the
OSNR closed-form principle applied to the cortical learning rule.}
\end{table}

\subsection{The gradient-free cortex, scale-up, and application fronts}

The synthesis: a fully gradient-free agent that builds its own architecture (NEAT), learns deep
features from pixels by a per-block local error sweep (no global backward), and accumulates tasks
with zero order-invariant forgetting (closed-form Gram), at competitive-to-SOTA accuracy and the
$10^5$-neuron scale target, across vision, language, control, and games.

\begin{table}[h]
\centering
\small
\begin{tabular}{@{}p{2.9cm}p{4.3cm}p{4.8cm}p{2.0cm}@{}}
\toprule
Experiment (\#) & Setup & Key numbers & Verdict \\
\midrule
Federated cortex (\#63--65) & closed-form PC area $+$ Gram memory, digits/CIFAR continual & 2-task digits CF 0.967, A-after-B 0.978 vs backprop 0.009; CIFAR 5-task front-end$+$Gram 0.649, gen-PC$+$ 0.655; backprop 0.189 & no-backprop $+$ no-forget holds \\
Shallow ceiling diagnosed (\#67,\#70) & richer front-end; head depth & richer 0.683 (order-independent, exact); depth0--3 PC 0.681--0.685, even backprop plateaus $\sim$0.69 & 0.68 is fixed-front-end ceiling, not no-backprop \\
Ceiling broken (\#71,\#72) & deep conv from pixels by per-block LOCAL sweep $+$ Gram & local-from-pixels 0.842 (backprop 0.864); deep-local$+$Gram \textbf{0.889} 5-task continual (reverse-order identical), backprop 0.189 & competitive no-backprop, no-forget \\
Self-constructing agent (\#73,\#74) & NEAT evolves conv arch, local weights, Gram & seed [32] 0.429 $\to$ [32,64,192] 0.693; full agent \textbf{0.864} 5-task continual (order-independent) vs backprop 0.193 & builds own arch, no backprop, no forget \\
SOTA-class vision (\#75,\#84) & deeper CNN $+$ aug, per-block local sweep, CIFAR-10 & local \textbf{0.9024} $\ge$ backprop 0.8855 (same net); multi-seed local \textbf{0.8994$\pm$0.0004} $>$ backprop 0.8824$\pm$0.0024 (3/3) & local matches-or-beats backprop, locked \\
Multimodal deep / scale (\#77,\#78) & deep no-bp vision $+$ GPT-2; neuron accounting & vision lifted 0.653$\to$\textbf{0.888}, language 0.652 retained; base substrate 246k neurons/img, 0.855 in 12 ep, $<$1\% of 128\,GB, 7k img/s & $10^5$-neuron target met \\
Games from pixels (\#79,\#82) & local-sweep perception $+$ (known/learned) dynamics $+$ MPC & CartPole from pixels 236/300 ($\sim$3k frames); Catch learned-dynamics R$^2$ 0.997, catch-rate \textbf{0.965} ($\sim$2.7k transitions) & DQN recipe, no backprop, no millions of frames \\
Transformers / real GPT (\#83,\#85) & per-block local sweep on attention & induction-head \textbf{1.000} $=$ backprop; real char-GPT (4 blocks) val bpc \textbf{1.946} vs backprop 1.933, coherent generation & PC $=$ backprop on the full architecture family \\
Atari Pong (\#86,\#90) & local-sweep conv policy on real ALE; BC then DAgger (relabel policy-visited states with predictive teacher) & BC $-$21.0 (distribution shift); DAgger lifts $-$21$\to$\textbf{$-$8.8} over 4 rounds (teacher $-$5.4, random $-$20.8), all no global backward & gradient-free policy plays real Pong from pixels \\
Online video (\#81) & deep local perception $+$ online Gram, day$\to$night shift & online Gram DAY 0.922 retained vs online backprop head 0.880 (forgets) & autopilot-style no-forget perception \\
\bottomrule
\end{tabular}
\caption{The gradient-free cortex and application fronts. One agent self-constructs its
architecture, learns deep features from pixels with no global backward pass, and accumulates skills
with structural (order-invariant) zero forgetting; backprop continual collapses to $\sim$0.19 on the
same streams. The learning rule was never the limit---architecture, augmentation, and scale were.}
\end{table}

\subsection{Standardized benchmark campaign}\label{app:benchmarks}

To test the substrate against the field's canonical benchmarks rather than bespoke setups, we ran a
standardized campaign across continual learning, language, scientific computing, and single-task
vision. Each result is gradient-free (no global backward pass) and reproduced by a single script.
The pattern is consistent with the theory: on \emph{specifiable-structure} tasks (PDEs) the
operator-matched closed-form solve wins by orders of magnitude; on generic perception the local
sweep matches or beats backprop on the same architecture; on continual streams the closed-form
Gram memory is competitive with the no-gradient SOTA with exactly zero forgetting.

\begin{table}[h]
\centering
\small
\begin{tabular}{@{}p{3.0cm}p{4.4cm}p{5.0cm}p{1.6cm}@{}}
\toprule
Benchmark (\#) & Setup & Key numbers & Verdict \\
\midrule
Split-CIFAR-100 CIL (\#87) & frozen ViT-B/16-in21k, random-proj $+$ closed-form Gram, 10$\times$10 & final \textbf{0.900}, avg-inc 0.936, order-invariant (zero forgetting), no gradient; RanPAC $\sim$0.92 & competitive, zero-forget, gradient-free \\
Split-ImageNet-R CIL (\#88,\#89) & same recipe, 200 classes, 10$\times$20 & final 0.669 (avg-inc 0.690), order-invariant; below RanPAC $\sim$0.78 (gap $=$ second-moment whitening; reimpl did not close it) & honest partial \\
enwik8 byte-GPT (\#91) & 4-block causal transformer, byte-level, per-block local sweep & TEST \textbf{2.24}\,bpc vs backprop 2.238 (identical), local FASTER (122\,s vs 416\,s); learned Wikipedia markup & local $=$ backprop on standard LLM corpus \\
PDEBench-formal 1D (\#92) & Advection/Burgers/Diff-React, OSNR closed-form vs trained FNO (74k\,p), nRMSE & extrap nRMSE: adv 0.004 vs 0.314, Burgers 0.042 vs 0.243, DR 0.000 vs 0.023; $\sim$400$\times$ faster, FLAT extrapolation & orders-of-magnitude win (matched regime) \\
CIFAR-100 single-task (\#93) & deep 4-block CNN $+$ aug, per-block local sweep & local \textbf{0.660} $>$ backprop 0.596 (same net), faster, monotonic; from-scratch (not SOTA $\sim$0.85) & gradient-free beats backprop, 100 classes \\
MuJoCo HalfCheetah (\#94--98) & deep dynamics model by local sweep $+$ short-horizon CEM-MPC, all no global backward & ridge model underfits (R$^2$ 0.64, return $-$60); deep local-sweep model R$^2$ \textbf{0.99}; short-horizon CEM-MPC \textbf{345$\pm$39} (8 ep) vs random $-$415, from $\sim$6--8k transitions & gradient-free continuous control: runs forward stably \\
\bottomrule
\end{tabular}
\caption{Standardized gradient-free campaign: continual learning, language,
scientific computing, vision, and continuous control.}
\end{table}

\begin{table}[p]
\centering
\footnotesize
\begin{tabular}{@{}p{3.0cm}p{4.4cm}p{5.0cm}p{1.6cm}@{}}
\toprule
Benchmark (\#) & Setup & Key numbers & Verdict \\
\midrule
Atari Freeway --- ES baseline (\#99) & OpenAI-ES (Salimans 2017) evolving a conv policy from pixels; \emph{standard method, not ours} & 21.0 vs random 0, from $\sim$176k env steps ($\sim$2.2M total); converges by gen 1 & gradient-free baseline FLOOR (reproduces ES-plays-Atari) \\
Atari Freeway --- ours (\#100) & local-sweep autoencoder $+$ latent dynamics $+$ discrete CEM-MPC, all no global backward & \textbf{21.3$\pm$0.5} vs random 0, from $\sim$\textbf{14k} env steps ($\sim$12$\times$ fewer than ES; $\sim$150$\times$ vs ES budget) & our-framework win, sample-efficient; Freeway easy (caveat \S\ref{app:mbrl}) \\
Atari Pong --- model-based (\#101) & reward-MPC; ball-tracking stress test & $-$21.0, reward-$R^2\approx0$ (reconstruction latent drops the $\sim$2$\times$3-px ball; sparse/delayed reward) & wrong tool for sparse reward (\S\ref{app:mbrl}) \\
Atari Pong --- local-sweep DQN (\#102--103) & value-based (Double, n-step) TD, credit assignment by local sweep, no global backward, Q-net on GPU (arch: Fig.~\ref{fig:arch-dqn}) & \textbf{wins from pixels, reproducibly: final greedy $+$18.5$\pm$0.8 over 3 seeds} (4/4 incl.\ seed0 positive; ceiling $+$21, random $-$20.7) & \textbf{reliable, not luck}; matches a DQN baseline (not SOTA) \\
Atari Breakout --- same config (\#105) & \emph{identical} stabilized DQN config, no game-specific tuning & final greedy \textbf{58.9} (curve 0$\to$31$\to$66 peak; random $\sim$1) over 1M steps & generalizes: the fix is principled, not Pong-overfit \\
\bottomrule
\end{tabular}
\caption{Standardized gradient-free campaign: pixel-control benchmarks;
remaining gaps are explicit.}
\end{table}

\subsection{Gradient-free model-based control: one recipe, state and pixels}\label{app:mbrl}

The control results (MuJoCo HalfCheetah \#94--98, Atari Freeway/Pong \#99--100) share a single
gradient-free recipe: learn a world model with the per-block local sweep, then plan with the
cross-entropy method (CEM-MPC). No global backward pass appears anywhere---not in perception, not in
the dynamics model, not in the planner. The only difference between the low-dimensional-state case
(MuJoCo) and the pixel case (Atari) is a front-end \emph{autoencoder} that supplies a latent state.

\paragraph{Components (all trained by the local sweep, Alg.~\ref{alg:mbrl}).}
\emph{(1) Encoder/latent (pixels only).} A convolutional autoencoder $E_\phi,D_\psi$ is trained on
collected frames by the per-block local error sweep on the reconstruction loss
$\lVert D_\psi(E_\phi(o))-o\rVert^2$; activations are detached between blocks and each block updates
from its own local target (no backprop through the stack). The encoder $E_\phi$ then maps a frame to a
$d_z{=}64$ latent $z$. For low-dimensional state (MuJoCo) this step is skipped and $z$ is the raw state.
\emph{(2) Latent dynamics + reward.} A deep MLP $f_\theta([z,a])\mapsto(\Delta z, \hat r)$ is trained by
the local sweep to predict the next-latent increment and the immediate reward (one-hot actions for
discrete control). \emph{(3) Planner.} CEM-MPC rolls candidate action sequences through $f_\theta$ in
latent space and keeps the elite by predicted return: a per-step Gaussian refined to its elite
mean/variance (continuous, MuJoCo) or a per-step categorical refined to its elite action frequencies
(discrete, Atari); the first action is executed, receding-horizon. \emph{(4) Aggregation.} Roll out the
planner, append the visited transitions, refit---the dynamics analogue of DAgger (\#90).

\begin{algorithm}[H]
\caption{Gradient-free model-based control (local-sweep world model + CEM-MPC)}\label{alg:mbrl}
\begin{algorithmic}[1]
\State Collect random transitions $\mathcal{D}=\{(o,a,r,o')\}$
\For{round $=1,\dots,R$}
  \If{pixels} train autoencoder $E_\phi,D_\psi$ on $\{o\}$ \textbf{by the local sweep} (recon loss); set $z=E_\phi(o)$
  \Else{} $z=o$
  \EndIf
  \State Train $f_\theta([z,a])\!\to\!(\Delta z,\hat r)$ on $\mathcal{D}$ \textbf{by the local sweep} (MSE) \Comment{no global backward}
  \State \textbf{CEM-MPC:} at each step, refine per-step action distribution by elite predicted return in $f_\theta$; execute first action
  \State Append planner-visited transitions to $\mathcal{D}$
\EndFor
\end{algorithmic}
\end{algorithm}

\paragraph{Results and honest scope.} On MuJoCo HalfCheetah a deep local-sweep dynamics model (reward
$R^2\,0.99$, vs $0.64$ for a closed-form ridge+random-feature model that underfits contact dynamics)
with short-horizon CEM-MPC runs the robot forward at $345\pm39$ (random $-415$); the horizon had to be
shortened ($25\!\to\!12$) to suppress model-exploitation over long plans (\#94--98). On Atari Freeway
the same pixel pipeline reaches $21.3\pm0.5$ (random $0$) from $\sim$14k environment steps, $\sim$12$\times$
fewer than the gradient-free Evolution-Strategies baseline (\#99, $21.0$ from $\sim$176k steps) and
$\sim$150$\times$ fewer than the ES total budget---the sample-efficiency edge of model-based planning.
We are explicit about the limits: Freeway's optimal policy is nearly trivial (press \textsc{up}), and the
reward head's $R^2$ degraded under aggregation while the score held, so Freeway demonstrates the pipeline
but not a high-quality \emph{pixel} world model. The stronger test, Pong (\#101), is a clean
\emph{negative}: the same pipeline scores $-21.0$ (random $-20.7$) with reward-$R^2\approx0$ from the first
round. The cause is diagnosed, not hand-waved---the encoder is trained by \emph{reconstruction}, whose
pixel-MSE is dominated by large static structures (paddles, background, score), so Pong's $\sim$2$\times$3-pixel
fast ball is dropped from the latent; with no ball in $z$, reward (a point) is unpredictable and the planner
is blind. This is the textbook failure of reconstruction-latent world models on small reward-relevant
features (why Dreamer-v2/v3 and TD-MPC2 use reward/value-predictive or self-supervised forward-predictive
latents). The fix is a task-aware latent---action-conditioned next-frame prediction, reward/value-predictive
encoding, or motion (frame-difference) input, all local-sweep-compatible---and the scaling of all of this to
the full 49-game suite is a cluster task (Track A3 in \texttt{RUNPOD\_TODO.md}).

\paragraph{Value-based control with local credit assignment (Pong).} Pong's reward is sparse and
delayed---the wrong regime for reward-seeking MPC but the home turf of \emph{value bootstrapping}, where TD
learning propagates the rare $\pm1$ backward into a dense learned value. We therefore ran a DQN (standard
$84\!\times\!84\!\times\!4\to$conv$_{32/64/64}\to$FC$_{512}\to Q$, replay buffer, target network,
$\epsilon$-greedy) but with credit assignment by the \emph{per-block local sweep} instead of backprop---the
TD error on the taken action is propagated block-locally, no global backward pass, Q-network trained on the
GPU. It \textbf{learns Pong from pixels}: greedy score $-21\!\to\!-9.3$ over $\sim$1.5M frames, action
distribution spreading across all six actions. This establishes that the gradient-free local rule performs
value-based RL---bootstrapped, non-stationary targets---on a real Atari game, not only supervised and
model-based learning. \emph{Reproducible}, not a lucky run: the initial setup was high-variance (1 of 5
single-machine runs reached $-9.3$), which we root-caused---not to the local rule (it computes the exact
backprop gradient) but to a replay buffer $10\times$ too small (correlated samples $\to$ Q-collapse to a
degenerate action) and 1-step TD too slow for the sparse reward. With the standard fixes (300k buffer,
n-step returns $n{=}3$, Double-DQN), three fresh seeds all converge to a winning score---final greedy
$+19.6/+17.8/+18.1$, mean $\mathbf{+18.5\pm0.8}$ (ceiling $+21$, random $-20.7$), every seed a smooth
monotone-trend climb with no collapse. An earlier flat $-21$ traced to a preprocessing bug of ours (a
missing frame-max de-flicker that erased Pong's flickering ball), not the learning rule. Honest scope: this
is algorithmically a (Double, n-step) DQN---the contribution is the no-global-backward \emph{local} rule
doing value-based RL to a reproducible near-ceiling score, not a new RL algorithm or a SOTA result; the
full 49-game suite is deferred to RunPod (Track A1).

\paragraph{Stable recursive self-improvement: two knobs make a self-training loop compound instead of collapse.}
We frame recursive self-improvement (RSI) as a dynamical system in capability-space: the binding constraint is the
\emph{stability} of the closed self-training loop (a tiny labelled seed pseudo-labels an unlabelled pool, the model
learns from its own labels, repeat), not the per-step improvement. Backprop is allowed---it is the baseline that
collapses. On frozen DINOv2 features, naive backprop self-training collapses ($0.61\!\to\!0.32$, self-label error
$0.39\!\to\!0.68$) while the same loop anchored to a never-forget Gram memory is collapse-proof and self-improves
($0.318\!\to\!0.661$): the additive, order-invariant memory is the Lyapunov anchor. For the harder case---a
\emph{trainable} CNN representation self-improving from scratch on CIFAR-10---a fixed reference anchor either caps or
drags the net, but a slow EMA mean-teacher that \emph{tracks} the improving student (the dissipative drift-following
anchor) yields stable self-improvement. Two stability knobs are each necessary. (i) The EMA anchor must EMA the
batch-norm statistics \emph{consistently} with the weights; copying the fast student's BN stats onto slow-EMA weights
produces a mismatched teacher that looks like collapse (worse the slower the EMA: at momentum $0.95$ the teacher fell
$0.43\!\to\!0.22$)---a BN-consistency bug, not a property of slow anchoring. With it fixed, the EMA teacher climbs
\emph{monotonically} where naive is noisy or degrades: at $25$/class it compounds $0.183\!\to\!0.326$ over $30$ rounds
(still rising, the per-round gain \emph{growing} as the self-paced pool admitted grows $31\!\to\!362\!\to\!788$),
while naive peaks early and settles back at chance; at $100$/class it is smooth-monotone $0.430\!\to\!0.478$.
(ii) The confidence threshold is the curriculum: at $\text{conf}{=}0.95$ the loop admits only reliable labels and
compounds, but relaxing to $0.90$ floods the loop with the teacher's lower-confidence errors ($21$--$49$k of the
$\sim$$50$k pool admitted) and reproduces the classic confirmation-bias collapse ($\text{EMA peaks}\sim\!0.20\!\to\!0.13$).
Honest scope: absolute accuracies are modest (small CNN; $25$/class reaches $0.326$, $39\%$ of the $0.827$ supervised
ceiling) and EMA mean-teaching is established semi-supervised learning---the contribution is the framing (RSI as a
stable dynamical system, slow anchor $+$ confidence curriculum as the stabilizers) and the precise characterization
that stable self-improvement is a BN-consistency-plus-curriculum property of the anchor.

\paragraph{One frozen encoder (or several) + one closed-form memory = a cross-domain, cross-modal lifelong brain.}
The closed-form never-forget memory composes into a concrete instantiation of ``attach our framework to any frozen
model and get a lifelong-learning brain.'' On a single frozen DINOv2 ViT-S encoder we stream three very different
\emph{datasets} as a task-free sequence---CIFAR-100 (objects) $\to$ Flowers-102 $\to$ FGVC-Aircraft (planes),
unified into $302$ classes, evaluated by predicting over \emph{all} seen classes with no task-id. A finetuned head
catastrophically forgets (final average $0.035$, forgetting $+0.42$); the closed-form memory retains every domain
with near-zero forgetting ($+0.011$) at average $0.754$, beating nearest-class-mean ($0.720$, the edge on
fine-grained aircraft, $0.498$ vs $0.368$, from feature decorrelation). One honest pitfall surfaced and was fixed:
plain accumulation is \emph{swamped} by dataset-size imbalance (CIFAR's $50$k examples vs Flowers' $2$k drive the
ridge to fit the large set and ignore the small ones, $0.000$ on Flowers/Aircraft---imbalance, not forgetting, since
CIFAR stays at $0.764$); class-balanced accumulation ($G=\sum_i w_i z_i z_i^\top$, $w_i=1/n_{c_i}$) is the correct
default for imbalanced multi-domain streams. The \emph{strong} form crosses modalities: two different encoders
(DINOv2 for images, BGE-base for text) feed one shared memory through per-modality random feature maps into a common
space, on an interleaved vision/language stream (CIFAR-100, banking77 intents, Flowers, dbpedia topics, Aircraft;
$393$ classes). One memory spans both modalities with near-zero forgetting ($+0.006$, average $0.803$) where the
finetuned head collapses ($0.069$, $+0.27$); the per-modality random maps occupy near-orthogonal subspaces, so
vision and language do not interfere. The contribution is the architecture---any encoder, any modality, one
gradient-free buffer-free closed-form lifelong memory at $\sim$zero forgetting---not a new accuracy record (it tracks
the frozen encoders' ceilings, per the teacher-scaling law).

\paragraph{Bulletproofing the closed-form memory: it beats replay, learns new classes instantly, and is private against exemplar attack.}
Three receipts harden the continual-learning claim, all on frozen DINOv2 features. (i) \emph{Beats buffered replay.} On Split-CIFAR-100
class-incremental learning, the gradient-free never-forget memory reaches $0.868$ at average forgetting $0.050$ versus a $2000$-exemplar
replay head at $0.801$/$0.183$ and naive online SGD at $0.335$/$0.703$; on the harder cross-modal stream a \emph{fair} replay (balanced
half-current/half-buffer batches) still collapses to $0.188$ against the memory's $0.803$, because $2000$ exemplars over $393$ classes
($\sim$5/class) cannot rehearse a high-dimensional head while the memory accumulates every example exactly as sufficient statistics. (ii)
\emph{Instant few-shot class-incremental.} A new class is one additive update ($O(1)$, no epochs); cross-domain FSCIL (CIFAR-100 base, then
Flowers/Aircraft $K$-shot) keeps base retention exactly constant ($0.773$--$0.775$, zero forgetting) while new-class accuracy scales with
shots ($0.663$ at $1$-shot to $0.773$ at $10$-shot). Finetuning faces an unwinnable stability--plasticity dilemma: few steps underfit the
new classes (new $0.000$), more steps erase the base (base $0.000$). (iii) \emph{Privacy.} A membership-inference attack cannot distinguish
the parametric memory from a gradient head (attack AUC $\sim\!0.53$ for both---privacy comes from frozen features plus regularisation, not the
closed form per se), but both are near-private against the \emph{exemplar} alternative: a replay/RAG store of raw features is trivially
de-anonymised (AUC $1.0$). The memory's privacy edge is therefore over the exemplar methods it replaces, exactly the buffer it does without.

\paragraph{The deep frontier: anchoring a never-forget memory to a \emph{plastic} representation.}
Every lifelong-brain result above freezes the encoder; the open problem is retaining old tasks while the representation keeps learning, where
the memory's stored sufficient statistics go stale as features drift. Isolating it (a trainable MLP on frozen DINOv2 features, Split-CIFAR-100,
$5\times20$), a live plastic representation forgets catastrophically ($0.291$, forgetting $0.844$) because the stored statistics no longer match
the drifted features. The fix is the same slow anchor that stabilised recursive self-improvement: an EMA-slow copy of the representation feeds
the memory, and as the EMA momentum increases ($0.90\to0.95\to0.99$) retention climbs monotonically ($0.645\to0.765\to0.837$, forgetting
$0.373\to0.060$), the $0.99$ case nearly matching the frozen upper bound ($0.885$) while the representation still trains on every task---the
momentum is a clean stability--plasticity knob. Two honest bounds. First, freezing still edges EMA when the encoder is adequate (confirmed on
CIFAR-100 \emph{and} fine-grained Aircraft, $0.705$ vs $0.541$): plastic-plus-EMA is the controllable fallback for when a frozen encoder is
insufficient, not a universal improvement, and demonstrating EMA${>}$frozen needs a genuinely out-of-domain stream. Second, the EMA anchor needs
a good base to anchor toward: from a random initialisation a slow EMA simply preserves the random start (a from-scratch CNN gives EMA
$\approx$ frozen-random $\approx$ chance), the same lesson as the gradient-free self-improvement negative---it resolves the staleness catch-22
given a usable representation but cannot manufacture one. The validated recipe therefore remains: freeze a strong pretrained encoder and let the
closed-form memory carry the lifelong, cross-domain, cross-modal continual learning on top.

\paragraph{A unified agent brain: one closed-form memory that both perceives and acts.}
The lifelong-brain wins are all perception; a real agent must also act. We give a single never-forget memory three faculties in one interleaved
stream---vision (DINOv2 CIFAR-100), language (BGE banking77), and control (behavior-cloning a gradient-free CartPole expert found by random
linear-policy search, return $500/500$)---through per-faculty random feature maps, and after the stream evaluate all three, the control faculty
by \emph{deploying} the memory as a policy in the live environment. One memory perceives (vision $0.689$, text $0.550$) and acts at expert level
(CartPole $500/500$) with zero forgetting, while a sequentially-finetuned head catastrophically forgets every earlier faculty (vision/text at
chance $0.03/0.04$) and even underperforms on control ($418$). The design that makes this work---diagnosed by first finding closed-form control
underperforming in a naive shared memory ($234$), then isolating that closed-form ridge clones the policy nearly as well as cross-entropy SGD
when given its own features ($478$ vs $492/500$, so the deficit was interference, not least-squares)---is \emph{fully independent per-faculty
blocks}: each faculty gets its own bias, random-feature coordinates, and regularization (perception $\lambda{=}100$ class-balanced, control
$\lambda{=}1$ full-mass) inside one block-structured solve. Shared coordinates corrupt the joint Gram matrix (an early version gave vision
$0.000$); a shared bias is swamped by the high-mass control faculty; class-balancing crushes the two-action control faculty. With independent
blocks all three reach their standalone ceilings at once. The contribution is architectural: a single gradient-free, buffer-free memory spanning
perception \emph{and} action with structural zero-forgetting---the closed-form approach extends to decision-making, not only classification.

\paragraph{Add to remember, subtract to forget: exact machine unlearning falls out of the additive structure.}
The same additivity that makes the memory order-invariant and zero-forgetting also makes the right-to-be-forgotten problem trivial. Because
$G=\sum_t Z_t^\top Z_t$ and $B=\sum_t Z_t^\top Y_t$ are sums over tasks, a task is removed \emph{exactly} by subtracting its contribution
($G\!-\!=\!Z_k^\top Z_k$, $B\!-\!=\!Z_k^\top Y_k$) followed by one re-solve---$O(1)$ in the forgotten task's size, gradient-free, no access to
the retained data. On Split-CIFAR-100 (frozen DINOv2 ViT-B, 10 tasks), unlearning a task this way yields a model bit-identical to one retrained
from scratch without it ($\max|W_{\text{subtract}}-W_{\text{scratch}}|\approx1.5\times10^{-5}$, retained accuracy $0.872$ in both), whereas an SGD
head finetuned on the remaining data carries no such guarantee (its weights still encode the removed task; exact removal would require full
retraining). Add to remember, subtract to forget---both exact, both $O(1)$ per task---a privacy/compliance capability gradient training cannot match.

\paragraph{The limits of self-improvement: self-reference plateaus, verification compounds the readout, only teaching improves the representation.}
Asking directly what makes an online model \emph{self-improve}, we mapped which information source lets the loop compound. (i) \emph{Self-labels
(no new information)} degrade or plateau everywhere---a model training on its own predictions can only sharpen what its features already separate.
(ii) We first hoped our stability machinery would tame online RL, but it is an honest negative: a closed-form gated memory is no more stable than
gradient online-TD on CartPole (gradient final $256$/late-std $84$ vs memory $220$/std $144$), because RL targets are \emph{non-stationary}
(bootstrapped and policy-dependent) whereas our memory's advantage requires \emph{stationary} targets---the deadly triad is a target problem the
memory just faithfully fits. (iii) This predicted the right source: \emph{verification}. A verifier of reliability $v$ (checking correctness is
cheaper than generating the answer) supplies new \emph{and} stationary information. Added to the very fixed-feature never-forget loop that degrades
on self-labels ($v{=}0.5$: $0.32{\to}0.29$), it compounds monotonically to the supervised ceiling as $v$ rises ($v{=}0.9\to88\%$, $v{=}1.0\to
0.51{=}102\%$)---a threshold effect ($v$ must be $\gtrsim0.9$). But verification breaks only the \emph{readout} plateau: with a trainable CNN it
does \emph{not} break the representation plateau ($v{=}1.0$ max $0.456$ $=$ $v{=}0.5$ max $0.455$, both $\sim55\%$ of ceiling), because
verified-correct labels concentrate on already-correctly-classified examples and carry no new \emph{representational} signal (and below a
viability threshold---a near-chance start---even a perfect verifier cannot bootstrap, as there is nothing correct to certify). (iv) What breaks the
representation plateau is \emph{teaching}: external true labels on pool examples lift the CNN from $0.46$ to $0.63$--$0.66$ ($\sim77\%$ of ceiling)
where verification stalls at $0.46$, and hard-vs-random selection is a minor secondary effect---the decisive factor is new information about examples
the model \emph{cannot already handle}. \textbf{Conclusion:} self-improvement compounds only up to the information already latent in model-plus-data;
verification (new correctness information) compounds the knowledge/readout layer to its ceiling, but improving the representation itself requires
external teaching---it cannot be reached by self-reference or by verifying the model's own outputs. A precise, honest boundary on ``recursive self-improvement.''
(v) \emph{Deployability caveat.} The verifier must be \emph{external}: a self-derived cheap verifier---ensemble agreement / query-by-committee over
random feature-subsets of the same model---does not substitute for it (max $0.372$ vs.\ self-training $0.375$ vs.\ oracle $0.477$), because members
sharing the same features have errors \emph{correlated} with the model's own (they agree-but-wrong on feature-confusable classes), so agreement
$\approx$ confidence, not correctness. Verification-based self-improvement therefore works in domains with a genuine external checker---code (unit
tests), math (execution/proof), simulation---not from a model's own self-consistency; there, never-forget memory $+$ external verifier gives
gradient-free compounding to the ceiling, with zero forgetting and exact unlearning. (vi) \emph{Instantiated with a real LLM.} A frozen
\texttt{qwen2.5vl:7b} on $3\times3$-digit multiplication (a genuine weak spot, $0$-shot $0.31$), with EXECUTION as the external verifier and a
never-forget memory of execution-verified worked solutions retrieved as few-shot context, self-improves to $0.49$ ($+18$ points) over $120$
problems---no weight updates, no labels, only the execution signal. A control isolates verification as the cause: showing the model's \emph{own
unverified} solutions (including wrong ones) as exemplars \emph{hurts} to $0.13$, well below both verified ($0.49$) and $0$-shot ($0.31$), because
it imitates the wrong procedure. Execution-filtered exemplars help decisively; unfiltered ones harm. This is the constructive realisation of the
map: a genuine external verifier plus a never-forget memory yields real, label-free, gradient-free self-improvement of a frozen model.

Full per-experiment detail, exact configurations, multi-seed statistics, and the code for every row
above live in \texttt{RESULTS\_SCOREBOARD.md} and the corresponding
\texttt{bio\_growth/closed\_form\_neat\_*.py} and \texttt{bio\_growth/pde\_*.py} scripts.

\paragraph{Independent robotic-system campaign (Gates 192--211).}
The KUKA transfer first produces a representation-positive result: the operator
map beats the published linear baseline but a compact MLP is more accurate.
Its initial transaction analysis paired adjacent blocks incorrectly. Gate 209's
input-only audit identifies repeated programs $(0,3),(1,4),(2,5)$; on these
pairs sparse improves every repeat and retains 99.23\% of pooled dense gain with
3.189$\times$ fewer labels, though one pair initially misses a frozen retention
clause. Gate 210 then corrects the same pairing error in source-only ridge
selection, changing the ridge from 100 to 0.01 and raising sparse retention to
97.45\% of dense gain with every pair above 89\%. All mechanism clauses pass.
The earlier cross-program harm statement is retracted and archived as a
protocol error; both corrections are post-exposure.
The subsequent real Nano-Drone protocol commits source and independent-repeat
receipts before confirmation and passes all twelve clauses with $2.500\times$
fewer labels. Frozen ablations attribute the gain to exponential actuator
memory and reject the broad cardinal expansion. A matched MPS MLP loses
accuracy but exposes a slow literal compiler; exact packed validation,
sufficient-statistic composition, and Gram-only admission then reproduce the
program in 0.935 s, $13.35\times$ faster than neural optimization. A final
direct/recursive control formally fails three-group dominance while isolating
large long-horizon attitude/rate gains. Finally, an adequate pinned ASIA
reproduction with matched sparse personalization beats the compact operator in
all four error groups, while the operator is 1,020$\times$ smaller and compiles
65.9$\times$ faster. Gates 200--208 then implement neural proposal followed by
exact operator consolidation. One-execution projection compresses but does not
transfer; typed local-cardinal recurrence repairs translation; and additive
pseudo-transition Grams from two executions produce the best compact cumulative
errors $1.646/5.368/2.624/15.957$. Generic/discovered interaction expansion and
recency forgetting fail their gates. Protocols, phase seals, outputs, and
negative results are retained in the Gate 192--211 Markdown records and
corresponding industrial-breakthrough scripts.
Gate 211 packages the resulting provenance rule into a typed evidence capsule:
all six true KUKA Gram compositions are exact, all twelve wrong-program
compositions are rejected, and verified symmetric-circulant normals invert
$55\times$ faster at 1,024 coordinates with $4,096\times$ storage reduction at
4,096 coordinates. The complete repository suite passes 46/46 tests.
Gate 212 then transfers the frozen API and an additive cardinal history model
to a previously unused measured Wiener--Hammerstein circuit. Its source-only
gate fails (40.220 mV versus 39.958 mV for the polynomial control), and the
official target therefore remains sealed. The committed failure motivates
operator-coordinate discovery rather than a denser cardinal grid.
Gate 213's generic rank-4 Jennrich decomposition then has no real
well-conditioned solution and produces no predictive score. Gate 214 instead
ties the shifted nonlinear coordinate through the disclosed physical
factorization; source RMS falls to 1.967 mV and every clause passes. Gate 215
commits the full-source states before opening the official target once. Its
input-only cardinal program reaches 1.569 mV versus 43.338 mV linear and 8.486
mV factorized cubic, with a 3,920-byte executable. The result is prospective
but remains above the published 0.241-mV deep-encoder frontier.

Gates 216--225 are post-target-exposure source-only diagnostics. Grid density,
longer FIR memory, nominal inverse-Chebyshev transfer, lifted stable ARX,
output-error refitting, and exact front/output pole variable projection fail to
cross the 0.90-mV band. Gates 226--228 construct the complete quadratic and
cubic stable-pole Hermite jets. Quadratic reaches 0.844451 mV; cubic regresses.
The overlap-streamed branch compiler reproduces quadratic RMS to
$1.11\times10^{-13}$ relative error at 584.8 MiB rather than 881.5 MiB peak.
Gate 229 admits one cubic atom only after gains repeat on both validation
halves, reaching 0.800018 mV. Gate 230 exposes exactly three independent
numerator tangent directions after the scaling mode is removed and admits one
at 0.596536 mV. Gate 231 admits a second at 0.454572 mV. Gate 232 rejects the
last: it improves aggregate RMS to 0.420820 mV and the later half by 14.1\% but
harms the earlier half by 0.81\%. All these refinements are source-only and do
not constitute a second prospective official-target claim.

Gates 235--237 admit three repeatable transfer-function curvature atoms and
reach 0.363553 mV; Gate 238 rejects all six remaining exposed atoms. Gate 239
then refits the frozen 16-branch topology on all source samples and commits the
complete causal artifact before target execution. It reaches 0.338035 mV on
source and 0.367337 mV on the official target, 76.59\% below Gate 215, with
15,200 static bytes and 24,136 bytes including carried state. The target had
already been opened, so this is retrospective transfer evidence rather than a
new prospective benchmark claim.

Gates 240--243 extend output memory to 96 taps and replace coefficient-space
compression with exact functional-Gram factorization. Rank 10 reaches 0.331004
mV, reproduces dense execution to $2.77\times10^{-16}$, and transfers after
full-source refit at 0.332097 mV. Gates 244--249 diagnose a colored residual but
reject recent-input tails, autonomous stable ARX, exponential-reproduction
edges, a second nonlinear chart, and recurrent exponential-channel memory.
Gate 250 replays the six previously rejected finite atoms under the changed
temporal metric; exactly one numerator-zero atom is now admitted on both
segments, reaching 0.288277 mV. Gate 251 transfers its full-source refit at
0.304471 mV, narrowly missing the frozen 0.300-mV clause. Gates 252--254 reject
cutoff retuning, higher-rank accuracy growth, and alternating edge--temporal
refits; float32 functional rank 17 is prediction-neutral and occupies 12,948
static bytes. Gates 255--258 test pure and hybrid cardinal cubic Hermite edges.
Their aggregate RMS reaches 0.272181/0.273758 mV, but both harm one of four
contiguous blocks, so zero-forget rejects them and closes circuit engineering.

Gate 259 seals the Bouc--Wen benchmark, three independently phased source
programs, and all code before fitting. A 1,624-byte physical Gram law improves
official input-only multisine/sweep RMS from 1.559830/1.197468 to
0.158032/0.124041 mm; all prospective clauses pass. A tensor-cardinal residual
is rejected at every nonzero admission step. Gate 260 differentiates the full
recurrent simulator into three output-trajectory tangents. Four exact Gram
cycles, admitted on two separate source programs, recover the disclosed
hysteresis coefficients within 0.006\% and reach source RMS
$4.04\times10^{-5}/4.42\times10^{-5}$ mm. Its source artifact and official
input-only trajectories are committed before retrospective scoring at
0.002212/0.005213 mm. The first Gate-260 run completed computation but failed
JSON serialization; the type-only repair and identical rerun are both recorded.

Gates 261--265 remove disclosed coefficients and then the equation pair. Gate
261 jointly fits six physical parameters. Gates 262--264 retain their negative
greedy, singleton, and local-pair receipts. Gate 265 enumerates all 55 pair
topologies, calibrates the best ten on program 0, selects on program 1, and
opens program 2 only once; it recovers the exact signed pair. Gates 266--267
retain two strict sensitivity-equivalence failures together with the complete
finite-difference convergence tables. Gate 268 performs a matched end-to-end
race from the same initializer and passes. Gate 269 records a faster
backtracking endpoint but fails its validation-count clause. Gate 270 emits
analytic atom gradients for the full dictionary and repeats the Gate-265
selection without loading its learned artifact. All charters, source hashes,
cycle receipts, nonfinite rejections, timings, and frozen programs are stored
with the corresponding scripts and JSON outputs.

Gate 271 seals four device files and a truth manifest in commit \texttt{71b6ad31}
before its fitting executable. It freezes each selected program before moving
to the next device, then opens third-program prefixes for routing and suffixes
for scoring. Gate 272 seals the fifth device in \texttt{b7634aa4}, fixes the adequacy
threshold from the earlier bank, executes abstention before fitting, and
records all pre/post routes and hashes. Gate 273 seals a continuous switched
trajectory in \texttt{edcb9c09} before its rolling detector. Its complete per-window
score/route trace, latencies, tail errors, model hashes, stream hash, runtime,
and resource checks are retained. Generator modules are never imported by
their corresponding fitting/detection executables.

Gates 274--276 freeze a device-0 wear ladder and two active-probe designs. The
passive detector crosses its adequacy threshold between 0.25 and 0.5\% wear;
generic E-optimal and wear-directional RMS probes formally fail, at 1.153x and
1.467x passive separation. Gate 277 freezes a nominal-only matched detector
before generating 100 nominal and 100 worn trials; all 200 decisions are
correct. Gate 278 commits one noisy adaptation response, updater, repaired
artifact, validation record, and evaluator in that order; untouched RMS falls
98.46\%.

Gates 279--281 retain three strict near-failures. Unknown six-dimensional
$\pm0.5\%$ updates all finish below 0.001 mm but one small parent improves only
70.46\%. A 24-device $\pm1\%$ panel achieves at least 92.4\% reduction but two
final scores exceed 0.003 mm. Same-observation relinearization on 12 $\pm2\%$
faults puts every final score below 0.000666 mm, but marginally worsens two
already-excellent first steps. Each charter, generator, observation, updater,
artifact set, validation set, evaluator, and outcome is a separate chronological
commit.

Gate 282 freezes a 5\% active-residual acceptance margin before its new
$\pm3\%$ panel. Sixteen observations generate step-one, step-two, and selected
artifacts; all are committed before broadband validation exists. The evidence-
only selector accepts 13 second steps. Every frozen clause passes, including no
unseen regression versus step one, 95.71\% minimum parent repair, 0.000830-mm
maximum RMS, bitwise parent replay, stable hashes, runtime, and memory. The
complete machine-readable chronology resides in the Gate 274--282 JSON outputs
and Markdown charters/reports.

Gate 283 freezes 24 anonymous local/structural cases and routes them before
labels open; the exact 12/12 split and all admitted validation checks pass.
Gates 284--288 retain successive exhaustive, shared-dictionary, top-two,
dual-probe, and triple-probe receipts, including each formal failure and
runtime. Gate 289 freezes a fourth maximin capsule, four responses, compiler,
artifacts, validation generator, outcomes, and evaluator in separate commits;
every clause passes, although only one case actually accepts a four-capsule
refinement.

Gate 290 freezes its charter and six anonymous devices before controller code.
The controller's six decisions and executables are committed before the hidden
three-drift/three-structural ordering and broadband records are generated. It
routes 6/6 and passes every transfer, resource, hash, topology, and evidence-
economy clause. The initial invocation stops before processing cases because of
a sealed-array key mismatch; the one-line repair is a separate commit.

Gate 291 commits two programs and two rollback receipts before labels reveal
that these are exactly the two in-dictionary and two out-of-dictionary cases.
It formally fails one known-law transfer ceiling. Gate 292 separately freezes
six cases and two mandatory full-capsule proposals per case. Four accept an
update, two retain the preceding state, and the gate fails the per-case update
and runtime clauses although all withheld accuracy clauses pass. The Markdown
charters/reports and JSON records preserve these negative outcomes without
threshold changes.

Gate 293 authorizes the exogenous-force operator before sealing four new cases.
Its compiler, four artifacts, validation generator, outcomes, and evaluator are
separate commits. Discovery and accuracy pass; the first evaluator formally
fails its runtime clause. Gates 294--295 retain exact but formally negative
value-only and vector-fused deployment attempts. Gate 296 freezes an inlined,
precomputed vector kernel before execution and passes exactness, scientific
score, hash, memory, absolute-time, and relative-speed clauses. None of these
implementation repairs retroactively changes the earlier gate labels.

Gate 297 freezes its charter and Cholesky-risk designer before generating 20
candidate waveforms; its selected probe and complete score table are committed
before any response exists. Gate 298 then freezes four new response sets, the
two-branch compiler, eight artifacts, the validation generator, broadband data,
and evaluator in chronological commits. Its selected branch wins all four
independent comparisons and every preregistered clause passes. The report notes
that this panel does not isolate VOI from minimum-eigenvalue design.

Gates 299--300 retain two outcome-free non-separation results: narrowed task risk
and damping-coordinate variance both choose the generic 35--50-Hz probe. Gate
301 freezes an enlarged transient language before evaluation and records the
first divergent design. Gate 302 then commits fresh evidence, two matched
branches, and both branches' hashes before generating two validation families;
its coefficient-identification result passes while the overall task gate fails.
Gate 303 freezes and selects a constrained hybrid before any new responses.
Gate 304 repeats the full evidence--compiler--artifact--validation chronology on
new systems and retains its mixed broadband success and formal transient/
coefficient failure. No threshold is revised after outcomes.

Gate 305 freezes a five-version minimax Gram design and retains its failure to
improve the nominal mixture. Gate 306 freezes full-transaction simulation and
records a nonfinite structural-initialization failure without interpreting its
NaN ranking. Gate 307 commits a one-change neutral-initialization wrapper before
replay; it chooses the generic endpoint but fails the frozen novelty clause.
Gate 308 freezes that acquisition rollback, a new panel, both matched branches,
their hashes, and only then two validation families. Its 19.92\% worst-risk
advantage narrowly misses the 20\% threshold, and the exact endpoint receipt
also fails. Both clauses remain unchanged in the report and JSON.

Gate 309 freezes an eight-system acquisition-only panel, literal generic
endpoint, 16 fixed-topology branch artifacts, and only then eight validation
pairs. All acquisition-policy clauses pass, while one marginal evidence residual
keeps the overall gate false. An initial validation invocation stops before
sealing because a four-entry frequency schedule is exhausted; the eight-entry
input-only repair is a separate commit. Gate 310 is explicitly retrospective:
its evidence-only router and one-branch escalation are frozen before replay, and
its passing receipt does not relabel Gate 309.

Gate 311 freezes an end-to-end two-level controller before eight new systems.
Its acquisition decision passes after artifacts are committed and validation is
generated, but the one-extra-step model contract fails on two branches. Gate 312
is explicitly retrospective: after validation exposure it demonstrates that a
further certificate-monotone step closes both branches, without relabeling Gate
311. Gate 313 then freezes the bounded two-step rule, generator, compiler,
validation generator, and evaluator. Adaptation evidence is sealed first; final
compiler artifacts and hashes are committed second; only then is the eight-case
validation panel generated, sealed, and opened. Four routed branches close in
one step and every precommitted end-to-end clause passes.

The first Gate-313 compiler invocation ends after its parent phase when its
execution session expires. Audit also identifies a dormant tuple/subtraction
typo in the escalation branch that had not run. The interruption record and
one-line repair are committed before a complete deterministic replay; evidence,
programs, thresholds, routing, budgets, and validation code remain unchanged.

Gate 314 freezes an external MuJoCo environment receipt and retains its unsafe
open-loop termination. Gate 315 freezes a feedback-filtered nominal selector,
then formally fails because the snapshot omits Gymnasium's cumulative time-limit
counter. Its output and corrected interpretation are committed. Gate 316 changes
only that snapshot coordinate; identical controller, candidates, perturbation,
seed, and thresholds then pass exact replay, observability, restoration, time,
and memory clauses. No learning result is inferred from this environment receipt.

Gate 317 next seals 2,048 adaptation-only transition pairs and freezes linear
and cardinal residual artifacts while validation does not exist. Its compiler
fails an independent-solve clause because the raw additive cardinal dictionary
has five partition-of-unity gauges. Gate 318 freezes a contrast projection and
replays the same evidence: the compiler passes every algebra, rank, size, time,
and memory clause, and its artifact is committed before any validation generator
is authorized.