Operator-spline theory: consolidated research manuscript
Historical source. Some claims in older records were subsequently corrected. The associated article states the adopted interpretation. This record preserves the original source alongside its rendered reading view.
Rendered archival TeX
This is an HTML reading rendition of the local TeX record. Mathematical notation is rendered with KaTeX; archived figures are included when their source assets are part of this collection.
Operator-Spline Neural Representations study how known continuous structure can be represented and manipulated through discrete coefficients. For admissible constant-coefficient operators, polynomial and exponential B-splines provide compact local generators, operator-null-space reproduction and analog-to-digital filtering identities. Hermite generators provide explicit value and derivative coordinates. This document develops those continuous–discrete connections and records the implementation and empirical conditions under which they are useful.
The central computational objects are continuous function, derivative and cross-basis Grams. Fixed generators and grids permit their reuse across coefficient updates; projection, differential energies and adjoints then become structured linear algebra. Periodic equal-spacing settings admit circulant or block-circulant realizations, whereas unequal-grid cross-Grams, finite-boundary operators and data-dependent statistics need their actual structure respected. The implementations address canonical endpoint handling, stable reciprocal-root factorizations, positive spectra, banded/bordered solves, exact polynomial-piece integration and local matrix/Horner execution. Grid refinement preserves a function only with the required subspace embedding; changing a feature family does not in general preserve sufficient statistics for unseen directions.
Recent extensions apply this calculus to compact causal physical models. For fixed spline directions and an unclipped regression model, a stable first-order actuator permits a joint constrained quadratic fit after absorbing its pole-dependent scale into static spline coefficients. For monotone fitted laws, explicit saturation also admits finitely many ordered data regions with constrained quadratic subproblems; batched quadratic lower bounds prune most solves while retaining numerical optimality checks. This is not a claim that the clipped objective is globally convex or that every fit meets a fixed real-time bound. Combining bounded-function products with stable exponential response sums yields an exact infinite-response inner product and a compact source-bank descriptor. This is exponential response calculus around a cardinal cubic model, not an exponential B-spline neural activation experiment or a new universal-policy theorem.
A thirty-pool temporal-task transfer experiment provides a further boundary: cardinal support planning solves 161/480 composed missions versus MLP 242/480, with sequential reuse matching both. Exact Bernstein specialization preserves audited development outputs while reducing planning time by 2.10–2.46×, but informative state acquisition and useful composition do not follow from correct local calculus. This experiment uses cubic fields, not the full exponential/Hermite physical-operator toolbox.
A measured CT extension combines exact rectangular field queries, finite-domain hierarchical product matrices, incremental likelihood information and full Gaussian block-design solves. Its strongest classical control answers 47 of 72 late questions without confident mistakes, but fails the fixed coverage target and gains no measurement saving from targeted acquisition. Exact-output memory compilation does not by itself establish useful uncertainty or a spline-specific inspection advantage.
The statistical and behavioral claims remain conditional. Generalized increments filter continuous innovations and need not be independent. Operator matching alone does not guarantee lower prediction risk or sparse support adaptation by ridge. Pooled fixed-feature statistics preserve a fitting objective, not every old prediction or private datum; protected supports and immutable archived programs provide different preservation contracts. Physical identification also does not imply useful closed-loop recovery: the accompanying controlled experiments include both gains and failed restoration or comparator criteria. The resulting foundation is an auditable set of representation, calculus and compilation mechanisms, with explicit boundaries on observability, model mismatch, numerical conditioning, memory and downstream performance.
tableofcontents
Introduction
Coordinate networks have become a standard tool for representing continuous signals and fields. A typical implicit neural representation (INR) maps coordinates to field values through a multilayer perceptron, fθ:x↦y, where the weights θ are optimized by stochastic gradient descent. SIRENs improve high-frequency representation by using sinusoidal activations, while PINNs add differential-equation residuals to the training objective.
The OSNR thesis is that this approach is structurally misaligned for physical systems governed by known operators. It is inspired by the operator-based spline signal-processing program of Unser, Blu, Vetterli, and collaborators [unser1993bspline1,unser1993bspline2,unser2005cardinal1,unser2005cardinal2,vetterli2002fri]. If a field is constrained by L{s}=r, then the representation should be built from the operator L itself. The role of learning or numerical inversion should be reduced to coefficient recovery in an already appropriate continuous function space.
OSNR therefore replaces black-box nonlinear layers by operator-matched spline dictionaries. The spline coefficients play the role of the latent representation. The architecture is not a generic neural network with a different activation; it is a continuous-domain inverse-problem engine presented in neural-representation form, aligned with continuous-domain inverse-problem representer theorems [gupta2018continuous,debarre2019hybrid,debarre2021composite]. The same principle also suggests a deterministic alternative to the trunk side of branch-trunk neural operator models such as DeepONet and MIONet [lu2021deeponet,jin2022mionet]: the coordinate-to-field map can be a spline synthesis operator with exact calculus rather than a learned MLP.
The empirical record tests this thesis rather than establishing universal superiority. S[sec:grown-topology] covers closed-form identification, grown-topology control and continual-learning comparisons; S[sec:matched-vs-sindy] compares operator-matched identification with SINDy and neural ODEs, including a negative blind-discovery control. S[sec:pde-vs-fno] studies PDE comparisons with neural operators, regime changes and partially known physics. Their measured gains depend on the information supplied, comparator, accuracy and computational accounting: knowing an operator family is not by itself a sufficient condition for an advantage. The later exact battery-response experiment, for example, admits a nearly equally fast training-free classical control; the real battery-aging fit loses to its matched polynomial comparator. S[sec:ssp-view] develops the sparse-stochastic-process connection under its stated assumptions, not a universal statistical dominance theorem. S[sec:rsi] records recursive-improvement experiments, while subsequent claim audits distinguish pooled fitting statistics, protected function supports, immutable archives and retained task performance. None alone establishes unrestricted self-improvement without forgetting or drift. The constructive result is a tested representation and calculus toolbox; independent-task benefit against strong controls remains the capability test.
Mathematical Background
Cardinal spline representation
The classical cardinal spline model represents a continuous signal as [unser1993bspline1,unser1993bspline2] s(t)=k∈Z∑c[k]β(t−k), where β is a spline generator and c[k] are discrete coefficients. The essential point is that a continuous function space is controlled by a discrete sequence. This makes exact digital processing of continuous objects possible, provided the generator is stable and the correct coefficient-domain filters are used.
Linear differential operators and null spaces
Let L be a constant-coefficient differential operator with characteristic roots or poles α=(α1,…,αN). The null space is NL={s:L{s}=0}=span{eαmt}m=1N with polynomial factors included for repeated roots. For the Helmholtz operator L=dx2d2+k2, the poles are α=±jk, and the null space consists of sinusoidal modes.
Cardinal exponential splines
Cardinal exponential splines are compactly supported spline generators matched to α [unser2005cardinal1]. Their Fourier-domain form is βα(ω)=m=1∏Njω−αm1−eαm−jω. Integer shifts of βα generate a stable spline space under appropriate Riesz conditions. The basis reproduces exponential polynomials and supports exact operator calculus in the coefficient domain.
Given an operator L with pole vector α and physical knot spacing T, an OSNR dictionary is a matrix of samples An,k=βα(Txn−k), augmented when necessary by explicit null-space columns. The continuous field is represented by s(xn)≈(Ac)n.
Exponential Hermite splines
Cardinal E-splines attach the operator to a scalar coefficient stream. For higher-order boundary value problems, OSNR requires a vector-valued cardinal system whose coefficients store not only function values but also derivatives. The second-order exponential-polynomial Hermite generator of Schmitter, Badoual, Uhlmann, Fageot, and Unser [schmitter2016hermite] provides exactly this structure.
Let Φ(t)=[ϕ0(t)ϕ1(t)ϕ2(t)]⊤. The Hermite spline expansion is f(t)=k∈Z∑(c0[k]ϕ0(t−k)+c1[k]ϕ1(t−k)+c2[k]ϕ2(t−k)), with coefficient vectors c[k]=c0[k]c1[k]c2[k]=f(k)f′(k)f′′(k) for exactly interpolated data. The interpolation conditions are ϕp(r)(k)=δprδk,p,r∈{0,1,2},k∈Z. On the interval [0,1], each channel has the exponential-polynomial form ϕp(t)=Ap+Bpt+Cpt2+Dpt3+Epejω0t+Fpe−jω0t,ω0=M2π. The constants Ap,…,Fp are determined by the six Hermite endpoint constraints ϕp(r)(0)=δpr,ϕp(r)(1)=0,r=0,1,2. The negative branch is fixed by the Hermite parity ϕ0(−t)=ϕ0(t),ϕ1(−t)=−ϕ1(t),ϕ2(−t)=ϕ2(t), so the generator is compactly supported on [−1,1]. This construction is C2, reproduces polynomials up to cubic degree, and reproduces the trigonometric modes sin(ω0t) and cos(ω0t) through the coefficient samples of the function and its first two derivatives.
For OSNR Tier 3, the crucial change is semantic: Dirichlet, Neumann, and curvature boundary data become direct coefficient assignments. A clamped boundary at knot kb, for example, is imposed by setting c[kb]=000, rather than by adding a soft loss term. The same vector-valued structure also produces the block-circulant Hermite Gram system derived in Appendix [app:hermite-block-gram].
OSNR Architecture
Master core apparatus
The current OSNR research code separates into three production tiers. Each tier has a different admissible function space, solver topology, and verification target.
Tier 1 targets smooth physical fields that lie in, or close to, the null space of a known linear operator. Examples include Helmholtz waves, damped oscillators, and linear constant-coefficient PDE components. The solver pipeline is:
identify L and its poles α;
construct calibrated E-spline or null-space dictionaries;
recover coefficients through a stable linear solve or circulant FFT inversion;
compute differential quantities through operator identities rather than autograd.
Tier 2: adaptive sparse continua
Tier 2 targets composite fields with sparse discontinuities or shocks: s=ssmooth+ssparse. Uniform grids are not sufficient for non-bandlimited discontinuities at sub-grid coordinates. The sparse tier must first identify finite-rate innovations, then adapt the dictionary to those coordinates:
estimate innovation locations using FRI or matrix-pencil methods;
snap sparse knots to the recovered coordinates;
solve a cross-Gram-coupled sparse-plus-smooth inverse problem;
debias with scale-invariant ridge stabilization.
Tier 3: higher-order neural operators
Tier 3 targets families of PDE solutions rather than a single fitted field. Neural operators such as DeepONet and MIONet learn maps between function spaces by pairing branch networks, which encode input functions or boundary data, with trunk networks, which encode query coordinates [lu2021deeponet,jin2022mionet]. Spline-PINN shows a related but distinct path: a CNN predicts Hermite spline coefficients, and a continuous Hermite spline layer evaluates PDE residuals without finite-difference losses [wandel2022splinepinn]. OSNR adopts the continuous Hermite idea but removes the black-box coordinate trunk where the governing operator and boundary calculus are known.
Let Φ(t)=[ϕ0(t)ϕ1(t)ϕ2(t)]⊤ be a second-order Hermite generator compactly supported on [−1,1]. The generalized higher-order Hermite construction of Schmitter, Badoual, Uhlmann, Fageot, and Unser stores value, slope, and curvature data at each knot and interpolates all three channels exactly [schmitter2016hermite]: ϕ0(r)(k)=δr0δk,ϕ1(r)(k)=δr1δk,ϕ2(r)(k)=δr2δk,r=0,1,2. The synthesized field is f(t)=k∈Z∑c0[k]ϕ0(t−k)+c1[k]ϕ1(t−k)+c2[k]ϕ2(t−k), where c[k]=c0[k]c1[k]c2[k]=f(k)f′(k)f′′(k) for exactly interpolated data. The Hermite generator in [schmitter2016hermite] is piecewise polynomial-exponential, C2, compactly supported, and reproduces cubic polynomials as well as trigonometric functions. This is the missing boundary mechanism for a neural-operator OSNR tier: Dirichlet, Neumann, and curvature constraints are coefficient assignments, not soft loss penalties.
For a PDE solution operator S:u↦v, the branch-side network or analytic encoder should output Hermite coefficient tensors {c[k;u]}k∈Z, while the trunk side is replaced by deterministic Hermite synthesis. In multidimensional domains, tensor products of one-dimensional Hermite generators give mixed value/derivative channels, matching the construction used by Spline-PINN for continuous PDE residuals [wandel2022splinepinn]. Unlike Spline-PINN, OSNR evaluates continuous energies and cross-correlations in coefficient space through the block-Gram calculus derived in Appendix [app:hermite-block-gram], avoiding Monte Carlo residual quadrature whenever the operator and boundary model admit exact inner products.
Continuous-Discrete Calibration
The most important implementation condition is the cardinal coordinate map, which preserves the shift-invariant structure required by spline filtering calculus [unser2005cardinal1,unser2005cardinal2]. vk(x)=Tx−k. Here T is the physical knot spacing. For a domain [0,D] and a compact generator of support order q, a stable finite dictionary uses T=M−qD. Then vk(x+T)=vk(x)+1, so one physical knot step maps to one cardinal interval.
If the map is implemented as vk(x)=x+bk with biases distributed over a fixed interval while M changes, increasing M decreases the relative shift between adjacent columns without changing physical support. The resulting dictionary columns become nearly collinear, and the Gram matrix A⊤A develops near-zero eigenvalues.
Let adjacent atoms be sampled as ϕ(x+bk) and ϕ(x+bk+1). If bk+1−bk=O(1/M) while the support width of ϕ is fixed, then a first-order expansion gives ϕ(x+bk+1)=ϕ(x+bk)+O(1/M). Thus adjacent columns converge to each other as M increases. The Gram matrix approaches rank deficiency and the pseudoinverse amplifies roundoff along small singular directions.
This explains why a derivative residual can be zero while reconstruction fails. If the derivative dictionary is defined algebraically by Ad2=−k2A, then the PDE residual Ad2c+k2Ac is identically zero regardless of whether A is a well-conditioned reconstruction basis.
2D tensor-product Hermite expansion
The 2D Tier 3 engine is obtained by tensorizing the one-dimensional second-order Hermite streams. Let hix(x),hjy(y),i,j∈{0,1,2}, denote the value, slope, and curvature Hermite generators along the two axes. The tensor-product basis functions are hi,j(x,y)=hix(x)hjy(y),i,j∈{0,1,2}. For a grid node (k,ℓ), OSNR stores a nine-stream local state ck,ℓ=[f∂xf∂xxf∂yf∂xyf∂xxyf∂yyf∂xyyf∂xxyyf](k,ℓ)⊤. The synthesized field is f(x,y)=k,ℓ∑i=0∑2j=0∑2ci,j[k,ℓ]hix(Txx−k)hjy(Tyy−ℓ). This formula is the tensor-product analogue of the Schmitter et al. Hermite generator and the multidimensional counterpart of the periodic inner-product calculus of Badoual, Schmitter, and Unser.
The practical consequence is that the spatial calculus of 2D physical fields is a forward coefficient operation. In the stream-function formulation for incompressible flow, vx=∂yaz,vy=−∂xaz, and therefore ∇⋅v=∂x∂yaz−∂y∂xaz=0 up to the commutation error of the discrete finite-difference stencil. The nonlinear transport term is evaluated as (v⋅∇)v=[vx∂xvx+vy∂yvxvx∂xvy+vy∂yvy], and viscous diffusion as Δv=[∂xxvx+∂yyvx∂xxvy+∂yyvy]. All operators appearing in these expressions are evaluated by shift-invariant finite-difference ladders on the Hermite coefficient streams during the forward pass. No backward-mode automatic differentiation tape is constructed; the measured PyTorch autograd graph allocation in all Tier 3 validations is therefore 0.00 bytes.
Autograd-Free Differential Calculus
For an exponential spline with pole vector α, applying a first-order operator (D−αm) reduces the order of the spline [unser2005cardinal1,delgadogonzalo2012exponential]: (D−αm)βα(t)=βα∖αm(t)−eαmβα∖αm(t−1). This is the finite-difference ladder. Differential fields can be evaluated by filtering coefficients or by applying deterministic lower-order dictionary maps. For the Helmholtz null space, dx2d2s(x)=−k2s(x) inside the smooth spans, with boundary innovations handled separately.
This removes the need for backward-mode automatic differentiation in Tier 1 PDE residual evaluation. The differential operator is encoded in the basis and coefficient algebra.
Circulant and FFT Solvers
When the dictionary is shift-invariant and periodized, the Gram matrix is circulant. This is the coefficient-domain form of the exact periodic inner-product calculus for spline curves and functions [badoual2016inner,badoual2018periodic]: G=g0g1⋮gM−1gM−1g0⋮gM−2⋯⋯⋱⋯g1g2⋮g0. Such matrices are diagonalized by the discrete Fourier transform: G=F−1ΛF. Solving Gc=b reduces to c[ℓ]=λℓb[ℓ]. This converts O(M3) dense solves into O(MlogM) FFT operations.
Tomographic Radiance Fields as Spline Inverse Problems
The failed direct ray-kernel DL3DV experiment clarifies an important modeling boundary. A nontrivial view-synthesis scene is not naturally a smooth map from ray origin and direction to RGB. The physically shared quantity is a latent field in space, observed through line or ray measurements. The spline tomography literature gives the appropriate replacement model: represent the unknown continuous field by shifted basis functions, push those basis functions through the forward projector, and solve the coefficient inverse problem with fast adjoint and normal operators [nilchian2013fast,mccann2016fast,donati2018multiscale,haouchat2025generalized]. Jin et al. make the complementary point that inverse problems whose normal operators are convolutional admit physics-aware direct inversions before any learned artifact-removal stage [jin2017deep]; in OSNR, the direct inverse is the primary object, and any neural residual must remain secondary.
For a first linearized density or opacity stage, write σ(x)=k∑a[k]φσ(x−Λk),g=Ha+η, where H samples line or ray integrals of the spline density field. The regularized inverse update is a⋆=argamin21∥W1/2(Ha−g)∥22+λR(a), with W encoding ray confidence or frequency reliability, as in weighted phase-retrieval formulations [bostan2016variational]. For quadratic R(a)=∥La∥22, the normal equation is (H⊤WH+λL⊤L)a=H⊤Wg. The computational opportunity is the McCann–Donati normal-operator identity. For shift-invariant basis functions and locally stationary projection blocks, H⊤H acts as a discrete convolution: (H⊤Ha)[k]=(a∗r)[k], where r is the sampled autocorrelation of the projected basis function. Thus the expensive repeated normal-operator application inside conjugate gradients or ADMM becomes a Fourier-domain multiplication. Multiscale basis functions then provide a controlled coarse-to-fine path that is robust to pose and angular uncertainty [donati2018multiscale].
Fast forward projection and sparse acquisition variants can be imported from the same lineage: Arcadu et al. use Fourier regridding with minimal oversampling for efficient forward projectors, while Donati et al. show how randomized STEM sampling can be coupled to regularized tomographic recovery [arcadu2016forward,donati2017compressed]. This does not make full NeRF rendering linear. The volume-rendering equation contains transmittance C(r)=∫T(t)σ(r(t))c(r(t),d)dt,T(t)=exp(−∫0tσ(r(s))ds), so exact RGB fitting remains nonlinear in σ. The OSNR route is therefore staged: first recover coarse density/support with a tomographic spline inverse solve; then refine the density multiscale; then solve color/radiance coefficients on the recovered support; and finally apply visibility-weighted nonlinear corrections. Positivity, support constraints, total-variation, Hessian, and sparse-innovation priors enter naturally through the constrained ADMM machinery developed for spline tomography [nilchian2013constrained,nilchian2015spline].
The controlled runner apps\_industrial\_breakthrough/spline\_tomographic\_radiance\_solver.py validates only the linearized operator claim. It builds a periodic synthetic density field, samples 48 discrete projection directions, constructs H⊤H explicitly once from a delta impulse, and then replaces all subsequent normal-operator applications by FFT convolution. On a 96×96 field, the convolutional normal operator matches explicit H⊤H with relative error 2.4462×10−7. One explicit normal-operator application costs 2.4956 ms, while the FFT version costs 0.0969 ms. Solving the same ridge-regularized inverse problem by conjugate gradients takes 188.1746 ms with explicit normals and 4.9666 ms with FFT normals, yielding a reconstruction PSNR of 22.3578 dB under noisy sparse projections. This is not yet a NeRF result; it is the isolated mathematical validation that the spline-tomographic normal operator can be diagonalized as the literature predicts.
The next controlled runner, apps\_industrial\_breakthrough/spline\_ray\_operator\_validation.py, validates the more fundamental Haouchat-style requirement: the forward ray operator and its adjoint must be matched before any real-scene radiance experiment is meaningful. On a 36×36 coefficient grid with 2688 parallel rays, the script constructs two explicit small operators for auditability: a pixel basis and a quadratic tensor-product spline basis. The spline data are generated by the spline operator itself, and both models solve the same noisy inverse problem with conjugate gradients. The adjoint identity ⟨Hc,p⟩=⟨c,H⊤p⟩ holds to 1.5672×10−15 relative error for the spline operator and 2.7427×10−16 for the pixel operator. At the same coefficient count, the spline inverse reconstructs the continuous rendered target at 54.0869 dB, while the pixel model reaches only 27.0125 dB. This is an intentionally controlled operator test: it proves that the coefficient-domain ray basis and adjoint are now correctly formulated, not that the full DL3DV visibility problem is solved.
The basis choice itself is not incidental. Following the exponential-spline construction of Delgado-Gonzalo, Thevenaz, and Unser [delgadogonzalo2012exponential], a one-dimensional cardinal exponential B-spline associated with poles α=(α1,…,αN) has Fourier-domain form βα(ω)=m=1∏Njω−αm1−exp(αm−jω). The two-dimensional smooth tier then uses the tensor-product generator φαx,αy(x,y)=βαx(x)βαy(y), so the basis can reproduce the local modes implied by the operator rather than merely interpolate samples. The controlled sweep apps\_industrial\_breakthrough/exponential\_spline\_basis\_sweep.py tests this precision-first hypothesis on the same 1920 rays and 784 coefficients while changing only the tensor-product basis. The data are generated by a damped-harmonic exponential spline with poles (−λ,−λ+jω,−λ−jω), and all candidate operators reuse the identical ray geometry. The matched damped-harmonic basis reaches 56.3837 dB, compared with 52.1420 dB for a monotone exponential-decay basis, 50.0620 dB for an undamped harmonic basis, 49.9084 dB for a quadratic polynomial spline, and 28.7579 dB for a pixel box basis. This isolates the important design rule for the next radiance-field stage: the highest precision should come from tensor-product operator splines whose poles are matched to the expected local dynamics, with sparsity and compression applied after the operator basis is correct.
The follow-up runner apps\_industrial\_breakthrough/exponential\_spline\_pole\_sweep.py turns this from a hand-picked basis comparison into a deterministic pole-selection problem. It fixes the target field, ray geometry, coefficient count, noise level, ridge parameter, and conjugate-gradient budget, then sweeps damped-harmonic exponential splines of orders 2, 3, and 4 over λ∈{0.18,0.28,0.35,0.42,0.56,0.72} and periods {8,10,12,16}. The target is generated by the order-3 pole set (−0.42,−0.42+j2π/10,−0.42−j2π/10). The pole sweep correctly ranks that matched operator first at 55.2077 dB on 1280 rays and 576 coefficients. The best order-4 candidate reaches 52.4859 dB, and the best order-2 candidate reaches only 42.1900 dB. This is the first automated evidence that pole selection is a meaningful OSNR model-selection axis: increasing support/order blindly does not dominate, while matching the operator poles controls reconstruction precision.
The support/regularity runner apps\_industrial\_breakthrough/exponential\_spline\_support\_regularization\_sweep.py then isolates the opposite regime: a field generated by the shortest first-order Green-matched spline with pole α=−0.42. The one-pole basis has support length 1 and no continuity guarantee, but it exactly matches the local Green mode. It reaches 65.7918 dB with operator density 0.0362 and an average of 20.86 active coefficients per ray. The smooth order-3 mixed basis [0,0,α] reaches only 27.1299 dB and has density 0.1084 with 62.44 active coefficients per ray. This confirms a second design rule: if the modeled object is a Green response or sparse innovation, the shortest matched spline can be both more accurate and more localized than a smoother high-order basis. Regularity should be introduced because the signal class requires it, not by default.
Finally, apps\_industrial\_breakthrough/exponential\_spline\_basis\_selection\_map.py evaluates the full pole-multiset selection problem. It tests seven target regimes against fourteen candidate bases: first-order Green response, repeated real poles, two distinct real poles, zero-augmented smooth operators, damped oscillators, pure polynomial splines, and support-4 mixed bases. Across all regimes, the exact pole multiset ranks first. This is the strongest evidence so far that OSNR basis design should be formulated as operator pole selection rather than degree selection. Order controls support and regularity; pole multiplicity and location control the reproduced null-space modes. Higher order improves asymptotic approximation power for smooth functions, but it is not a substitute for matching the operator that generated the signal.
The final synthetic step, apps\_industrial\_breakthrough/exponential\_spline\_operator\_inference.py, removes access to rendered target PSNR during model selection. Each candidate basis is fitted on 75% of the rays and scored on held-out rays using a normalized validation residual plus small support and density penalties. This exposes a practical distinction between generative pole matching and predictive operator selection. Under sparse-ray training, the oracle PSNR basis is sometimes a smoother support-4 model rather than the exact generating basis, because the added regularity improves interpolation across unobserved rays. The validation score still selects the oracle basis in four of seven regimes and stays within 0.2039 dB of oracle in all cases, with mean PSNR loss 0.0465 dB. Thus the operational rule becomes: use the differential operator poles as the first prior, then choose among nearby pole augmentations by held-out measurement prediction rather than training residual.
We then stress-test the same idea in a mixed local-operator field with apps\_industrial\_breakthrough/exponential\_spline\_local\_operator\_adaptation.py. The synthetic field assigns different pole multisets to different spatial regions and compares a single global basis, a held-out-ray local block selector, and an oracle local label map. The oracle local solve reaches 33.5291 dB, while the best global held-out model reaches 28.6559 dB. This proves that local operator adaptation has substantial headroom. However, a blind ray-only greedy block selector reaches only 28.8074 dB with 25% block-label accuracy. Global line measurements make small independent block labels weakly identifiable unless the selection objective includes stronger spatial priors, localized measurements, or joint segmentation/coefficient optimization.
The sparse-innovation remedy is tested in apps\_industrial\_breakthrough/exponential\_spline\_fri\_region\_adaptation.py. Instead of selecting arbitrary blocks, the method first reconstructs a global proxy, computes derivative-energy profiles, extracts sparse FRI-style transition proposals, expands them into a small boundary lattice, and then jointly scores region geometry and region pole choices on held-out rays. The raw derivative peaks locate approximate boundaries at x=(−1.4149,1.4149) and y=3.0319; held-out refinement moves them to x=(−3.4149,3.4149) and y=4.0319, yielding 100% region-label agreement at the coefficient-grid resolution. The resulting FRI-region adaptive model reaches 34.2000 dB, outperforming both the global model and the nominal oracle-region pole assignment. This is the first successful local adaptation mechanism: sparse innovation proposals supply the missing spatial prior that held-out global rays alone could not provide.
The robustness sweep apps\_industrial\_breakthrough/exponential\_spline\_fri\_region\_stress\_sweep.py repeats the experiment over three projection-angle budgets and three additive noise levels. To keep the sweep diagnostic rather than combinatorial, it uses a compact region-basis candidate set containing the physical oracle family, the single-case selected family, a smoother support-4 family, and the best global family; the exhaustive 54 region-basis search remains available as an optional mode. Across all nine stress cases, the FRI-region model improves over the best global held-out basis. The gain increases with measurement density, from a mean +1.1667 dB at 12 angles to +5.0277 dB at 24 angles, while the recovered region accuracy rises from 89.50% to 100.00%. This confirms that the sparse-innovation step is not a one-off artifact: as the inverse problem receives enough projections to identify the transition set, local operator adaptation becomes reliably beneficial.
We then deliberately break the axis-aligned assumption with apps\_industrial\_breakthrough/exponential\_spline\_fri\_curve\_adaptation.py. The target field is generated by three smooth transition curves: two vertical pole-boundary curves and one top-interface curve. A row/column FRI-style derivative tracker fits low-order curve proposals from the global proxy and refines them by held-out rays. This harder test exposes the current bottleneck. The global held-out basis reaches 32.2114 dB, while the true curved local operator assignment reaches 41.5032 dB, proving that curved local operators have large headroom. However, the detected curved selector reaches only 31.1969 dB despite 91.00% region-label agreement and boundary RMSE 1.4606. The failure is not the absence of local operator advantage; it is the scoring layer. Under curved imperfect labels, the held-out ray residual prefers smoother surrogate pole assignments rather than the physical local pole map. The next algorithmic step is therefore joint geometry–basis–coefficient refinement, or a region-contrastive validation score that prevents the local operator assignment from collapsing to a globally smooth surrogate.
The joint refinement runner apps\_industrial\_breakthrough/exponential\_spline\_fri\_curve\_joint\_refinement.py implements this next correction. It expands the curve-offset lattice near the best FRI proposal, fits coefficients for each candidate local operator assignment, and augments the held-out residual with an edge-consistency contrast term that rewards reconstructions whose gradient energy concentrates on the proposed sparse transition curves. This converts the curved selector from a negative result into a partial recovery: the selected joint model reaches 34.7216 dB, a +2.5102 dB gain over the global held-out basis. The best candidate present in the searched family reaches 35.8310 dB, while the true-curve oracle remains at 41.5032 dB. Thus the scoring fix is directionally correct but incomplete. The remaining gap now separates two effects: curve localization error and the limited pole-assignment candidate family. This is the cleanest current target for further algorithmic work.
Finally, apps\_industrial\_breakthrough/exponential\_spline\_geometry\_reproduction.py isolates the more geometric point raised by the exponential-spline curve and surface literature [delgadogonzalo2012exponential]. The target is a closed harmonic curve with modes up to order three, represented from only twelve parameter samples and eight control points. The matched compact harmonic E-spline uses the pole set {0,±j2π/M,±j4π/M,±j6π/M} and reaches dense curve RMSE 5.4457×10−6. With the same number of control points and the same samples, a generic cubic polynomial spline reaches only 1.0256×10−2 RMSE, and a piecewise-linear polygon reaches 3.8693×10−2 RMSE. This is a small but important result: if the expected geometry is known to be harmonic, elliptic, spherical, cylindrical, or otherwise parametrizable by a known exponential-polynomial family, then OSNR should place that family directly in the geometric span rather than recover it indirectly through a generic volumetric grid. For NeRF-like scenes this suggests a patch-based route: segment or initialize object surfaces with geometry-reproducing parametric E-splines, then fit texture/radiance on those surfaces. For PDE domains it suggests an even cleaner route: represent both boundary geometry and the solution field in operator-matched spline spaces.
The follow-up optimizer apps\_industrial\_breakthrough/eggroll\_spline\_shape\_optimizer.py tests whether this geometric advantage can be used when the target shape is not known in closed form. Motivated by Schmitter and Unser's continuous-domain shape projectors and functional PCA construction [schmitter2018landmark], the experiment represents a closed spline curve by 16 control points but restricts the learned geometric search to an 8-dimensional continuous shape subspace. This is the analogue of replacing an arbitrary coordinate-field parameter vector by a learned spline-shape chart. The stochastic search component is motivated by the EGGROLL low-rank evolution-strategy result [sarkar2026eggroll]: rank-one perturbations can be evaluated as hardware-friendly low-rank updates, but the experiment separates this hardware trick from the geometric prior itself.
The result is deliberately diagnostic. Full Gaussian ES over all 32 control coordinates reaches dense RMSE 2.9097×10−2 after 67,200 forward evaluations, while rank-one EGGROLL-style perturbations applied directly to the raw control matrix reach 3.3517×10−2. Thus low-rank noise alone does not solve geometry discovery. When the same evaluation budget is spent inside the Schmitter-style spline subspace, the error drops to 6.9268×10−3; rank-one EGGROLL perturbations inside that subspace reach a comparable 7.4862×10−3. The algebraic continuous-subspace projection oracle reaches 3.5689×10−4 in 0.293 ms, exposing the remaining optimization gap. The practical implication is precise: the promising route is not blind evolution over arbitrary OSNR coefficients, but variable projection. Use stochastic low-rank search only for nonlinear geometry, visibility, and knot variables; solve the linear radiance or texture coefficients algebraically once a candidate geometry is proposed.
The next controlled runner, apps\_industrial\_breakthrough/eggroll\_adjoint\_variable\_projection.py, implements that variable-projection step explicitly. A one-dimensional spline boundary partitions a 40×40 radiance field into two continuous regions. For each candidate boundary, the code constructs the ray operator H(φ) by projecting masked smooth atoms, eliminates the linear radiance coefficients by the adjoint normal equation c⋆(φ)=(H(φ)⊤H(φ)+λI)−1H(φ)⊤y, and scores the resulting field on acquisition directions not used in the coefficient solve. This is the minimal inverse-problem analogue of a NeRF geometry/radiance separation: nonlinear geometry is searched, while linear radiance is solved in closed form.
The experiment clarifies both the opportunity and the bottleneck. The mean-geometry variable-projection baseline reaches 16.8968 dB on held-out projections. Full Gaussian ES over raw boundary controls improves to 18.9082 dB, rank-one EGGROLL over raw controls reaches 19.9034 dB, and rank-one EGGROLL in the six-dimensional spline subspace reaches 20.1043 dB. However, the true-geometry variable-projection oracle reaches 33.1874 dB with a field RMSE of 3.3521×10−2. Thus the adjoint variable-projection mechanism is working, but stochastic boundary discovery remains underidentified from the current projection residual alone. This is an important negative constraint for the NeRF/SIREN campaign: the next improvement must add stronger geometry evidence, such as FRI edge measurements, silhouette consistency, epipolar visibility constraints, or a learned continuous shape prior. More low-rank perturbation budget alone is unlikely to close the oracle gap.
The Haouchat-matched follow-up apps\_industrial\_breakthrough/haouchat\_matched\_variable\_projection.py performs that correction. Instead of using a primitive projection mask, it builds the inner ray operator from the same quadrature-evaluated tensor-product spline basis used in the 52–56 dB matched-ray experiments. The target data are generated by a damped-harmonic exponential spline ray operator. Candidate geometries still define a boundary-dependent masked coefficient dictionary, but the forward map is now Hφ=HβLΦ(φ), where HβL is the Haouchat-style ray projector for the selected tensor-product basis and Φ(φ) applies the boundary-dependent coefficient atoms. The stochastic outer loop is also given a weak FRI-like edge observation of the boundary, so the score combines held-in projection residual and edge consistency. This is the first experiment in this line that combines all three ingredients: matched spline rays, adjoint variable projection, and sparse boundary evidence.
The result changes the interpretation sharply. With the matched damped-harmonic operator, the true-geometry projection oracle reaches 96.9888 dB on held-out rays, confirming that the inner ray/inverse model itself is not the limiting factor. The edge-aware rank-one spline-subspace EGGROLL search reaches 42.2339 dB, up from the mean-geometry baseline of 33.3709 dB, with boundary RMSE reduced from 9.5169×10−2 to 2.7066×10−2. The exponential-decay candidate also benefits from edge evidence, improving from 40.6125 dB without the edge term to 41.0644 dB with it. This supports the current thesis: OSNR does not need a dense NeRF-style MLP to represent the radiance once the operator is matched; the hard remaining problem is physically constrained geometry and visibility discovery.
The follow-up apps\_industrial\_breakthrough/haouchat\_fri\_edge\_variable\_projection.py removes the remaining artificial part of the edge-aware score. Instead of injecting a boundary hint directly from the true geometry, it forms a measurement-derived FRI proxy: training rays are backprojected through the matched adjoint, row-wise derivatives of the normalized adjoint image are localized, and the resulting peak track is smoothed into a candidate boundary. This is not a complete multidimensional FRI surface solver, but it is a measurement-only sparse-transition proposal. In the default run, the synthetic edge hint has boundary RMSE 1.3676×10−2, while the adjoint-derived edge track has RMSE 1.9117×10−2.
Using this measurement-derived edge evidence, the matched damped-harmonic model reaches 41.5632 dB on held-out rays, compared with 33.3709 dB for mean geometry and 96.9888 dB for the true-geometry oracle. The cleaner synthetic edge hint reaches 45.9174 dB in the same runner. The gap between 41.56 dB and 45.92 dB is useful: it quantifies the price of deriving geometry evidence from measurements rather than providing it externally. The experiment therefore validates the direction without hiding the remaining work. The next real-scene version should replace row-wise adjoint peaks by multi-view epipolar FRI proposals and visibility-aware surface clustering.
The next controlled refinement, apps\_industrial\_breakthrough/haouchat\_multiray\_fri\_edge\_projection.py, replaces the row-wise adjoint peak estimate by a multi-ray residual selection loop. The adjoint boundary is used only as initialization. Each boundary control point is perturbed over a small local offset lattice, and candidates are scored by the held-in variable-projection residual plus curvature and anchor penalties: S(φ)=∥HβLΦ(φ)c⋆(φ)−yscore∥22+η∥Δ2φ∥22+ρ∥φ−φadj∥22. This converts the crude measurement edge into a ray-consistent FRI boundary proposal without accessing the hidden target geometry.
The improvement is large. The row/adjoint edge has boundary RMSE 1.9117×10−2; the multi-ray refined edge has RMSE 1.4848×10−3, better than the noisy synthetic edge control (1.3676×10−2). With the matched damped-harmonic operator, direct variable projection on the multi-ray FRI boundary reaches 66.1488 dB on held-out rays and field RMSE 2.5205×10−3. The true-geometry oracle remains 96.9888 dB, so the experiment is still controlled rather than a final SOTA benchmark, but it shows that the geometry bottleneck can be attacked algebraically by ray-consistent sparse innovation refinement. Interestingly, applying the stochastic EGGROLL search after this refined geometry is worse (40.3169 dB), so the current best path is not more random search but better deterministic boundary proposal.
The robustness sweep apps\_industrial\_breakthrough/haouchat\_multiray\_fri\_robustness\_sweep.py repeats the deterministic part of the experiment across two boundary variants, projection-angle counts {12,18,24}, and additive noise levels {0,3×10−4,10−3}. The mean-geometry baseline averages 32.5922 dB over the 18 cases; direct row/adjoint edge projection averages 38.0604 dB; multi-ray FRI variable projection averages 64.0770 dB; and the true-geometry oracle averages 97.9883 dB. The multi-ray method remains above 61.76 dB in every tested case. This confirms that the 66.15 dB result is not a one-off numerical accident, but a stable consequence of selecting boundary offsets by held-in ray residuals.
Quantity
Explicit normal
FFT convolution normal
Single H⊤H application
2.4956 ms
0.0969 ms
CG inverse solve
188.1746 ms
4.9666 ms
Normal-operator relative error
multicolumn(2)c2.4462×10−7
Reconstruction PSNR
multicolumn(2)c22.3578 dB
Controlled synthetic validation of the spline-tomographic normal-operator identity. The experiment isolates the linear inverse-problem component needed before returning to nonlinear radiance-field rendering.
Quantity
Pixel basis
Quadratic spline basis
Matched-adjoint relative error
2.7427×10−16
1.5672×10−15
CG inverse solve
19.5137 ms
21.9203 ms
Coefficient RMSE
4.1936×10−2
7.5915×10−3
Rendered reconstruction PSNR
27.0125 dB
54.0869 dB
Controlled validation of the matched spline ray operator Hφ and adjoint Hφ⊤ on 2688 rays and 1296 coefficients. The result establishes the correct operator foundation before porting the method to DL3DV camera rays with visibility weights.
Basis profile
PSNR
Adjoint error
CG solve
Coefficient RMSE
Pixel box
28.7579 dB
1.8997×10−16
4.8182 ms
1.8724×10−1
Quadratic spline
49.9084 dB
3.0959×10−15
4.1478 ms
1.8928×10−1
Exponential decay
52.1420 dB
2.0145×10−16
4.7577 ms
1.1161×10−2
Harmonic exponential
50.0620 dB
2.2573×10−16
4.1359 ms
1.5208×10−2
Damped-harmonic exponential
56.3837 dB
0.0000
3.9541 ms
6.8398×10−3
Tensor-product basis sweep for a matched ray inverse problem on 1920 rays and 784 coefficients. The target field is generated by the damped-harmonic exponential spline; all candidate bases reuse the same rays, regularization, and conjugate-gradient inverse solve.
Best candidate class
Pole profile
PSNR
CG solve
Coefficient RMSE
Order 2
λ=0.35, period 16
42.1900 dB
2.1103 ms
1.6017×10−1
Order 3
λ=0.42, period 10
55.2077 dB
1.9365 ms
7.5642×10−3
Order 4
λ=0.35, period 16
52.4859 dB
2.2178 ms
4.9702×10−2
Deterministic pole-selection sweep for damped-harmonic tensor-product exponential splines. The target is generated by the order-3, λ=0.42, period-10 operator; the matched pole set ranks first across the tested grid.
Pole multiset
Regularity
PSNR
CG solve
Density
Active/ray
[α]
C−1
65.7918 dB
2.1925 ms
0.0362
20.86
[0]
C−1
29.1917 dB
2.1154 ms
0.0362
20.86
[0,α]
C0
26.8630 dB
2.1370 ms
0.0724
41.71
[α,α]
C0
26.5164 dB
2.1558 ms
0.0724
41.71
[0,0,α]
C1
27.1299 dB
2.1973 ms
0.1084
62.44
[α,α,α]
C1
26.8975 dB
2.0574 ms
0.1084
62.44
Support-versus-regularity sweep for a first-order Green target with α=−0.42 on 1280 rays and 576 coefficients. The shortest matched basis wins because the target is Green-like rather than smooth.
Target regime
Best candidate
Best PSNR
Matched PSNR
Support
Regularity
Green [α]
[α]
52.9697 dB
52.9697 dB
1
C−1
Repeated [α,α]
[α,α]
45.6373 dB
45.6373 dB
2
C0
Two-real [α,β]
[α,β]
44.6760 dB
44.6760 dB
2
C0
Smooth [0,0,α]
[0,0,α]
49.4038 dB
49.4038 dB
3
C1
Damped oscillator
[−λ,−λ±jω]
48.4246 dB
48.4246 dB
3
C1
Polynomial [0,0,0]
[0,0,0]
49.9511 dB
49.9511 dB
3
C1
Mixed [0,0,α,α]
[0,0,α,α]
52.2576 dB
52.2576 dB
4
C2
Pole-multiset basis-selection map on 768 rays and 400 coefficients. The exact pole multiset ranks first in every tested target regime, confirming that operator matching and pole multiplicity are distinct from simply increasing spline order.
Target regime
Oracle basis
Inferred basis
Oracle PSNR
Inferred PSNR
Green [α]
[α]
[α]
35.3689 dB
35.3689 dB
Repeated [α,α]
[0,0,α,α]
[0,0,0,0]
33.1575 dB
33.1202 dB
Two-real [α,β]
[0,0,α,α]
[0,0,0,0]
32.2104 dB
32.0064 dB
Smooth [0,0,α]
[0,0,α,α]
[0,0,α,α]
34.9124 dB
34.9124 dB
Damped oscillator
[0,0,α,α]
[0,0,α,α]
34.6249 dB
34.6249 dB
Polynomial [0,0,0]
[0,0,0,0]
[0,0,α,α]
34.9254 dB
34.8408 dB
Mixed [0,0,α,α]
[0,0,α,α]
[0,0,α,α]
36.1621 dB
36.1621 dB
Measurement-driven operator inference with a held-out ray split. The selected basis is chosen without rendered target access; it remains within 0.2039 dB of the oracle PSNR basis across all tested regimes.
p0.22linewidthp0.34linewidthp0.27linewidthp0.10linewidth@
Model
PSNR
Selection signal
Block-label accuracy
Global held-out basis
28.6559 dB
Held-out rays
n/a
Local greedy basis
28.8074 dB
Held-out rays
25.00%
Oracle local labels
33.5291 dB
Ground-truth region labels
100.00%
Local operator adaptation stress test on a mixed pole-multiset field. The oracle gap confirms that local bases can matter, while the weak greedy-label recovery identifies the next algorithmic bottleneck.
Model
Boundary source
Region accuracy
PSNR
Global held-out basis
none
n/a
28.6559 dB
Blind greedy blocks
fixed grid
25.00%
28.8074 dB
Oracle region labels
ground truth
100.00%
33.5291 dB
FRI-region adaptation
derivative peaks + held-out refinement
100.00%
34.2000 dB
FRI-guided local operator adaptation. Sparse-innovation boundary proposals convert the weak blockwise selection problem into a region-level operator-selection problem and recover the local-basis advantage.
Projection angles
Mean global PSNR
Mean FRI PSNR
Mean gain
Region accuracy
12
28.2410 dB
29.4078 dB
+1.1667 dB
89.50%
18
28.7692 dB
31.9501 dB
+3.1809 dB
86.00%
24
28.9903 dB
34.0180 dB
+5.0277 dB
100.00%
FRI-region stress sweep averaged over additive noise levels {0,5×10−4,2×10−3}. The sparse-transition prior becomes more valuable as the ray geometry provides enough measurements to localize region boundaries.
Model
Boundary model
Region accuracy
PSNR
Global held-out basis
none
n/a
32.2114 dB
FRI curved selector
fitted curves
91.00%
31.1969 dB
Oracle curved labels
ground-truth curves
100.00%
41.5032 dB
Curved-interface local operator adaptation. The oracle gap confirms large local-operator headroom, while the selected model identifies the next bottleneck: held-out ray residuals alone are not sufficient to choose physical pole assignments under imperfect curved segmentation.
Model
Selection signal
Region accuracy
PSNR
Global held-out basis
held-out rays
n/a
32.2114 dB
FRI curved selector
held-out rays
91.00%
31.1969 dB
Joint curve refinement
residual + edge contrast
92.25%
34.7216 dB
Best searched candidate
hidden PSNR oracle
n/a
35.8310 dB
Oracle curved labels
ground-truth curves
100.00%
41.5032 dB
Joint curved-interface refinement. Adding edge-consistency contrast to the measurement score recovers a useful local-operator gain, but the remaining oracle gap shows that curved sparse-innovation geometry and pole assignment still need joint refinement.
Geometry basis
Parameters
Dense RMSE
Max error
Matched harmonic E-spline, L=3
16
5.4457×10−6
9.2024×10−6
Generic cubic polynomial spline
16
1.0256×10−2
1.9079×10−2
Piecewise-linear polygon
16
3.8693×10−2
1.6177×10−1
Geometry-reproduction benchmark for a closed harmonic curve from twelve parameter samples. The matched exponential-spline pole set places the target geometry in the span; generic polynomial and polygonal bases require more parameters to reach the same precision.
Optimizer
Search dimension
Dense RMSE
Runtime
Structural compression
Full Gaussian ES on controls
32
2.9097×10−2
46.708 ms
0.00%
Rank-one EGGROLL on controls
32
3.3517×10−2
39.477 ms
0.00%
Schmitter spline-subspace ES
8
6.9268×10−3
35.385 ms
75.00%
Rank-one EGGROLL in spline subspace
8
7.4862×10−3
35.693 ms
75.00%
Continuous subspace projection oracle
8
3.5689×10−4
0.293 ms
75.00%
Low-rank stochastic geometry search on a continuous spline shape family. The result separates the EGGROLL hardware mechanism from the Schmitter-style geometric prior: raw rank-one perturbations are not enough, while low-dimensional continuous spline shape coordinates produce the large error reduction.
Controlled spline-shape discovery benchmark. Red points are sparse noisy observations, gray curves mark the target where shown, and black curves show the recovered continuous spline shape for each optimizer.
Method
Search dim.
Held-out PSNR
Field RMSE
Boundary RMSE
Search time
Mean geometry + variable projection
0
16.8968 dB
2.2344×10−1
1.4287×10−1
0.00 ms
Full Gaussian ES controls
18
18.9082 dB
1.7413×10−1
3.9663×10−1
7132.59 ms
Rank-one EGGROLL controls
18
19.9034 dB
1.5539×10−1
3.4238×10−1
7333.54 ms
Rank-one EGGROLL spline subspace
6
20.1043 dB
1.7349×10−1
3.2396×10−1
7177.82 ms
True geometry projection oracle
6
33.1874 dB
3.3521×10−2
0.0000
0.00 ms
Adjoint variable-projection geometry benchmark. Each candidate boundary defines H(φ), radiance coefficients are eliminated by a ridge normal solve, and quality is measured on held-out projection directions. The oracle gap shows that coefficient elimination is not enough; physical boundary evidence must be strengthened.
Variable-projection radiance benchmark. Left panel is the target field; subsequent panels show recovered fields for mean geometry, full ES, rank-one control EGGROLL, rank-one spline-subspace EGGROLL, and true-geometry oracle. Red curves mark the recovered boundary.
Basis
Method
Held-out PSNR
Field RMSE
Boundary RMSE
Polynomial quadratic
Mean geometry
33.3249 dB
3.8306×10−1
9.5169×10−2
Polynomial quadratic
Edge-aware EGGROLL
41.7015 dB
3.8036×10−1
2.3912×10−2
Polynomial quadratic
True-geometry oracle
48.9766 dB
3.7926×10−1
0.0000
Exponential decay
Mean geometry
33.3627 dB
6.6277×10−2
9.5169×10−2
Exponential decay
Edge-aware EGGROLL
41.0644 dB
3.3052×10−2
2.4452×10−2
Exponential decay
True-geometry oracle
52.4398 dB
5.1297×10−3
0.0000
Damped harmonic
Mean geometry
33.3709 dB
6.6046×10−2
9.5169×10−2
Damped harmonic
EGGROLL, no edge term
39.9315 dB
4.0312×10−2
5.0651×10−2
Damped harmonic
Edge-aware EGGROLL
42.2339 dB
2.9005×10−2
2.7066×10−2
Damped harmonic
True-geometry oracle
96.9888 dB
3.2065×10−5
0.0000
Haouchat-matched variable projection. The inner operator is a quadrature-evaluated tensor-product spline ray projector, while the outer loop searches boundary geometry. Matching the damped-harmonic exponential basis restores the very high oracle ceiling, and adding sparse edge evidence moves the stochastic search into the 42 dB held-out regime.
Haouchat-matched variable projection for the damped-harmonic basis. Panels show the target, mean geometry, no-edge EGGROLL, edge-aware EGGROLL, and true-geometry oracle.
Basis
Method
Held-out PSNR
Field RMSE
Boundary RMSE
Polynomial quadratic
Measurement FRI edge
41.2085 dB
3.8043×10−1
2.5332×10−2
Polynomial quadratic
True-geometry oracle
48.9766 dB
3.7926×10−1
0.0000
Exponential decay
Measurement FRI edge
40.5459 dB
3.4998×10−2
4.8217×10−2
Exponential decay
True-geometry oracle
52.4398 dB
5.1297×10−3
0.0000
Damped harmonic
Mean geometry
33.3709 dB
6.6046×10−2
9.5169×10−2
Damped harmonic
Synthetic edge hint
45.9174 dB
2.1096×10−2
1.9775×10−2
Damped harmonic
Measurement FRI edge
41.5632 dB
3.3397×10−2
4.2715×10−2
Damped harmonic
True-geometry oracle
96.9888 dB
3.2065×10−5
0.0000
Measurement-derived FRI edge variable projection. The sparse boundary proposal is estimated from adjoint backprojection and derivative peak localization, not from the hidden target geometry. It recovers most of the useful edge-aware gain but remains below the cleaner synthetic edge hint, identifying multi-view FRI geometry extraction as the next bottleneck.
Measurement-derived FRI edge variable projection for the damped-harmonic basis. Panels show target, mean geometry, synthetic-edge EGGROLL, measurement-edge EGGROLL, and true-geometry oracle.
Basis
Method
Held-out PSNR
Field RMSE
Boundary RMSE
Polynomial quadratic
Multi-ray FRI variable projection
48.6908 dB
3.7930×10−1
1.4848×10−3
Polynomial quadratic
True-geometry oracle
48.9766 dB
3.7926×10−1
0.0000
Exponential decay
Multi-ray FRI variable projection
51.9659 dB
5.6990×10−3
1.4848×10−3
Exponential decay
True-geometry oracle
52.4398 dB
5.1297×10−3
0.0000
Damped harmonic
Mean geometry
33.3709 dB
6.6046×10−2
9.5169×10−2
Damped harmonic
Row/adjoint FRI edge
37.7852 dB
4.7983×10−2
6.8656×10−2
Damped harmonic
Multi-ray FRI variable projection
66.1488 dB
2.5205×10−3
1.4848×10−3
Damped harmonic
EGGROLL after multi-ray edge
40.3169 dB
3.6051×10−2
3.9707×10−2
Damped harmonic
True-geometry oracle
96.9888 dB
3.2065×10−5
0.0000
Multi-ray FRI edge refinement. Local boundary offsets are selected by held-in ray residuals after variable projection. The direct refined geometry nearly closes the oracle gap for the polynomial and exponential-decay bases and raises the matched damped-harmonic profile to 66.15 dB without synthetic edge injection.
Multi-ray FRI edge refinement for the damped-harmonic basis. The direct multi-ray FRI geometry, not the subsequent EGGROLL search, gives the dominant quality gain.
Method
Mean PSNR
Minimum PSNR
Mean boundary RMSE
Mean geometry
32.5922 dB
31.5426 dB
1.0261×10−1
Row/adjoint edge
38.0604 dB
32.5913 dB
3.8393×10−2
Multi-ray FRI variable projection
64.0770 dB
61.7645 dB
1.9800×10−3
True-geometry oracle
97.9883 dB
83.1976 dB
0.0000
Robustness sweep for multi-ray FRI edge refinement over 18 cases: two boundary variants, three projection-angle budgets, and three noise levels. The multi-ray FRI geometry remains consistently high quality and closes most of the gap between crude adjoint peaks and the oracle.
Representative robustness-sweep preview for the base boundary at the largest projection budget.
Sparse Stochastic Innovation Models
The sparse stochastic framework defines a process by [unser2014sparse1,unser2014sparse2] L{s}=w, where L is a whitening operator and w is white innovation noise. Gaussian w produces dense least-sparse processes; non-Gaussian Levy noise produces sparse or impulsive innovations. The operator controls correlation and physics, while the Levy measure controls sparsity.
The discrete-domain theory shows that matched B-spline filters convert continuous innovations into discrete generalized increments [unser2014sparse2]. MAP and MMSE estimators for these priors are developed in [bostan2013sparse,amini2013bayesian,kamilov2013mmse]. For OSNR, this means sparse parameters should not be arbitrary dense neural weights. They should be coefficient-domain innovations induced by the correct operator.
Controlled SPDE validation: advection–diffusion with Levy innovations
To convert the sparse stochastic theory into a PINN/weather-facing experiment, we implemented apps\_industrial\_breakthrough/spde\_operator\_spline\_benchmark.py. The controlled PDE is a periodic one-dimensional advection–diffusion–reaction model over a space–time block,
Lu=(∂t+a∂x−ν∂xx+λ)u=w(x,t),
where w is not restricted to be Gaussian. Following the Unser–Tafti sparse process model, the operator L fixes the correlation and propagation physics, while the innovation law determines the forcing morphology. We test four innovation profiles: a smooth periodic source, a Gaussian stochastic source, a compound-Poisson sparse impulse source, and a mixed weather-like source containing smooth waves, Gaussian background, and sparse jump events.
On the periodic grid, eqref(eq:spde-advection-diffusion) has the Fourier-domain symbol
L(ωt,ωx)=jωt+jaωx+νωx2+λ,
so the OSNR state-free solve is the diagonal complex division
u(ωt,ωx)=∣L(ωt,ωx)∣2+ϵL(ωt,ωx)w(ωt,ωx).
This is the SPDE analogue of an operator-matched exponential spline solve: the Green structure is built into the inverse operator, and no neural coordinate residual or automatic-differentiation tape is required. A low-pass Fourier reconstruction is included as a spectral-bias baseline; it mimics what happens when a smooth model family cannot carry non-Gaussian sparse innovations.
Profile
OSNR PSNR
Low-pass PSNR
Innovation RMSE
Event error
Hard zeros
Solve time
Smooth periodic, 1282
138.0840 dB
85.4051 dB
1.2801×10−4
n/a
0.00%
0.1383 ms
Gaussian SPDE, 1282
90.9460 dB
38.7232 dB
2.9458×10−4
n/a
0.00%
0.1370 ms
Poisson sparse, 1282
69.9296 dB
34.4731 dB
3.6954×10−3
0.0000 px
99.78%
0.1319 ms
Mixed Levy weather, 1282
99.5862 dB
60.4309 dB
3.7530×10−3
7.0438 px
0.00%
0.1414 ms
Poisson sparse, low diffusion
62.1205 dB
30.7575 dB
9.2609×10−3
0.0000 px
99.41%
0.1353 ms
Mixed Levy, low diffusion
80.6083 dB
47.5993 dB
9.2813×10−3
3.2817 px
0.00%
0.1472 ms
Poisson sparse, 1922
71.4692 dB
34.3896 dB
2.4829×10−3
0.0000 px
99.83%
0.5018 ms
Mixed Levy weather, 1922
104.2661 dB
63.6906 dB
2.5052×10−3
5.1575 px
0.00%
0.5070 ms
Controlled SPDE operator-spline benchmark for advection–diffusion with Gaussian and sparse Levy innovations. The OSNR solve is the direct FFT inversion of eqref(eq:spde-fft-solve); the low-pass row is a smooth spectral-bias baseline.
SPDE profile comparison on the 1282 benchmark. Each row shows the target field, the OSNR state-free FFT reconstruction, and the low-pass smooth baseline. The sparse and mixed rows expose why a Gaussian/smooth-only surrogate is not enough for weather-like fronts and impulses.
Table [tab:spde-operator-spline] gives the current interpretation. The state-free operator solve is essentially exact for all four innovation laws and remains below one millisecond even at 1922. The pure compound-Poisson case recovers event coordinates exactly at the tested grid resolutions, validating the sparse innovation view. The mixed case is more realistic and more difficult: the field reconstruction remains excellent, but raw top-K event localization degrades because the smooth and Gaussian components overlap the sparse impulses in the recovered innovation. This is not a failure of the operator inverse; it identifies the next algorithmic requirement. A weather-grade OSNR solver should add the same sparse-plus-smooth oblique innovation sieve used elsewhere in this paper, but now applied to Lu rather than to the field u itself.
We therefore added an explicit innovation-domain sieve to the same benchmark. Given the recovered innovation w~=Lu~, the sieve first estimates a smooth background wsm=Gσ∗w~ and then extracts sparse events from the residual
wsp=T(w~−wsm),
where T is either a known-cardinality top-K selector or an adaptive median-absolute-deviation threshold. This is not a field smoother; it acts after applying the physical operator and is therefore an innovation prior in the sense of Unser and Tafti. On the 1282 mixed Levy/weather case, raw top-K localization has mean event error 7.0438 px. The known-cardinality sieve reduces this to 0.0000 px. The same result holds for the low-diffusion stress case, where raw localization is 3.2817 px, and for the 1922 case, where raw localization is 5.1575 px. The adaptive MAD sieve with threshold 8 also recovers the mixed-weather events exactly without being told the number of events; it selects 36 sparse sites in the mixed case and yields 0.0000 px event error. In the pure Poisson case it selects a larger sparse support (602 sites at 1282) because the Gaussian smoothing residual leaves a local halo around each impulse, but nearest-event localization is still exact. Thus the next refinement is amplitude/support debiasing, not event detection. The practical weather implication is encouraging: OSNR can solve the stochastic PDE block globally and then separate sparse front/impulse innovations from smooth meteorological background in the physically meaningful residual domain.
We also tested the immediate nonlinear extension in apps\_industrial\_breakthrough/forced\_burgers\_spde\_benchmark.py. The model is a periodically forced viscous Burgers equation,
ut+uux−νuxx=fsmooth(x,t)+fsp(x,t),
where fsp is a sparse set of localized Gaussian events. A high-resolution spectral RK4 rollout is treated as the reference trajectory; compressed OSNR rollouts retain only a fixed number of Fourier/operator modes. The key inverse-problem distinction is that the forcing innovation must be estimated by applying the nonlinear physical operator to the observed trajectory,
f~(x,t)=ut+uux−νuxx,
not by thresholding the difference between a coarse rollout and the reference. The latter is mostly a truncation and phase-defect diagnostic. The former is the nonlinear analogue of the operator-domain innovation extraction used in the linear SPDE experiment.
We therefore report both sparse support recovery and compact event-atom recovery. The point sieve thresholds f~−Gσ∗f~ and measures whether each true event overlaps the recovered sparse support. The weak-form atom score integrates f~ against anisotropic Gaussian test functions matched to the injected event scale and then applies non-maximum suppression. We also apply a local centroid debiasing step around each detected atom. This approximates
ηm=⟨f~,φm⟩,
where φm is a compact adjoint/test atom. On clean synthetic forcing, direct operator-domain support recovery is sharper than the weak score; the weak form is expected to become more useful once observations are noisy or irregular.
Coarse modes
Rollout PSNR
Support error
Atom error
Centroid error
Defect ratio
8
39.5805 dB
0.0000 px
1.0357 px
0.7917 px
0.2290
18
68.7114 dB
0.0000 px
1.0357 px
0.7917 px
0.0098
32
108.0643 dB
0.0000 px
1.0357 px
0.7917 px
0.0001
Forced Burgers SPDE diagnostic after correcting the inverse-problem residual. Applying the nonlinear operator to the observed trajectory recovers every sparse forcing support location at the tested grid resolution. The atom-center error is about one pixel because the injected events are finite-width Gaussian blobs and overlapping events shift local maxima; local centroid debiasing reduces this to 0.7917 px. The defect ratio reports the norm of the coarse-rollout phase/truncation defect relative to the physical innovation norm.
Forced Burgers SPDE diagnostic at 18 retained modes. The corrected operator innovation f~=ut+uux−νuxx exposes the sparse forcing structure directly. The support map recovers the event locations, while the weak-form score produces compact event atoms within about one pixel.
The conclusion is important for the weather/PINN program. Linear operator-matched SPDEs are already a home-turf win for OSNR: exact global solves, sparse Levy innovations, and sub-millisecond runtime. The corrected nonlinear Burgers diagnostic shows that sparse forcing can also be recovered when the physical operator is applied in the right domain. A denser stress case with 56 injected events still yields 0.0000 px support error and 0.8912 px centroid error. The remaining bottleneck is not event detection but support and amplitude debiasing for finite-width/overlapping events, especially under noisy or partially observed fields. The next layer should estimate sparse innovations through an adjoint weak form,
⟨f,φm⟩=⟨ut+uux−νuxx,φm⟩,
with test functions φm matched to the operator and the expected front scale, plus a local centroid/amplitude debiasing step. An operator-splitting scheme that alternates deterministic nonlinear advection with a sparse forcing inverse problem is the natural production path before claiming weather-grade nonlinear SPDE recovery.
Direct PINN home-turf challenger.
We added a more direct PINN-facing control in apps\_industrial\_breakthrough/pinn\_operator\_home\_turf\_challenger.py. The benchmark is a periodic two-dimensional Helmholtz/Poisson problem,
(−Δ+λ)u(x,y)=f(x,y),
where u is a mixed-frequency smooth field and f is obtained by applying the known operator. The OSNR path solves the field by a single FFT-domain division. The baseline is a SIREN-style coordinate PINN trained with Adam on data samples and automatic-differentiation residual collocation. On the quality-first 1282 run with λ=6, the full OSNR solve reaches 141.7829 dB PSNR and RMSE 1.9704×10−7 in 0.1695 ms on CPU. A compressed low-mode OSNR profile retaining only 8.3557% of Fourier bins still reaches 136.3778 dB in 0.2935 ms. The SIREN PINN baseline, after 1800 epochs, reaches only 21.1458 dB and RMSE 2.1203×10−1 after 109.112 s. This is not a noisy external-data claim; it is a clean operator-known PINN control. It demonstrates the central home-turf point: when the differential operator and boundary topology are known, structural inversion gives both higher accuracy and roughly 6.44×105 lower training latency than residual-learning the same field.
Profile
PSNR
RMSE
Time
Active coefficients
OSNR full spectral solve
141.7829 dB
1.9704×10−7
0.1695 ms
100.00%
OSNR low-mode solve
136.3778 dB
3.6712×10−7
0.2935 ms
8.3557%
SIREN PINN, 1800 epochs
21.1458 dB
2.1203×10−1
109.112 s
dense MLP
Direct PINN home-turf challenger on a periodic 1282 Helmholtz/Poisson field. The OSNR rows are measured FFT/operator inversions; the SIREN PINN row is measured Adam training with autograd residual collocation.
PINN home-turf visual panel. The full and low-mode OSNR inversions are visually indistinguishable from the target at the displayed scale, while the trained SIREN PINN remains visibly over-smoothed after the measured optimization budget.
The scale follow-up at 2562 confirms that the coefficient fraction improves with resolution when the operator spectrum is compact. With the same low-mode budget, OSNR reaches 143.7870 dB in 0.6588 ms for the full solve, and 136.8069 dB in 1.1293 ms while retaining only 2.0889% of Fourier bins. A 900-epoch PINN baseline on the same field reaches 19.4427 dB after 59.916 s. At 5122, the same low-mode budget retains only 0.5222% of Fourier bins and still reaches 136.4969 dB in 3.5055 ms; the full solve reaches 143.3450 dB in 2.3255 ms, while a 300-epoch PINN baseline reaches 19.0573 dB after 19.664 s. The purpose of these rows is not to claim a universal neural-operator benchmark victory; they isolate the regime where PINN residual learning is structurally the wrong computational tool.
We ran an additional observation-noise stress test to separate robust atom detection from brittle support thresholding. Gaussian observation noise is added to the trajectory before evaluating the nonlinear operator. At 0.1% relative observation noise, raw pointwise support thresholding misses many events (7.8618 px support error), but ranked atom selection from the same operator residual remains accurate (0.7801 px after centroid refinement). Mild pre-operator smoothing restores support overlap (0.0357 px) but blurs atom centers (1.4350 px). At 0.5% noise, the best tested atom setting uses σ=0.75 pre-smoothing and reaches 0.7975 px centroid error, while binary support thresholding is unreliable. This confirms the correct noisy-weather design: detect a ranked set of operator-domain event atoms first, then run local amplitude/support debiasing rather than relying on a global hard threshold.
Operator-symbol identification by variable projection
The preceding SPDE experiments assume that the differential operator is known. The next weather/PINN question is whether OSNR can also learn a compact operator from data without falling back to a dense coordinate network. We therefore added apps\_industrial\_breakthrough/operator\_pole\_identification\_benchmark.py. The controlled model is the same advection–diffusion–reaction family
(∂t+a∂x−ν∂xx+λ)u=w,
but now the coefficients (a,ν,λ) are treated as unknown operator parameters. In Fourier space,
w−jωtu=(jaωx+νωx2+λ)u,
so the unknown operator coefficients enter linearly once the observed field and innovation are transformed. OSNR therefore identifies the operator by one complex ridge least-squares solve over selected frequency bins,
θ=argθ=(a,ν,λ)min∥D(u)θ−(w−jωtu)∥22+ϵ∥θ∥22.
This is a variable-projection step: linear field coefficients remain solved by the operator inverse, while the low-dimensional operator symbol is recovered directly from the data. A backpropagation baseline optimizes the same three parameters by Adam through the spectral residual for 800 steps.
Profile
Method
a^
ν^
λ^
PSNR
Time
Clean, all bins
OSNR LS
0.730001
0.021008
0.168636
76.9176 dB
1.4462 ms
Clean, band 24
OSNR LS
0.729999
0.021000
0.170013
116.0950 dB
0.2236 ms
Clean
Adam residual
0.729998
0.037428
0.038625
17.9194 dB
206.4690 ms
0.5% noise, band 24
OSNR LS
0.729844
0.019821
0.366630
40.2696 dB
0.1575 ms
0.5% noise, band 12
OSNR LS
0.729962
0.020995
0.171116
56.3422 dB
0.1897 ms
0.5% noise
Adam residual
0.723947
0.030782
0.044012
20.8344 dB
205.5888 ms
Operator-symbol identification for the advection–diffusion–reaction family with true parameters (a,ν,λ)=(0.73,0.021,0.17). The OSNR row uses a single complex least-squares solve in the Fourier/operator domain; the baseline uses iterative backpropagation through the same residual. Conservative spectral fitting bands suppress derivative-amplified observation noise.
Noisy operator-identification result at 0.5% observation noise with a conservative fitting band. The recovered operator reconstructs the state at 56.3422 dB after one structured coefficient solve.
Table [tab:operator-pole-identification] is the first explicit operator-learning result. In the clean case, a frequency band of 24 modes recovers all three coefficients to near machine precision and improves the reconstruction from 76.9176 dB to 116.0950 dB by avoiding ill-conditioned bins. With 0.5% observation noise, fitting too many frequencies corrupts the reaction estimate because derivative operators amplify high-frequency noise. Tightening the band to 12 modes restores the coefficients to sub-percent relative error and yields 56.3422 dB, while the Adam residual baseline remains near 20.8 dB after 800 gradient steps. The lesson is directly relevant to weather data: unknown physics should be learned as a compact, stability-constrained operator symbol with explicit spectral/noise control, not as an unconstrained dense coordinate network.
We then tested the harder field-only variant in apps\_industrial\_breakthrough/blind\_operator\_sparsity\_identification.py. Here w is hidden: the search chooses the operator whose residual Lθu is most compressible as a low-pass smooth field plus a fixed number of sparse atoms. This is closer to unsupervised weather-model discovery, but it exposes an identifiability boundary. On four independent trajectories sharing the same true operator, the blind compressibility score selects (a^,ν^,λ^)=(0.91,0.021,0.26) instead of (0.73,0.021,0.17), even though the sparse event locations are recovered exactly. A support-projected oracle that masks the true sparse event neighborhoods but does not know their amplitudes also fails to recover the reaction coefficient. The reason is structural: from u alone, a wrong operator can be absorbed into a different smooth forcing background, so sparse-plus-smooth compressibility is not a unique operator identifier.
Setting
a^
ν^
λ^
Event error
Time
Blind sieve, 4 clean trajectories
0.910000
0.021000
0.260000
0.0000 px
2596.29 ms
Blind sieve, 4 trajectories, 0.2% noise
0.910000
0.009000
0.260000
0.0000 px
2594.86 ms
Support-projected oracle, clean
0.570370
0.014276
−14.202470
oracle support
5.86 ms
Support-projected oracle, 0.2% noise
0.314185
0.000886
1.900169
oracle support
5.96 ms
Blind operator-discovery diagnostic. Sparse event geometry can be recovered from Lθu, but field-only sparse-plus-smooth compressibility does not uniquely identify the true operator because operator mismatch can be reinterpreted as smooth forcing.
Blind operator-sparsity diagnostic. The selected residual preserves sparse event locations but corresponds to the wrong operator, demonstrating that fully blind field-only operator discovery needs additional physical anchors.
This negative result is useful. It says the breakthrough lane is not arbitrary unsupervised PDE discovery from a single scalar field. The credible path is semi-blind operator learning: use measured innovations, multiple observed state channels, conservation laws, boundary/flux constraints, or assimilation windows to anchor the smooth forcing ambiguity, then recover the compact operator symbol by structured least squares or variable projection.
The first semi-blind anchor test follows this prescription. In apps\_industrial\_breakthrough/anchored\_operator\_identification.py, only a random subset of the forcing samples is revealed. The field u is observed everywhere, but the operator coefficients are fitted from the pointwise equations
at the anchor sites. All derivatives are evaluated analytically by spectral/operator columns, and the three unknown coefficients are recovered by a real ridge least-squares solve.
Setting
Anchors
a^
ν^
λ^
PSNR
Clean, 0.25% anchors
41
0.729829
0.021105
0.143701
49.6630 dB
Clean, 0.50% anchors
82
0.729892
0.021072
0.159291
58.4940 dB
Clean, 5.00% anchors
819
0.729956
0.021002
0.172572
69.6683 dB
0.2% noise, no denoise, 5.00% anchors
819
0.732336
0.008352
2.355422
29.6130 dB
0.2% noise, band 12, 5.00% anchors
819
0.729782
0.020778
0.183559
53.7874 dB
Semi-blind operator identification from sparse forcing anchors. Clean operator recovery is accurate with very few forcing samples. Under observation noise, derivative columns require spectral denoising; with a band-12 field prefilter, 5% anchors recover the operator to below 8% worst relative error and reconstruct the state at 53.7874 dB.
Semi-blind noisy operator identification with sparse forcing anchors. A small set of pointwise forcing measurements breaks the field-only ambiguity exposed in Table [tab:blind-operator-sparsity].
Table [tab:anchored-operator-identification] is a more realistic weather/PINN direction than fully blind scalar discovery. It shows that a small number of physical anchors can make compact operator identification well-posed again. The nonmonotone noisy rows also identify the next engineering layer: anchors should be selected by leverage or derivative-energy criteria rather than uniformly at random.
We tested the simplest version of that idea by selecting anchors with the largest normalized derivative-column energy. This naive leverage rule is not sufficient. In the clean case it reaches only 52.1897 dB at 5% anchors, below the random-anchor 69.6683 dB result. With 0.2% observation noise and band-12 denoising, it degrades to 37.4777 dB at 5% anchors because high-leverage points are also the points where derivative noise is most amplified. The active-anchor rule must therefore combine derivative leverage with noise sensitivity and spatial diversity; selecting the largest rows of the design matrix is too brittle.
A follow-up robustification adds a trimmed ridge solve: after the first anchor fit, the largest pointwise residuals are discarded and the operator is refit on the lowest-residual fraction. Random anchors with trimming do not improve the best 5% noisy row, but diverse leverage plus trimming uncovers a useful low-anchor operating point. With only 0.5% anchors under 0.2% observation noise, diverse-trimmed anchors estimate (a,ν,λ)=(0.731760,0.020692,0.175461), corresponding to only 3.2124% worst relative parameter error and 49.5094 dB reconstruction. The best state PSNR still comes from denser random anchors, but the active-trimmed result shows that carefully chosen anchors can reduce physical measurements by an order of magnitude while preserving an accurate compact operator.
Nonlinear Burgers operator identification and held-out forecasting
The next SOTA-facing PINN target is nonlinear forecasting rather than static reconstruction. We implemented apps\_industrial\_breakthrough/burgers\_operator\_identification\_forecast.py, which treats viscous Burgers dynamics
ut+cuux=νuxx
as a compact operator-identification problem. From the observed training window, OSNR forms analytic derivative columns (ut,uux,uxx) and solves the two unknown coefficients (c,ν) by a tiny ridge system over sparse anchor samples. The learned operator is then rolled forward over the held-out future window. This is the nonlinear analogue of the semi-blind anchor experiments above, but the validation target is future prediction, which is the quantity that PINNs and neural operators usually report.
Setting
Anchors
c^
ν^
Future PSNR
ID time
Clean, 0.25% anchors
35
1.003308
0.0045239
70.7236 dB
0.4064 ms
Clean, 0.50% anchors
69
0.998324
0.0044914
77.6890 dB
0.2062 ms
Clean, 1.00% anchors
138
0.999828
0.0044959
88.6117 dB
0.1814 ms
Adam residual, clean
all
0.997542
0.0044901
74.8301 dB
82.15 ms
0.1% noise, 0.50% anchors
69
0.995605
0.0044539
66.3736 dB
0.2458 ms
0.2% noise, 1.00% anchors
138
0.995355
0.0044588
66.8131 dB
0.3160 ms
Adam residual, 0.2% noise
all
0.979668
0.0044440
56.8246 dB
81.62 ms
Wrong prior
n/a
0.75
0.00225
30.7955 dB
n/a
Nonlinear Burgers operator identification and held-out forecasting. True parameters are (c,ν)=(1.0,0.0045), the training window is 45% of the timeline, and the future window contains the remaining 90 frames. OSNR identifies the operator in sub-millisecond time from sparse anchors and forecasts the future without a neural training loop.
Noisy Burgers operator-ID forecast at 0.2% observation noise. OSNR recovers the nonlinear operator from 1% training-window anchors and forecasts the held-out future at 66.8131 dB.
This is the strongest nonlinear PINN-facing result so far. The clean 1% anchor row reaches 88.6117 dB future PSNR with a 0.1814 ms identification solve, while the Adam residual fit is about 450× slower and reaches only 74.8301 dB. Under 0.2% observation noise, OSNR still reaches 66.8131 dB from 1% anchors, outperforming the Adam residual fit by about 10 dB. The wrong-prior row shows that the forecast is not trivially easy: incorrect physics collapses to 30.7955 dB. This is the first result that directly combines nonlinear coefficient discovery, held-out forecasting, sparse physical measurements, and a clear optimization-speed gap.
The scalar Burgers forecast is a necessary control, but the weather/PINN claim requires a coupled multi-field nonlinear system. We therefore implemented apps\_industrial\_breakthrough/nonlinear\_shallow\_water\_operator\_forecast.py. The state is q=(η,u,v) and the operator family is
The unknown physical vector is θ=(H,β,g,f,r,ν,μh), covering mean depth, nonlinear transport strength, gravity, Coriolis coupling, damping, momentum viscosity, and height diffusion. OSNR forms the full analytic derivative library from the observed training window and solves a scaled sparse-anchor linear system for all seven coefficients at once. The fitted operator is then advanced over the held-out future window using the same spectral RK4 physics core. The comparison baseline fits the same residual equations by Adam over all training rows, so the quality comparison is not against a weak interpolant but against the standard differentiable residual-minimization path used by PINN-style methods.
Setting
Anchors
β^
g^
Max param. err.
Future PSNR
ID time
482, clean, 0.25%
795
0.7178
0.8590
1.8214%
68.7348 dB
14.24 ms
482, clean, 1.00%
3180
0.7181
0.8590
1.6975%
68.7314 dB
13.24 ms
Adam residual, 482 clean
all
0.7186
0.8590
1.8572%
68.6608 dB
830.30 ms
482, 0.1% noise
795
0.7091
0.8594
21.347%
66.3867 dB
16.93 ms
Adam residual, 0.1% noise
all
0.7171
0.8587
17.495%
66.0641 dB
1093.02 ms
482, 0.2% noise
795
0.7218
0.8599
34.417%
65.6597 dB
17.96 ms
Adam residual, 0.2% noise
all
0.7156
0.8585
32.218%
63.4251 dB
1095.57 ms
642, clean, 0.10%
713
0.7191
0.8594
1.8368%
72.6945 dB
24.00 ms
Adam residual, 642 clean
all
0.7191
0.8593
1.1991%
72.5755 dB
1103.88 ms
Wrong prior, 642
n/a
0.3240
1.0750
n/a
27.7993 dB
n/a
Coupled nonlinear shallow-water operator identification and held-out forecasting. True parameters are (H,β,g,f,r,ν,μh)=(1.0,0.72,0.86,0.58,0.065,0.006,0.004). OSNR uses sparse derivative anchors; Adam optimizes the same residual over all training rows.
Coupled nonlinear shallow-water forecast at 121×642. With only 0.1% sparse derivative anchors, OSNR identifies the seven-parameter nonlinear operator and forecasts the held-out future at 72.6945 dB. The wrong-prior forecast falls to 27.7993 dB, confirming that the high score is not a trivial smoothness artifact.
This is the first multi-field nonlinear weather-core forecast result in the project. It preserves the key advantage seen in Burgers: the residual landscape can be collapsed into a small structured operator solve instead of optimized by thousands of neural/PINN gradient steps. On the 642 run, OSNR uses only 713 anchor equations out of the full derivative library and identifies the operator in 24.00 ms, while the Adam residual fit takes 1103.88 ms. Both methods converge to similar coefficients in the clean case because the library is correct, but OSNR reaches the solution in one scaled linear solve with about a 46× identification-speed advantage and no neural training loop. Under observation noise, derivative bias still affects the weak damping/diffusion terms, but the future forecast remains above 65 dB and stays ahead of Adam in the tested 0.2% setting. The next moonshot is therefore not another scalar PDE; it is sparse/partial observation data assimilation for this same coupled nonlinear operator family.
Adaptive sparse sensors for nonlinear shallow-water assimilation
We then cross-pollinated the weather station-placement result with the coupled nonlinear shallow-water core. The runner apps\_industrial\_breakthrough/nonlinear\_shallow\_water\_adaptive\_sensor\_assimilation.py keeps the same no-backprop pipeline: sparse sensors reconstruct the observed training window by a closed-form Fourier-dictionary ridge solve, the seven-parameter nonlinear operator is identified by sparse least squares, and the state is rolled into the held-out future. The only changed variable is where the sparse sensors are placed. Sensor policies are computed from the training window only, never from held-out future frames. We compare random points, raw training-window variance/gradient/leverage scores, lattice-plus-score hybrids, residual-innovation hybrids, and centered/phase-shifted coverage policies.
Observation setting
Modes
Sensor policy
Sensors
Future PSNR
Clean
4
lattice 3%
123
19.842 dB
Clean
4
lattice+hybrid 10% budget, 3% total
123
19.822 dB
Clean
3
centered lattice 1.5%
61
19.852 dB
Clean
3
centered lattice 2%
82
19.942 dB
Clean
3
offset-best lattice 2%
82
19.958 dB
Clean
3
random 20%
819
19.920 dB
1% sensor noise
3
centered lattice 2%
82
19.948 dB
1% sensor noise
3
random 20%
819
19.906 dB
Adaptive sparse-sensor placement for nonlinear shallow-water assimilation and forecasting. All rows use the same 121×642 trajectory, closed-form sparse-window assimilation, sparse least-squares operator identification, and no backpropagation. Mode 3 centered/offset coverage reaches dense-random forecast quality with 10× fewer observations. The offset-best policy chooses the best phase among 16 centered lattice shifts by training-window reconstruction PSNR only.
Adaptive shallow-water sparse-sensor forecast panel for the mode-3 coverage-geometry run. Centered/offset lattice rows preserve the large-scale future height field with 2–3% sensors, while dense random placement needs about 20% sensors to reach the same forecast band.
This result is useful because it is positive and diagnostic. The naive high-information policies are not winners: variance, gradient, and hybrid-diverse placement overconcentrate sensors in active regions and can make the Fourier reconstruction ill-conditioned. The follow-up lattice-plus-information experiment confirmed the same boundary: at 3% total sensors with mode 4, lattice+gradient, lattice+hybrid, and lattice+residual 10% allocation reach 19.786, 19.822, and 19.807 dB, all below the pure lattice row at 19.828 dB. The real improvement is coverage geometry plus basis order. Mode 3 is the sparse bias-variance sweet spot; modes 1–2 underfit and modes 5–6 are underconstrained at low sensor counts. A centered lattice at 2% sensors reaches 19.942 dB clean future PSNR, above the same-run 20% random reference at 19.920 dB; with 1% sensor noise, the same 2% centered lattice row reaches 19.948 dB versus noisy random 20% at 19.906 dB. The low-count sweep shows the transition: 0.5% centered sensors fail (15.628 dB), 1% is not yet dense-random quality (19.158 dB), 1.5% approaches it (19.852 dB), and 2% crosses it.
We then tested whether this was a single-trajectory phase artifact. The script now exposes initial roll and amplitude controls, and the offset-best policy chooses the best of 16 lattice phases by training-window reconstruction only. Across four robustness variants, offset-best 2% sensors remains at or above random 20%: roll (7,11) gives 19.946 versus 19.942 dB, roll (13,5) gives 19.955 versus 19.919 dB, amplitude scale 1.25 gives 19.948 versus 19.894 dB, and a changed dynamics profile (β,g,f)=(0.9,0.78,0.45) gives 19.981 versus 19.932 dB. Thus adaptive station placement is not only a terminal weather-assimilation trick. In a coupled nonlinear weather-core forecast, a coverage-aware sensor topology can reduce observations by 10× while preserving future forecast quality, using only algebraic assimilation and operator identification.
Finally, we removed the known-family assumption after sparse sensing. The runner apps\_industrial\_breakthrough/nonlinear\_shallow\_water\_sparse\_sensor\_library\_discovery.py first reconstructs the training window from sparse sensors, then fits the 24-column shallow-water library by sequential thresholded least squares, and forecasts from the assimilated last state. This is a harder test because the solver must reject decoy columns and no longer receives the seven-parameter operator family.
p0.17linewidthp0.13linewidthc c c c p0.15linewidth@
Observation setting
Policy
Sensors
Threshold
Support (TP,FP,FN)
Future PSNR
Reference
Clean
offset-best 2%
82
0.003
(8,0,5)
19.960 dB
known-family 19.962 dB
Clean
random 20%
819
0.005
(7,0,6)
19.888 dB
known-family 19.902 dB
1% sensor noise
offset-best 2%
82
0.003
(8,0,5)
19.965 dB
known-family 19.964 dB
1% sensor noise
random 20%
819
0.003
(7,0,6)
19.885 dB
known-family 19.912 dB
Sparse-sensor governing-equation discovery after Fourier assimilation. The discovered support is counted against the 13 true library columns and 11 decoys. Offset-best 2% sensors recover an 8-term true subset with zero decoys and match or exceed dense-random 20% forecast quality.
The sparse-library result changes the interpretation. The 2% offset-best row does not fully recover all weak nonlinear/damping terms, but it recovers the dominant conservative, pressure, Coriolis, and diffusion operators with no decoys and forecasts within about 0.002 dB of the known-family coefficient fit. A lower threshold 0.001 recovers 9 true terms with one decoy and reaches 19.964 dB clean, but the zero-decoy 8-term threshold is the cleaner scientific claim. Thus the sensor topology is not merely helping a fixed PDE prior; it preserves enough operator information for sparse governing-equation discovery from partial observations.
The next step was to remove a flaw in the pointwise library: it thresholds every equation-specific term and every decoy on the same normalized scale, even though the physically meaningful shallow-water operators are shared typed groups. We therefore added a weak-form typed library in nonlinear\_shallow\_water\_sparse\_sensor\_weakform\_discovery.py. Each window enforces q(tb)−q(ta)≈∫tatbLj(q(t))dt and projects the balance onto low Fourier test modes. The 13 true pointwise terms are then tied into seven shared physical groups (H,β,μh,g,f,r,ν), while the 11 nuisance columns receive a larger typed-selection threshold. This is not a neural loss or a backward pass: it is a weak-form operator balance followed by weighted sequential thresholded least squares.
p0.17linewidthp0.14linewidthc c c c p0.17linewidth@
Observation setting
Policy
Sensors
(τ,λdecoy)
Support (TP,FP,FN)
Future PSNR
Reference
Clean
random 2%
82
(2×10−6,50)
(7,1,0)
unstable
known-family 16.055 dB
Clean
offset-best 2%
82
(5×10−7,20)
(7,0,0)
19.961 dB
weak dense 19.961 dB
Clean
random 20%
819
(10−6,50)
(7,0,0)
19.909 dB
weak dense 19.909 dB
1% sensor noise
offset-best 2%
82
(5×10−7,20)
(7,0,0)
19.956 dB
weak dense 19.956 dB
1% sensor noise
random 20%
819
(10−6,50)
(7,0,0)
19.910 dB
weak dense 19.910 dB
Typed weak-form sparse-sensor governing-equation discovery. Support is counted over seven shared physical operator groups and eleven decoys. The decoy multiplier λdecoy applies only to nuisance columns. Offset-best 2% sensors recover the complete seven-group shallow-water operator with zero decoys and match the dense weak-form solve, while using 10× fewer observations than random 20%.
This closes the support-recovery gap left by Table [tab:nonlinear-shallow-water-sparse-sensor-library]. With typed weak-form rows, offset-best 2% sensors recover all seven physical groups with zero false positives in both the clean and 1% sensor-noise settings. The same row is also forecast-competitive: 19.961 dB clean and 19.956 dB noisy, above the corresponding typed random-20% rows (19.909 and 19.910 dB). Random 2% still fails despite selecting most physical groups, which confirms that the result is not merely a threshold artifact. The station geometry must preserve a well-conditioned weak operator balance; once it does, typed OSNR selection can recover the full coupled shallow-water operator from sparse partial observations without backpropagation.
The typed weak-form result also passed the first robustness sweep. At 1% offset sensors, the selector already recovers (7,0,0) support but only reaches 19.505 dB, so full support and forecast-quality crossing are separate requirements. At 2% offset sensors, all four trajectory variants recover (7,0,0): roll (7,11) gives 19.926 dB, roll (13,5) gives 19.948 dB, initial scale 1.25 gives 19.946 dB, and changed dynamics (β,g,f)=(0.9,0.78,0.45) gives 19.978 dB. The random-20% typed weak-form rows for the same variants are 19.942, 19.916, 19.944, and 19.977 dB, respectively. Thus complete support recovery is robust in the tested variants; the 10× forecast-quality advantage holds in three of four variants and narrowly fails on roll (7,11).
We then tested whether the same typed weak-form mechanism learns a reusable operator rather than a trajectory-specific correction. The multi-trajectory runner trains one shared operator from sparse-observed variants \base, roll (7,11), scale 1.25\ and forecasts unseen roll variants (13,5) and (5,17) from their held-out states. Offset-best 2% sensors recover full support and reach a mean unseen-trajectory PSNR of 51.100 dB; random 2% is unstable even with nearly full support. Dense random 20% also recovers full support and reaches 52.232 dB, while offset-best 20% reaches 62.902 dB. This is the first sparse-observation cross-trajectory operator-learning result in this section. It is not a 10× dense-quality win at 2% observations, but it shows that the typed OSNR weak-form solver can learn a shared coupled operator from partial observations and transfer it to unseen initial conditions without backpropagation.
A follow-up sensor-fraction and placement sweep showed that the cross-trajectory coefficient bottleneck is primarily geometric. With offset-best placement and the same typed solver, 3% sensors already recover (7,0,0) and reach 53.282 dB, exceeding the random-20% result; 5% reaches 57.847 dB; 10% drops to 52.865 dB; and 20% reaches 62.902 dB. Thus adding sensors is not monotone unless the station geometry remains well conditioned for the weak-form operator rows. A placement sweep at 3–10% found that pointwise saliency policies (variance, gradient, and hybrid-diverse additions) consistently introduce decoy groups and degrade transfer. The best sparse clean-support result is a centered lattice with 5% sensors and a stronger decoy multiplier: it recovers (7,0,0) and reaches 58.829 dB on the two unseen trajectories. If the two derivative-decoy groups are allowed, the same 5% centered lattice reaches 60.775 dB, but we treat this as a numerical correction rather than a clean governing-equation discovery. The practical conclusion is that operator-identifiability-balanced station geometry is more important than generic high-activity station placement.
We then made the geometry test explicit by adding row-conditioning and fixed lattice-phase sweeps. Unit-design, unit-joint, and clipped row normalizations all made the solver worse: they activated most decoys and collapsed transfer to roughly 22–30 dB. This negative result is important because the weak-form row magnitudes carry physical operator information; flattening them destroys the balance rather than improving conditioning. In contrast, fixed quarter-phase lattice placement is a productive control variable. At 5% sensors on the three-training-trajectory protocol, the best clean fixed phase (0.75,0.75) reaches 60.365 dB with exact (7,0,0) support, and the result validates on fresh roll and scale variants with mean 60.365 dB. Expanding the training set to eight sparse-observed trajectories raises the same clean 5% phase result to 60.945 dB. Most importantly, a focused sensor curve with this phase shows that 7.5% sensors per training trajectory recover exact support and reach 65.072 dB on four fresh test variants, exceeding the same-protocol 20% phase reference of 64.149 dB. The curve remains nonmonotone: 10% drops to 55.014 dB and 15% to 59.252 dB. Thus the current lesson is sharper than ``more sensors'': sparse OSNR operator learning can beat denser observation budgets when station geometry is phase-balanced for the weak operator, but station-count increases can still harm coefficient estimation if they alias the weak-form rows.
We also audited whether the phase can be selected without looking at the final test variants. Simple training-only proxies failed: physical-column condition number, physical–decoy coherence, dense weak residual, leave-one-training-trajectory weak residual, leave-one theta stability, and an inner training-window rollout score did not rank the best phases. A standard validation split, however, does. Selecting the phase on validation variants \roll (3,9), scale 0.75\ chooses (0.25,0.75), which then transfers to disjoint test variants \roll (11,4), scale 1.40, roll (19,2), scale 1.30\. The validation-selected 7.5% phase recovers exact support and reaches 65.161 dB on that disjoint test set, while the same phase with 20% sensors reaches 64.283 dB. This is the cleanest current sparse cross-trajectory result: the station geometry is selected on validation data, the test variants are unseen, and the learned seven-group operator still beats the denser observation budget without backpropagation.
Finally, we tested whether the nonmonotone phase-lattice curve could be repaired by replacing the lattice with low-discrepancy or jittered station families. It could not. On the same validation-selected test protocol, Sobol phase stations activated 7–10 decoys and produced unstable forecasts at 7.5%, 10%, and 15% sensors. Jittered phase lattices were stable but much weaker: 46.774 dB at 7.5%, 44.902 dB at 10%, and 58.205 dB at 15%. The unjittered phase lattice remains the best clean geometry, with 65.161 dB at 7.5%. Thus the current station rule is not generic space filling; it is a Fourier-compatible phase-balanced sampling rule.
We then promoted the validation split from phase selection to joint phase/count selection. The training set stayed fixed at eight sparse-observed variants \base, roll (7,11), scale 1.25, roll (13,5), roll (5,17), roll (2,19), scale 0.90, scale 1.10\. The validation variants were again roll (3,9) and scale 0.75. The grid searched lattice phases (0.25,0.75), (0.75,0.25), (0,0.5), (0.75,0.75), (0,0.75), and (0.25,0.5) at sensor fractions 5%, 6.25%, 7.5%, 8.75%, 10%, 12.5%, and 15%, with threshold 5×10−7, decoy multiplier 20, no row normalization, window 7, stride 2, and a 75% inner training-window diagnostic split. Validation selected the 8.75% lattice with phase (0.75,0.75): it recovered exact (7,0,0) support and reached 67.996 dB on the two validation variants. Without changing any hyperparameter, the selected row transferred to disjoint test variants \roll (11,4), scale 1.40, roll (19,2), scale 1.30\, reaching 67.841 dB with exact support. Same-run references were 65.432 dB at 6.25%, 65.161 dB for the previous 7.5% phase (0.25,0.75) row, and 64.283 dB for the same 20% phase (0.25,0.75) row. Thus held-out validation can now select both observation count and phase, and the selected sparse geometry uses only 358 stations per training trajectory while outperforming 819-station dense-phase references.
A final fine phase/count refinement around this winner exposed a sharper resonance. We searched 8.125%, 8.4375%, 8.75%, 9.0625%, and 9.375% sensors with phases (0.62,0.62), (0.62,0.75), (0.75,0.62), (0.75,0.75), (0.75,0.87), (0.87,0.75), (0.87,0.87), (0.62,0.87), and (0.87,0.62). The adjacent count bands 8.125% and 8.4375% were poor despite exact support, reaching only about 53–54 dB; 9.0625% recovered to 68.520 dB at phase (0.75,0.75), but the validation winner was again 8.75%, now with phase (0.75,0.87). This row reached 70.774 dB on validation and transferred to the disjoint test variants at 69.803 dB, with exact (7,0,0) support and a 1.31 ms sparse solve. The same fine-test run reproduced the old 8.75% phase (0.75,0.75) result at 67.841 dB and showed that moving the winning phase to 9.0625% drops to 65.941 dB. The result is therefore not a generic phase preference. It is a count-specific Fourier sampling geometry that materially improves coefficient accuracy while keeping the observation budget at 358 stations per training trajectory.
To check whether the fine geometry was overfitting the two validation variants, we ran a broader fresh-variant audit with roll shifts (1,23), (23,1), (31,17), (17,31) and amplitude scales 0.60, 1.60, 0.50, and 1.75. The selected 8.75% phase (0.75,0.87) row reached 70.166 dB across these eight variants with exact support. Same-count controls were 67.882 dB for phase (0.75,0.75) and 63.047 dB for phase (0.25,0.75), while 20% references reached only 64.044, 64.104, and 65.107 dB for the three tested phases. This robustness audit strengthens the interpretation: the selected sparse station geometry generalizes across unseen roll and amplitude perturbations and beats substantially denser station budgets because it better identifies the weak operator coefficients, not because it sees more observations.
Because the resonance was phase-sharp, we then ran a local phase-only refinement at the fixed 8.75% count. The validation grid swept x phases 0.70,0.72,0.75,0.78,0.80 and y phases 0.84,0.87,0.90,0.93 around the previous winner. Most rows were much weaker even with exact support; for example x=0.80 remained below 60 dB and (0.75,0.93) dropped to 66.390 dB. The validation winner was (0.75,0.90) at 73.568 dB. Tested on the union of the four disjoint variants and the eight broad-audit variants, this row reached 73.379 dB with exact support and a 1.17 ms solve. On the same 12-variant audit, (0.75,0.87) reached 70.045 dB and (0.75,0.75) reached 67.868 dB. This is the current best clean shallow-water result: validation-selected sparse station geometry with 358 observations per training trajectory beats both same-count neighboring phases and all tested 819-station references by a large margin.
One more one-percent refinement around (0.75,0.90) saturated rather than improved the result. Sweeping x∈{0.73,0.74,0.75,0.76,0.77} and y∈{0.88,0.89,0.90,0.91,0.92} at the same 8.75% count again selected (0.75,0.90); (0.75,0.91) tied it because the rounded station set is effectively equivalent. Nearby rows drop quickly: (0.75,0.89) gives 72.614 dB, (0.75,0.92) gives 68.431 dB, x=0.76–0.77 with y=0.90–0.91 gives 70.367 dB, and x=0.73–0.74 remains near 62–63 dB. Thus the station-design frontier appears locally saturated at this lattice resolution; the next improvement must come from a different station family, a richer validation criterion, or a stronger operator/library model rather than sub-percent phase nudging.
We next tested whether more sparse-observed training trajectories improve the shared operator. They do not automatically help. On a fresh test set \roll (9,27), roll (27,9), roll (15,29), roll (29,15), scale 0.70, scale 1.50, scale 0.40, scale 1.90\, the current eight-training-variant row with 8.75% phase (0.75,0.90) reaches 73.381 dB. Adding four more sparse-observed training variants \roll (1,23), roll (23,1), scale 0.60, scale 1.60\ while keeping the same per-trajectory station budget and solver drops the same fresh-test mean to 70.316 dB. The support remains exact, but the coefficient vector shifts, especially in the nonlinear and damping terms. Thus the next operator-learning lever is not simply more trajectories; training variants must be selected or weighted so that assimilation bias from scale-extreme trajectories does not distort the shared weak-form coefficients.
The isolating controls confirm that the degradation is not caused by one family alone. Adding only the two extra roll variants to the eight-variant training set gives 71.321 dB on the same fresh test set; adding only the two scale-extreme variants gives 70.131 dB. Both retain exact support, but both move the coefficients away from the high-PSNR eight-variant estimate. This suggests that the original eight sparse-observed trajectories already form a good coefficient-calibration design for this station phase. Additional trajectories should enter only through validation-selected weights or subset selection, not by unweighted concatenation.
We implemented that weighting hook in the runner as --train\_variant\_weights, multiplying each variant's weak-form rows and targets by the square root of its weight before the closed-form solve. Downweighting the four rejected variants improves over unweighted concatenation but still does not beat the eight-variant subset: weights 0.25, 0.10, 0.03, and 0.01 on the four added variants yield 72.490, 73.035, 73.280, and 73.346 dB, respectively, on the same fresh test set, versus 73.381 dB for weight zero. Thus the validation-selected action for these candidates is rejection. The useful research conclusion is that the no-backprop operator learner can support neuromodulatory-style reliability weights, but the first weighted audit says the next gain requires discovering better candidate trajectories or operator features, not softly retaining known harmful variants.
We then audited three alternative explanations before changing the operator model. First, a threshold/decoy-pressure sweep around the 8.75% phase (0.75,0.90) frontier used thresholds 10−7, 2×10−7, 5×10−7, 10−6, and 2×10−6 with decoy multipliers 5, 10, 20, 50, and 100. Low decoy penalties admitted false positives and dropped validation to about 66 dB, but every exact-support row gave the same 73.568 dB validation score; applying the validation-selected 10−7, multiplier-50 row to the 12-variant fresh audit reproduced 73.379 dB. Thus the frontier is not limited by the sparse threshold once decoys are suppressed.
Second, we implemented phase-preserving score-mixed station policies such as lattice\_phase75\_90\_gradient01. These keep the tuned phase lattice as the backbone and replace only 1–5% of the station budget with diverse high-gradient, high-variance, or hybrid-score sites. This fairer saliency audit was decisively negative. At the fixed 352–358 station scale, generic score-mixed policies tied to the untuned lattice collapsed to 49.713 dB or worse, and even the phase-preserving variants degraded monotonically: gradient replacement at 1%, 2%, 3%, and 5% gave 68.468, 65.134, 62.046, and 52.616 dB; variance replacement gave 55.350, 51.061, 46.917, and 42.942 dB; hybrid replacement gave 58.324, 54.872, 50.958, and 45.586 dB. The conclusion is that pointwise saliency is not an adequate station objective for this weak operator learner. The lattice points themselves carry Fourier conditioning, and replacing even a few of them damages the coefficient estimate despite exact support in several rows.
The positive improvement came from exact decimation of the phase lattice. Scanning the integer station counts 350 through 361 at phase (0.75,0.90) found a new validation winner at 352 stations, i.e. sensor fraction 352/4096=0.0859375. This row recovers exact (7,0,0) support and reaches 74.442 dB on the validation variants, compared with 73.568 dB for the previous 358-station row and 73.267 dB for the complete 19×19 grid with 361 stations. Nearby counts are sharply worse: 350–351 give about 66 dB, 353–354 give 68.795–69.939 dB, 356 gives 66.661 dB, and 359–360 give 70.570–72.844 dB. A local phase refinement at the 352-station count confirmed (0.75,0.90), with (0.75,0.91) tied by an effectively equivalent rounded station set. On the 12-variant fresh audit, the validation-selected 352-station row reaches 73.925 dB with exact support and a 1.21 ms sparse solve, improving the previous 358-station fresh frontier of 73.379 dB while using fewer observations. This is now the cleanest sparse shallow-water operator-learning result in the manuscript: progress came not from more data, saliency replacement, or threshold tuning, but from validation-selected Fourier-compatible station decimation.
We added an exact --sensor\_counts option and widened the decimation sweep to counts 320–380 at the same phase. This exposed an even sharper sparse resonance at 334 stations, i.e. 334/4096=0.08154296875 observations per training trajectory. The 334-station row reaches 75.851 dB on the validation variants with exact (7,0,0) support, while nearby counts again fluctuate strongly: 328–330 sit near 71–72 dB, 332 activates false support, 335 gives 71.424 dB, 336 activates two false positives, and the entire 362–380 side-20 band stays below 68 dB except for false-support rows. A fresh 12-variant audit of the locked 334-station row reaches 75.841 dB with exact support and a 1.18 ms sparse solve. Local phase refinement at count 334 again selects (0.75,0.90), with (0.75,0.91) tied by the rounded station set. This supersedes the 352-station checkpoint: the validation-selected operator now improves the broad fresh audit by 2.462 dB over the previous 358-station frontier while using 6.7% fewer observations.
We also checked whether the same phase contains an even lower-count resonance. A validation sweep over exact station counts 220–319 at phase (0.75,0.90) was negative. The best row in that band is 318 stations at only 68.349 dB, and most rows sit near 55–66 dB, with occasional false-support failures such as counts 232, 235, 240, 297, and 306. Thus the current sparse optimum is not simply ``as few stations as possible.'' For this Fourier dictionary and weak-form window, the useful resonance appears to start near the high end of the side-19 decimation family, with 334 stations as the current validated minimum-quality sweet spot.
The next audit asked whether the 334-station geometry was limited by the Fourier assimilation basis itself. Holding the training variants, validation variants, station count, phase (0.75,0.90), window, threshold, and decoy pressure fixed, we swept the reconstruction basis from modes 2 through 6. Mode 2 still selected exact support but underfit the observed window and biased the nonlinear coefficient, reaching only 52.464 dB validation PSNR. The previous mode-3 row reached 75.851 dB validation and 75.841 dB on the locked 12-variant fresh audit. Mode 4 gives a small but clean improvement: it reaches 76.474 dB on validation, exact (7,0,0) support, and a 1.23 ms sparse solve; the disjoint 12-variant audit reaches 76.444 dB with exact support and a 1.30 ms solve. Modes 5 and 6 regress to 75.486 and 74.903 dB, respectively, despite exact support. Repeating the exact-count sweep 320–380 under mode 4 again selects 334 stations; count 352 rises to 74.894 dB but stays below the 334-station row, and the side-20 band remains weaker or false-support. The current interpretation is therefore a two-axis resonance: the best sparse operator learner is not maximal observation count or maximal basis bandwidth, but the mode-4, 334-station, phase-balanced Fourier geometry.
We then revisited the weak temporal projection itself. The committed rows used a window of 7 frames, stride 2, and test-mode radius 3. At the locked mode-4, 334-station geometry, shortening the window is a major coefficient-calibration lever. With stride 2 and test-mode radius 3, validation PSNR rises from 76.474 dB at window 7 to 76.131 dB at window 6, 78.346 dB at window 5, 78.244 dB at window 4, 79.204 dB at window 3, and 79.582 dB at window 2, all with exact (7,0,0) support. The projection radius is sharp: at window 3, radius 2 collapses to 59.003 dB and radius 4 drops to 69.496 dB; at window 5, radius 2 and 4 give 59.126 and 70.595 dB. The lower boundary and stride controls also reject a trivial ``shorter is always better'' rule: window 1 gives 78.707 dB, while window 2 with stride 1 and 3 gives 79.189 and 78.717 dB. The selected weak setting is therefore window 2, stride 2, radius 3. On the locked 12-variant fresh audit this reaches 79.509 dB, exact support, and a 1.26 ms sparse solve, improving the previous mode-4 fresh frontier by 3.065 dB and the old 358-station frontier by 6.130 dB. The inferred vector (H,β,g,f,r,ν,μh)=(0.99980,0.72348,0.86033,0.58257,0.06370,0.005998,0.004022) is now close to the true operator across all seven groups. The lesson is precise: long weak windows were smearing the sparse-assimilated trajectory balance; a short, non-overdense weak window better matches the local truncation and assimilation error scale.
Re-sweeping station counts after the weak-window correction shows that this is not a new count-search problem. With mode 4, window 2, stride 2, radius 3, phase (0.75,0.90), and the same validation variants, counts 300–360 again select 334 stations at 79.582 dB. The low-count side remains far below the frontier: the best sub-320 count is 318 at 68.956 dB. The old 352-station checkpoint improves from 74.894 dB to 76.376 dB under the shorter weak window, and the previous 358-station phase row rises to 74.607 dB, but both remain clearly below 334. Secondary bumps such as count 328 at 72.560 dB and count 343 at 73.083 dB do not change the ordering. Thus the weak-window correction improves coefficient calibration at fixed geometry, while the station-count resonance itself remains locked.
We also ported exact count and phase-station support into the sparse weak-form library-discovery runner, then audited the harder raw and grouped libraries at the locked geometry. This uses the same mode-4, 334-station, phase (0.75,0.90), window-2 weak system, but asks the selector to reject nuisance terms rather than assuming the seven physical groups. The raw typed weak solve at threshold 5×10−7 and decoy penalty 5 recovers all 13 physical columns with zero decoys, active support (13,0,0), in 0.54 ms. Its grouped typed counterpart recovers all seven shared physical groups with zero decoys, active support (7,0,0), in 0.56 ms. Raising the threshold to 5×10−6 prunes weak true terms, and thresholds 5×10−5 or larger over-prune the operator. The forecast values in this single-trajectory discovery runner remain near 20 dB because the rollout starts from the sparse-assimilated state whose reconstruction PSNR is only about 21 dB; this row should therefore be read as a support-identifiability result, not as the high-quality multitrajectory forecast frontier above.
We then inserted the same raw-vs-grouped choice into the high-quality multitrajectory forecast protocol. This separates support recovery from long-horizon transfer. At the locked mode-4, 334-station, phase (0.75,0.90), window-2, stride-2, radius-3 setting, the raw typed library again recovers exact support (13,0,0) at threshold 5×10−7, but its two-variant validation forecast is only 73.676 dB. The grouped typed operator recovers (7,0,0) and reproduces the 79.582 dB frontier. Thus the seven-group collapse is not merely a reporting convention. Enforcing the shared physical coefficients (β,g,f,r,ν) across their equation-specific columns is a strong structural regularizer for rollout quality, even when the ungrouped raw support is exactly correct.
We also repeated the station and weak-projection controls under the improved window-2 setting. A 20-policy local phase grid with x∈{0.65,0.70,0.75,0.80,0.85} and y∈{0.80,0.85,0.90,0.95} again selects phase (0.75,0.90) at 79.582 dB; the nearest strong neighbor is (0.75,0.80) at 79.179 dB, while many exact-support phases fall into the 60–69 dB range. Phase-preserving score replacement remains decisively negative at the locked count: replacing only 2–15% of the (0.75,0.90) lattice by gradient, variance, hybrid, or leverage stations never improves the frontier. The best mixed row is gradient-2% at 66.530 dB, and larger gradient replacements introduce decoys or drop below 54 dB; variance, hybrid, and leverage replacements are similarly weaker. A targeted low-amplitude train-weighting control also loses: weights (1,1,0.5,1,1,1,2,1) on \base, roll (7,11), scale 1.25, roll (13,5), roll (5,17), roll (2,19), scale 0.90, scale 1.10\ reach 79.153 dB, below the equal-weight row. Finally, the radius sweep closes the weak-test-mode axis for window 2: radius 1 diverges, radius 2 gives 58.987 dB, radius 3 gives 79.582 dB, radius 4 gives 69.786 dB, and radius 5 gives 65.683 dB. The selected radius is therefore not arbitrary; it is the unique tested projection scale that balances sparse-assimilation bias and weak-form identifiability.
As a no-backprop control-theory follow-up, we added optional sparse-sensor rollout calibration to the multitrajectory runner via --theta\_calibration\_steps. Starting from the weak-form coefficient vector, the routine performs coordinate search over the seven physical parameters and accepts changes that reduce a held-out training-tail loss measured only at the observed sparse station locations. This is biologically and control-theoretically plausible in the sense that it uses forward rollouts and local observation residuals, not reverse-mode differentiation or full-field labels. On the current frontier, however, it is a hard negative result. With steps 1%, 0.3%, and 0.1%, the sparse sensor-tail loss decreases only from 0.2092457 to 0.2092395, while validation PSNR collapses from 79.582 dB to 57.751 dB. The calibrated vector moves to (0.99780,0.73364,0.85514,0.58257,0.06281,0.005914,0.003970). The conclusion is useful: direct sparse-sensor replay is an overfitting objective for this problem. The weak-form grouped solve generalizes because it optimizes an operator balance, not because it best replays a short sparse observation tail.
We then attacked the same bottleneck from the reconstruction and weak-row side. First, temporal smoothing of the assimilated training sequence before weak integration is negative: centered binomial-3, binomial-5, and box-3 smoothing reduce validation to 54.490, 50.174, and 51.966 dB. The smoothed sequences have nearly the same reconstruction PSNR as the raw assimilated fields, but their weak coefficients are biased. Second, concatenating additional weak views is also negative. At the locked mode-4, 334-station, phase (0.75,0.90) geometry, a 50/50 window-2/window-3 weak-row mixture reaches only 79.341 dB, an 80/20 mixture reaches 79.470 dB, and a 90.9/9.1 window-2/window-4 mixture reaches 79.344 dB. A lightly weighted radius-4 test-mode view also degrades the diagnostic validation run. Thus the selected window-2, radius-3 weak projection is not merely underdetermined; adding nearby valid projections injects biased rows rather than averaging out the reconstruction error.
The next reconstruction audit clarifies the direction. Raising the Fourier assimilation dictionary from mode 4 to modes 5, 6, and 7 lowers the training-window reconstruction PSNR and drops validation to 76.850, 75.503, and 71.546 dB, even though exact (7,0,0) support is still recovered. In contrast, applying a rectangular low-pass filter to the mode-4 assimilated fields before forming the weak rows gives a tiny but clean improvement when the keep radius is 3. Filter radius 2 underfits and reaches only 76.787 dB; radius 4 is effectively the unfiltered baseline at 79.583 dB. Radius 3 reaches 79.587 dB on validation and 79.519 dB on the 12-variant fresh audit with exact support and a 1.19 ms sparse solve. The coefficient vector is (0.99980,0.72336,0.86033,0.58258,0.06371,0.005998,0.004022). This supersedes the unfiltered same-ridge fresh row at 79.511 dB and the previous unfiltered frontier at 79.509 dB, but only by about 0.01 dB. The scientific value is therefore diagnostic rather than headline: the weak learner is now limited by aliasing and sparse-reconstruction bias at the operator-balance level.
We then made this anti-aliasing more operator-specific. The runner now supports --weak\_filter\_application and --assim\_spatial\_filter\_shell\_weight. The selected row keeps the weak target and all linear operator columns on the unfiltered mode-4 assimilated sequence, replaces only the true quadratic flux/advection columns by their mode-3 low-pass values, and retains the first excluded Fourier shell with weight 0.05. This is a term-local weak-form filter: it does not smooth the rollout state, does not alter the sparse station observations, and does not use held-out future fields. On the validation variants this nonlinear\_terms row reaches 79.588 dB with exact (7,0,0) support; on the locked 12-variant fresh audit it reaches 79.521 dB, exact support, and a 1.24 ms sparse solve, with coefficients (0.99980,0.72336,0.86033,0.58258,0.06371,0.005998,0.004022). Filtering nonlinear terms and decoys gives the same validation score; filtering all columns with the same shell reaches only 79.519 dB fresh. Finally, the same nonlinear-only filter does not make higher assimilation bandwidth safe: mode 5 and mode 6 validation runs fall to 76.693 and 75.301 dB despite exact support. The interpretation is sharper: the main aliasing source is the quadratic product library, but high-bandwidth assimilated states also bias the linear weak balance and coefficient calibration. The next material step should therefore be an operator-aware station or anti-aliasing objective that keeps the stable mode-3 nonlinear products while preserving the useful mode-4 linear state information.
A decomposition audit then checked whether the filter should act after products are formed or only on one nonlinear physical channel. Product-column filtering is too late: filtering the already-formed nonlinear product columns reaches only about 79.583 dB validation, essentially the unfiltered row. Momentum-advection-only state filtering is also negative at 79.581 dB. Mass-flux-only state filtering is the strongest validation row, reaching 79.591 dB when only the continuity-equation β column is formed from the mode-3 filtered state. However, this validation gain does not transfer: the hard mass-flux row reaches 79.520 dB on the 12-variant fresh audit, and shell weights 0.05 and 0.10 also reach only 79.520 dB. Thus the fresh frontier remains the all-nonlinear state-prefiltered row above. The useful conclusion is methodological: two validation variants can over-rank continuity-specific anti-aliasing, so the next selector must use a broader validation design or a physically derived anti-aliasing criterion rather than a two-trajectory validation score alone.
We therefore widened the selector itself before running further filter searches. The broad validation set contains eight additional roll and amplitude variants, roll9\_27, roll27\_9, roll15\_29, roll29\_15, scale070, scale150, scale040, and scale190, while the eight sparse-observed training variants remain fixed. Under this V8 selector, the unfiltered row scores 79.513 dB, all-column keep-3 filtering scores 79.521 dB, all nonlinear-state filtering scores 79.523 dB, the all-nonlinear shell-0.05 row scores 79.523 dB, mass-flux-only hard filtering scores 79.522 dB, and mass-flux-only shell-0.05 filtering scores 79.522 dB, all with exact (7,0,0) support. The broader selector therefore chooses the same all-nonlinear shell-0.05 rule that transferred best to the 12-variant fresh audit, rather than the continuity-only row that won the narrow two-variant validation. This is the current robust selection rule: keep the mode-4 assimilated state for targets and linear operators, form all true nonlinear state products from the mode-3 state with a 0.05 first-shell taper, and validate across both phase rolls and amplitude extremes.
Two follow-up geometry audits closed the obvious remaining local axes. Re-sweeping exact station counts 318,328,334,343,352,358 under the V8 selector and the nonlinear shell rule again selects 334 stations. The tested rows all recover exact support, but coefficient calibration is sharply count-dependent: the V8 means are 68.898, 71.748, 79.523, 73.055, 75.713, and 74.287 dB. A nearby phase grid at count 334 is even sharper. The nine policies lattice\_phase70\_85 through lattice\_phase80\_95 all recover exact support, but only lattice\_phase75\_90 reaches the frontier. The other phase rows range from 58.980 to 68.973 dB. Thus the station objective is not ``recover the seven groups''; it is to preserve the Fourier-compatible weak-row geometry that calibrates the nonlinear and damping coefficients.
We also tested whether the weaker high-amplitude validation rows could be fixed by adding amplitude-extreme trajectories to the training weak solve. On a disjoint V8b holdout, the fixed eight-trajectory training set scores 79.536 dB. Adding scale070 and scale150 with equal weights drops to 76.522 dB; giving those added variants only 0.1 weight still drops to 79.317 dB. A new whole-trajectory --train\_block\_normalization control was added to test block-level scaling without row-wise physics destruction. Target-RMS block normalization on the same augmented set activates all 11 decoys and drops to 65.019 dB, while the dense true-support coefficient row is still only 77.358 dB. The result is negative but useful: broad amplitude coverage is not automatically a regularizer for this sparse weak system. The next structural lever should be an operator-domain anti-aliasing or station-design criterion, not more amplitude variants or block/row rescaling.
The first useful station-design criterion came from the operator domain rather than from reconstruction or row conditioning. We added --operator\_reference\_mode dense\_train, which fits a dense weak-form reference operator on the training window only and scores each sparse station geometry by the relative distance between its sparse-derived θ and this dense-training θ. The dense-training reference for the current eight-trajectory set is (1.00037,0.71970,0.86033,0.58005,0.06454,0.006010,0.004011), within 4.59×10−4 relative error of the known simulator coefficients. This training-only objective selects lattice\_phase75\_90 in the nine-phase grid and selects the same count-334, phase-75_90 row in a local 3×3 count/phase grid. The selected row has operator-reference error 0.00283 and V8 PSNR 79.523 dB; the next closest local candidate is count 352, phase 75_90 with error 0.01088 and 75.713 dB. This is not a new quality frontier, but it is a more principled path to station placement: design sparse observation geometry to reproduce the dense training operator, then evaluate transfer on disjoint futures.
We then stress-tested that rule as an actual selector over a broader 30-candidate grid: station counts 318,328,334,343,352,358 crossed with phases 70\_90, 75\_85, 75\_90, 75\_95, and 80\_90. On the V8 holdout, both held-out PSNR and the training-only dense-operator objective select count 334, phase 75\_90, with 79.523 dB and operator-reference error 0.00283. The top held-out alternatives are count 352, phase 75\_90 at 75.713 dB and count 358, phase 75\_90 at 74.287 dB. Repeating the same grid on a disjoint V8b holdout gives the same top row: 79.536 dB for count 334, phase 75\_90; the next held-out rows are count 352, phase 75\_90 at 75.834 dB and count 358, phase 75\_90 at 74.347 dB. The caveat is that operator-reference error is a successful top-1 selector here, not yet a calibrated total ordering: for example, count 358, phase 75\_85 has the second-lowest reference error but only about 68 dB. The next station-design objective should therefore keep dense-operator matching as the primary constraint, but add a training-only stability term that penalizes geometries whose coefficient match is fragile under small phase, count, or trajectory perturbations.
A follow-up diagnostic showed that the failure mode is not local instability but coefficient scaling. The original dense-operator score is an absolute relative L2 distance, so it is dominated by the large coefficients (H,β,g,f) and can underweight small but rollout-sensitive coefficients such as damping and viscosity. We therefore added --operator\_reference\_metric with absolute\_l2, relative\_l2, and relative\_linf options. The relative\_linf metric scores the maximum coefficient-wise relative error to the dense training operator. On the same V8 30-candidate grid, relative\_linf again selects count 334, phase 75\_90, with score 0.01294 and 79.523 dB. The row count 358, phase 75\_85, which was second-best under absolute L2, is demoted to score 0.04446 because its small-coefficient errors are large. This makes the selector more physically balanced: it still does not perfectly rank all held-out PSNR values, but it removes the most obvious scale artifact in dense-operator matching.
Finally, a broader relative-ℓ∞ selector search crossed eight counts, 300,318,328,334,343,352,358,372, with the full 3×3 phase neighborhood from 70\_85 through 80\_95. Across all 72 rows, both held-out PSNR and the balanced dense-operator score again select count 334, phase 75\_90, with 79.523 dB and score 0.01294. The best non-frontier held-out row is count 352, phase 75\_90 at 75.713 dB; count 372 never exceeds 67.353 dB. This closes the obvious count/phase lattice search. The next lever should be the weak-form estimator itself: multi-window test functions, coefficient-balanced row construction, or model-bias correction for the nonlinear product columns.
We then held the selected station geometry fixed and tested whether concatenated weak views could reduce estimator bias. They did not. The baseline single view, window 2 and test radius 3, remains 79.523 dB with balanced operator score 0.01294. Equal-weight windows (1,2,3) drop to 79.334 dB; equal-weight windows (2,3,4) drop to 78.918 dB; a conservative 90/10 mix of windows (2,3) still drops to 79.476 dB. Spatial test-mode concatenation is worse: modes (2,3) drop to 71.673 dB and modes (3,4) drop to 74.947 dB. This also validates the balanced score: the (1,2,3) row has a smaller absolute operator distance than the baseline, but worse balanced relative error and worse rollout. Thus the useful weak-form view is now sharply identified as window 2, test radius 3; the remaining bias is not fixed by naive multi-scale row concatenation.
Finally, we tested whether the remaining gap is fundamental or merely a calibration-objective problem. A diagnostic forward-only coordinate search was run from the current sparse coefficient vector, using full-field training-tail rollout residuals over the eight training variants and no reverse-mode differentiation. This is not a sparse-only claim, because the calibration objective uses full-field training tails; it is a boundary diagnostic. The result is decisive: the sparse vector (0.99980,0.72336,0.86033,0.58258,0.06371,0.005998,0.004022) gives 79.523 dB on the disjoint V8 holdout, while the tracked full-field calibration runner nonlinear\_shallow\_water\_theta\_calibration\_diagnostic.py reaches (1.00000,0.72076,0.85999,0.58002,0.06489,0.006001,0.004000), 104.088 dB on V8, and 104.319 dB on the disjoint V8b holdout. The accepted moves mainly correct β, f, r, H, and μh toward the simulator values. In contrast, rolling out the dense-training weak-form reference vector, although closer to the simulator coefficients, gives only 78.583 dB. Therefore the 79.5 dB frontier is not a support, station-count, or weak-window ceiling; it is a coefficient-calibration objective ceiling.
The follow-up calibration audit isolates what information is missing. If the calibration starts from the exact full training-tail state but observes only the 334 sparse station values over the tail, the same forward-only coordinate search reaches the same vector and the same V8/V8b values, 104.088/104.319 dB. Sparse station values therefore contain enough coefficient information once the hidden state at the calibration start is correct. The failure is the state anchor: using the sparse-assimilated tail-start state makes even a full-field tail loss collapse to 55.875 dB, and optimizing either the sparse station tail, the assimilated pseudo-field tail, or the weak residual of the assimilated tail drives the coefficients in the wrong direction. Dense-training weak residual and dense-operator relative-ℓ∞ objectives are also not sufficient by themselves: they reduce their training losses but reach only 78.997 and 79.294 dB on V8.
We then tested sparse or partial state-anchor repairs. A low-mode forward-sensitivity correction fitted from station history is too weak: mode-2, mode-3, and mode-4 anchors improve full-tail PSNR by only 0.00029, 0.00043, and 0.00084 dB, respectively, and the calibrated mode-2 row still collapses to 63.281 dB. A stronger but more privileged dynamics anchor, obtained by propagating the known full initial training state to the calibration split with the sparse OSNR operator, gives a much better training-tail state (75.778 dB tail PSNR and station loss 5.98×10−7). Unconstrained station-tail calibration from this anchor still overfits, dropping to 73.021 dB after 39 accepted moves, but a one-accepted-move trust-region update raises only the damping coefficient, r:0.0637088↦0.0643459, and improves V8/V8b to 80.123/80.139 dB. A two-move variant immediately accepts an overlarge Coriolis correction and falls to 76.831 dB. Thus the first legitimate direction beyond the 79.5 dB row is not blind coordinate search; it is trust-region, state-aware station calibration.
The next selector audit made that trust-region rule explicit in the tracked diagnostic runner. The new one-step selector enumerates all single-coordinate candidates at the configured step sizes, ranks them using training-side objectives only, and evaluates V8/V8b only after the selected move is fixed. Station-tail loss alone is a negative control: it selects a depth decrease, H:0.9998017↦0.9988019, because that gives the largest station-tail improvement, but the held-out scores collapse to 74.987/74.999 dB. Adding a dense training-tail weak-residual consistency gate changes the selected move. Among candidates that improve both the model-anchor station tail and the dense training weak residual, the selector chooses the small Coriolis correction f:0.5825801↦0.5808324; only after that training-only selection do the disjoint diagnostics evaluate to $82.194/82.221$ dB on V8/V8b. This is a larger lift than the previous damping-only trust move, but it is still diagnostic rather than sparse-only, because the auxiliary gate uses dense training-tail weak rows. Multi-accept variants do not improve the result: accepting a subsequent viscosity move gives 82.050/82.079 dB, and an additional dense-operator relative-ℓ2 gate still keeps the best state at the first accepted f move. The useful conclusion is sharper: sparse station replay supplies candidate moves but cannot select them safely by itself; a training-side operator-consistency gate can reject destructive station overfits and select a real coefficient correction. The next publishable step is to replace the dense weak-residual gate with a station-observable or assimilated operator-consistency surrogate while preserving the one-move trust-region discipline.
That replacement attempt is now also informative. We added station-observable selector objectives based on sparse-assimilated weak residuals, station finite-difference RHS residuals, one-step station-increment replay, held-out station splits, and separate higher-mode gate reconstructions. None is a safe substitute for the dense weak gate. The sparse weak gate selects the same destructive H move and gives 74.987/74.999 dB; station-RHS and station-one-step primaries select β:0.7233627↦0.7161291 and slightly reduce V8/V8b to 79.489/79.509 dB; held-out station one-step replay selects g:0.8603310↦0.8517277 and collapses to 55.525/55.525 dB; higher-mode sparse weak gates at modes 5 and 6, with and without temporal smoothing, still admit the destructive H move. A clean sparse weak-solve hyperparameter audit over shell weights 0,0.025,0.075,0.10 and sensor ridges 10−5,10−4 also fails to move the frontier, staying at 79.520–79.523 dB. The only positive replacement so far is structural rather than learned: restrict the selector to momentum coefficients (f,r,ν) and to small trust steps (0.003,0.001). With the same model-anchor station primary, this selects the same f:0.5825801↦0.5808324 move and reaches 82.194/82.221 dB without dense weak rows; on two fresh eight-variant audit lists it moves 79.521/79.524 dB to 82.191/82.194 dB. This is a cleaner diagnostic than the dense-gated selector, but it is still not a sparse-only claim because the selector primary uses the model-propagated full initial training state. When the primary is changed to the fully sparse assimilated station tail, the same small-trust momentum selector chooses ν:0.0059976↦0.0059796 and drops to 78.351/78.360 dB. The next real problem is therefore state anchoring: station observations contain the coefficient signal, but the current sparse-assimilated calibration-start state distorts the selector enough that station-local objectives prefer wrong coefficient directions.
The state-anchor follow-up gives the first clean sparse calibration lift. Instead of using the privileged full initial training state, we propagate only the sparse-assimilated initial frame to the calibration split with the learned sparse OSNR operator, then run the same one-move small-trust momentum selector on station-tail loss. This sparse-model anchor is still a low-PSNR field in full space (mean start/tail PSNR 21.069/20.386 dB over the training variants), but it is dynamically consistent with the learned operator and improves the station-tail selector geometry. With primary objective sparse\_model\_station\_tail, coordinate subset (f,r,ν), and steps (0.003,0.001), the selector chooses f:0.5825801↦0.5808324 and moves V8/V8b from 79.523/79.536 dB to $82.194/82.221$ dB, without dense weak rows, without full-field calibration tails, and without the full-initial-state model anchor. The same selected move transfers on two fresh eight-variant audit lists, 79.521/79.524↦82.191/82.194 dB. The constraints are sharp: allowing the large 0.01 step without validation oversteps to f=0.5767543 and falls to 76.540/76.547 dB; allowing a second small momentum move falls to 80.905/80.925 dB; allowing all coordinates even at small trust selects β:0.7233627↦0.7255328 and falls to 79.213/79.248 dB.
We then added a training-variant split to the sparse-model station objective. The selector can now use sparse\_model\_station\_tail\_fit as the primary loss and require improvement on sparse\_model\_station\_tail\_val. This removes the hand-coded step scale within the momentum subspace: with candidate steps (0.01,0.003,0.001), the validation split rejects the destructive 0.01 overstep and selects the same small f:0.5825801↦0.5808324 move, preserving 82.194/82.221 dB. A second split-validated refinement from that point, using only f and smaller steps (0.001,0.0003,0.0001,0.00003), accepts f:0.5808324↦0.5806582 and improves V8/V8b to $82.278/82.305$ dB; a fresh disjoint two-list audit gives 82.274/82.278 dB. A subsequent (f,r,ν) pass finds no eligible move. However, the split does not learn the coordinate mask: all-coordinate split validation at small trust selects H and collapses to 66.444/66.443 dB, all-coordinate large-trust split validation selects β and gives 78.174/78.270 dB, and adding a sparse weak-residual gate selects g and collapses to 66.058/66.061 dB. The publishable statement is therefore narrow but stronger than before: a sparse station-derived dynamic anchor, a physics-motivated momentum coordinate mask, and a training-split trust rule give a reproducible +2.75 dB held-out lift over the 334-station exact-support frontier. The next missing piece is a station-observable coordinate-confidence rule, likely based on operator-block sensitivity or adjoint/Fisher geometry, that rejects the compensatory H,β,g moves without using dense labels or held-out futures.
Sparse governing-equation discovery for nonlinear shallow water
The preceding experiment assumes that the correct seven-column operator family is known. The more ambitious physics-learning problem is to discover the governing equation itself from a larger nonlinear library. We implemented apps\_industrial\_breakthrough/nonlinear\_shallow\_water\_library\_discovery.py, which expands the residual library to 24 candidate columns: the 13 true equation-specific terms for height and both velocity channels, plus 11 decoys including raw fields, quadratic field products, and misplaced height-gradient terms. OSNR applies a sequential thresholded least-squares solve on only 0.25% of the training residual rows. The discovered support is then collapsed back into the shared physical vector (H,β,g,f,r,ν,μh) and used for held-out future forecasting.
Setting
Support (TP,FP,FN)
Future PSNR
Discovery time
Adam support
Adam PSNR
Adam time
Clean, STLS threshold 5⋅10−4
(13,0,0)
72.7835 dB
1.0432 ms
(10,11,3)
43.8385 dB
1587.90 ms
0.1% noise, band-18
(13,0,0)
69.2189 dB
1.4878 ms
(10,10,3)
33.8352 dB
1632.81 ms
0.2% noise, band-16
(12,0,1)
64.5354 dB
1.6853 ms
(12,0,1)
66.7134 dB
1634.47 ms
Oracle true-library LS, clean
n/a
72.6779 dB
0.2817 ms
n/a
n/a
n/a
Sparse nonlinear shallow-water governing-equation discovery from a 24-term overcomplete library. OSNR recovers the exact clean PDE support and remains robust at 0.1% observation noise. The Adam library baseline uses the same anchor rows and 20,000 optimization steps with an ℓ1 penalty.
Sparse governing-equation discovery for the coupled nonlinear shallow-water core. The discovered PDE recovers the clean future at 72.7835 dB after selecting all true terms and no decoys from the overcomplete library.
This is the first result in the project that begins to look like a genuine PINN/SINDy-class breakthrough rather than only a fast solver. In the clean case, OSNR selects every true nonlinear PDE term and rejects every decoy, then slightly outperforms the oracle true-library forecast. The comparable Adam library run is not just slower; after 20,000 gradient steps it keeps 11 false-positive decoys, misses 3 true terms, and loses almost 29 dB of forecast quality. The wall-clock ratio for discovery is about 1522× in favor of OSNR. At 0.1% observation noise, the same sparse support is still recovered exactly and the speed ratio remains above 1000×. At 0.2% noise, spectral denoising preserves zero false positives but one weak damping term drops below threshold; the optimized Adam library catches up in forecast quality only after paying the full 1.6 s optimization cost. This identifies the next hard technical layer: noise-aware thresholding or group sparsity for weak physical terms, not larger neural networks.
Canonical PDE discovery: Burgers and Kuramoto–Sivashinsky
To reduce the risk that the shallow-water result is viewed as a repository-specific construction, we added apps\_industrial\_breakthrough/canonical\_pde\_discovery\_benchmark.py. It evaluates two standard equation-discovery controls: viscous Burgers, ut=−uux+νuxx, and the chaotic Kuramoto–Sivashinsky equation, ut=−uux−uxx−uxxxx. Both are discovered from the same 10-term library {u,u2,ux,uux,u2ux,uxx,uuxx,u3,uxxx,uxxxx}, using only 128 sampled residual rows. The Kuramoto–Sivashinsky trajectory is generated with the standard ETDRK4 spectral integrator; the discovery stage is independent of that generator and sees only the sampled field values.
Equation
OSNR support
OSNR PSNR
OSNR time
Adam support
Adam time
Burgers, clean
(2,0,0)
43.1187 dB
0.3690 ms
(2,0,0)
1265.03 ms
Kuramoto–Sivashinsky, clean
(3,0,0)
12.2751 dB
0.0972 ms
(3,0,0)
1244.63 ms
Canonical PDE discovery controls. Support is reported as (TP,FP,FN) against the known governing equation. Adam uses the same sampled residual rows and 20,000ℓ1-regularized optimization steps.
Canonical Burgers and Kuramoto–Sivashinsky discovery from a shared overcomplete library. Both OSNR and Adam recover the clean support, but OSNR does it via a millisecond-scale sparse solve rather than a long gradient-optimization loop.
The canonical control confirms that the sparse operator-discovery mechanism is not confined to the shallow-water generator. On Burgers, OSNR recovers exactly {uux,uxx} and is about 3428× faster than the Adam library optimizer. On Kuramoto–Sivashinsky, OSNR recovers exactly {uux,uxx,uxxxx} and is about 12809× faster. The chaotic KS forecast PSNR is naturally low over the held-out horizon because small coefficient and phase errors amplify quickly; for this control, support recovery and coefficient recovery are the meaningful scientific-discovery metrics. At 0.1% direct observation noise, both OSNR and Adam pick decoys under simple pointwise derivative regression, which confirms that the next publishable robustness layer must be weak-form or group-sparse denoised discovery rather than more gradient steps.
Weak-form canonical PDE discovery under observation noise
The pointwise canonical experiment exposes the correct failure mode: differentiating noisy data directly creates spurious high-frequency library columns. We therefore implemented the weak-form variant in apps\_industrial\_breakthrough/canonical\_pde\_weakform\_discovery.py. Instead of regressing ut at individual grid points, OSNR integrates the PDE over temporal windows and projects the resulting balance onto low-frequency spatial Fourier test functions: u(tb)−u(ta)=∫tatbΘ(u(t))ξdt. This is the operator-spline analogue of weak-form PDE discovery: the test functions absorb observation noise before sparse regression sees the library.
Equation/noise
OSNR support
OSNR PSNR
OSNR time
Adam support
Adam PSNR
Adam time
Burgers, 0.1%
(2,0,0)
49.8283 dB
0.4965 ms
(2,0,0)
47.9135 dB
1560.30 ms
Burgers, 0.5%
(2,0,0)
49.3129 dB
0.2580 ms
(2,0,0)
40.7179 dB
1537.55 ms
Burgers, 1.0%
(2,0,0)
48.8199 dB
0.2103 ms
(2,0,0)
45.2210 dB
1539.72 ms
KS, 0.1%
(3,0,0)
12.6155 dB
0.2870 ms
(3,0,0)
12.7014 dB
1549.60 ms
KS, 0.5%
(3,0,0)
12.5214 dB
0.3026 ms
(3,0,0)
11.9874 dB
1560.78 ms
KS, 1.0%
(3,0,0)
12.6433 dB
0.3148 ms
(3,0,0)
11.7123 dB
1529.31 ms
Weak-form canonical PDE discovery under observation noise. The support tuple is (TP,FP,FN). Both methods use the same weak rows and library, while Adam uses 20,000ℓ1-regularized optimization steps.
Weak-form Burgers and Kuramoto–Sivashinsky discovery under noisy observations. The weak operator rows recover the correct governing support through 1% noise while avoiding the pointwise derivative decoys observed in the direct regression control.
This is the strongest canonical scientific-ML result so far. The weak-form OSNR solver recovers the exact Burgers and KS support at every tested noise level up to 1%. On Burgers, it is also materially more accurate than Adam in forecast quality, improving the 0.5% noise row by 8.5950 dB. On KS, both methods recover the same support, but OSNR reaches the solution roughly 4,858× to 5,400× faster for the threshold-0.002 profile. This converts the earlier ``mostly speed'' canonical result into a robustness result: integral operator rows eliminate noisy derivative decoys while preserving millisecond-scale discovery.
The next CFD-facing step is a genuinely two-dimensional incompressible flow operator rather than a scalar one-dimensional PDE. We implemented apps\_industrial\_breakthrough/navier\_stokes\_weakform\_discovery.py, which generates periodic vorticity trajectories and discovers the vorticity equation ωt=−uωx−vωy+νΔω,u=ψy,v=−ψx,−Δψ=ω. The discovery stage is not told the two-term equation. It sees a 12-term library containing the advective term, the Laplacian, and ten decoys built from raw vorticity, velocity, first derivatives, and nonlinear products. As in the canonical weak-form experiment, OSNR integrates over time windows and projects the balance onto low-frequency two-dimensional Fourier test functions, ω(tb)−ω(ta)=∫tatbΘ(ω(t),u(t),v(t))dt, then applies a scaled sequential thresholded solve. The Adam control optimizes the same weak rows for 20,000ℓ1-regularized steps. A first high-resolution attempt at 1602 with the coarse timestep became numerically unstable, so the retained scaled run tightens the timestep and reference substepping rather than hiding the CFL boundary.
Grid/noise
OSNR support
OSNR coefficients (cadv,ν)
OSNR PSNR
OSNR time
Adam support
Adam time
962, clean
(2,0,0)
(0.9999865,0.0015000)
135.3978 dB
0.6477 ms
(2,11,0)
1958.15 ms
962, 0.1%
(2,0,0)
(0.9999592,0.0014999)
121.0461 dB
0.5487 ms
(2,11,0)
2013.22 ms
962, 0.5%
(2,0,0)
(1.0007806,0.0015010)
99.6485 dB
0.3822 ms
(2,8,0)
1966.74 ms
1282, clean
(2,0,0)
(0.9999944,0.0015000)
146.9002 dB
0.7231 ms
(2,11,0)
2322.75 ms
1282, 0.1%
(2,0,0)
(1.0000714,0.0015000)
126.0550 dB
0.4781 ms
(2,11,0)
2324.12 ms
Weak-form two-dimensional Navier–Stokes vorticity discovery. The support tuple is (TP,FP,FN) relative to the two true terms {−uωx−vωy,Δω}. Adam is the same weak-library regression optimized by backpropagation, not a full neural Navier–Stokes model.
Scaled 1282 Navier–Stokes weak-form discovery. OSNR identifies the exact advection–diffusion vorticity operator and forecasts the held-out future from the recovered coefficients, while the gradient-optimized sparse regression admits many decoys.
This result is the first high-impact two-dimensional CFD discovery benchmark in the repository. It is still a controlled periodic vorticity system, not a direct DeepMind weather-model comparison. The important claim is narrower and stronger: for a known candidate library and noisy observations, weak-form OSNR recovers the exact incompressible Navier–Stokes vorticity support and coefficients at 962 and 1282 resolution, remains stable through 0.5% noise in the 962 run, and solves the sparse operator identification in less than a millisecond. The Adam control uses the same rows and library but remains thousands of times slower and selects many decoy terms. The next step is therefore to move from periodic vorticity discovery to partial-observation assimilation and forced/stochastic Navier–Stokes, where the sparse innovation machinery can be tested on genuinely unknown forcing rather than only coefficient recovery.
We then tested the more realistic assimilation problem in apps\_industrial\_breakthrough/navier\_stokes\_sparse\_forcing\_assimilation.py. The governing operator is assumed known, but the trajectory is driven by hidden sparse spatiotemporal forcing: ωt=−uωx−vωy+νΔω+f(t,x,y),f(t,x,y)=r=1∑Rarφt(t−τr)φx(x−xr,y−yr). This is closer to weather and flow data assimilation than coefficient discovery: the unknowns are localized forcing events, not just scalar PDE coefficients. OSNR first denoises the observed trajectory spectrally, applies the known Navier–Stokes operator to form the innovation residual f(t+21)=Δtω(t+Δt)−ω(t)−[−uωx−vωy+νΔω]t+1/2, then runs a weak three-dimensional matched atom detector: the residual is convolved with the separable spatial–temporal Gaussian test function associated with the forcing atom before non-maximum suppression. A final small least-squares amplitude debiasing step fits the active atoms to the residual. The comparison baselines are an unforced rollout and a smooth low-pass residual forcing field.
Run
Events
Noise
Event error
Recall@3
Traj. PSNR
Assimilation time
962×121
48
0.2%
0.7416
97.92%
81.9812 dB
225.76 ms
1282×161
96
0.2%
1.2556
94.79%
83.0769 dB
851.00 ms
1282×161
96
0.5%
2.6500
87.50%
79.5815 dB
883.36 ms
1282×161
96
1.0%
5.2118
73.96%
74.5790 dB
905.40 ms
1922×201
160
1.0%
5.0134
78.75%
79.1944 dB
4086.68 ms
1922×201
240
1.0%
3.8576
82.50%
77.9266 dB
6200.45 ms
1922×201
240
2.0%
4.6710
77.50%
70.9027 dB
6158.36 ms
1922×201
240
3.0%
4.8846
77.08%
67.3531 dB
6367.19 ms
1922×201
240
5.0%
8.1801
60.83%
61.7619 dB
6334.94 ms
Forced two-dimensional Navier–Stokes sparse innovation assimilation. Event error is measured in joint (t,x,y) grid units against the injected forcing centers. The weak three-dimensional atom detector keeps the sparse residual useful through 5.0% observation noise and improves trajectory reconstruction over both unforced and low-pass residual rollouts.
Forced Navier–Stokes sparse innovation assimilation on the 1922×201 stress run with 240 hidden events at 3.0% observation noise. The recovered weak-form sparse atoms preserve the assimilated trajectory substantially better than an unforced model and better than a smooth low-pass residual forcing field.
The 1282 dense-event case recovers 94.79% of the hidden forcing events within three grid units at 0.2% noise and reaches 83.0769 dB trajectory PSNR, compared with 77.8847 dB for low-pass forcing and 71.1900 dB for the unforced operator. After replacing the point detector with the weak three-dimensional matched atom score, the same dense case remains useful at 0.5% and 1.0% noise: at 1.0% noise it recovers 73.96% of events within three grid units and reaches 74.5790 dB, compared with 69.8860 dB for the low-pass residual and 67.8846 dB for the unforced operator. The larger 1922×201 stress run with 240 hidden events at 1.0% noise recovers 82.50% of events within three grid units and reaches 77.9266 dB, compared with 72.5998 dB for the low-pass residual and 70.5923 dB for the unforced model. With scale-adjusted weak atom smoothing, the same 240-event stress run remains ahead of the low-pass residual at 2.0%, 3.0%, and 5.0% observation noise. At 3.0% noise, it recovers 77.08% of forcing events and improves trajectory quality by 3.8774 dB over low-pass; at 5.0% noise, it still recovers 60.83% of events and keeps a 2.0572 dB trajectory advantage. This is a meaningful step beyond coefficient discovery: sparse OSNR innovations can assimilate unknown localized forcing in a nonlinear two-dimensional flow under noisy observations. The remaining bottleneck is now external benchmark standardization and heavy-overlap amplitude calibration, not basic sparse forcing recovery.
To compare against a trained coordinate-field alternative, we added apps\_industrial\_breakthrough/navier\_stokes\_neural\_forcing\_baseline.py. The neural baseline receives the same innovation residual as OSNR and fits a Fourier-feature MLP gθ(t,x,y) with AdamW, using a sample distribution biased toward high residual magnitude so that sparse events are not hidden by uniform sampling. The learned forcing is then rolled through the same Navier–Stokes solver. On the 1282×161 case with 96 hidden events and 3.0% observation noise, a 5-layer, 128-hidden-unit Fourier MLP trained for 5,000 steps reaches only 42.9701 dB trajectory PSNR. OSNR reaches 64.6633 dB from the same residual, while the low-pass residual baseline reaches 60.8749 dB. The recovery step takes 0.9165 s for OSNR versus 52.1354 s for the neural training loop, a measured 56.9× speed advantage before rollout. The conclusion is not that this small MLP is a definitive neural SOTA baseline; rather, it isolates the key mechanism: dense coordinate-field training smooths or misallocates sparse innovations, while the operator-sparse residual directly preserves the hidden forcing events.
Forced Navier–Stokes sparse forcing recovery against a trained Fourier-feature neural residual field. The neural field is trained directly on the same residual observations but remains much less accurate in the downstream flow rollout.
Finally, we converted the high-noise forced-flow result into a replicated stress suite in apps\_industrial\_breakthrough/navier\_stokes\_high\_noise\_suite.py. The protocol repeats the 1922×201, 240-event experiment across three independent random seeds and reports aggregate gains over low-pass residual assimilation. At 3.0% observation noise, OSNR reaches a mean trajectory PSNR of 67.0929 dB versus 63.4397 dB for low-pass, a mean gain of 3.6532 dB with a worst-seed gain of 3.3038 dB. At 5.0% observation noise with the high-noise weak-atom smoothing profile, OSNR reaches 61.6653 dB versus 59.6858 dB for low-pass, a mean gain of 1.9795 dB with a worst-seed gain of 1.7582 dB. This establishes that the forced-flow advantage is not a single-seed artifact.
Noise
Seeds
Mean recall
Mean OSNR
Mean low-pass
Worst gain
3.0%
3
73.33%
67.0929 dB
63.4397 dB
3.3038 dB
5.0%
3
59.44%
61.6653 dB
59.6858 dB
1.7582 dB
Replicated high-noise forced Navier–Stokes sparse innovation assimilation at 1922×201 with 240 hidden forcing events. The reported gain is OSNR trajectory PSNR minus low-pass residual trajectory PSNR.
Replicated high-noise forced-flow suite. OSNR remains ahead of smooth low-pass residual assimilation across all tested seeds at 3.0% noise and after retuning the weak atom scale at 5.0% noise.
To probe whether the effect survives across a broader operating envelope, we added the ``destroyer'' matrix apps\_industrial\_breakthrough/navier\_stokes\_destroyer\_protocol.py. It evaluates 24 forced-flow cases across four grid families (962, 1282, 1602, 1922), event counts from 48 to 240, noise levels from 1.0% to 5.0%, and two random seeds per configuration. OSNR wins 22/24 cases against the low-pass residual baseline, with mean trajectory gain +4.2384 dB and mean event recall 74.24%. At 3.0% noise, OSNR wins all 12/12 cases with mean gain +2.9985 dB. At larger grids (1282, 1602, and 1922), OSNR wins every tested case; the only two losses occur in the smallest 962 grid with the densest 96-event, 5.0% noise setting, where event overlap exceeds the available spatial resolution.
Slice
Cases
Wins
Mean gain
Minimum gain
All destroyer cases
24
22
+4.2384 dB
−0.7136 dB
1.0% noise
4
4
+13.6026 dB
+12.7441 dB
3.0% noise
12
12
+2.9985 dB
+0.7201 dB
5.0% noise
8
6
+1.4161 dB
−0.7136 dB
1282–1922 grids
16
16
+4.3012 dB
+1.7768 dB
Destroyer forced Navier–Stokes sparse assimilation matrix. Gains are OSNR trajectory PSNR minus low-pass residual trajectory PSNR.
Destroyer matrix summary. Bars show mean OSNR gain over low-pass residual forcing for each grid/event/noise configuration; labels show mean event recall.
External PDEBench/FNO weather and fluid assimilation audits
Test 28 stabilizer boundary.
To move beyond internally generated forced-flow fields, we audited the hosted prediction tensors from the external pdebench-fno-audit/fno-predictions artifact. The target case is Test 28, a 5122 incompressible Navier–Stokes vorticity–Poisson benchmark. Each chunk stores FNO vorticity predictions ω^, target vorticity ω, and the published velocity-space nRMSE obtained by solving the Dirichlet Poisson problem −Δψ=ω,v=(∂yψ,−∂xψ), then comparing velocity fields after the first ten input frames. We implemented the same DST-I Poisson recovery in apps\_industrial\_breakthrough/pdebench\_fno\_test28\_stabilizer.py and verified that the recomputed FNO velocity nRMSE on chunk 00 matches the stored metric to within expected numerical drift.
The OSNR diagnostic applies a deterministic spectral-viscosity operator to the FNO vorticity field, ω^osnr=F−1[MK(kx,ky)Fω^], where MK is a compact rectangular frequency support. This is a blind post-processing stabilizer: it does not use held-out targets. We also tested a diagonal spectral transfer calibrated from five samples and applied to the remaining five samples; this is reported only as an assimilation diagnostic because it uses calibration targets.
Profile
Vorticity nRMSE
Velocity nRMSE
FNO artifact, chunk 00
1.390145
0.243828
OSNR spectral viscosity, K=64
0.772164
0.243647
OSNR best vorticity filter, K=16
0.542349
0.244357
Diagonal spectral calibration, held-out
0.497082
0.266637
Oracle replace low modes, K=8
–
0.057275
Oracle replace low modes, K=32
–
0.023451
Oracle replace high modes, K=32
–
0.243831
External PDEBench/FNO Test-28 stabilizer audit on chunk 00. The OSNR spectral operator strongly suppresses vorticity outliers, but the official velocity-space metric changes only marginally and aggressive vorticity filtering can hurt velocity. The calibrated diagonal transfer is not a blind forecast result and is included to expose the metric boundary.
External PDEBench Test-28 vorticity panel for a held-out chunk sample. Spectral OSNR filtering removes large high-frequency FNO vorticity spikes, but this does not automatically translate into a large improvement in the benchmark velocity-space nRMSE.
This external audit is a useful boundary result. It confirms that operator-spline spectral structure can repair raw vorticity instability in a real hosted FNO artifact, but it also prevents an overclaim: the official velocity metric is dominated by low-frequency phase and Poisson-integrated velocity structure. The oracle rows make this precise. Replacing only the lowest spectral modes of the FNO prediction with the target reduces velocity nRMSE from 0.243828 to 0.057275 at K=8 and 0.023451 at K=32, whereas replacing high modes while leaving the low modes unchanged barely moves the metric. A SOTA-facing improvement on this benchmark therefore requires a velocity-aware low-mode dynamics corrector or Poisson-adjoint training objective, with sparse OSNR machinery reserved for high-frequency vorticity stabilization.
We tested three follow-up low-mode correction families on held-out chunk 02 after calibrating on chunks 00–01. A diagonal vorticity-space transfer improved held-out vorticity nRMSE from 1.3716 to 1.0479 but worsened velocity nRMSE from 0.2303 to 0.2491. Direct velocity-space diagonal and mean-residual transfers also worsened the held-out metric, reaching best velocity nRMSE 0.2436. Finally, a compact MPS-trained low-resolution velocity CNN fit the training loss but evaluated at 0.2488 nRMSE on chunk 02. These negative results are informative: the low-mode error is not a stationary spectral bias and not solved by a small framewise image corrector. It is a sample-specific dynamical phase error. The next external benchmark attempt must either learn a genuine temporal low-mode evolution operator from the input history or select a benchmark where sparse/operator innovations, not phase drift, dominate the published metric.
Sparse-station assimilation.
We therefore reframed the external FNO artifacts as sparse-station data assimilation problems, which is closer to operational weather and fluid monitoring. In apps\_industrial\_breakthrough/pdebench\_weather\_sparse\_station\_assimilation.py, the FNO forecast is treated as a neural dynamical prior. At each forecast step, a small number of station observations are used to solve either a closed-form per-channel affine correction or a low-rank DCT residual correction. On the three external Test 29 four-channel forecast configurations, this post-processing consistently improves the FNO forecast. The strongest case, M01\_Eta01, drops from mean nRMSE 0.00461 to 0.000947 with a rank-8 DCT residual fit from 512 stations, a 79.5% relative reduction. The M10\_Eta01 case drops from 0.00551 to 0.00159 (71.1% reduction), while the harder M10\_Eta001 case drops from 0.01204 to 0.00908 (24.6% reduction). We also tested the harder 5122×101 vorticity artifact in apps\_industrial\_breakthrough/pdebench\_vorticity\_sparse\_station\_assimilation.py. Per-step affine station calibration reduces mean full-window vorticity nRMSE from 1.1802 to 0.6872 with 1024 sparse stations, a 41.8% relative reduction, but the velocity-space audit above shows that affine-only vorticity calibration is not the right final correction family.
External forecast artifact
FNO mean nRMSE
Sparse-station OSNR nRMSE
Relative gain
Test 29 M01\_Eta01
0.004611
0.000947
79.5%
Test 29 M10\_Eta01
0.005507
0.001594
71.1%
Test 29 M10\_Eta001
0.012038
0.009076
24.6%
Test 28 vorticity chunk 00
1.180217
0.687232
41.8%
External PDEBench/FNO sparse-station assimilation. The neural forecast is kept as the dynamical prior, while operator/dictionary corrections are solved from sparse observations without backpropagation.
The stronger Test 28 result comes from making the station correction operator-aware in the published metric. The runner apps\_industrial\_breakthrough/pdebench\_vorticity\_dct\_station\_assimilation.py fits, at each forecast step, an affine vorticity calibration followed by a rank-32 DCT residual from 1024 contemporaneous vorticity stations. Unlike the affine-only station correction, this low-mode residual directly repairs the Poisson-integrated velocity structure. On all three available Test 28 chunks, the fixed profile improves both vorticity and the recomputed velocity metric:
Chunk
FNO velocity
DCT-station velocity
FNO vorticity
DCT-station vorticity
00
0.243826
0.191702
1.389743
0.429227
01
0.247869
0.188312
1.421356
0.408249
02
0.230292
0.189971
1.371190
0.435880
Mean
0.240662
0.189995
1.394096
0.424452
External PDEBench/FNO Test 28 DCT station assimilation. Metrics are computed after the first ten input frames. The correction uses 1024 lattice stations, affine calibration, and a rank-32 DCT residual solve at each forecast step. Mean velocity nRMSE drops by 21.0% and mean vorticity nRMSE drops by 69.5% across the three local chunks.
This per-frame result was the first Test 28 improvement in the project that moved the official velocity-space metric rather than only suppressing raw vorticity outliers. The next run added temporal structure to the station adapter. In apps\_industrial\_breakthrough/pdebench\_vorticity\_temporal\_osnr\_station\_rescue.py, chunk 00 selects the profile and chunks 01–02 are held out. For each trajectory, the frozen FNO rollout remains the neural dynamical prior, sparse contemporary vorticity stations are observed at each future frame, and one separable spatiotemporal OSNR residual is fitted over the whole forecast window, ω^(y,x,t)=atω^FNO(y,x,t)+bt+p,q,r∑cpqrϕp(y)ϕq(x)τr(t), where ϕ are spatial DCT atoms and τ is a temporal DCT basis over the post-input frames. The metric is again the Poisson-recovered velocity nRMSE plus raw vorticity nRMSE.
Method on held-out chunks 01–02
Stations/frame
Velocity nRMSE
Gain
Vorticity nRMSE
Gain
Frozen FNO prior
0
0.238631
–
1.396631
–
Previous per-frame DCT, rank 32
1024
0.189142
20.7%
0.422065
69.8%
Temporal OSNR, rank 32×8
2048
0.112993
52.6%
0.349838
75.0%
Temporal OSNR, rank 32×12
4096
0.106656
55.3%
0.328363
76.5%
Temporal OSNR, rank 48×12
4096
0.103622
56.6%
0.327553
76.5%
External PDEBench/FNO Test 28 temporal OSNR station rescue. The final row uses 4096/262144=1.5625% of grid sites per frame and fits one spatiotemporal residual over each forecast trajectory. Hyperparameters are selected on chunk 00 and reported on held-out chunks 01–02.
This became the strongest closed-form Test 28 external FNO result in the workspace. Across all three local chunks, the rank-48×12 temporal adapter reduces mean velocity nRMSE from 0.240406 to 0.103368 and mean vorticity nRMSE from 1.394469 to 0.330206. It is still an assimilation result, not a blind forecast: the method uses contemporary sparse measurements. That distinction is important, but it is also exactly the operational setting where station, buoy, radar, and satellite observations are available and a neural forecast acts as the dynamical prior. The broader mechanism is now clearer than in the first Test 28 audit: OSNR can act as a closed-form, low-rank test-time correction layer on top of a frozen neural PDE forecaster, and the same plug-in idea already transferred from external Darcy sparse assimilation to time-dependent vorticity forecasts.
We then ran the harder SOTA-facing comparison: a trained sparse neural assimilator under the same station protocol. The runner apps\_industrial\_breakthrough/pdebench\_vorticity\_neural\_sparse\_assimilation\_baseline.py trains on chunks 00–01 and reports held-out chunk 02. Its inputs are the frozen FNO vorticity frame, the same-station affine calibration, sparse target and residual maps, a station mask, coordinates, and forecast time. A compact 770,241-parameter U-Net predicts a 1282 residual that is upsampled to the full 5122 grid before both vorticity and Poisson velocity metrics are computed. This is not a no-backprop OSNR result; it is the competent trained sparse neural baseline that the closed-form adapter must be compared against.
Method on held-out chunk 02
Stations/frame
Velocity nRMSE
Gain vs FNO
Vorticity nRMSE
Frozen FNO prior
0
0.230295
–
1.371558
FNO + temporal OSNR
256
0.198508
13.8%
0.551600
Sparse neural assimilator
256
0.073335
68.2%
0.191289
FNO + temporal OSNR
512
0.147680
35.9%
0.461464
Sparse neural assimilator
512
0.031716
86.2%
0.154792
FNO + temporal OSNR
1024
0.134717
41.5%
0.416209
Sparse neural assimilator
1024
0.019973
91.3%
0.139996
Sparse neural assimilator, repeat seed
1024
0.021120
90.8%
0.140282
FNO + temporal OSNR
4096
0.110239
52.1%
0.338437
Sparse neural assimilator
4096
0.015279
93.4%
0.155159
PDEBench Test 28 trained sparse neural assimilation on held-out chunk 02. The 1024-station row uses only 1024/262144=0.390625% of grid sites per frame and is stable under one seed repeat. The neural model is trained with backpropagation and is included as the relevant SOTA-facing sparse-assimilation comparator.
This shifts the Test 28 frontier. The 1024-station neural row reduces velocity error by about 91% against the frozen FNO and by about 85% against the matched closed-form FNO+OSNR row. The 4096-station row reaches the best velocity value, 0.015279, but the 1024 row is the cleaner observation-efficiency result. The fixed FNO-tuned temporal OSNR adapter does not transfer unchanged onto the trained neural prior: at 1024 stations it worsens the neural velocity row from 0.019973 to 0.075423, and at 4096 from 0.015279 to 0.048526. This negative adapter result is useful. Once the neural model has learned the low-frequency station-conditioned correction, the next OSNR layer must be selected specifically for a neural prior, with identity/gating/high-ridge/low-rank candidates or an orthogonalized residual space. Reusing the FNO prior's adapter is not valid.
The active-station follow-up then exposed that the original closed-form gap was mostly geometric. We extended the same runner with centered, interior, space-filling, and rounded phase-lattice station policies while keeping the train/test split, neural architecture, epochs, and Poisson velocity metric fixed. The old lattice includes boundary-heavy samples; a centered or interiorized lattice spends the same budget on Fourier-compatible interior coverage. Table [tab:pdebench-vorticity-active-stations] shows the result on held-out chunk 02.
Station policy
Stations
FNO+OSNR velocity
FNO+OSNR vorticity
Neural velocity
Neural vorticity
Old edge lattice
512
0.147680
0.461464
0.031716
0.154792
Centered lattice
512
0.036161
1.289455
0.042066
1.085325
Space filling
512
0.142717
0.461791
0.031003
0.176815
Old edge lattice
1024
0.134717
0.416209
0.019973
0.139996
Centered lattice
1024
0.020786
1.281390
0.034668
1.077794
Best rounded phase (0.50,0.25)
1024
0.020563
1.285525
0.038600
1.078815
Space filling
1024
0.118799
0.350398
0.017219
0.176978
Old edge lattice
2048
0.119788
0.364925
0.020052
0.147939
Space filling
2048
0.108844
0.336009
0.016828
0.212521
Old edge lattice
4096
0.110239
0.338437
0.015279
0.155159
Space filling
4096
0.106388
0.340152
0.015478
0.221071
PDEBench Test 28 active station-geometry audit on held-out chunk 02. Regular interior station geometry nearly closes the velocity gap between closed-form temporal OSNR and the trained sparse neural assimilator at 1024 stations, but does not repair raw vorticity. Space-filling stations improve the trained neural velocity curve at 1024–2048 stations but worsen vorticity and do not beat the old 4096-station neural velocity frontier.
The best closed-form phase row reduces FNO velocity nRMSE from 0.230295 to 0.020563 using only 1024/262144=0.390625% contemporary station sites per frame. This is an 84.7% reduction relative to the old same-budget edge-lattice OSNR row and is only about 3% worse than the trained 1024-station neural row. It also beats one repeat seed of that neural row (0.021120). The caveat is just as important: the same regular/phase lattice rows leave vorticity near 1.28, so the win is a Poisson-velocity low-mode correction, not full vorticity reconstruction. Space-filling gives the complementary behavior: the neural model reaches new 1024- and 2048-station velocity-efficiency rows, 0.017219 and 0.016828, but with worse vorticity and no improvement over the old 4096-station velocity frontier. A local phase refinement around (0.50,0.25) was locally saturated and quantized by integer-grid rounding. A naive one-mask hybrid that concatenates 50% or 75% centered/interior lattice stations with space-filling fill points was decisively negative: the best 1024-station hybrid neural velocity was only 0.131524, and closed-form hybrid velocity was worse than the frozen FNO. Thus the next Test 28 step should be a vorticity-aware two-geometry or two-head adapter: keep Fourier-compatible interior stations for the low-mode velocity correction, add a separate high-frequency/vorticity residual mechanism, and gate any OSNR residual against the trained neural prior rather than reusing the FNO-prior adapter blindly.
The two-head follow-up made this decomposition explicit. Because the benchmark velocity is recovered by a DST-I Poisson solve, the useful fusion basis is not a generic DCT split but the same sine basis that diagonalizes the reported metric. The runner pdebench\_vorticity\_two\_head\_frequency\_adapter.py keeps the phase-lattice OSNR head for low modes, adds a separate space-filling OSNR vorticity head for high modes, selects the DST cutoff and scalar weight on chunk 01, and reports chunk 02. Table [tab:pdebench-vorticity-dst-two-head] summarizes the resulting Pareto rows.
Method
Station observations/frame
Velocity nRMSE
Vorticity nRMSE
FNO
0
0.230295
1.371558
Phase low head
1024
0.020563
1.285525
Space-filling high head
1024
0.118799
0.350398
DST two-head
1024+1024
0.023612
0.192818
DST two-head, larger high head
1024+2048
0.023712
0.195319
Cached neural comparator
1024
0.017219
0.176978
Corrected neural repeat
1024
0.017219
0.176978
Neural + DST high-pass
1024
0.017060
0.113794
Corrected neural repeat
2048
0.016828
0.212521
Neural + DST high-pass
2048
0.016403
0.097739
Corrected neural repeat
4096
0.015478
0.221071
Neural + DST high-pass
4096
0.014927
0.097201
Corrected neural repeat, seed 20260611
1024
0.020466
0.177627
Neural + DST high-pass, seed 20260611
1024
0.020210
0.095205
PDEBench Test 28 DST two-head and neural-prior high-pass adapters on held-out chunk 02. Separate station counts indicate distinct low-mode and high-mode observation sets for the closed-form rows. Each neural high-pass row uses the same deterministic space-filling station set as its neural comparator, selects the DST split on chunk 01, and improves both reported metrics on chunk 02.
The closed-form DST row is the first Test 28 adapter in the workspace that substantially improves both sides of the earlier closed-form tradeoff: vorticity drops from the standalone space-filling value 0.350398 to 0.192818, while velocity remains close to the phase-lattice value (0.023612 versus 0.020563). Increasing the high-frequency head to 2048 stations does not improve the fused row, so the bottleneck is not simply high-head station count. The neural-prior row is the clean frontier. The first repeat used the correct station geometry but the wrong statistics-sampling seed; after matching the baseline convention, the runner exactly reproduces the cached 1024-station neural row. A validation-selected DST high-pass OSNR layer with cutoff 80 and weight 0.2 then improves held-out velocity from 0.017219 to 0.017060 and raw vorticity from 0.176978 to 0.113794, using the same deterministic space-filling station set. The follow-up ladder strengthens the claim. At 2048 stations, the vorticity-aware selector chooses cutoff 128 and weight 0.1, improving the neural row from 0.016828/0.212521 to 0.016403/0.097739. At 4096 stations, cutoff 128 and weight 0.05 improve 0.015478/0.221071 to 0.014927/0.097201. A second 1024-station seed repeats the pattern, improving 0.020466/0.177627 to 0.020210/0.095205. This reverses the earlier negative neural+OSNR result, where an FNO-tuned residual damaged the trained neural prior: the useful adapter is neural-prior-specific and orthogonalized into the DST high-frequency space.
Gate row
Fit/gate stations
Choices
Neural vel/vort
Gated vel/vort
1024, seed 20260610
768/256
112:10,128:0
0.017219/0.176978
0.016922/0.096622
2048, seed 20260610
1536/512
112:10,128:0
0.016828/0.212521
0.016403/0.099564
4096, seed 20260610
3072/1024
112:6,128:4
0.015478/0.221071
0.014927/0.098693
1024, seed 20260611
768/256
112:10,128:0
0.020466/0.177627
0.020210/0.097617
No-leakage station-heldout gate for the Test 28 neural-prior DST adapter. The gate chooses identity, cutoff 112, or cutoff 128 from held-out station residuals only, then refits the chosen correction on all available stations for the reported field. Identity is never selected in these runs.
The gate table removes the remaining hand-picked-selector weakness. The decision uses only contemporary station values: 25% of stations are withheld from the gate fit, candidate residuals are scored on those stations, and the selected candidate is then refit on the full station set. This standard cross-validation pattern is operationally different from using full-field validation metrics. The gate is conservative relative to the full-field oracle, which would choose cutoff 128 in all four runs; it often chooses cutoff 112 instead. The price is a small vorticity gap versus the oracle, but the no-leakage rows still improve both velocity and vorticity over the trained neural prior at every tested budget and on the second seed. A fit-only ablation that permanently discards the gate stations damages the velocity metric, so the deployable protocol is station-heldout selection followed by all-station refit.
The more aggressive follow-up removes the assimilation advantage entirely. The blind diffusion-refiner runner sees only the first ten true vorticity frames of each held-out Test 28 trajectory at inference time. It never reads the FNO prediction, never observes future stations, and uses the cached FNO tensor only as an evaluation comparator. The model treats forecasting as an iterative refinement-time PDE, uk+1=P(uk+ηFθ(uk,history,τ,k)), where P is either the identity or a DST spectral-viscosity projection. After the first run showed a clean failure mode–excellent vorticity but weaker Poisson velocity–we added a light Poisson-weighted DST coefficient loss and then a differentiable low-mode Poisson-velocity loss.
Blind Test 28 from-initial-history refiner. The OSNR/DST rows do not
use FNO predictions or future observations at inference time. They are trained
from chunks 00/01 and evaluated on held-out chunk 02; the FNO row is a frozen
external comparator.
The direct velocity objective tightened the single-seed frontier from 0.252706 to 0.246451 while preserving the vorticity win. More importantly, seed diversity exposed an ensemble effect rather than a single lucky run. A three-seed uniform average nearly closes the FNO velocity gap, and the five-seed uniform ensemble becomes the first blind from-initial-history Test 28 row in this project to beat the hosted FNO comparator on both reported metrics: velocity improves from 0.230295 to 0.228610 (0.73%), while vorticity drops from 1.371558 to 0.366712 (73.3%). We then froze an explicit validation-locked model-selection protocol in pdebench\_vorticity\_blind\_locked\_ensemble.py: train candidate seeds on chunk 00, select a uniform seed subset on validation chunk 01 with the predeclared score velocity plus 0.02 times vorticity, refit only the selected seeds on chunks 00/01, and evaluate chunk 02 once. That protocol selects seeds 20260614/11/15 and reaches 0.233002 velocity and 0.370837 vorticity on held-out chunk 02. The locked row therefore confirms the blind raw-vorticity result–73.0% lower vorticity than FNO–but it does not yet confirm the exploratory velocity edge, missing FNO velocity by 1.18%. The scientific signal is sharp but narrower than the frontier row: without FNO input or future observations, OSNR/DST refinement has a defensible no-leakage mechanism for collapsing raw vorticity, while the Poisson-velocity win still requires a better locked selector or dynamics model.
We then tested whether the missing velocity margin was simply a selector issue. The weighted locked runner, pdebench\_vorticity\_blind\_locked\_weighted\_ensemble.py, builds validation-only low-resolution prediction quadratics and searches a convex 0.05 simplex grid with the predeclared score mean velocity plus 0.02 times mean vorticity plus 0.25 times velocity p90 plus 0.05 times maximum velocity. Allowing four or five active seeds selects weights (0.20,0.15,0.20,0.45) on seeds 20260614/12/11/15 and reaches 0.231505 velocity, 0.370884 vorticity on chunk 02. Forcing all five seeds active selects weights (0.20,0.05,0.10,0.15,0.50) and reaches the best strict locked velocity row so far: 0.231374 velocity and 0.371718 vorticity. This narrows the locked velocity gap from 1.18% to 0.47% relative to FNO, but still does not beat the FNO velocity comparator. The vorticity result remains stable at roughly 72.9% lower error. Thus the next velocity gain is unlikely to come from selector-only sweeps; it likely requires a stronger cross-mode or horizon-conditioned dynamics model.
The negative controls are also informative. A no-backprop per-mode DST ridge dynamics runner reaches only 0.352961 velocity and 0.555332 vorticity at r128/K64; the validation-selected cross-mode kernel dynamics follow-up is worse still at 0.868662 velocity and 0.887061 vorticity; naive 256-grid scaling reaches only 0.320404 velocity and 0.462010 vorticity after rescaling the Poisson loss; and two 20-epoch schedules improve training loss without improving held-out velocity (0.270341 and 0.267655). Thus the current blind lesson is specific: iterative learned refinement plus light Poisson-weighted OSNR/DST structure and seed diversity are the useful branch, while per-mode closed-form spectral extrapolation, naive cross-mode kernels, naive resolution scaling, and longer single-seed training do not close the single-seed velocity gap.
This is not a claim of beating end-to-end global weather systems such as GraphCast or GenCast. It is a more precise and immediately defensible claim: operator-spline station assimilation can dramatically improve external neural PDE forecasts when a small number of contemporary observations are available. That setting is scientifically meaningful because real forecasting systems already assimilate sparse stations, buoys, sondes, radar, and satellite products; the OSNR contribution is a closed-form, low-rank correction layer that can sit on top of a neural forecaster without retraining it.
We also tested a local compact-kernel variant in apps\_industrial\_breakthrough/pdebench\_weather\_rbf\_station\_assimilation.py. Here the station residual after affine calibration is interpolated through a Gaussian RBF kernel, which better matches spatially local forecast-error structure than a global DCT basis at low station counts. On M01\_Eta01, affine+RBF improves the best result further from 0.000947 to 0.000819, an 82.2% reduction from the FNO baseline. On M10\_Eta01, RBF reaches 0.001640 and already obtains a 52.9% reduction with only 16 stations, while the global DCT correction remains slightly better at the highest station budget. On M10\_Eta001, DCT remains the better choice. The conclusion is algorithmic rather than cosmetic: the assimilation layer should select its correction dictionary from the forecast-error geometry. Smooth global biases favor low-rank DCT; localized residual structure favors compact RBF/operator-spline kernels.
The follow-up hybrid runner apps\_industrial\_breakthrough/pdebench\_weather\_hybrid\_station\_assimilation.py combines the two correction families in one sequential closed-form layer: affine calibration, rank-8 DCT residual fitting, and station-centered Gaussian RBF residual interpolation. A low-budget targeted sweep shows that blindly mixing dictionaries can underperform the best single family because station equations are split across redundant atoms. At the higher 512-station random budget, however, the hybrid layer improves all three external forecasts: M01\_Eta01 drops to 0.000707 (84.7% reduction), M10\_Eta01 drops to 0.001438 (73.9% reduction), and M10\_Eta001 drops to 0.008584 (28.7% reduction). This became the random-station reference for the adaptive placement tests. The design rule is now clearer: use a single compact dictionary at very sparse station counts, but switch to a hybrid global–local operator dictionary once the observation budget is high enough to identify both smooth bias and localized forecast residuals.
Adaptive Test 29 station placement.
The next experiment, apps\_industrial\_breakthrough/pdebench\_weather\_adaptive\_station\_assimilation.py, tests whether station placement can reduce the observation budget. The correction dictionary is kept fixed and only the station policy changes. Forecast-gradient, DCT-leverage, and two-stage pilot-residual policies are not reliable: they oversample high-variation or high-error regions and leave the RBF/DCT normal equations poorly covered. A farthest-point space-filling policy is better because it improves interpolation coverage and dictionary conditioning. The first adaptive sweep nearly matched the 512-random hybrid with 256 stations. A targeted follow-up then retuned only the RBF length scale and showed that 256 space-filling stations are enough to beat the 512-random reference on all three external Test 29 files: M01\_Eta01 reaches 0.000705, M10\_Eta01 reaches 0.001377, and M10\_Eta001 reaches 0.007649. The same policy at 288 stations improves all three again, and 384 stations gives the strongest current Test 29 assimilation results: 0.000646, 0.001267, and 0.007274, respectively. Thus the useful high-information criterion for this external artifact is not local forecast-error magnitude; it is conditioned spatial coverage for the global–local operator dictionary. The result is a concrete observation-efficiency gain: half as many stations now beat the former 512-station hybrid reference, without retraining the FNO and without backpropagating through the assimilation layer.
For reproducibility, this weather result should be read as a frozen-neural-prior plus deterministic assimilation architecture, not as a newly trained neural network. The neural component is the external FNO prediction tensor already present in the hosted audit artifact. The OSNR code never changes the FNO weights and does not run reverse-mode differentiation. Each Test 29 file contains arrays preds and targets with shape 10×128×128×21×4: ten held-out forecast samples, a 1282 spatial grid, 21 forecast frames, and four weather/PDE channels. For each sample n, time index τ, and channel c, the layer treats the FNO forecast pc(x) as a dynamical prior and receives contemporary station observations yc(xi) at a selected set S of grid sites. The correction is the three-block operator AS,r,ℓ(p,y)=RS,ℓ(DS,r(CS(p,y),y),y), where CS is per-channel affine calibration, DS,r is a low-rank DCT residual solve, and RS,ℓ is a compact Gaussian-RBF station residual solve. The first block solves (ac,bc)=arga,bmini∈S∑(apc(xi)+b−yc(xi))2,qc(0)(x)=acpc(x)+bc. The DCT block builds a tensor-product dictionary Φr∈RHW×r2 with normalized atoms ϕky,kx(i,j)=Zky,kx−1cos(Hπ(i+1/2)ky)cos(Wπ(j+1/2)kx),0≤ky,kx<r, and solves one ridge system shared across channels, Bc=(ΦS⊤ΦS+λI)−1ΦS⊤(yc(S)−qc(0)(S)),qc(1)(x)=qc(0)(x)+Φr(x)Bc. The local residual block then places Gaussian atoms at the same station sites, Kij=exp(−2ℓ2∥xi−xj∥22),αc=(KSS+λI)−1(yc(S)−qc(1)(S)), and evaluates qc(2)(x)=qc(1)(x)+KxSαc on the full grid. All reported rows use λ=10−3 and double-precision NumPy linear algebra. The metric is per-sample relative L2 error over all forecast pixels, times, and channels, nRMSEn=∥Yn∥2+10−20∥Y^n−Yn∥2, with the table reporting the mean over the ten samples.
The station-placement rule in the winning rows is also deterministic given the script seed. The grid is the integer lattice {0,…,127}2. For space\_filling, the runner draws one initial grid point from np.random.default\_rng(seed) and then repeatedly adds the point whose squared distance to the already selected set is largest. For each file and hyperparameter row the seed is 20260531+1009ifile+131m+17r+⌊ℓ⌋, because the winning sweeps use only the space\_filling strategy. Here m is the station count, r is the DCT rank, and ℓ is the RBF length scale. The exact runs that produced Table [tab:pdebench-weather-adaptive-stations] are:
The runner uses only numpy and PIL; the adaptive station-placement results in this table were executed on CPU, not on CUDA or MPS. This is acceptable for the claim being made because the layer is a small closed-form assimilation operator rather than a trained network. GPU acceleration would mainly speed up the dense station-kernel solves and repeated grid evaluations; it is not responsible for the reported accuracy.
We next implemented a GPU-ready no-backprop OSNR network runner, apps\_industrial\_breakthrough/pdebench\_weather\_gpu\_osnr\_network.py, to test whether the assimilation block should become a trained operator network rather than a fixed dictionary layer. The architecture keeps the same frozen FNO prior and station-conditioned affine/DCT/RBF heads, but adds a learned residual POD dictionary fitted from training samples by closed-form SVD. No PyTorch autograd graph is constructed; the entire run uses torch.inference\_mode(), and the script selects MPS/CUDA when the environment exposes those backends. The first bounded probe used training samples 0–3, held-out samples 4–5, forecast frames 10–13, rank-8 DCT atoms, RBF length scale 16, and 128 or 256 space-filling stations. The local environment reported cuda\_available=false and mps\_available=false, so this probe executed on CPU despite the GPU-capable code path. The result is informative but not a new headline: on M01\_Eta01, 256 stations with eight learned POD atoms per channel slightly improves the probe nRMSE from 0.000665 to 0.000663; on M10\_Eta01 and M10\_Eta001, the learned atoms slightly worsen the pure DCT/RBF head. Thus the next serious no-backprop weather-network target is not generic residual PCA; it is a trained station policy, operator gate, or dictionary-selection controller that preserves the conditioned spatial coverage responsible for Table [tab:pdebench-weather-adaptive-stations].
A subsequent MPS audit exposed the true high-budget behavior of the same frozen-prior assimilation network. The Apple-MPS backend is visible in the project venv through the runpy invocation path, and the confirmed run reports device=mps, torch\_version=2.12.0, and autograd=disabled. The protocol is deliberately stricter than the earlier all-sample adaptive table: for each Test 29 file, samples 0–3 fit any closed-form residual statistics, while samples 4–9 are held out; all 21 forecast frames and all four channels are evaluated. The best rows use no learned POD atoms, no joint residual atoms, sequential affine+DCT+RBF correction, rank-8 DCT atoms, RBF ridge 5×10−4, and deterministic space-filling stations. The only swept variables in the headline rows are the number of stations and the RBF length scale. At 2048 stations, which is 12.5% of the 1282 grid per frame, the layer reached the first strong external-weather frontier: M01\_Eta01 dropped to 3.19×10−4, M10\_Eta01 to 4.90×10−4, and M10\_Eta001 to 1.40×10−3. We then pushed the dense local operator harder under the same MPS memory guard. At 3072 stations the same architecture reached 2.57×10−4, 3.84×10−4, and 7.53×10−4, respectively. At 4096 stations, or 25% of the 1282 grid per frame, it reached the current external Test 29 frontier: 2.21×10−4, 3.46×10−4, and 6.66×10−4. Relative to the previous 2048-station frontier, the 4096-station rows improve the three held-out means by 30.8%, 29.4%, and 52.5%, respectively. The cost is the expected dense-kernel bottleneck: the 4096 rows use a 320 MB station kernel and take about 185 s per file on MPS.
MPS high-budget sparse-station scaling on external PDEBench/FNO Test 29 weather artifacts. The protocol holds out samples 4–9, evaluates all 21 forecast frames and four channels, and keeps the FNO forecast frozen. All rows use torch.inference\_mode(), MPS tensors, rank-8 DCT correction, no learned residual POD atoms, no backpropagation, RBF ridge 5×10−4, and a deterministic space-filling station policy. The 4096-station frontier gives relative held-out mean reductions of 95.6%, 94.9%, and 94.9%, respectively, but the dense RBF block now costs about 185 s per file on MPS.
This MPS sweep also falsified several tempting station-selection and operator-network variants. Residual-energy, forecast-gradient, observability, neuro-coverage, dynamic-lattice, dynamic-attention, reward-gated, time-local POD, global POD, and joint-atom variants did not beat the best conditioned space-filling baseline on the full held-out protocol. We also tested two stronger control-style placement ideas. A two-stage dynamic pilot-innovation policy first spends a station subset on a topographic pilot, diffuses the observed innovation magnitude through an RBF field, and then places the remaining stations from that inferred regional error signal. On a bounded M10\_Eta001 smoke it underperformed plain space-filling. A training-loss station-seed search is valid and mildly useful: with 512 stations on the hard held-out file it improves the best mean from 0.005390 to 0.005359, but the gain is too small to explain the frontier. The useful mechanism is therefore not merely ``look where the forecast is large,'' ``follow dopamine-like innovation,'' or ``add more learned atoms.'' It is the numerical conditioning of a global–local operator dictionary under sparse contemporary observations. The next weather-scale research target is a subquadratic or partitioned station solver, because the current dense RBF block scales like O(m3) in the station count and becomes the runtime bottleneck exactly where accuracy is best.
A subsequent sub-512 MPS audit confirms this conclusion under the same strict held-out protocol. With rank-8 DCT, no learned POD or joint atoms, RBF ridge 5×10−4, and all 21 frames evaluated, a 256/384 station sweep over the three Test 29 files finds best means 0.000791, 0.001301, and 0.006318. The first is not a new M01\_Eta01 frontier, but the latter two beat the old 512-random hybrid references for M10\_Eta01 and M10\_Eta001. A focused hard-file sweep then shows that 320 stations already reduce M10\_Eta001 to 0.007178, below the old 512-random 0.008584 reference, while 448 space-filling stations with length scale 9 reach 0.005784. Training-loss station-seed search helps some rows (M01\_Eta01 and M10\_Eta01) but hurts or misses the best M10\_Eta001 rows. Thus the high-information station-placement rule is still conditioned coverage plus length-scale matching, not residual hot-spot chasing.
We then tested the subquadratic solver implied by this bottleneck. A purely spatial partition-of-unity RBF head fails: on M10\_Eta001, side-2 local blocks over 448/640/768 stations give best mean 0.009167, and adding a 128-station coarse global scaffold before the local residual solve gives best mean 0.009540. Both are much worse than the dense 448-station 0.005784 result, so the weather residual is not a set of independent local patches. It requires the global covariance geometry of the RBF kernel. Replacing the dense station kernel by a global inducing-center RBF dictionary is the useful compression. With all station observations retained as regression rows but only 384 globally distributed RBF centers as columns, rank-8 DCT, no POD or joint atoms, length scale 7, and torch.inference\_mode() on MPS, the three-file run reaches 0.000661, 0.001029, and 0.004823 at 1024 stations for M01\_Eta01, M10\_Eta01, and M10\_Eta001; increasing to 512 inducing centers, 1536 stations, and length scale 5 improves these to 0.000596, 0.000883, and 0.003766; increasing again to 768 centers and 2048 stations gives 0.000494, 0.000690, and 0.002736; using 1024 centers, 3072 stations, and length scale 4 gives 0.000408, 0.000561, and 0.002092; using 1536 centers, 4096 stations, and length scale 3 gives 0.000329, 0.000444, and 0.001368. A rank-2048 inducing profile with 6144 stations and length scale 2.5 reaches 0.000240, 0.000383, and 0.001034 with relative reductions from the frozen FNO of 95.2%, 94.3%, and 92.1%. The recorded peak kernel/design estimate for this row is 192.0 MB versus 528.0 MB for a dense 6144-station kernel, and the row evaluates in about 58.8 s on MPS. We also fixed the runner so that inducing mode no longer materializes the unused dense station kernel before the low-rank solve; the subquadratic memory estimate now matches the executed branch. A first inducing-center audit shows that center geometry, not just center count, matters: on the hard file at rank 2048, 6144 stations, and length scale 2.5, the historical farthest-point station prefix reaches 0.001040, independent global space-filling centers reach 0.001059, station-stride centers reach 0.001253, and late station-tail centers collapse to 0.005219. A second score-weighted audit keeps a coverage scaffold and spends the remaining centers on training residual geometry: on the hard file, 75% coverage plus residual centers improves to 0.000966, while 50% coverage gives 0.000998, 90% gives 0.001023, and a broader neuro-score field gives 0.000976. The all-file rank-2048 residual-center validation reaches 0.000253, 0.000366, and 0.001012 for M01\_Eta01, M10\_Eta01, and M10\_Eta001. Retuning its length scale over 2,2.25,2.5,2.75,3 selects ℓ=2.75 on the hard file with 0.000957; all-file validation at this scale reaches 0.000255, 0.000361, and 0.000993. A bottleneck isolation then rejects merely adding observations at fixed rank: rank 2048 with 8192 stations gives only 0.000972 on M10\_Eta001, worse than the 6144-station row. Increasing the inducing dictionary instead is decisive. With rank 3072, 6144 stations, 75% coverage plus residual centers, and ℓ=2.5, the MPS all-file validation reaches 0.000208, 0.000306, and 0.000634 for M01\_Eta01, M10\_Eta01, and M10\_Eta001, with FNO reductions of 95.8%, 95.5%, and 95.2%. The kernel/design estimate is 300.0 MB versus 528.0 MB for the dense 6144-station kernel. Rank 4096 remains viable under a 512 MB guardrail; a hard-file sweep over ℓ=2,2.25,2.5 selects ℓ=2 with mean 0.000540, and all-file validation reaches 0.000183, 0.000290, and 0.000534 with FNO reductions of 96.3%, 95.7%, and 95.9%. Its kernel/design estimate is 416.0 MB versus 528.0 MB dense and each row costs about 202 s on MPS. A rank-4096 center-allocation audit keeps ℓ=2 and varies only the center score: 50% residual coverage gives 0.000547 on M10\_Eta001, 87.5% residual coverage gives 0.000552, and a neuro-score center field at 75% coverage gives 0.00054037, fractionally worse than the 75% residual-energy row at 0.00054017. Ridge retuning then shows a shallow cross-file tradeoff: λrbf=2×10−4 gives the best hard-file row at 0.000531 but worsens M10\_Eta01 to 0.000292, λrbf=10−4 regresses the hard file to 0.000539, and the best single-profile macro mean is λrbf=3×10−4, reaching 0.000182, 0.000290, and 0.000532 across M01\_Eta01, M10\_Eta01, and M10\_Eta001; λrbf=4×10−4 lands slightly worse at 0.000182, 0.000290, and 0.000533. Intermediate compression points preserve the same geometry with smaller dictionaries: rank 3584 chooses ℓ=2.25 and gives 0.000192, 0.000295, and 0.000564 at 357.0 MB and about 159 s per row, while rank 3840 also chooses ℓ=2.25 and improves the tradeoff to 0.000189, 0.000292, and 0.000552 at 386.25 MB and about 180 s per row. Finally, we added an --eval\_split hook to audit profile selection without held-out leakage. Training-split losses over λrbf∈{2,3,5}×10−4 select 5×10−4 for M01\_Eta01, 2×10−4 for M10\_Eta001, and 5×10−4 for M10\_Eta01; this misses the tiny held-out M01\_Eta01 optimum but selects both M10 held-out winners, giving held-out values 0.000183, 0.000290, and 0.000531, better in macro mean than any single ridge. Thus the scalable weather solver should be low-rank global RBF/Nystr\"om geometry with enough inducing capacity and a balanced global/residual center budget, not spatial partitioning or additional stations at a fixed insufficient rank; beyond rank 4096, the next question is compression or per-regime regularization because memory approaches dense parity.
After this profile-selection audit, we retuned the rank-4096 frontier more finely. A control-style, station-held reward-gain branch first tested whether a terminal dopamine-like scalar could improve the correction magnitude. The runner was patched so that reward-gated assimilation uses the same inducing RBF operator as the frontier path; with 10% held stations and candidate gains {0.75,0.9,1.0,1.1,1.25}, the hard M10\_Eta001 file selected average gain 0.94 but worsened to 0.000542, so scalar terminal gain is a negative branch. The positive branch is earlier operator regularity. A fine length-scale audit at 6144 stations, rank-4096, residual-center coverage 0.75, and λrbf=2×10−4 over ℓ∈{1.5,1.75,2,2.125,2.25} selects ℓ=2.125; all-file validation reaches 0.000182, 0.000526, and 0.000289 for M01\_Eta01, M10\_Eta001, and M10\_Eta01. Retuning λrbf at this length scale gives a hard-file knee at 5×10−5: the all-file single-profile row reaches 0.0001816, 0.0005214, and 0.0002914 with the same 416 MB inducing design estimate. A no-leakage train-split selector between 2×10−4 and 5×10−5 at ℓ=2.125 selects the hard-file and M10\_Eta01 held-out winners, missing only the tiny M01\_Eta01 preference; the train-selected held-out profile is approximately 0.000182, 0.000521, and 0.000289. Applying the same lower ridge to the rank-3840 compression point shows that compressed profiles prefer a slightly broader kernel: ℓ=2.25, λrbf=5×10−5 reaches 0.0001857, 0.0005369, and 0.0002919 at 386.25 MB, improving the earlier rank-3840 tradeoff while remaining below rank-4096 quality. The practical conclusion is that the next weather lever is per-regime operator regularity and compression, not terminal scalar reward gain.
The next lower compression point confirms the rank curve. Rank 3584 with the same lower ridge and ℓ=2.25 reaches 0.0001896, 0.0005505, and 0.0002958 at 357 MB and about 159 s per row. This improves the old rank-3584 row, but rank 3840 is the cleaner quality–memory compromise before the rank-4096 frontier.
Pushing the opposite direction gives the current quality frontier while still respecting the memory guard. Rank 4608 remains below the dense 6144-station kernel, with a 477 MB design estimate. A hard-file sweep at λrbf=5×10−5 selects ℓ=2.25 with M10\_Eta001 mean 0.0005048. The all-file validation reaches 0.0001805, 0.0005048, and 0.0002837 for M01\_Eta01, M10\_Eta001, and M10\_Eta01, respectively, at about 251 s per row. This improves the rank-4096 single-profile frontier on every Test 29 weather file while remaining below dense-kernel memory.
The edge-of-guardrail rank-4800 profile improves the frontier again while still staying below dense memory: the design estimate is 500.39 MB versus 528 MB dense. A hard-file sweep over ℓ∈{2.125,2.25,2.375} selects ℓ=2.25 and reaches 0.0004994 on M10\_Eta001, the first confirmed sub-0.0005 hard weather result in this campaign. The all-file validation reaches 0.0001797, 0.0004994, and 0.0002822 for M01\_Eta01, M10\_Eta001, and M10\_Eta01, at about 271 s per row. This is the current quality frontier under the sub-dense memory guard.
The edge-rank check uses rank 4864, which still fits under the guard at 508.25 MB. With ℓ=2.25, λrbf=5×10−5, and a 75% inducing-center coverage scaffold, the all-file validation reaches 0.0001797, 0.0004973, and 0.0002820 for M01\_Eta01, M10\_Eta001, and M10\_Eta01, at about 278 s per row. The matching train-split audit gives 0.0001188, 0.0006095, and 0.0003040, supporting selection of the edge-rank profile from training trajectories. A fine scaffold retune at the same rank and memory then improves the hard frontier. The first hard-file sweep over coverage fractions 0.50,0.625,0.875,1.00 gives 0.0005017, 0.0004967, 0.0005087, and 0.0005129. A matched-station-seed sweep around the knee gives 0.0004966, 0.0004950, 0.00049323, 0.00049266, and 0.0004939 for coverage 0.5625,0.600,0.625,0.650,0.6875. Promoting the 0.650 scaffold to all three files gives the current sub-dense quality frontier, 0.00017950, 0.00049266, and 0.00028253, still at 508.25 MB versus 528 MB dense and about 277 s per row. This improves all three files versus the 0.625 row and improves the hard file by about 0.92% relative to the old 75% scaffold. The no-leakage train-split audit for the 0.650 scaffold gives 0.00011875, 0.00061174, and 0.00030438, so raw train loss would still prefer the 75% scaffold for the hard and M10\_Eta01 files. Thus the 0.650 row is a real held-out frontier, but robust train-selectable per-regime control remains unsolved; the remaining rank headroom before dense parity is only a few MB, so further progress should come from algorithmic compression or better no-leakage controllers rather than raw rank escalation.
The first post-frontier compression audit therefore varied the inducing-center allocation instead of the raw rank. At rank 3840, 6144 stations, ℓ=2.25, λrbf=5×10−5, rank-8 DCT, and no POD/joint atoms, pure residual-energy centers with no global coverage scaffold validate at 0.0001845, 0.0005313, and 0.0002962 for M01\_Eta01, M10\_Eta001, and M10\_Eta01. This slightly improves the rank-3840 macro mean over the 75% coverage scaffold row (0.0001857,0.0005369,0.0002919) at the same 386.25 MB design estimate, but worsens M10\_Eta01. Pure global space-filling centers are worse on the hard file (0.0005612), broad neuro-score centers are also worse (0.0005413), and an intermediate 25% coverage scaffold reaches only (0.0001853,0.0005359,0.0002945). At rank 4864, the pure-residual endpoint worsens the hard file to 0.0004998 versus the 75% scaffold frontier 0.0004973. Thus center allocation is not a universal scalar setting: lower-rank compression benefits from more aggressive innovation-driven centers, while the edge-rank frontier still needs a coverage scaffold for conditioning. A no-leakage selector audit exposed the next bottleneck. Raw train-split losses at rank 3840 give cov0 (0.0001202,0.0006838,0.0003135) and cov75 (0.0001208,0.0006750,0.0003132); this selects the held-out winners for M01\_Eta01 and M10\_Eta01 but misses the hard file. A one-sample inner validation split, fitting on samples 0–2 and scoring sample 3, gives cov0 (0.0001339,0.0008521,0.0003611) and cov75 (0.0001343,0.0008527,0.0003614), selecting cov0 for all three and therefore missing the held-out M10\_Eta01 scaffold preference. A two-sample inner validation split, fitting on samples 0–1 and scoring samples 2,3, gives cov0 (0.0001387,0.0010027,0.0004333) and cov75 (0.0001393,0.0009815,0.0004329); it selects cov75 for the hard file even though held-out hard prefers cov0. We also implemented a station-held center-coverage gate that scores candidate center profiles on held observed stations inside each forecast slice and then refits the selected profile on all observed stations. At rank 3840 with candidates {0,0.75} and 10% held stations, M10\_Eta001 reaches 0.0005348 with mean selected coverage 0.268, better than static cov75 but worse than static cov0; M10\_Eta01 reaches only 0.0003009 with mean selected coverage 0.298, worse than both static profiles. A cheap train-only proxy audit over residual-score entropy, grid coverage, station coverage, and center nearest-neighbor spacing is also inconclusive: cov0 and cov75 have nearly identical coverage means (about 0.81 grid units) and nearest-neighbor statistics. We then patched the dynamic station-placement path so that dynamic gradient, neuro-attention, and pilot-innovation policies use the same low-rank inducing RBF solver as the static frontier. The full hard-file pilot-innovation row at 6144 stations and rank 3840 reaches only 0.0005426 in 706 s, worse than static cov0 (0.0005313). A short prefix scan over frames 0–4 was invalid because those frames have zero baseline forecast error in the artifact. On the meaningful frame-5–9 prefix, dynamic neuro attention with a 75% lattice scaffold improves the local comparator from 0.0003793 to 0.0003679, but the promoted full late-horizon hard-file run rejects the signal: static neuro coverage 0.90 reaches 0.0006612 in 137 s, whereas dynamic neuro coverage 0.75 reaches 0.0006768 in 491 s and coverage 0.90 reaches 0.0007066 in 459 s. The compression result is real, but robust per-regime controller selection remains open; raw training loss, tiny validation windows, simple geometry proxies, repeated high-rank station-held gates, and the current greedy dynamic station selector are too noisy or too expensive for this knob.
We also ran the harder replacement test: remove the frozen FNO prior and forbid target-time stations. The new runner apps\_industrial\_breakthrough/pdebench\_weather\_blind\_osnr\_dynamics\_mps.py identifies a blind dynamics model from observed trajectories only and then rolls held-out samples forward from their history frames. Its cell has two coupled no-backprop components. The local PDE library uses current fields, velocity memory, first derivatives, Laplacians, biharmonic terms, Laplacian velocity, quadratic products, self-advection, and cross-channel products, with coefficients solved by ridge systems. A low-mode spectral liquid cell adds DCT coefficients, finite differences, stable pole traces {0.25,0.50,0.75,0.90,0.97}, and tanh coefficient features. All fits run under torch.inference\_mode() on MPS; no FNO prediction, no future observation, and no autograd graph are used. The result is a clear boundary rather than a win. A global fit with three history frames is unstable and worse than persistence. A sample-online fit with per-sample history clamps and five observed frames becomes stable but only improves persistence by 0.30%, 0.42%, and 0.41% on M01\_Eta01, M10\_Eta001, and M10\_Eta01, while remaining 2.63×, 21.20×, and 12.06× worse than FNO. With eight history frames the persistence gains rise only to 0.83%, 0.88%, and 0.88%, with FNO gaps of 1.45×, 20.53×, and 5.75×. Thus the current weather result is assimilation, not autonomous GraphCast/GenCast replacement. Replacing the entire neural forecaster requires substantially more trajectory diversity or a stronger endogenous PDE/ODE neural architecture; the present local library mostly learns a small homeostatic correction to persistence.
The follow-up autonomous-weather loop raised the history to 16 frames and changed the local cell rather than only retuning ridge constants. The temporal feature set adds previous-frame state, velocity derivatives, previous Laplacian, saturating nonlinearities, and current–velocity products. The memory2 feature set adds a second-order temporal state: older frame, prior velocity, acceleration, acceleration derivatives/Laplacian, older Laplacian, and acceleration cross-products. Finally, a partitioned operator layout fits separate PDE matrices on a 2×2 spatial grid but shrinks them toward the shared trajectory-level operator, Wb=(1−s)Wshared+sWlocal,b,s=0.25, which preserves global conditioning while allowing regional deviations. Table [tab:pdebench-blind-weather-autonomous] reports the current best rows. The important result is not yet a weather-model victory: the autonomous cell beats frozen FNO only on the two easy regimes where persistence is already very strong, and it remains about 9.5× worse than FNO on the hard M10\_Eta001 regime. The positive scientific signal is narrower but real: second-order memory plus light regional shrinkage improves the hard no-FNO row from 0.130594 to 0.121345, and a finer 8×8 partition with a global observed-history rollout selector improves the hard row further to 0.115310. This selector is not trained on held-out future frames: it scores candidate ridge/update settings on training-sample observed-history tails, then applies the chosen setting to held-out histories. Its selected settings are (λ,γ)=(300,0.70) for M01\_Eta01, (500,0.75) for hard M10\_Eta001, and (300,0.70) for M10\_Eta01. The partition-resolution audit also sets a boundary: 4×4 reaches hard 0.116181, 8×4 reaches 0.115653, 8×8 fixed reaches 0.115466, but 16×8 and 8×16 regress to about 0.11587 while doubling runtime; 8×8 shrinkage s=0.20 and s=0.40 regress to 0.115805 and 0.116198. Negative controls were decisive. A multiscale smoothed-context dictionary worsened hard error to 0.241980, pure 2×2 blocks worsened to 0.124196, temporal blocks worsened to 0.126452, explicit coordinate augmentation of the memory2 cell reached only 0.116311, compact per-channel global moment modulators reached only 0.116603 and a damped version 0.117145, nearby 8×8 shrinkage s=0.25 and s=0.35 reached only 0.115562 and 0.115525, a p90-robust forced row reached 0.115537, and top-3/top-5 history-tail ensembles reached only 0.115345/0.115357 on the hard file. Tail-length validation also closed around the current setting: observed-history rollout tails of 2, 3, 5, and 6 steps reached 0.118317, 0.118317, 0.115710, and 0.115949, so the four-step tail remains best. A wider channel-coupled cross-operator dictionary was especially diagnostic: its lower-ridge run improved the observed-history tail score from 0.283297 to 0.280035 but worsened held-out future error to 0.117344; the stronger-ridge run reached 0.117920. The observed-history selector can therefore be fooled by high-capacity local dictionaries. Low-mode spectral blending also remained dangerous: an ungated conservative spectral search would choose blend 0.20 and degrade hard held-out error to 0.135724. The same mean/p90 spectral gate now blocks that row because the observed-tail mean gain over zero blend is only 4.68%<5%, restoring the zero-blend 0.115310 frontier. Thus the next autonomous-weather step should not be more spectral blending, coordinate tags, compact global moments, scalar shrinkage tuning, tail-length retuning, top-k row averaging, or wider local dictionaries without a stronger validation guard; it should use a batched/subquadratic partitioned solver and a structurally different endogenous low-mode/state mechanism with no-leakage model selection.
Autonomous no-FNO cell, history 16
M01\_Eta01
M10\_Eta001
M10\_Eta01
Macro mean
FNO artifact
0.006846
0.012127
0.009781
0.009585
Persistence
0.001090
0.206921
0.008352
0.072121
Temporal shared OSNR
0.000861
0.130594
0.006783
0.046079
Memory2 shared OSNR
0.000897
0.123513
0.007313
0.043907
Memory2 2×2 shrink OSNR
0.000905
0.121345
0.007127
0.043126
Selector: temporal easy, 2×2 hard
0.000861
0.121345
0.006783
0.043330
Memory2 8×8 history-tail global OSNR
0.000851
0.115310
0.007224
0.041128
Autonomous no-backprop weather replacement audit on external PDEBench/FNO Test 29 artifacts. The model observes only the first 16 true frames of each held-out trajectory and predicts the remaining five frames. It does not read FNO predictions, target-time stations, or future labels at evaluation time. All rows use per-trajectory closed-form ridge identification under torch.inference\_mode() on MPS. The last row uses the formal history\_tail\_global selector: one ridge/update setting is chosen from training-sample observed-history rollout error for each file, then applied to held-out histories.
External forecast artifact
FNO
256 sf tuned
288 sf
384 sf
512 random ref.
Test 29 M01\_Eta01
0.004611
0.000705
0.000686
0.000646
0.000707
Test 29 M10\_Eta01
0.005507
0.001377
0.001334
0.001267
0.001438
Test 29 M10\_Eta001
0.012038
0.007649
0.007639
0.007274
0.008584
Adaptive station placement on external PDEBench/FNO Test 29 artifacts. All adaptive rows use farthest-point space-filling placement with the same closed-form affine+DCT+RBF assimilation layer; only the station count, DCT rank, and RBF length scale are selected by the bounded sweep. The 512-station random hybrid reference is beaten by 256 tuned space-filling stations on all three files, and the 384-station row gives the current best Test 29 means.
Direct no-backprop PDEBench sequence continuation
The next non-Fashion test removes both the frozen FNO prior and target-time stations, but keeps a short true history of the same trajectory. The runner apps\_industrial\_breakthrough/pdebench\_1d\_direct\_spectral\_forecaster\_mps.py is a no-backprop direct sequence forecaster for the 1D PDEBench audit artifacts. For a target tensor Y∈RN×X×T×C, history length h, and normalized spectral matrix ΨK∈RX×K, it forms coefficient tokens an,t=ΨK⊤yn,t∈RKC. The per-sample input feature is zn=[1,an,0:h−1,Δan,0:h−2,an,h−1,aˉn,std(an),tanh(0.5an,h−1)], flattened over time, modes, and channels. The future coefficients are obtained by one closed-form ridge solve, Wλ,K=(Z⊤Z+λI)−1Z⊤Ah:T−1,A^h:T−1=ZWλ,K. The low-mode field is reconstructed by ΨKA^, and a high-frequency identity residual from the last observed frame is added with validation-selected decay ρ: y^n,t=ΨKa^n,t+ρ(yn,h−1−ΨKΨK⊤yn,h−1),t≥h. No FNO prediction, no future observation, and no reverse-mode graph are used by the model. The updated runner also validates a small no-backprop expert set—direct spectral, persistence, and linear extrapolation—so the reported ``selected'' row is the validation-only choice among physically simple history-conditioned mechanisms. Runs use torch.inference\_mode() on Apple MPS through the project venv/bin/python -c "... runpy.run\_path(...)" route, because direct script execution can hide the MPS backend in this environment.
The latest autonomous extension adds an explicit physical-cell feature grid while keeping the default base model unchanged. For each low-mode coefficient history, the liquid feature set forms causal pole traces pt(α)=αpt−1(α)+(1−α)at,α∈{0.10,0.30,0.55,0.75,0.90,0.98}, then appends acceleration, last velocity, bounded softsign/tanh coordinates, pole innovations ah−1−ph−1(α), low-mode quadratic products, and leading-mode spline hinges max(a−κ,0) with validation-selected spline width. A mixed base,liquid grid lets validation route each PDE file to the simpler DCT history model or to the richer liquid/spline cell. A separate autoregressive mode trains the same closed-form map on all sliding history windows and rolls forward without backprop; it is used only as a long-horizon diagnostic below.
PDEBench audit file
FNO future nRMSE
no-backprop selected
Gain vs FNO
Selected expert
test\_01
0.004911
0.004824
1.8%
base spectral
test\_02
0.001513
0.076136
−4931.1%
liquid spectral
test\_03
0.001801
0.001412
21.6%
persistence
test\_07
0.003583
0.002440
31.9%
base spectral
test\_08
0.004866
0.003091
36.5%
liquid spline
test\_09
0.005087
0.009067
−78.2%
liquid spectral
test\_10
0.004532
0.264953
−5746.4%
autoreg spectral
test\_11
0.003982
0.003134
21.3%
base spectral
test\_13
0.000451
0.001221
−171.0%
persistence
test\_16
0.000976
0.000938
3.9%
liquid spline
MPS no-backprop direct sequence continuation on scalar 1D PDEBench/FNO audit artifacts after the liquid/pole/spline routing extension. The model observes the first eight true frames and predicts all remaining frames (thirteen frames for the 21-step scalar rows; thirty-three frames for the 41-step test\_10 row). All rows use training samples 0–849, validation samples 850–899, and held-out test samples 900–999. The result is a genuine non-Fashion no-backprop hit but not a universal replacement: six scalar rows beat the cached FNO future error, several failures improve materially, and test\_10/test\_13 remain unresolved.
The boundary is equally important. The liquid/pole/spline extension improves, but does not solve, the coupled three-channel artifacts. The old direct runner gave test\_05=0.112862 and test\_06=0.052174; the mixed liquid/spline grid with validation top-3 averaging improves these to 0.092994 and 0.046179, respectively, versus FNO 0.003861 and 0.006197. The coupled gain is real (about 17.6% and 11.5% relative error reduction against the old no-backprop direct rows), but the remaining FNO gap is still too large for a SOTA claim. The next autonomous audit asked whether this was a validation-split, temporal-decoder, or locality problem. It was not. A shuffled validation split on the same scalar runner still failed to rescue the selector: on test\_13, FNO is 0.000451 while validation selects persistence at 0.001221 although the direct row is 0.001416; on test\_03, validation again selects persistence at 0.001412 while the direct row is 0.001413. Adding compressed future-time DCT decoding through --future\_time\_ranks gives test\_10=0.271725 at rank 8, still worse than the previous autoregressive hard-row frontier near 0.264953, and a six-row scalar audit selects no temporal compression at all. A global RBF/kernel memory over observed histories is worse again on the hard row, reaching only 0.300106. Finally, the new local runner apps\_industrial\_breakthrough/pdebench\_1d\_local\_direct\_forecaster\_mps.py predicts each spatial point from a periodic local stencil, velocity, local derivatives, liquid pole states, spline hinges, global low-mode context, and optional online per-trajectory PDE coefficient tokens. Its best hard-row run reaches 0.272812 on test\_10; on the coupled rows it reaches 0.096646 and 0.085312, improving persistence but trailing the earlier global spectral/liquid rows. We then made the proposed regime idea explicit: the same local runner can fit separate ridge operators for quantile bins of local shock, curvature, velocity, or transport score, and can also identify per-trajectory local PDE-library coefficients from the observed frames and roll them forward autonomously. On the full test\_10 hard-row confirmation, validation still selects the original local map at 0.282656 validation and 0.273743 held-out test; the best shock-regime row is slightly worse at 0.283498 validation, and the best online PDE rollout is much worse at 0.319173 validation. The conclusion is now stronger: the missing mechanism is not simply random validation leakage, a compressed future basis, kernel memory, local finite-difference tokens, quantile shock partitioning, or stepwise local PDE coefficient rollout. The current direct solver needs a genuinely conservative flux/discontinuity operator or a different endogenous architecture, not another ridge feature expansion.
We also added apps\_industrial\_breakthrough/pdebench\_2d\_direct\_spectral\_forecaster\_mps.py, which extends the architecture to 2D tensor-DCT tokens, an,t,k,ℓ,c=i,j∑Ψi,k(H)Ψj,ℓ(W)Yn,i,j,t,c, and uses the same closed-form future-coefficient solve plus identity residual expert gate. On test\_27 with history 12, rank 32, train/validation/test split 70/10/20, the 2D direct model improves as rank grows but still reaches only 0.009609 against FNO 0.001856; porting the 1D liquid pole/quadratic feature grid to the 2D tensor-DCT runner leaves validation on the original base row. On test\_26, the same 2D grid is worse: direct spectral reaches 0.554028 and validation falls back to linear extrapolation at 0.080726 versus FNO 0.002635. Additional negative controls on test\_10 show that higher DCT rank, explicit real Fourier bases, more observed history, sliding-window autoregressive ridge training, global shift/transport extrapolation, and projected PDE-library features all plateau near 0.265–0.267, far from FNO. The same mixed base/liquid/PDE feature grid beats persistence but remains far from FNO on additional vector rows: test\_17=0.037810 versus FNO 0.001431, test\_19=0.111576 versus FNO 0.008693, and test\_20=0.041286 versus FNO 0.004741. Thus the current publishable claim is narrow and precise: no-backprop OSNR-style spectral sequence continuation can beat cached FNO on several scalar PDEBench dynamics and improve hard failures, but shock-like long-horizon rows, coupled systems, and 2D fields require a stronger endogenous architecture than one global ridge map from history tokens to future coefficients. The next operator-learning target should combine this direct spectral solver with local conservation-law blocks, channel-coupled low-mode poles, and validation-stable regional expert gates.
The same hosted FNO artifact contains five static Darcy-flow rows, Tests 21–25. Since the artifact exposes only FNO predictions and targets, not the Darcy coefficient input fields, we do not claim blind PDE solving in this audit. Instead, apps\_industrial\_breakthrough/pdebench\_darcy\_compression\_audit.py measures a representation question: how compactly can an OSNR-style global operator dictionary encode the target solution fields compared with the error of the trained FNO prediction?
For each target field u, we retain a compact rectangular set of Fourier/operator coefficients, u^K=F−1[MK(kx,ky)Fu], and report relative L2 nRMSE. Table [tab:pdebench-darcy-compression] shows that the Darcy targets are extremely compressible. On the hardest row, Test 21, the hosted FNO prediction has nRMSE 0.2670, while OSNR target representation reaches 0.2039 with only 0.096% of spectral bins, 0.0730 with 0.385%, and 0.0396 with 0.865%. For the easier Darcy rows, FNO is already strong, but OSNR still passes the FNO error level with a small coefficient budget: K=8 for Test 23 and K=16 for Tests 24–25.
PDEBench row
FNO nRMSE
First OSNR K beating FNO
Coefficient fraction
OSNR nRMSE
Darcy Test 21
0.267029
2
0.096%
0.203930
Darcy Test 22
0.117649
4
0.385%
0.075482
Darcy Test 23
0.027689
8
1.538%
0.027178
Darcy Test 24
0.011536
16
6.154%
0.009421
Darcy Test 25
0.009500
16
6.154%
0.009427
External PDEBench Darcy target-representation compression audit. This table measures compact representation of the target solution fields, not blind prediction from Darcy coefficients.
PDEBench Darcy Test 21 compression panel. A tiny low-frequency OSNR coefficient set captures the smooth elliptic solution structure more accurately than the hosted FNO prediction error, but this is a target-representation result rather than a complete Darcy solver.
This result identifies a strong but precise opportunity. For elliptic PDE outputs, the operator-spline basis has excellent compression power on real external benchmark tensors. To turn this into a SOTA solving claim, the next experiment must ingest Darcy coefficient fields and solve or learn the coefficient-to-solution operator directly; target compression alone is not enough.
We began this coefficient-to-solution step using the PhysArena/PDEBench Darcy Parquet mirror, which exposes both the diffusion coefficient and the flow target. A simple finite-volume CG solve of −∇⋅(a∇u)=0.01 with homogeneous Dirichlet boundaries recovers the spatial shape of many samples very accurately after an oracle scalar alignment: on a 120-sample pilot, oracle-scaled relative error averages 0.0473 with median 0.0243. However, the required scalar varies strongly with coefficient geometry. A CNN trained on coefficient images to predict this scalar improves the median held-out error but leaves large outliers, with robust training still giving mean error about 0.305 versus oracle 0.060. A larger hybrid attempt using 3000 samples and an MPS residual CNN from (a,uCG,x,y) to u also failed, increasing held-out error to 1.239 compared with 0.317 for the raw CG shape and 0.055 for oracle-scaled CG. We therefore do not yet claim a Darcy solver win. The next step is to recover the exact dataset forcing/normalization convention from the Plaid metadata or learn in constrained coefficient/operator space rather than using an unconstrained image residual network.
The direct operator is nevertheless already a strong solver in a clearly defined regime. In apps\_industrial\_breakthrough/pdebench\_darcy\_operator\_regime\_audit.py, we evaluate 1000 real coefficient fields from the same external shard and stratify by the fraction of low-conductivity cells. Table [tab:pdebench-darcy-operator-regime] reports the result. When low-conductivity inclusions occupy less than 25% of the domain, the raw finite-volume CG solve reaches mean nRMSE below 6.2×10−4. For the intermediate 25–50% regime, it remains useful with mean nRMSE 0.0654 and median 0.0236. The failure begins once low-conductivity cells dominate the grid, which is exactly where the simple arithmetic-face cell-centered stencil diverges from the external generator's apparent discretization.
We then stress-tested that failure with apps\_industrial\_breakthrough/pdebench\_darcy\_operator\_stress\_audit.py. The goal was to determine whether the high-inclusion collapse was a shallow implementation issue. A generalized face-transmissibility sweep over power means p∈{−4,−2,−1,−0.5,0,0.5,1,2,4} and boundary scales {0.5,1,2} did not fix the regime: the best training-screen mean was 0.3300 and the high-inclusion mean remained 0.5664. A held-out jump-aware residual correction using local features (u,∣∇a∣u,∣∇a∣,Δu,au,1a<0.5u,1) was unstable, increasing held-out mean error from 0.2931 to between 4.525 and 7.259 depending on ridge strength. Finally, high-contrast coefficient convention tests rejected simple metadata mistakes: the original coefficient field gave oracle-scaled high-contrast error 0.1327, while inverse coefficients, vertical flips, horizontal flips, and rotations worsened to 0.4069, 0.2717, 0.3174, and 0.3096, respectively. This negative result is useful. It says that the next Darcy improvement should not be another scalar transmissibility tweak or unconstrained residual network; it should recover the exact generator discretization or move to a constrained multiscale/interface operator that preserves ellipticity.
We also tested the more realistic hybrid idea: keep the direct operator solve, but train a small neural module only for the unknown correction. In apps\_industrial\_breakthrough/darcy\_neural\_osnr\_pde\_layer.py, a compact encoder observes the coefficient field and the direct CG solution. Three heads are compared on held-out Darcy fields: a direct low-resolution residual decoder, an OSNR/PDE-layer residual decoder that passes the predicted source through a screened Poisson inverse before upsampling, and a scalar normalization head. The correction is gated by the coefficient regime: low/mid-inclusion samples keep the direct operator output, while high-inclusion samples use the learned correction. On a 500-sample smoke split with 350 training samples and high-regime specialization, the base held-out mean nRMSE is 0.2876. The gated OSNR/PDE residual reduces this to 0.2516, the gated direct residual to 0.2536, and the gated scale head to 0.2332. In the hardest low-conductivity-dominant bin, the base mean error 0.8923 drops to 0.7202 with the scale head. This is not a final SOTA solver, but it is a real hybrid lesson: the current external Darcy gap is mostly a hidden sample-dependent normalization/interface convention, and a physics-gated neural correction is useful only when it respects the regimes already solved by the operator.
We then pushed directly on that convention gap with two additional real-data calibration experiments. First, apps\_industrial\_breakthrough/pdebench\_darcy\_interface\_calibration\_search.py evaluated 21 positive symmetric face-transmissibility variants on a 360/240/120 train/test split. The best learned calibration used a maximum-face law and reduced the held-out mean from 0.3715 to 0.3416, with the high-inclusion bins improving from 0.5535 to 0.4752 and from 0.8539 to 0.7386. More importantly, its oracle-scaled held-out error was only 0.0822, showing that the operator family can produce a much better shape than our deployable scalar calibration recovers. Second, we tested whether the missing scalar is easily recoverable from coefficient geometry. A hand-feature kNN calibration in apps\_industrial\_breakthrough/pdebench\_darcy\_geometry\_scale\_knn.py failed, increasing gated mean error from 0.2876 to 0.3678. Supervised scalar-head variants that explicitly regress the oracle alignment scalar also failed to beat the earlier 0.2332 gated scale result. These negative results narrow the remaining problem: the bottleneck is not generic network capacity or a simple geometry-to-scale map, but recovery of the external generator's discretization/normalization law or a richer constrained multiscale elliptic operator.
The same diagnosis suggests a different, stronger problem formulation: sparse-sensor PDE assimilation. In real deployments one often has a small number of pressure/head/flow probes but cannot afford dense field acquisition. In apps\_industrial\_breakthrough/pdebench\_darcy\_sparse\_sensor\_assimilation.py, the coefficient field defines the elliptic OSNR/PDE shape, and m point observations determine the remaining amplitude by the closed-form least-squares scalar
sm=∑q∈ΩmuCG(q)2+ϵ∑q∈ΩmuCG(q)uobs(q).
On a larger 1000/700/300 real PDEBench Darcy split, the blind direct operator has mean nRMSE 0.3262. A single interior sensor reduces the error to 0.0889, four sensors reduce it to 0.0681, and 32 sensors reach 0.0597, close to the full-field oracle scalar ceiling 0.0585. In the hardest low-conductivity-dominant bin, the same four-sensor assimilation reduces mean error from 0.8642 to 0.1460, while 32 sensors reach 0.1304 against an oracle of 0.1288. This result is substantially stronger than the learned residual attempts: it uses no neural training, preserves the elliptic operator, and converts a failed blind coefficient-to-solution setting into a practical sparse-observation reconstruction problem.
Method
Mean held-out nRMSE
Hard-bin nRMSE
Blind direct CG operator
0.3262
0.8642
1 sparse sensor
0.0889
0.2030
4 sparse sensors
0.0681
0.1460
16 sparse sensors
0.0606
0.1327
32 sparse sensors
0.0597
0.1304
Full-field oracle scalar
0.0585
0.1288
Real PDEBench Darcy sparse-sensor assimilation on 300 held-out coefficient fields. A few point observations close most of the blind-operator amplitude gap without neural training.
Low-conductivity fraction
Samples
Raw CG nRMSE
Oracle-scaled nRMSE
0–0.10
11
0.000366
0.000054
0.10–0.25
86
0.000615
0.000271
0.25–0.50
411
0.065379
0.009759
0.50–0.75
351
0.471048
0.074186
0.75–1.00
140
0.876031
0.114951
External PhysArena/PDEBench Darcy coefficient-to-solution operator regime audit on 1000 real coefficient fields. This is a true solver experiment, not target compression. The direct operator is highly accurate in low/mid contrast regimes and fails when low-conductivity inclusions dominate.
External Darcy coefficient-to-solution example in the low-inclusion regime. The direct operator solve reproduces the target flow without neural training.
Weather-core validation: coupled rotating shallow water
The next weather-facing benchmark is apps\_industrial\_breakthrough/shallow\_water\_operator\_spline\_benchmark.py. This moves beyond scalar advection–diffusion and Burgers equations to a coupled three-component linearized rotating shallow-water core. The unknown state is q(t,x,y)=(η,u,v)⊤, where η is the height anomaly and (u,v) are horizontal velocities. The periodic operator is
Here H is mean depth, g is gravity, f is the Coriolis parameter, r is velocity damping, and μ is a small height-relaxation gauge that removes the resonant zero-frequency mass mode. In Fourier space, every (ω,kx,ky) bin is a dense 3×3 complex linear system. OSNR solves the full coupled block by batched frequency-bin inversion, not by fitting a neural coordinate model or by stepping a recurrent simulator.
The forcing combines smooth planetary-wave structure with sparse localized height/vorticity impulses. The sparse impulses are treated as storm/front innovations in the operator domain. A low-pass spectral reconstruction is included as a smooth surrogate baseline. Table [tab:shallow-water-weather-core] reports the current results.
Profile
Grid
OSNR PSNR
Low-pass PSNR
Atom error
Solve time
Main, 24 events
32×642
33.2500 dB
29.0210 dB
0.0000 px
13.09 ms
Dense, 48 events
32×642
34.9617 dB
27.1203 dB
0.3363 px
13.08 ms
Heavy, 96 events
32×642
34.1593 dB
26.2245 dB
1.2163 px
12.85 ms
0.5% noise, raw
32×642
33.0138 dB
29.0210 dB
0.0000 px
13.08 ms
1.0% noise, smoothed
32×642
31.8288 dB
29.0210 dB
0.0000 px
13.32 ms
2.0% noise, smoothed
32×642
31.8257 dB
29.0210 dB
0.0000 px
13.84 ms
Scale stress
48×962
34.5164 dB
28.9019 dB
0.0711 px
40.55 ms
Scale stress, 96 events, 0.5% noise
48×962
32.3981 dB
27.7452 dB
0.4625 px
37.83 ms
Linearized rotating shallow-water weather-core benchmark. OSNR solves the coupled height/velocity operator by batched 3×3 frequency-bin inversions. The low-pass row is a smooth spectral-bias baseline. Atom error measures sparse storm/front innovation recovery from the height forcing channel after operator-domain residual extraction.
Coupled shallow-water OSNR benchmark at 48×962. The figure shows target height, OSNR reconstruction, low-pass baseline, recovered height innovation, sparse event atoms, and velocity magnitude. Unlike the scalar tests, the solve couples height and both velocity components through Coriolis and pressure-gradient terms.
This is the first result in the manuscript that begins to resemble a real weather core. It is still not a direct GraphCast/GenCast/WeatherNext comparison: those systems operate on global ERA5-scale atmospheric states and are trained on decades of data. The scientific significance is narrower but important. OSNR can invert a physically coupled, multi-variable periodic atmospheric operator in milliseconds, preserve sparse front/storm innovations, and outperform a smooth low-pass surrogate under noise. The degradation is graceful: even the 96-event, 0.5% noisy 48×962 stress case remains above 32 dB with sub-pixel atom error. The next hard step is to leave the linearized core and introduce nonlinear advection, partial observations, and data assimilation windows while preserving this operator-domain sparse innovation advantage.
We also tested a first partial-observation assimilation variant in which only η is treated as observed. A naive geostrophic lift fails badly because the synthetic state contains wave and forced components outside static balance. A dynamic momentum lift performs much better: given the observed η, it solves the two Fourier-domain momentum equations for (u,v) while assuming small direct velocity forcing. This height-only lift reaches 33.2250 dB on the 32×642 case and 35.1743 dB on the 48×962 scale case, so balanced field reconstruction remains plausible from partial observations. However, sparse event localization degrades to 5.8550 px and 10.9152 px, respectively. The lesson is precise: partial state assimilation can recover smooth balanced dynamics, but front/storm innovation recovery needs an explicit sparse assimilation stage rather than a purely balanced velocity closure.
Coupled weather-core operator identification
The scalar operator-identification experiments show that field-only discovery is underdetermined, while sparse physical anchors make the problem well posed. We repeated the same idea on the coupled shallow-water core in apps\_industrial\_breakthrough/shallow\_water\_operator\_identification.py. The unknown parameter vector is now θ=(μ,H,r,f,g), corresponding to height relaxation, mean depth, velocity damping, Coriolis coupling, and gravity. Given observed (η,u,v) and sparse samples of the forcing channels, the pointwise equations are linear in θ:
Thus the multi-channel operator is recovered by one real ridge least-squares solve over analytic derivative columns. The important correction is the continuity-column term ux+vy; omitting vy makes H unidentifiable.
Setting
Anchors
μ^
H^
r^
f^
g^
PSNR
Clean, 1.0% anchors
3933
0.1817
0.9689
0.0822
0.7988
0.9999
25.7477 dB
0.1% noise, σ=0.5, 0.5% anchors
1965
0.2108
0.9924
0.0865
0.8036
0.9999
31.3868 dB
0.2% noise, σ=0.75, 0.5% anchors
1965
0.2435
1.0068
0.0929
0.8017
1.0001
32.5423 dB
0.2% noise, σ=0.75, 1.0% anchors
3933
0.2162
1.0086
0.0886
0.7973
1.0002
31.4024 dB
0.2% noise, σ=0.75, 2.0% anchors
7863
0.2179
1.0048
0.0933
0.7966
1.0000
34.2686 dB
Coupled shallow-water operator identification with true parameters (μ,H,r,f,g)=(0.2,1.0,0.08,0.8,1.0). Multi-channel derivative columns identify the physical operator from sparse forcing anchors.
Coupled shallow-water operator identification at 0.2% observation noise with σ=0.75 pre-smoothing. Multi-channel physics anchors recover the operator and reconstruct the weather-core state without training a coordinate network.
This result is stronger than the scalar anchor test in two ways. First, the multi-channel structure anchors the coupling constants f and g very tightly. Second, even with observation noise, small forcing-anchor fractions recover the coupled state above 31 dB. The remaining weak parameter is μ, the artificial height-relaxation gauge, because it is weakly excited relative to the wave and forcing terms. This suggests a practical design rule for weather-grade OSNR: learn physically meaningful coupling symbols from multi-channel states, and treat gauge/damping terms with explicit priors or assimilation-window constraints.
FRI-Guided Adaptive Sparse Tier
Uniform sparse dictionaries fail at sub-pixel discontinuities because a step located at τ∈/TZ cannot be represented by a finite block of rigid grid atoms without tail error or leakage. FRI theory instead recovers the innovation coordinate first [vetterli2002fri,dragotti2007moments].
For a stream of K weighted Diracs, w(t)=k=1∑Kakδ(t−τk), moments satisfy mℓ=∫tℓw(t)dt=k=1∑Kakτkℓ. The annihilating filter or matrix-pencil method recovers the roots τk. Once τk are known, sparse step atoms are snapped exactly to those locations. This changes the sparse tier from an approximation grid into an adaptive representation of the true innovation geometry.
Hybrid Sparse-Plus-Smooth Decomposition
Let As=Asmooth,Ax=Asparse. The naive alternating update cs=cargmin∥y−Axz−Asc∥22 is not wrong by itself, but it is incomplete if implemented as if the two bases are orthogonal. The full block normal equations contain cross terms: [As⊤AsAx⊤AsAs⊤AxAx⊤Ax][cscx]=[As⊤yAx⊤y]. The cross-Gram matrix Across=As⊤Ax must appear directly in the right-hand side of block updates: (As⊤As)cs=As⊤y−Acrossz,(Ax⊤Ax+ρI)cx=Ax⊤y−Across⊤cs+ρz−u. This prevents smooth atoms from absorbing sparse shocks and prevents sparse atoms from chasing smooth energy. It is the finite-dimensional expression of the hybrid-spline coupling described by Debarre, Aziznejad, and Unser [debarre2019hybrid,debarre2021composite].
Failure Modes and Corrections
Uncalibrated knot grids
Defect: using weight 1 and spreading biases over a fixed interval. Correction: use vk(x)=x/T−k.
Sequential smooth-first fitting
Defect: solving the smooth component first lets the smooth basis approximate discontinuities through oscillatory combinations, producing Gibbs residuals. Correction: use joint or cross-Gram-shielded updates.
Joint coherent dictionaries
Defect: concatenating coherent dictionaries and applying naive ADMM can allocate smooth energy into sparse atoms and vice versa. Correction: explicitly include cross-Gram blocks and stabilize each block solve.
Partition-of-unity trap
Spline systems reproduce constants. A contiguous block of step atoms also reproduces a constant over an interval. This is not merely high coherence; it is an identifiability collision. Let χm(x)=1x≥τm be step atoms sorted by their knot locations. On any interval tiled by an active adjacent block, a difference or finite linear combination of these atoms can reproduce an indicator plateau 1[τa,τb)(x)=χa(x)−χb(x). Inside the plateau support, this function is exactly constant. At the same time, valid cardinal and exponential spline spaces satisfy partition-of-unity conditions and reproduce the global constant mode. Hence, after restriction to a local active support, the sparse step block and the smooth E-spline block contain indistinguishable constant directions.
During active-support debiasing, Af=[As∣Ax,active], the columns can become locally indistinguishable. Then Af⊤Af is singular or nearly singular. This is the numerical origin of the observed coefficient explosions when unregularized lstsq was applied to a joint smooth-plus-step active set.
Correction: use a scale-invariant Tikhonov solve, consistent with ADMM/proximal regularization views of Bayesian denoising [nguyen2018regularizers], cf=(Af⊤Af+γI)−1Af⊤y, with γ=ϵmean(diag(Af⊤Af)). In the production sparse core, ϵ=10−6. This makes the ridge invariant to the absolute scaling of the dictionary and shifts the zero singular directions by an amount proportional to the local Gram energy. The FRI stage further reduces the degeneracy by snapping sparse knots to physical innovation coordinates before debiasing, so the active sparse columns describe true shock interfaces rather than a diffuse uniform-grid approximation. In the batched Sprint 3 validation, the resulting ridge-stabilized debiasing matrices remain bounded with maximum condition number 18.1639<19.0 while preserving 118.78 dB mean reconstruction precision and 97.7% hard-zero sparse parameters.
Benchmarks
The current prototype benchmark scripts demonstrate:
calibrated Helmholtz coefficient recovery with zero autograd graph construction;
cascading derivative evaluation with machine-precision PDE residuals;
FFT/circulant inversion at O(MlogM) complexity;
adaptive FRI sparse shock recovery with snapped knots and ridge-stabilized debiasing;
batched multi-edge sparse core execution;
second-order Hermite neural-operator coefficient recovery through block-circulant 3×3 Fourier solves;
2D tensor-product Hermite fluid simulation through block-circulant 9×9 Fourier solves;
first SOTA-style SIREN comparison on a multi-edge non-bandlimited silhouette.
Benchmark environment
All timings in this draft were measured on a local Apple Silicon workstation. The benchmark scripts were executed as standalone Python processes from the repository virtual environment using torch.no\_grad() for every numerical solve. No PyTorch backward pass was constructed in any benchmark.
ll@
Component
Configuration
CPU / SoC
Apple M4 Max
Memory
128 GiB unified memory
Architecture
arm64
Operating system
macOS 15.7.4, build 24G517
Python
3.14.3
PyTorch
2.12.0
NumPy
2.4.6
LaTeX compiler
Tectonic 0.16.9
Hardware and software environment used for the current OSNR benchmark run.
Numerical results
Table [tab:benchmarks] records the current measured outputs of the repository benchmark scripts after the calibrated-grid, cross-Gram, FFT, FRI, Hermite block-Gram, and ridge-stabilization corrections. The scripts live under the repository's benchmark and source directories. The timings are wall-clock processing times reported by the scripts and should be interpreted as prototype measurements rather than final library-level performance claims.
p0.28linewidthp0.32linewidthp0.32linewidth@
Tier
Mechanism
Official package validation profile
Tier 1 steady-state
Pole-locked trigonometric E-splines plus circulant Fourier division
Measured benchmark results for the current OSNR prototype scripts.
Interpretation
The Helmholtz and cascading-derivative experiments isolate the Tier 1 deterministic case. Their high PSNR values and machine-precision PDE residuals confirm that, once the cardinal grid is calibrated, coefficient recovery in an operator-matched basis can replace iterative PINN-style optimization for this synthetic null-space task. The circulant experiment verifies the expected O(MlogM) path when periodized shift-invariant structure is available.
The sparse experiments test the Tier 2 adaptive case. The single-edge FRI shock tracker recovers the discontinuity at x=0.834200, snaps the sparse atom to that location, and obtains 124.84 dB PSNR with 99.2% hard-zeroed sparse coefficients. The batched adaptive sparse core extends this to batch size B=3 with K=3 innovations per signal, achieving mean PSNR 118.78 dB, minimum PSNR 116.73 dB, maximum edge localization error 7.627232e−14, and ridge-stabilized condition number below 19.
The Hermite operator experiment tests the Tier 3 coefficient-space trunk. A batched value/slope/curvature payload is passed through the exact Hermite block-circulant Gram system and recovered by independent 3×3 Fourier-domain solves. The boundary knots are overwritten with clamped value, slope, and curvature vectors, giving a structural boundary residual of exactly zero without a boundary loss term. The measured PSNR of 157.60 dB confirms that the Hermite trunk can act as a deterministic neural-operator synthesis layer under the matched periodic benchmark model.
Nonlinear PDE validation: Burgers PINN versus OSNR-liquid rollout
The first nonlinear PINN-facing CFD experiment is apps\_industrial\_breakthrough/burgers\_liquid\_pinn\_challenger.py. It solves the viscous Burgers equation ut+uux=νuxx,x∈[−1,1),ν=0.01/π, with periodic boundary conditions and initial condition u(x,0)=−sin(πx). A high-substep pseudospectral rollout provides the reference trajectory on a 128×101 space-time grid. The evaluation is intentionally split in time: the first 70% of frames are available for fitting or residual correction, while the final 30 frames are held out as a future prediction window.
The comparison has three profiles. The first is a compact SIREN-style coordinate PINN trained with Adam on data samples, initial-condition samples, periodic boundary consistency, and the Burgers residual computed by backward-mode automatic differentiation. The second is a truncated OSNR spectral rollout that keeps only a fixed number of Fourier/operator modes and evolves the known PDE directly without training. The third adds a tiny exact liquid residual in coefficient space. The liquid cell is trained only on the residual Fourier coefficients in the training window and uses the closed-form update xk+1=γk(xk−ω+∑sfs(ck)∑sAsfs(ck))+ω+∑sfs(ck)∑sAsfs(ck). Here ck contains the truncated spectral coefficients and normalized time. This is a deliberately hybrid experiment: the spectral rollout remains no-autograd, while the liquid residual uses a small training graph to learn truncation-error compensation.
First viscous Burgers CFD/PINN challenger. Metrics are computed on the held-out future window for the first two columns. The liquid rows train a tiny coefficient-space residual corrector; they are reported as hybrid OSNR-liquid profiles rather than zero-autograd deterministic solves.
Burgers space-time heatmaps for the quality-first 32-mode run. Left to right: reference trajectory, truncated OSNR spectral rollout, OSNR plus exact liquid coefficient residual, and SIREN PINN.
The follow-up capacity sweep, apps\_industrial\_breakthrough/burgers\_liquid\_capacity\_sweep.py, separates three effects: retained operator modes, liquid hidden-state count, and forecast horizon. The result is not that arbitrarily larger liquid networks replace resolution. Instead, the dominant lever is still the operator basis. Increasing the spectral support from 8 to 48 retained modes raises the zero-training future PSNR from 11.6562 to 38.2908 dB on the 70%/30% train/future split. The liquid residual is most valuable when the basis is deliberately compressed: at 16 modes and a 50%/50% split, it raises future PSNR from 19.5291 dB to 22.2635 dB with 128 liquid states. At 48 modes the same scaling gives only a sub-dB correction because little truncation error remains.
Burgers OSNR-liquid capacity sweep. The liquid cell is useful as a compact coefficient-space truncation-error corrector, especially under aggressive mode budgets. Once the operator basis is sufficiently resolved, additional liquid capacity yields diminishing returns.
Best capacity-sweep Burgers profile. Left: reference trajectory. Right: 48-mode OSNR rollout with 128 liquid states on the 70%/30% train/future split, reaching 38.8746 dB future PSNR.
The spline-native liquid implementation apps\_industrial\_breakthrough/burgers\_operator\_spline\_liquid.py then replaces the generic sigmoid gates by compact cubic B-spline conductance banks and solves the liquid neuron ODE with an exponential Green update over Gauss–Legendre nodes inside each time cell. This is now treated as a rejected prototype rather than the final architecture: cubic B-spline gates are not matched to the liquid neuron operator, and numerical quadrature reintroduces the approximate integration step that OSNR is meant to remove. The ablation is still useful because it shows that merely making the gates ``spline-shaped'' is insufficient. A correct operator-spline liquid layer must derive its basis from the neuron ODE itself.
First operator-spline liquid Green-solver ablation. The implementation uses spline-parametric conductance and forcing fields and solves the neuron ODE by exponential Green steps, but this initial parameterization is slower and less accurate than the generic sigmoid liquid control. This table identifies the next mathematical bottleneck: the spline-liquid state must be tied more directly to modal residual coefficients or initialized by a coefficient-space linear solve.
Composite Operator-Spline Liquid Networks
Motivation: liquid dynamics as an operator equation
Liquid time-constant networks and closed-form continuous-time networks model hidden states as continuous-time ODEs whose coefficients are modulated by the input and by the state itself [hasani2020ltc,hasani2022cfc,cantini2025exact]. In scalar form, a liquid neuron can be written as x˙(t)=−[wleak+f(I(t),x(t);θ)]x(t)+f(I(t),x(t);θ)A. The productive OSNR interpretation is not to regard this as a black-box recurrent layer. It is a first-order operator equation with a known leak component and an innovation term. Splitting the deterministic leak from the nonlinear synaptic feedback gives Lx(t)=(D+wleak)x(t)=s(t),s(t)=f(I(t),x(t);θ)(A−x(t)). The operator L=D+wleak has Green function g(t)=e−wleaktH(t), so the isolated liquid neuron is exactly an exponential-memory system. The correct OSNR liquid layer should therefore use exponential/operator splines matched to D+wleak, not polynomial splines inserted as generic nonlinear activations.
This viewpoint changes the computational target. Standard neural ODE, ODE-RNN, and latent-ODE implementations propagate states through numerical solvers and differentiate through either the solver trace or an adjoint system [chen2018neuralode,rubanova2019latentode]. PINNs impose related continuous constraints by adding automatic-differentiation residuals at collocation points [raissi2019pinn]. Standard numerical LNN implementations likewise propagate states sequentially by an ODE solver or by a constrained closed-form approximation. OSNR instead asks whether the entire hidden trajectory can be represented in an operator-matched spline space x(t)=k∑c[k]βL(t−k), where βL is generated by the liquid leak operator. In the shift-invariant, fixed-conductance case, the coefficient-domain normal equations inherit a Toeplitz/circulant temporal structure and can be diagonalized by a one-dimensional FFT. In the state-dependent case, the exact global system is nonlinear; the mathematically controlled route is to isolate the nonlinear part as an innovation process s(t) and solve the leak-filtered liquid trajectory exactly for a proposed innovation representation.
Transfer-function bridge to state-free sequence models
The strongest modern precedent for the OSNR liquid formulation is the transfer-function view of state-space sequence models. A linear time-invariant state-space model x˙(t)=Ax(t)+Bu(t),y(t)=Cx(t)+Du(t) has Laplace-domain response Y(s)=[C(sI−A)−1B+D]U(s)=H(s)U(s). Parnichkun et al. parameterize this dual representation directly as a rational transfer function and evaluate sequence blocks by FFT-based state-free inference, avoiding materialization of a hidden state tensor across the whole sequence [parnichkun2024statefree]. This is conceptually aligned with OSNR's block-circulant operator calculus: when the governing operator is shift invariant, the recurrent scan can be replaced by a frequency-domain multiplication or division.
For the scalar liquid leak operator, the transfer function is the first-order rational filter Hleak(s)=s+wleak1. Thus, once an innovation trajectory s(t) has been specified, the liquid state satisfies X(s)=Hleak(s)S(s),x(t)=g∗s,g(t)=e−wleaktH(t). On a uniform periodized grid this becomes a single FFT solve, while on nonuniform intervals it becomes the closed-form Green update derived below. The distinction is essential: Parnichkun's state-free result applies to linear transfer operators; OSNR does not claim that the original nonlinear liquid conductance is globally diagonalized. The OSNR claim is that the nonlinear term can be represented as a structured continuous innovation field and that the leak-filtered liquid trajectory can then be recovered by the exact transfer operator.
This reframes the proposed architecture as a Liquid State-Space Spline Operator. Compared with pure rational transfer-function layers, the OSNR addition is the continuous sparse-plus-smooth innovation sieve: FRI atoms capture non-bandlimited temporal events, and smooth exponential/DCT modes capture low-frequency drift. This is the part needed for impact-like events, irregular measurements, and causal regime changes emphasized in continuous-time sequence papers [rubanova2019latentode,lechner2020odelstm,vorbach2021causal]. Compared with CfC and exact recursive LTC formulas, the OSNR contribution is not merely another cell update; it is a block trajectory representation that can use transfer-function diagonalization for uniform components and exact exponential Green kernels for local nonuniform components.
Dual-continuum liquid innovation model
The liquid innovation s(t) need not be dense. Sequential signals often combine smooth trends with abrupt events: contact impacts, gait transitions, sensor dropouts, arrhythmia-like spikes, or regime switches. Continuous-time circuit policies and causal navigation experiments show that structured continuous dynamics can improve robustness and interpretability, but they still train recurrent neural circuits by gradient-based rollout rather than solving the operator algebraically [lechner2020ncp,vorbach2021causal]. OSNR therefore models the innovation side as a sparse-plus-smooth continuum, Lx(t)=ssparse(t)+ssmooth(t). The sparse tier is a temporal FRI model, ssparse(t)=r=1∑Rarφ(t−τr), where matrix-pencil/TLS moment recovery estimates the nonuniform event times τr. The smooth tier is represented by a low-frequency orthonormal dictionary such as a DCT or, more strictly, by an exponential-spline residual dictionary matched to the leak-filtered temporal statistics. A cross-Gram shielding step is required exactly as in the image/video model: Across=Dsmooth⊤Asparse, so smooth coefficients do not absorb sharp liquid events and sparse atoms do not duplicate slow drift. This is the liquid-network analogue of the OSNR sparse-plus-smooth decomposition.
Global FFT solve under fixed leak
For a uniform temporal grid and fixed leak wleak, the sampled operator L=D+wleak is shift invariant under periodic or circulant boundary closure. Let c be the coefficient vector of the hidden trajectory and let s be the sampled innovation coefficients. The discrete operator relation has the form Lc=s, where L is Toeplitz/circulant up to boundary treatment. With circulant closure, c[ω]=D[ω]+wleaks[ω], and all frequency bins are solved concurrently by torch.fft.fft. For multiple liquid channels, this becomes either independent scalar divisions or small block solves when channels are coupled. If the coupling is constant, the frequency-bin update is c[ω]=(D[ω]I+Λ−W)−1s[ω], which is the spline-operator analogue of a rational transfer-function SSM. This is the non-iterative OSNR alternative to sequential liquid rollout, but it is exact only for the linear leak-filtered solve once the innovation sequence has been specified or estimated.
Closed-form operator-spline liquid cell
The liquid residual layer should be formulated as an operator-spline ODE solver, not as a generic recurrent neural network with spline activations. For a scalar liquid state, x˙i(t)=−λi(t)xi(t)+bi(t),λi(t)>0, the governing operator on a local time interval is Li,n=D+λi,n,t∈[tn,tn+1], after freezing or spline-predicting the conductance rate λi(t) on that cell. The correct basis is therefore the Green/operator spline of D+λi,n, not a polynomial cubic spline. Let h=tn+1−tn and write local time as τ=t−tn∈[0,h]. The exact variation-of-constants formula is xi,n+1=e−λi,nhxi,n+∫0he−λi,n(h−τ)bi,n(τ)dτ. To make this integral algebraic, the forcing is represented in an exponential-polynomial spline space on the same interval, bi,n(τ)=m=1∑Mρqi,n,meρmτ, where the poles ρm are chosen from the residual dynamics to be corrected: ρ0=0 for constant forcing, real negative poles for dissipative memory, imaginary pairs ±jω for oscillatory modes, and repeated poles when polynomial-exponential terms are required. Substitution gives the closed-form kernel Kλ,ρ(h)=∫0he−λ(h−τ)eρτdτ=λ+ρeρh−e−λh,λ+ρ=0. The removable singular case is handled by the analytic limit Kλ,−λ(h)=he−λh. Thus the exact operator-spline liquid update is xi,n+1=e−λi,nhxi,n+m=1∑Mρqi,n,mKλi,n,ρm(h). This is the central closed-form expression for the OSNR liquid cell.
The conductance and forcing coefficients must also live in coefficient space. Let cn denote the OSNR PDE coefficients after the deterministic operator step, and let rn denote the unresolved innovation or truncation residual to be modeled. Define a small set of operator-aligned sensors zn=[⟨ϕ1,cn⟩,…,⟨ϕS,cn⟩,⟨χ1,rn⟩,…,⟨χR,rn⟩], where ϕs and χr are modal or Hermite coefficient probes, not raw coordinate samples. The positive liquid rate is then λi,n=ωi+softplus(ℓ∑ai,ℓηℓ(zn)),ωi>0. Here ηℓ should be an exponential-spline dictionary matched to the coefficient process, for example modes generated by poles μℓ of an AR/CAR residual model. The interval forcing coefficients are qi,n,m=ℓ∑bi,m,ℓηℓ(zn). The liquid state contributes back to the PDE only through residual coefficient channels, cn+1=ΦOSNR(cn)+Bxn+1, where ΦOSNR is the deterministic operator-spline PDE step and B maps liquid states into the truncated/high-frequency innovation subspace. This prevents the learned liquid system from overwriting coefficients already explained by the physical operator.
Stability.
The closed-form update is contractive in the homogeneous part whenever λi,n>0: ∣e−λi,nh∣<1. If λi,n≥λmin>0 and ∣qi,n,m∣≤Qm, then ∣xi,n+1∣≤e−λminh∣xi,n∣+m∑Qm∣Kλi,n,ρm(h)∣. For residual poles with Re(ρm)≤0, the kernel is uniformly bounded over finite h. This gives a direct route to stable long-horizon rollout: enforce positive rates, bound the forcing coefficient functionals, and restrict the liquid-to-PDE map B to residual subspaces.
Algorithm.
The resulting implementation should follow this sequence.
Advance the physical field coefficients by the deterministic OSNR operator step: cn+1=ΦOSNR(cn).
Project the unresolved defect or modal state into operator-aligned sensors zn.
Compute positive rates λi,n and exponential forcing coefficients qi,n,m.
Update each liquid state with the closed-form kernel Kλ,ρ(h), using the analytic limit for λ+ρ=0.
Inject Bxn+1 only into the residual coefficient band, producing cn+1.
The learnable objects are therefore not generic recurrent weights: they are the sensor dictionary coefficients, the forcing coefficients qi,n,m, the positive rate functionals, and the residual injection map B. This is the mathematically defensible Operator-Spline Liquid Network target for the next implementation.
Repository algorithm target
The first principled implementation should be a separate module rather than another residual experiment. The intended script, apps\_industrial\_breakthrough/liquid\_transfer\_operator.py, should implement the following deterministic sequence.
Generate or ingest a continuous-time sequence tensor with explicit sample times, initially a Walker2d-like kinematic track with shape B×C×T.
Fit the smooth innovation tier in an orthonormal or exponential-spline residual dictionary, with cross-Gram shielding against the sparse tier.
Solve the leak-filtered hidden trajectory by the exponential transfer relation (D+wleak)x=s, using FFT diagonalization for the uniform component and closed-form Green updates for local nonuniform event cells.
Report trajectory RMSE, innovation sparsity, solve latency, and autograd allocation. The deterministic solver path should run under torch.no\_grad().
The benchmark comparison should be staged. First, compare against the existing generic sigmoid-liquid and rejected cubic-spline-liquid controls on synthetic sequences where the exact innovation structure is known. Second, compare the state-free transfer solve against recurrent CfC/LTC/Cantini-style exact cells on irregular synthetic sequences, separating numerical exactness from recurrent scan cost. Third, move to public liquid-network sequence benchmarks and compare against CfC, LTC, Neural ODE, ODE-RNN, and transfer-function SSM baselines under identical train/test splits. Only the first stage supports exact mathematical claims; the second and third stages are external efficiency and SOTA validation.
Controlled validation: state-free liquid transfer solve
The first controlled validation of this direction is apps\_industrial\_breakthrough/liquid\_transfer\_operator.py. The benchmark isolates the linear transfer claim from dataset, training, and irregular-sampling confounders. It evaluates (D+wleak)x=s on a 512-sample periodized grid with wleak=0.5 and float32 tensors. The state-free solver applies the rational transfer function H(iω)=iω+wleak1 by one-dimensional FFT division. This validates the same linear state-free mechanism used by transfer-function state-space layers, but with the innovation s represented by OSNR's sparse-plus-smooth model.
The validation uses three profiles. The first is a smooth periodic innovation made from a small number of Fourier modes, for which the analytic periodic Green response is known. The second is a sparse impulse innovation s(t)=r∑arδ(t−τr), where the event locations τr are recovered from Fourier moments by the regularized TLS matrix-pencil method. The third combines the sparse impulses with the smooth periodic component and estimates a joint sparse-plus-smooth frequency model. This last profile is a simple least-squares oblique separation, not yet the full ADMM cross-Gram solver.
p0.24linewidthrrrrrr@
Profile
Traj. RMSE
Innov. RMSE
Event err.
OSNR ms
Rec. ms
Sparse
smooth\_periodic
1.1000e−07
3.2166e−07
n/a
0.2142
19.4441
98.83%
sparse\_impulse
6.4509e−07
1.4450e−05
3.7253e−09
0.4932
21.1861
99.22%
mixed\_sparse\_smooth
1.1767e−05
2.8567e−03
4.0978e−08
0.3237
22.2715
93.30%
Controlled state-free liquid transfer validation on a 512-sample periodized sequence. The recurrent row is a periodic recurrent Green rollout reference, not a trained nonlinear LTC/CfC model. Event errors are measured in normalized time units.
Table [tab:liquid-transfer-validation] shows that the FFT transfer solve recovers the leak-filtered state at near-float32 accuracy for smooth and sparse controlled inputs, while reducing the measured block latency from roughly 19–22 ms for the recurrent Green reference to less than 0.5 ms. The mixed profile has a larger innovation error because sparse impulses have broadband Fourier support and the smooth dictionary is intentionally truncated; nevertheless, the leak-filtered trajectory error remains 1.1767e−05 with 93.30% structural sparsity. These results validate the state-free linear liquid transfer layer and the sparse-plus-smooth innovation separation under controlled periodic assumptions. They do not yet claim superiority over trained nonlinear LTC/CfC models on external datasets; that comparison is the next benchmark stage.
The mixed profile is intentionally harder than the separated profiles because Dirac atoms occupy all Fourier frequencies. After the event times are localized, the script solves a joint sparse-plus-smooth least-squares system over sparse atoms and retained low-frequency smooth modes. The remaining innovation error therefore measures dictionary cross-talk and smooth-mode truncation rather than failure of the transfer solve itself. The much smaller trajectory error indicates that the stable leak transfer function attenuates part of this residual mismatch before it reaches the liquid state trajectory.
Controlled validation: irregular causal Green evaluation
The second validation script, apps\_industrial\_breakthrough/liquid\_transfer\_irregular.py, removes the periodized uniform-grid assumption from the state evaluation stage. It samples a 512-point nonuniform time grid on [0,1], keeps wleak=0.5, and evaluates the causal Green response x(ti)=∫0tie−wleak(ti−u)s(u)du+τr≤ti∑are−wleak(ti−τr) directly at all irregular sample locations. Smooth forcing terms are integrated analytically on each interval; sparse events remain off-grid. The recurrent reference is an exact causal Green rollout over the same nonuniform intervals, while the approximate baseline is a zero-order-hold recurrent update that represents the kind of local forcing approximation used by simple closed-form recurrent cells. Event times are again recovered from exact sparse Fourier moments, so this remains a controlled operator validation rather than a noisy inverse problem.
p0.23linewidthrrrrrrr@
Profile
OSNR RMSE
Rec. RMSE
ZOH RMSE
Event err.
OSNR ms
Rec. ms
Sparse
smooth\_irregular
0.0000e+00
3.2506e−08
1.2631e−03
n/a
0.2467
27.6731
100.00%
sparse\_offgrid
3.6154e−07
4.7343e−07
4.7343e−07
6.1467e−08
0.6216
3.0790
99.22%
mixed\_irregular
3.6154e−07
1.0355e−06
1.2634e−03
6.1467e−08
0.2482
27.3737
99.22%
Controlled irregular causal liquid transfer validation. The OSNR column evaluates the closed-form Green response directly at nonuniform sample times. The recurrent column is an exact causal Green rollout; the ZOH column is a local zero-order forcing approximation.
Table [tab:liquid-irregular-validation] shows that the operator-spline Green evaluation preserves near-float32 agreement with the exact recurrent causal reference while avoiding the sequential scan over the full history. The zero-order recurrent approximation is accurate for pure off-grid impulses but loses roughly 10−3 RMSE on smooth forcing because it freezes the drive inside each irregular interval. This is the next step toward LNN relevance: the liquid response can be evaluated at irregular times by closed-form operator kernels, not only by a periodic FFT block. The remaining open problem is the harder one: estimating sparse-plus-smooth innovations from noisy irregular observations rather than from controlled moment access.
The third liquid validation script, apps\_industrial\_breakthrough/liquid\_innovation\_inverse.py, begins to address the inverse problem. Instead of giving the solver exact innovation moments, it observes noisy irregular samples of several leak-filtered liquid traces driven by the same hidden innovation. This multi-leak setting is intentional: a single scalar trace only identifies the interval containing an off-grid impulse from adjacent samples, while two or more leak rates identify the event location inside the interval through cross-leak residual ratios.
For leak rates {0.35,0.70,1.25} and a shared sparse-plus-smooth innovation, the script computes interval residuals rj,i=xj(ti+1)−e−wj(ti+1−ti)xj(ti). For an event τ∈(ti,ti+1], the sparse contribution obeys rj,ievent=ae−wj(ti+1−τ). Thus ratios across leak channels localize τ, after which a joint least-squares solve estimates event amplitudes and smooth forcing coefficients. This is still a controlled inverse problem: the smooth forcing dictionary and leak rates are known, and the detector is not yet a robust noisy-data estimator.
rrrrrrr@
Noise
Traj. RMSE
Innov. RMSE
FD innov.
Event err.
Events
Sparse
0.0e+00
1.4691e−07
8.8908e−07
4.7476e+01
1.7136e−07
4
99.54%
1.0e−05
2.1515e−05
6.7998e−05
4.7477e+01
4.7682e−05
4
99.54%
5.0e−05
6.4065e−05
1.8110e−04
4.7474e+01
1.9553e−04
4
99.54%
Controlled inverse liquid innovation recovery from irregular multi-leak observations. The finite-difference baseline estimates (D+w)x locally from one channel and is dominated by off-grid impulse discontinuities.
Table [tab:liquid-inverse-validation] shows that, in the noise-free setting, the multi-leak inverse recovers both the trajectory and the smooth innovation near float32 precision while localizing off-grid events to 1.7136e−07 normalized time units. At modest observation noise, the trajectory remains in the 10−5–10−4 RMSE range. The large finite-difference innovation errors confirm the expected failure mode of local derivative estimates on discontinuous off-grid events. The next algorithmic requirement is a noise-robust event detector and regularized sparse-plus-smooth inverse solve; without that layer, this result should be read as an identifiability and controlled recovery validation rather than a full real-world sequence benchmark.
Controlled validation: robust inverse noise sweep
The robustness follow-up, apps\_industrial\_breakthrough/liquid\_inverse\_robust\_sweep.py, compares the residual-ratio inverse against a matched-dictionary OMP variant and a practical local refinement. The OMP solver precomputes a candidate off-grid event dictionary over each irregular interval, alternates event selection with a joint sparse-plus-smooth refit, and reports the pursuit/refit latency after dictionary setup. This is a more robust but grid-quantized detector: it gives up some low-noise event precision in exchange for stability under larger observation noise. The local refinement keeps OMP's selected support but re-estimates each event time by a small closed-form multi-leak search inside the selected interval, followed by one global coefficient refit. This removes most of the useful quantization error without the multi-second cost of full variable projection. A ratio-after-OMP refinement is also implemented as an optional ablation, but it inherits the high-noise collapse of the ratio estimator and is not used as the default path.
rlrrrrr@
Noise
Method
Traj. RMSE
Innov. RMSE
Event err.
Latency ms
Sparse
0.0e+00
ratio
1.4691e−07
8.8908e−07
1.7136e−07
6.9665
99.54%
0.0e+00
OMP
9.0559e−05
5.6460e−07
4.1162e−04
5.4126
99.54%
0.0e+00
OMP-local
1.4395e−05
4.4244e−07
2.4807e−05
10.8437
99.54%
1.0e−04
ratio
2.0550e−04
6.0252e−04
5.3670e−04
5.6030
99.54%
1.0e−04
OMP
1.7557e−04
6.0251e−04
4.7795e−04
5.3485
99.54%
1.0e−04
OMP-local
8.9944e−05
6.0246e−04
2.1823e−04
11.3056
99.54%
5.0e−04
ratio
5.2915e−01
2.5590e+00
7.8826e−04
5.4806
99.54%
5.0e−04
OMP
2.6358e−04
3.0507e−03
5.3227e−04
5.3720
99.54%
5.0e−04
OMP-local
1.7237e−04
3.0505e−03
3.0643e−04
10.9159
99.54%
1.0e−03
ratio
5.4091e−01
3.5202e+00
8.8251e−04
5.6368
99.54%
1.0e−03
OMP
7.2456e−04
4.5331e−03
6.3139e−04
5.4304
99.54%
1.0e−03
OMP-local
7.8459e−04
4.5333e−03
7.6676e−04
10.5497
99.54%
Robust inverse noise sweep on the controlled multi-leak liquid problem after vectorized event-design construction. The ratio method is more accurate at very low noise but fails catastrophically at higher noise. OMP remains stable under large noise; OMP-local reduces candidate-grid quantization at low and mid noise with roughly doubled millisecond-scale latency.
Table [tab:liquid-robust-inverse-sweep] identifies the next engineering boundary. The state recovery problem is no longer limited by the Green transfer operator; it is limited by sparse event detection under noisy interval residuals. The matched-dictionary OMP variant removes the catastrophic high-noise failures of the residual-ratio method, and the OMP-local variant recovers much of the lost continuous timing precision at low and mid noise while remaining in the 10–11 ms range for the full 512-step controlled problem. At the largest tested noise, local refinement can overfit the noisy leak residuals, so the toolbox should expose both OMP and OMP-local as selectable estimators rather than treating refinement as uniformly dominant.
Operator-matched innovation routing with an inferred operator
The preceding experiments either specify the leak operator or fit a complete trajectory. We next ask a different algorithmic question: can the operator be inferred from contaminated observations and then used as an analytic conditional-computation gate? The mechanical benchmark runner generates separate training and test trajectories with an independent RK4 simulator. Each trajectory is a damped oscillator with sparse impulses at unknown off-grid times. A robust iteratively reweighted fit estimates the two-state flow map from noisy value/velocity jets. The interval innovation is then rn=xn+1−Fxn, and a fixed median/MAD threshold fitted on the unlabelled training residuals routes event intervals. Given a routed interval, the continuous event offset δ∈[0,T] is estimated by projecting rn onto exp(Aδ)b, where A=T−1logF and b=(0,1)⊤. No event label, event count, exact moment, pole, or location is supplied to the estimator.
llrrrr@
Setting
Method
AP
Unsupervised F1
Timing MAE/T
Flow error
linear, 40 dB
finite difference
0.7800
0.1996
0.3673
0.8188
linear, 40 dB
position-only AR(2)
0.4758
0.5245
n/a
n/a
linear, 40 dB
ordinary two-state
1.0000
0.9983
0.0426
0.0130
linear, 40 dB
robust inferred operator
1.0000
0.9994
0.0425
0.0005
linear, 30 dB
robust inferred operator
1.0000
0.9994
0.1249
0.0020
linear, 20 dB
robust inferred operator
0.9907
0.8553
0.2830
0.0115
strong Duffing, 30 dB
hybrid operator
1.0000
0.9963
0.1222
0.0800
Blind operator and off-grid innovation routing over 12 seeds. The
event gate is unsupervised; AP is threshold-free. Roughly 5.5% of intervals
contain events.
Table [tab:blind-operator-innovation-routing] supports the routing mechanism while isolating its limits. The inferred Hermite-jet residual gives essentially perfect non-oracle event separation at 30–40 dB and estimates the clean flow much more accurately than ordinary least squares. However, event detection is easy enough in this benchmark that ordinary two-state least squares also detects almost every event. More importantly, at 20 dB the timing error is 0.2830T, worse than the 0.25T expected from always choosing the interval midpoint. Thus interval detection and sub-sample timing are distinct claims: the latter must be confidence-gated under noise. The position-only AR control is substantially weaker and cannot determine timing; the derivative channel in the Hermite jet supplies genuine off-grid information.
The field-level benchmark tests whether the same principle survives a PDE and model mismatch. A 1024-point pseudo-spectral Strang-splitting solver generates periodic advection–diffusion–reaction trajectories with compact cubic B-spline sources at off-grid locations. Estimation sees only 256-point block averages. A trimmed unlabelled fit identifies transport parameters on a separate trajectory, after which only the largest 1% of test residuals are retained as conditional corrections. The true injected support occupies 0.283% of space–time points.
llrrr@
Setting
Predictor
Source AP
Base nRMSE
nRMSE after 1%
linear
identity
0.4323
0.1077
0.0813
linear
learned spectral map
0.8604
0.0506
0.0036
linear
blind inferred operator
0.8906
0.0505
0.0011
mild cubic
blind inferred operator
0.8866
0.0515
0.0011
strong cubic
learned spectral map
0.8586
0.0673
0.0057
strong cubic
blind inferred operator
0.8630
0.0672
0.0045
strong cubic
hybrid operator
0.8794
0.0671
0.0017
Eight-seed field innovation routing. The correction budget retains
the largest 1% of each one-step residual.
In the linear case, the blind fit recovers speed 0.72000001, diffusivity 0.00180000, and decay 0.07999999, from respective true values 0.72, 0.0018, and 0.08. Table [tab:pde-innovation-routing] shows that an operator residual is substantially more compressible than a raw temporal difference and also improves over an unconstrained learned spectral transition. Under strong cubic mismatch the pure linear operator degrades, but a four-feature closed-form local residual model lowers the 1%-budget error from 0.0045 to 0.0017. This is evidence for an analyze--\allowbreak annihilate--\allowbreak route--\allowbreak reconstruct algorithm: preserve the inferred transport operator and spend flexible capacity on its localized mismatch. It is not yet a SOTA claim; both studies are controlled, the mechanical study observes the full value/velocity jet, and the PDE metric is one-step correction rather than autonomous rollout.
The next test closes the loop and isolates the spline contribution. A sender observes each new 256-point field while the receiver retains only its previous reconstruction. Both apply the same predictor; a median-plus-six-MAD gate, fitted without labels on a separate trajectory, either sends no update or one amplitude/location packet. The coordinate control sends a Kronecker impulse. The spline codec sends a block-averaged cardinal cubic B-spline at one of four sub-cell phases. The corrected receiver state is fed into the next prediction for all 139 transitions, so errors are allowed to accumulate.
llrrr@
Setting
Predictor / packet
Trajectory nRMSE
Terminal nRMSE
Payload
linear
learned spectral / point
0.03971
0.06479
0.199%
linear
blind operator / point
0.01713
0.02242
0.202%
linear
blind operator / cardinal
0.00965
0.01100
0.209%
mild cubic
hybrid operator / point
0.01731
0.02299
0.202%
mild cubic
hybrid operator / cardinal
0.00976
0.01119
0.210%
strong cubic
learned spectral / point
0.09334
0.12578
0.200%
strong cubic
blind operator / cardinal
0.12662
0.11532
0.298%
strong cubic
hybrid operator / point
0.02882
0.03788
0.201%
strong cubic
hybrid operator / cardinal
0.02286
0.02240
0.209%
Eight-seed closed-loop innovation codec. Payload includes a float32
amplitude and the location bits, and is normalized by one dense float32 field.
At matched predictor and nearly matched payload, the cardinal packet reduces trajectory error relative to a point packet by 43.1% in the linear case, 43.1% for the hybrid under mild nonlinearity, and 21.3% under strong nonlinearity, winning all eight paired seeds in each comparison. This is the specific value of compact cardinal reproduction: one coefficient reconstructs the off-grid event footprint rather than one sampled coordinate. The factorization is necessary as well as the spline. A temporal-difference gate fails because smooth transport dominates its robust scale estimate. Under strong cubic feedback, the pure inferred linear operator false-triggers and loses to the learned spectral control; the small local mismatch model is what restores sparse routing.
A public-data follow-up uses the released PDEBench Test-17 FNO predictions. The frozen FNO is the neural prior; samples 900–999 and forecast frames 8–20 yield 3900 held-out, three-channel, 256-point residual fields. Equal-accounting packets encode the target-time FNO residual, with location and scale bits included. This is a residual codec/assimilation test, not a blind forecast improvement.
lrrrrr@
Codec
Packets
Payload
Field nRMSE
Frobenius nRMSE
Residual left
point
9
4.395%
0.004064
0.001242
69.80%
DCT
9
4.395%
0.001941
0.000500
36.18%
one-scale cardinal
9
4.395%
0.003810
0.001169
65.52%
multiscale cardinal OMP
8
4.199%
0.001460
0.000390
27.36%
Test 19: point
9
4.395%
0.018400
0.006024
59.65%
Test 19: DCT
9
4.395%
0.018944
0.006492
62.81%
Test 19: one-scale cardinal
9
4.395%
0.016993
0.005611
55.21%
Test 19: multiscale OMP
8
4.199%
0.012304
0.004138
42.05%
Public PDEBench FNO residual coding. The uncorrected FNO has field
nRMSE 0.004906 and sample-wise Frobenius nRMSE 0.001431.
Eight multiscale packets use fewer bits than nine DCT packets yet reduce paired sample error by 10.1% (bootstrap 95% interval 4.3–15.6%) and win on 77% of samples. The one-scale cubic control is weak and DCT wins at the smallest budgets: hierarchical scale selection is essential. A direct Python OMP costs approximately 477μs per field, versus 8.5μs for DCT. Replacing the per-field loop by batched FFT correlations and batched Gram solves reproduces its errors within 4.44×10−10 at 138μs per field, a 3.5× speedup. A refit-free batched pursuit reaches 36μs with a small accuracy loss. The remaining latency gap is an explicit engineering boundary.
The same scales and budgets transfer without tuning to PDEBench Test 19. Eight multiscale packets again use fewer bits than nine DCT packets, but now reduce paired sample error by 32.5% (bootstrap 95% interval 30.3–34.8%) and win on all 100 held-out samples. Per-field/Frobenius nRMSE is 0.012304/0.004138, versus 0.018944/0.006492 for DCT. Even the single-scale cardinal control beats DCT on Test 19, while the multiscale dictionary remains decisively better. The fast refit-free pursuit retains a 31.5% paired reduction at approximately 35μs per field.
Two stronger dictionary controls sharpen this result. We add an exact orthonormal Haar transform and a channel-specific PCA/KLT basis fit only on samples 0–899, then frozen for samples 900–999. With eight spline versus nine control packets, the Test-17 paired reductions are 14.0% (95% interval 8.4–18.9%) against Haar and 6.8% (1.1–12.3%) against learned PCA. On Test 19 they are 17.9% and 32.3%, respectively, with 100/100 wins. The effect therefore survives both a localized multiscale control and a training-only learned residual basis.
A separate channel audit gives each control ten packets against eight spline packets and includes a shared 16-bit block scale, amplitude quantization, coefficient noise, packet loss, and noisy sender residuals. Eight-bit payload is 2.051% of dense float32 for the spline versus 2.148% for controls; at four bits all methods use exactly 1.660%.
llrrr@
Set
Channel
Spline
learned PCA
paired reduction (95% CI)
17
float32 clean
0.001460
0.001794
1.4% (−4.9–7.3%)
17
int4 clean
0.001496
0.001878
7.4% (3.0–11.5%)
17
int2 clean
0.002837
0.003145
1.0% (−2.1–3.9%)
19
float32 clean
0.012304
0.018570
30.7% (28.5–32.9%)
19
int4 clean
0.012407
0.018601
30.3% (28.2–32.5%)
19
int2 clean
0.017709
0.020802
13.4% (12.5–14.3%)
19
10-dB coefficient SNR
0.014436
0.019466
23.6% (22.0–25.3%)
19
20% packet loss
0.015939
0.020264
18.7% (17.4–20.0%)
Quantized/impaired public residual packets. Entries are mean
per-field nRMSE; paired intervals bootstrap the 100 held-out samples.
The boundary is dataset-dependent rather than cosmetic. Test 19 wins all 100 paired samples against learned PCA in every tested clean, quantized, noisy, and loss-impaired condition. On heterogeneous Test 17, however, the lower clean aggregate nRMSE does not imply a resolved paired advantage over the stronger ten-packet PCA control: the clean, eight-bit, 5% loss, and 20-dB sender-noise intervals cross zero. Four bits is positive at exactly matched payload, while at two bits all Test-17 transform-comparison intervals cross zero and the spline payload is slightly larger. This establishes a low-rate boundary and supports a regime-dependent matched-residual claim, not universal codec dominance. Moreover all public experiments still encode the true target-time residual; predicted-residual or two-dimensional tests are the next gate.
We tested the predicted-residual gate directly. At time t a causal codec may use only previously revealed residuals. A cross-channel spectral AR selects order and ridge on samples 800–899 after candidate fits on 0–799, then refits on 0–899. Test 17 selects order four: eight delayed spline packets improve paired sample error by 18.25% (95% interval 15.47–21.10%) over no correction. This is a causal correction win but not a spline win: DCT and learned PCA are slightly better (spline reductions −0.51% and −0.64%), while spline beats Haar by only 0.63%. Test 19 selects order one; spline packets worsen no correction by 0.30% (interval −0.52–−0.08%) and lose to all three transforms. Thus spatial residual compressibility does not imply temporal predictability, and an assimilation result cannot be relabelled as a forecast result. A structurally different causal state, or genuine target-time sensor assimilation, is required.
The two-dimensional gate is more encouraging but remains regime-dependent. On public 128×128 PDEBench FNO residuals, we compare eight tensor-product multiscale cubic-cardinal packets with nine point, 2D-DCT, exact 2D-Haar, single-scale cardinal, and separable-KLT packets. The KLT row and column bases are fit per channel on samples 0–79 and frozen on held-out samples 80–99. Including two-dimensional locations and scale indices, spline payload is 0.0748% of a dense float32 field versus 0.0790% for controls.
llrr@
Set
Codec
Sample nRMSE
spline reduction (95% CI)
26
point / DCT / Haar
0.002315/0.002334/0.002245
19.1/18.7/16.7%
26
one-scale cardinal
0.002187
14.9%
26
learned separable KLT
0.001889
2.08% (0.55–3.60%)
26
multiscale cardinal
0.001850
–
26
adaptive spline/KLT atlas
0.001828
1.12% vs. spline
26
cheap preselector atlas
0.001848
2.19% vs. KLT
27
learned separable KLT
0.001462
−6.48% (−8.46–−4.51%)
27
multiscale cardinal
0.001562
–
27
adaptive spline/KLT atlas
0.001449
1.01% vs. KLT
27
cheap preselector atlas
0.001451
0.88% vs. KLT
Public 2D target-time FNO residual coding on 20 held-out samples.
Eight multiscale packets use fewer bits than nine controls.
On Test 26 the spline beats all fixed controls on 20/20 samples and learned KLT on 15/20, so the positive paired interval establishes a small but real 2D matched-dictionary gain. Untuned Test 27 supplies the counterexample: the spline still beats point, DCT, Haar, and one-scale cardinal atoms, but learned KLT wins decisively. Tensor-product cardinal atoms are therefore a strong compact prior for some 2D residual geometries, not a universal replacement for a learned covariance basis.
The fixed-basis boundary suggests an adaptive atlas. Since an assimilation encoder observes the residual being coded, it may evaluate both reconstructions and send one mode bit choosing the lower residual norm. Including this bit, payload is 0.0792%. The atlas selects spline on 70.0% of Test-26 fields but only 12.3% of Test-27 fields. It significantly improves both experts: on Test 26, 1.12% versus spline (interval 0.46–1.89%) and 3.22% versus KLT; on Test 27, 6.91% versus spline and 1.01% versus KLT (interval 0.46–1.70%). This converts the transfer failure into a useful design principle: route innovations among compact analytic and learned basis experts, treating operator splines as a specialized expert rather than a universal representation. The remaining systems question is whether a cheap preselector can preserve the gain without evaluating every encoder.
A training-only closed-form preselector answers that systems question positively in this audit. Thirteen cheap residual statistics—energy and tail ratios, periodic derivative energies, radial spectral fractions, and top-nine DCT/Haar energy—feed a ridge predictor of the spline/KLT error ratio. On Test 26 it ties the pure spline statistically and beats KLT by 2.19% (interval 0.91–3.45%); on Test 27 it beats spline by 6.79% and KLT by 0.88% (interval 0.40–1.52%). It selects spline on 84.3% and 8.6% of the respective fields. An actual conditional CPU path computes features and runs only the selected encoder, reproducing the reference reconstruction exactly. It costs 5.61 versus 6.13 ms/field for exhaustive selection on Test 26 and 1.09 versus 5.75 ms/field on Test 27, measured 1.09× and 5.26× speedups. Exact dual-mode search remains 1.03% and 0.13% more accurate, quantifying the price of preselection.
The atlas also survives coefficient quantization. With a shared 16-bit block scale and ten KLT packets against eight spline packets, exact-atlas sample nRMSE at int8/int4 is 0.001819/0.001828 on Test 26, versus 0.001868/0.001877 for KLT (both 2.57% paired reductions), and 0.001429/0.001431 on Test 27, versus 0.001439/0.001441 for KLT (0.81%/0.80%). Cheap preselection also beats KLT with positive intervals in all four cases. Atlas payload is 0.0452% at int8 and 0.0376% at int4, only one mode bit above KLT. The fixed Test-26 spline/KLT interval crosses zero after quantization; adaptive routing, rather than the spline alone, is the robust contribution. An optimized accelerator implementation remains open.
Rate–distortion atlas guarantee.
Let Em(r) denote the reconstruction produced by packet codec m for a residual field r. With equal or padded packet rates, the encoder-side mode decision m∗(r)=argm∈{1,…,M}min∥r−Em(r)∥22 requires only ⌈log2M⌉ mode bits and satisfies ∥r−Em∗(r)∥22≤minm∥r−Em(r)∥22 field by field. Consequently the summed squared error, and hence sample Frobenius error, cannot exceed any constituent codec. Unequal-rate operation replaces the objective by Dm+λRm. The guarantee is elementary but important: complementing an operator-matched spline expert with a learned covariance expert is safe at negligible rate, and empirical gains quantify whether their errors are truly complementary rather than redundant.
The complementarity transfers across all five public static 128×128 Darcy residual sets, using samples 0–899 for training-only construction and 900–999 for evaluation.
rrrrr@
Test
Spline mode
KLT nRMSE
Atlas nRMSE
reduction (95% CI)
21
18%
0.041215
0.040302
2.65% (1.48–4.01%)
22
22%
0.025046
0.024501
2.01% (1.06–3.08%)
23
42%
0.008331
0.007922
4.02% (2.68–5.66%)
24
55%
0.005932
0.005487
5.97% (4.21–7.93%)
25
62%
0.005612
0.005193
5.68% (3.99–7.58%)
One-bit spline/KLT atlas on held-out public static 2D residuals.
Together with dynamic Tests 26–27, the exact atlas beats KLT on all seven public 2D sets, with every paired interval positive. Spline mode use spans 12.3%–70.0%, direct evidence of regime-dependent complementarity. Cheap preselection retains a significant KLT gain on six of seven sets; Test 22 is unresolved (0.81%, interval −0.16–1.81%).
The two-expert atlas is nevertheless incomplete: on two Test-29 CFD regimes, DCT is substantially better than both spline and cross-regime KLT. We therefore apply the same rate–distortion rule to four experts—multiscale cardinal, training-only separable KLT, DCT, and Haar—at the cost of two mode bits. Rerunning Tests 21–27, the exact four-mode atlas improves on the strongest constituent by 5.85%, 3.03%, 4.44%, 6.41%, 4.40%, 1.13%, and 1.01%, respectively; every paired interval is positive.
The harder transfer protocol leaves out each of the three Test-29 four-channel CFD configurations in turn. KLT bases and the closed-form selector use only the other two configurations, while all ten samples of the third are held out.
lrrr@
Held-out regime
Best fixed expert
Four-mode atlas
gain (95% CI)
M01\_Eta01
DCT 0.001239
0.001157
5.17% (2.47–8.51%)
M10\_Eta001
spline 0.010867
0.010846
0.28% (0.11–0.49%)
M10\_Eta01
DCT 0.002264
0.002085
5.05% (2.13–9.68%)
Leave-one-regime-out Test-29 residual coding. Training-only
statistics come from the other two CFD configurations; the encoder still
observes each target-time assimilation residual.
Thus exact routing beats the strongest included expert with a positive paired interval on all ten public 2D datasets/regimes. This does not follow merely from averaging: field choices change sharply with physics. Spline accounts for 24.8%, 87.3%, and 22.2% of the three Test-29 regimes, while DCT accounts for 60.0%, 2.3%, and 69.1%. A multi-output version of the 13-feature selector runs only its predicted encoder and is 1.58×– 4.49× faster than exhaustive four-mode evaluation. Its boundary is equally clear: it is significantly better than the strongest fixed expert on six of ten public sets, unresolved on three, and 1.29% worse than spline on held-out M10\_Eta001. Exact encoder-side selection is robust; low-cost out-of-regime routing remains an open learning problem. Most importantly, the result rejects universal spline dominance: operator splines are a complementary analytic expert inside a compact adaptive atlas.
The cheap-router failure can be reduced by enforcing more of the transform calculus. If Um is an orthonormal codec and Im,K(r) indexes its K largest coefficients, Parseval gives ∥r−Em(r)∥22=∥r∥22−k∈Im,K(r)∑∣⟨r,um,k⟩∣2. Thus KLT, DCT, and Haar need no learned error predictor: the encoder chooses their minimum-distortion member exactly from retained energy. Only the nonorthogonal spline comparison remains unknown. Our guarded pilot computes one residual FFT and one matched-filter inverse FFT per cardinal scale, providing first-step OMP capture and top-correlation energies without the eight OMP iterations or least-squares refits. A binary ridge gate then chooses between the spline and the Parseval-best orthogonal transform using these features, residual morphology, and channel type.
Across Tests 21–27 and the three leave-one-regime-out Test-29 configurations, the guarded pilot is significantly better than the strongest fixed expert on nine of ten sets and statistically tied on the tenth. In the three OOD CFD rows it improves on DCT by 3.08% and 4.13% in the two DCT-dominant regimes, and ties spline at −0.004% (interval crosses zero) in the spline-dominant regime. This removes the old cheap router's significant 1.29% OOD loss. The guarded pilot improves the old router on nine of ten sets, with Test 27 unresolved, while measuring 1.13×–3.28× faster than exhaustive encoding. Exact search remains 0.04%–2.20% better. A feature-only four-output regressor is a decisive ablation: despite receiving the same 28 pilot features, it still loses 0.54% to spline in the hard regime. The robust gain therefore comes from decomposing the decision by Parseval and learning only the nonorthogonal comparison, not from adding features indiscriminately. Pilot transform coefficients are cached for the selected synthesis. Quantization also preserves the analytic decision: for retained coefficients c and quantized coefficients q, distortion is ∥r∥2−2⟨c,q⟩+∥q∥2. This value matches explicit KLT/DCT/ Haar reconstruction to 3.9×10−15 and selects the true best transform on every audited field. At int8/int4, guarded-pilot gains against the strongest fixed expert are 1.23%/1.06% on Test 26 and 0.79%/0.78% on Test 27, all with positive intervals, at 0.0454%/0.0378% dense payload.
Sparse-station operator-kernel atlas.
We next remove the encoder's full-field residual access. On the last target frame of each public Test-29 forecast, the method receives contemporary residual values at only S random pixels and reconstructs the complete 128×128 four-channel residual. This is sparse-observation assimilation, not blind forecasting. Periodic IDW is the equal-observation baseline. The spline expert uses the sampled columns of a periodized cardinal cubic interpolation operator; a Gaussian kernel supplies a smooth radial control. Kernel scales and ridges are selected on balanced fields from the other two CFD regimes. Within a field, 75% of stations fit each expert and 25% validate it. Training-only routing/audit splits select a switching margin by minimizing the worst source-regime squared-error ratio to IDW; the chosen expert is then refit on all S observations.
lrrrr@
Held-out regime
stations
IDW nRMSE
guarded atlas
gain (95% CI)
M01\_Eta01
512
0.001023
0.000972
3.12% (0.00–7.41%)
M10\_Eta001
512
0.004882
0.004686
3.48% (0.34–8.09%)
M10\_Eta01
512
0.001593
0.001572
0.67% (0.01–2.00%)
Leave-one-regime-out sparse-station Test-29 assimilation. All
hyperparameters and routing margins use only the other two CFD regimes; 512
stations are 3.125% of the spatial grid.
The guarded atlas therefore improves periodic IDW point estimates at the predeclared 512-station operating point in all three unseen regimes. The last two gains have strictly positive bootstrap intervals; the first interval touches zero and is unresolved. It routes 12.5%, 50.0%, and 15.0% of fields, respectively, to a cardinal or Gaussian expert, so the improvement is not a renamed IDW result. The wider density sweep is deliberately mixed: seven of nine point estimates at S∈{256,512,1024} favor the atlas, three have strictly positive bootstrap lower bounds, five are unresolved, and M10\_Eta01 at 256 stations loses 1.66% (interval −4.08–−0.00%). Thus current evidence supports moderate-density guarded interpolation, not uniform safety under extreme sparsity.
A compact-packet ablation is negative. Sensor-only OMP with eight cardinal or nine KLT/DCT/Haar coefficients loses substantially to IDW, which retains all station values. At 512 stations, cardinal nRMSE is 0.003960 versus 0.001023 on M01\_Eta01, and 0.006970 versus 0.004882 on M10\_Eta001. Sparse local support alone cannot identify atoms that receive no measurements. The successful spline mechanism is therefore a full-capacity periodized interpolation operator used behind a conservative router, rather than aggressive coefficient packetization.
The resolved low-density loss suggests that routing uncertainty, rather than expert capacity alone, is the immediate failure. We therefore add a cross-fitted agreement gate. Two disjoint validation-station folds must independently select the same kernel, and on both folds it must improve on IDW by a nonzero margin. Otherwise the gate abstains to IDW. The margin is chosen from {0.10,0.20,0.30} on the other two regimes by the same worst-group training audit; the accepted expert is finally refit on all stations.
Three independent station permutations, three held-out regimes, and three station densities yield 27 comparisons. The cross-fitted gate has 19 positive point estimates, six changes within 0.0001% of an exact tie, and two unresolved negative estimates. Eleven bootstrap lower bounds are strictly positive and none is a resolved loss. By comparison, the original gate has nine resolved gains and one resolved loss. Cross-fitting changes the worst point result from −1.659% to an unresolved −0.496%, while average gain decreases from 1.743% to 1.372%. At 512 stations, mean gains across the three layouts are 2.52%, 2.78%, and 0.23% for the three regimes. Thus fold agreement plus exact IDW abstention is an effective empirical safety device, but not a formal no-harm certificate: the remaining two negative point estimates, although unresolved, prevent that stronger claim.
newpage We next strengthen the spline expert without increasing its coefficient count. Let Kh1 and Kh2 be unit-diagonal, periodized tensor-product cardinal cubic kernels at two scales. Their direct-sum RKHS kernel is Kmulti=1+ηKh1+ηKh2,c=(Kmulti(X,X)+λI)−1y,η>0. The construction remains positive and has one coefficient per station, equal to a single-kernel interpolant. Scale pair, mixture weight, and ridge are minimax-selected on balanced fields from the other two regimes.
Across three station permutations, three densities, and three held-out regimes, the fixed pyramid improves the single cardinal kernel in 24 of 27 point comparisons, with 22 resolved gains. Hierarchy therefore improves the representation itself, but is unsafe alone: two low-density comparisons are resolved losses. On the base layout, the pyramid beats IDW by 4.56% at 512 stations and 14.43% at 1024 in M10\_Eta001, but loses 37.90% and 24.37% to IDW at 256 stations in M01\_Eta01 and M10\_Eta01, respectively.
Applying the same two-fold nonzero-margin abstention to this stronger expert gives the following means over three station layouts.
lrrr@
Held-out regime
256 stations
512 stations
1024 stations
M01\_Eta01
0.00%
1.49%
1.95%
M10\_Eta001
0.26%
2.08%
8.97%
M10\_Eta01
0.17%
0.18%
4.11%
Mean paired reduction versus periodic IDW for the cross-fitted
multiscale-cardinal pyramid, over three independent random station layouts.
Across all 27 comparisons, the routed pyramid has 18 positive estimates, seven numerical ties, two unresolved negatives, 12 resolved gains, and no resolved loss. Its mean gain is 2.134% and worst point estimate is an unresolved −0.082%. It lowers nRMSE relative to the earlier stable single-scale atlas in 16 of 27 cases and improves that method by 1.00% on average. The central result is consequently not ``more scales always win.'' A capacity-matched cardinal hierarchy creates a stronger analytic specialist; cross-fit abstention supplies its empirical robustness.
Sensor geometry is itself part of the sampling operator. We therefore repeat the 27-case audit with shifted periodic grids and with stratified layouts that place one sensor at a random subcell location in each grid cell. Relative to random stations, grid IDW reduces mean nRMSE by 11.2%, 12.1%, and 21.3% at 256, 512, and 1024 stations. Stratified IDW retains reductions of 7.6%, 8.2%, and 13.2%. The gain follows fill distance: averaged over three seeds, random fill radii are 13.54, 9.39, and 7.00 pixels; grid radii are 5.66, 4.47, and 2.83; stratified radii are 8.19, 6.63, and 4.24.
Coverage alone is insufficient. The normalized sampling mask of every grid has maximum non-DC Fourier magnitude one, the signature of exact reciprocal-lattice replicas. Random and stratified masks have maxima only 0.09–0.19. Consistent with this alias nullspace, the grid pyramid router has eight resolved gains but one resolved loss among 27 comparisons, reaching −2.06% in its worst case: validation on the same lattice cannot observe an off-lattice component. The stratified router has 19 positive estimates, eight ties, no negative estimate, nine resolved gains, and no resolved loss; its mean gain over the already stronger stratified IDW is 1.58%. Consequently the practical cardinal acquisition rule is to allocate one sensor per spline-scale cell to bound holes, then dither within cells to break coherent aliases. This is an empirical design rule rather than a universal optimality theorem.
The next experiment adapts the operator spectrum rather than the sampling grid. For channel c, we augment the polynomial pyramid by Kexp,c(x,y)=1+γcKmulti(x,y)+γcKhc(x,y)cos(ωc⊤(x−y)). The modulated term is the real sum of the two spectral shifts induced by the conjugate poles ±iωc. Since both the cardinal kernel and the stationary cosine kernel are positive semidefinite, their pointwise product is positive semidefinite by the Schur product theorem; the positive direct sum remains a valid kernel. It also retains one coefficient per station. Scale, frequency, weight, and ridge are selected separately for the four physical channels on balanced fields from the other two CFD regimes.
This gives an unusually clean positive/negative boundary. Used everywhere, the channelwise exponential pyramid beats its polynomial parent in only nine of 27 stratified held-out comparisons and loses in 18; it has two resolved gains, 13 resolved losses, and a mean relative change of −2.438%. Yet on the difficult M10\_Eta001 regime its mean gains over the polynomial pyramid are 0.830%, 0.596%, and 0.306% at 256, 512, and 1024 stations. Learned poles are therefore a specialized operator hypothesis, not a universally better spline degree.
We consequently place IDW, the polynomial pyramid, and the channelwise exponential pyramid in a cross-fitted operator atlas. Both station folds must clear a training-selected margin over IDW; exponential selection must also dominate polynomial selection on every fold, and routing below 7.5% of the field population triggers exact abstention. Over three regimes, three stratified layouts, and three densities, this guarded atlas has 19 positive results and eight exact abstentions, with no negative result, 12 resolved gains, and no resolved loss. Mean gain over stratified IDW is 1.829% and the best case reaches 9.231%. Its aggregate nRMSE is 0.429% lower than the polynomial-only router on average (16 wins, six ties, five losses), with a 1.016% average improvement in M10\_Eta001. Thus operator-pole adaptation is useful here only when evidence-gated; the negative fixed-basis result is as important as the atlas gain. The fieldwise audit localizes the mechanism: in M10\_Eta001, channels 1 and 2 average 18.78% and 21.49% reductions relative to IDW and route to the modulated expert on 52.2% and 53.3% of fields. Channel 3 selects zero modulation in all nine cases, so its nominal exponential routes are regularization-only. The gain is therefore specific to two channel dynamics, not generic added flexibility.
Cardinality also resolves the regular-grid hardware bottleneck exactly. On a rectangular periodic station lattice the pyramid Gram matrix is block circulant with circulant blocks. With the two-dimensional lattice DFT F, its coefficient solve is c=F∗Kgrid+λFy, and full-field synthesis is one FFT convolution after scattering c to the station lattice. No dense station-to-field matrix is formed. Across three regimes, three layouts, and three densities, the FFT implementation agrees with the dense solution to at worst 4.57×10−15. Median CPU speedups are 4.78×, 12.92×, and 19.72× at 256, 512, and 1024 stations. At 1024 stations the explicit Gram plus synthesis operators occupy about 136 MiB, versus 0.25 MiB for the kernel and its spectrum.
Dither makes the station Gram noncirculant, but does not destroy translation invariance of the much larger station-to-field synthesis map. We therefore retain the exact dense n×n irregular Gram solve, scatter its coefficients at the true station locations, and FFT-convolve on the output grid. This split diagonalization agrees with the original dense dithered implementation to at worst 3.42×10−15 across all 27 cases. Median speedups are 4.78×, 7.49×, and 6.54×, with runtimes 8.9, 11.7, and 28.3 ms. At 1024 stations it stores an 8 MiB Gram plus a 0.25 MiB spectrum rather than the 136 MiB Gram–synthesis pair, a 16.5× operator-memory reduction. The key is to preserve irregularity only where geometry requires it and diagonalize the globally stationary map.
This accelerator exposes the price of the anti-aliasing geometry above. Snapping stratified measurements to cell centers restores circulant structure but loses 4.20%, 10.86%, and 21.00% relative to the true dithered solve, with resolved losses in four of nine, nine of nine, and nine of nine cases. An exact irregular matvec can still scatter, FFT-convolve, and sample at the true stations. Preconditioned by the lattice inverse, it recovers the dense field to within 6.11×10−9, but needs 21–100 iterations and is 4–7× slower than optimized dense algebra at this scale. At 1024 stations, eight truncated iterations are 1.92× faster but lose 5.29% to dense and tie IDW; 16 iterations retain a 1.06× speedup and beat IDW by 4.89% on average, but still have five of nine resolved losses to dense. Snapping and iterative replacement of the Gram are therefore negative controls; the exact irregular solution is the dense-Gram/FFT- synthesis split above.
Continuously located sensors admit a further operator-derived bridge without a generic NUFFT. Write an off-pixel center as xj=nj+δj, where nj is its nearest grid point. For the cardinal pyramid K, full-grid synthesis becomes the truncated Hermite moment expansion j∑cjK(q−xj)≃∣α∣≤p∑[DαK∗j∑α!cj(−δj)αδnj](q). The continuous-coordinate station Gram is still assembled and solved exactly; only the much larger station-to-grid map is replaced. In two dimensions the order-p expansion needs (p+1)(p+2)/2 FFT convolutions, independent of the number of sensors.
We evaluate bilinearly sampled PDEBench residual measurements with random offsets up to 0.49 pixel over the same three regimes, three seeds, and three densities. The six-channel, second-order expansion has median relative synthesis error 3.40×10−4 and worst error 2.22×10−3 over 27 cases; its maximum absolute change in normalized reconstruction error is 3.71×10−7. Median synthesis speedups are 11.09×, 23.19×, and 45.13× at 256, 512, and 1024 sensors. Including the shared exact Gram solve, median speedups are 9.53×, 13.55×, and 10.92×. At 1024 sensors, six complex derivative spectra require 1.5 MiB rather than a 128 MiB dense synthesis matrix; including the common Gram gives approximately 9.5 versus 136 MiB of operator storage.
The controls expose the approximation mechanism. Nearest and bilinear coefficient gridding reach worst relative synthesis errors 0.357 and 1.992. A displacement-radius sweep gives error exponents 1.99 for first order and 3.01 for second order, as predicted by the Taylor remainder. Third order lowers the constant but its exponent remains 3.03, because the cubic B-spline is globally only C2 and knot crossings preclude a uniform fourth-order remainder. It therefore adds four FFT channels without material reconstruction benefit. Second-order Hermite moment gridding is the practical Pareto point. This closes the simulated off-pixel synthesis gate; validation on physical continuous-coordinate stations, rather than bilinearly sampled gridded fields, remains open.
The result is not specific to bilinear measurement formation. Replacing it by periodic Fourier upsampling and continuous sampling over the same 27 cases gives median/worst second-order synthesis errors 3.60×10−4/2.30×10−3 and a maximum absolute nRMSE change of 3.93×10−7. The near-identical envelope supports the translated- spline approximation mechanism rather than an accidental match to the pixel sampler.
Compact support also removes the dense storage assumption from the remaining continuous Gram. A periodic neighbor search assembles only nonzero cubic- cardinal interactions, after which one sparse factorization serves every field. Across 27 cases the sparse and dense matrices agree to at worst 3.68×10−16 and their coefficients to 9.83×10−14. On the fixed 1282 domain the Gram is 25% dense, uses 2.66× less raw storage, and gives median Gram-stage speedups 1.43×, 1.40×, and 1.19× at 256, 512, and 1024 sensors; one 512-sensor case is a 0.98× tie. Combined with second-order Hermite synthesis, median end-to-end speedups become 9.48×, 14.54×, and 12.74×.
This is an exact storage result but not yet an asymptotically fast sparse solver. On a 2562 diagnostic with 4096 continuous sensors, density falls to 6.25% and raw Gram storage improves 10.65×, yet generic sparse LU reaches only 1.03× dense parity because of fill-in. Unpreconditioned CG takes 466 iterations to reach relative residual 10−8. Thus the full operator need not be dense, but larger deployments require a cardinal multilevel or lattice-corrected preconditioner rather than generic sparse algebra.
The lattice-corrected preconditioner makes the scaling boundary constructive. Using the inverse BCCB operator of the underlying one-sensor-per-cell lattice reduces batched PCG from 466 to 100 iterations, although exact convergence is still slower than dense. At 4096 sensors, 32 iterations give 2.21× end-to-end speed with relative field error 4.998×10−3, while 64 iterations give 1.30× speed with error 9.99×10−6. Sixteen iterations are rejected: their 3.41× speed costs 11.3% field error. At 256–1024 sensors sparse direct factorization remains faster than the accurate truncated variants. The resulting solver policy is size dependent: sparse direct below the factorization crossover, 64-step lattice-PCG above it for high fidelity, and 32 steps only under an explicit 0.5% operator-error budget.
At this stage the liquid module is mature enough to be used as an OSNR toolbox component under controlled assumptions: exact Green/transfer evaluation is solved, irregular timing is supported, and sparse event recovery has both a fast robust estimator and a more precise local estimator. It is not yet a stand-alone SOTA learning claim against trained LTC/CfC/ODE-RNN models. The next learning benchmark should therefore use this module as a structured layer inside a small trainable hybrid system, while the main hard-core validation remains the CFD/PINN setting where analytic derivatives, hard boundary constraints, and FFT operator diagonalization are the central advantage.
The first learning-facing liquid benchmark tests whether the liquid toolbox is useful beyond deterministic reconstruction by constructing a train/test family of damped hybrid impact sequences. Each trajectory combines a smooth damped oscillatory component with sparse exponential impact responses. The model observes the first 60% of each noisy sequence and predicts the held-out future. This task is intentionally structured: the future is predictable when the prefix identifies the continuous operator, the smooth modes, and the sparse impact responses, but a generic neural sequence model must learn this structure from examples.
The benchmark uses 96 training sequences, 32 held-out test sequences, 128 time samples, three sparse impacts per sequence, and observation noise 10−3. Two gradient-trained baselines are included: a prefix MLP and a GRU encoder that maps the observed prefix directly to the future suffix. Three OSNR profiles are included. OSNR-head is a small neural head that maps the observed prefix to coefficients over a fixed normalized liquid dictionary and then renders the trajectory through the analytic basis. OSNR-support trains a classifier to imitate the OMP event support, selects a diverse top-K set of event atoms, and then solves amplitudes and smooth coefficients by ridge least squares on the prefix. OSNR-liquid performs no gradient training on this dataset. It fits each test prefix by a sparse-plus-smooth operator dictionary: damped Fourier/exponential smooth atoms plus a causal exponential event dictionary, selected by OMP and refit by ridge least squares, then extrapolated through the same closed-form liquid response.
lrrrr@
Profile
Future RMSE
Future PSNR
Train ms
Infer/Fit ms
ZOH hold
9.7691e−01
4.8727
0.0
0.000
MLP prefix
3.3785e−01
14.0954
241.4
0.044
GRU prefix
2.5258e−01
16.6219
10728.8
1.807
OSNR-head
3.2231e−01
14.5044
1010.7
0.112
OSNR-support
9.0795e+00
−14.4914
1121.2
1.075
OSNR-liquid
1.9521e−01
18.8600
0.0
61.654
Controlled hybrid impact sequence learning benchmark. Metrics are computed on the held-out future suffix across 32 test sequences. OSNR-head is a trained coefficient predictor over a fixed analytic liquid dictionary. OSNR-support is a trained event-support classifier followed by analytic coefficient refitting. OSNR-liquid is a per-sequence structured sparse-plus-smooth fit rather than a trained neural baseline.
Table [tab:liquid-hybrid-learning] gives three distinct lessons. First, on this operator-matched hybrid family, the per-sequence OSNR-liquid estimator improves future RMSE over the trained GRU by roughly 22.7% and improves future PSNR by 2.2381 dB, without gradient training. Second, the naive trainable coefficient head is not yet competitive with the GRU: it improves over the direct MLP but underperforms the recurrent encoder. Third, support imitation alone fails catastrophically. Even with diversity-constrained top-K selection, small support mistakes produce an ill-conditioned analytic extrapolation. This is an important boundary condition: the trainable hybrid should not try to classify sparse support independently from amplitude and trajectory fit. The next viable learned version should be a distillation/correction model around the OMP-selected OSNR fit, or a differentiable sparse solver layer with the support decision coupled to the reconstruction loss.
A sample-efficiency stress run reduces the gradient-trained training set while leaving the per-sequence OSNR-liquid fit unchanged. With only 8 training trajectories and 64 held-out test trajectories, the trained GRU reaches future RMSE 4.2430e−01 while OSNR-liquid reaches 2.1714e−01, a 48.8% reduction without dataset-level backpropagation. With only 4 training trajectories, the GRU reaches 3.9635e−01 and OSNR-liquid remains 2.1714e−01, a 45.2% reduction. This is the strongest liquid-network learning signal so far: not a public LTC/CfC benchmark victory, but a clean demonstration that a biologically inspired leak-filtered operator dictionary can replace gradient training when the sequence family is sparse-plus-smooth and operator matched. A multi-leak dictionary ablation was added to the runner, but the naive wider leak bank overfit the prefix and underperformed the fixed leak; future learned liquid hybrids should regularize leak selection or validate it inside the observed prefix rather than simply expanding the atom bank.
The liquid experiments above still treat the neuron mostly as a scalar leak-filtered state. The more radical cellular benchmark makes each artificial cell a multicompartment dendritic operator followed by a soma leak, and learning is a local sparse inverse problem rather than reverse-mode differentiation through a network.
For branch b with dendritic coordinate s∈[0,1], the controlled model starts from the cable-like PDE ∂tvb(s,t)=Db∂ssvb(s,t)−λbvb(s,t)+j∑θbjrj(t)δ(s−sbj), with cosine eigenmodes used as the exact reduced basis. The modal state obeys z˙bm(t)=−(λb+Dbπ2m2)zbm(t)+j∑θbjcos(πmsbj)rj(t), so every synapse generates an analytically known branch atom. The soma then integrates the branch root voltages by V˙(t)=−λsV(t)+b∑abvb(0,t). This gives a cellular analogue of OSNR: the dictionary atoms are not generic lags or learned hidden units; they are Green responses of dendritic diffusion plus soma leakage.
The learning rule is local. For each branch, dendritic probe traces at s∈{0,0.43,0.86} provide the local voltage/calcium-style observables. The branch solves θbminPbj∑θbjabj−yb22+α∥θb∥22,∥θb∥0≤Kb, by OMP plus ridge refitting. No global reverse pass, adjoint, or dataset-level backpropagation is used for the cellular row. A soma-only OMP ablation is also included: it sees only V(t), not the local branch probes, and therefore tests whether exact synapse placement is identifiable from the soma alone.
llrrrrr@
Train seq.
Profile
Test RMSE
PSNR
Train ms
Infer ms
Supp. F1
8
Delay-ridge local
5.7633e−04
29.0007
831.1
0.024
n/a
8
Dense cable ridge
4.9364e−04
30.3459
4106.4
0.043
n/a
8
GRU backprop
6.1679e−03
8.4113
7025.4
9.963
n/a
8
DOS-NC soma OMP
3.8017e−04
32.6145
4107.6
0.087
0.100
8
DOS-NC branch OMP
2.1202e−05
57.6867
5324.2
91.379
1.000
2
Dense cable ridge
4.5093e−03
11.1320
4151.8
0.025
n/a
2
GRU backprop
6.3951e−03
8.0972
6274.8
9.341
n/a
2
DOS-NC soma OMP
1.8702e−03
18.7764
4152.3
0.112
0.000
2
DOS-NC branch OMP
8.6729e−06
65.4508
5356.5
92.441
1.000
Controlled dendritic cellular operator learning. The teacher is a 4-branch dendritic cable cell with 6 cosine modes, 8 presynaptic traces, 10 sparse synapses, 160 time samples, 48 held-out test sequences, and observation noise 2e−3. DOS-NC branch OMP learns from local dendritic probe traces by sparse inverse solving; the GRU baseline uses ordinary backpropagation.
Table [tab:dendritic-cellular-operator-learning] is the first controlled validation of the proposed biological learning thesis. With only 8 training sequences, the branch-local dendritic operator learner reaches RMSE 2.1202e−05 on held-out soma voltage, about 291× lower than the backprop-trained GRU on the same split, and exactly recovers the teacher support. With only 2 training sequences, it remains at 8.6729e−06 RMSE with support F1 1.000. The soma-only OMP row is the critical ablation: it improves over generic temporal baselines but fails to identify the true synapses. This matches the biological premise. Local dendritic observables are not an implementation detail; they are the information channel that makes no-backprop synaptic learning identifiable.
Held-out dendritic cell trace for the 8-sequence benchmark. The branch-local DOS-NC curve is visually indistinguishable from the teacher soma, while the GRU and generic delay ridge baselines miss the operator-matched cellular dynamics.
This is still a controlled cellular-identifiability experiment, not a public LNN benchmark or a claim that arbitrary supervised learning can be replaced by local rules. Its significance is narrower and stronger: once the cell is modeled as a dendritic diffusion operator, local branch traces plus sparse inverse solving can learn the synaptic operator dramatically faster and more accurately than a small backprop-trained recurrent network in the matched regime. The next step is to stack these cells into layers where each branch receives its own local predictive or modulatory innovation, so the network-level credit signal is carried by local residual fields instead of exact reverse-mode gradients.
The next script, apps\_industrial\_breakthrough/dendritic\_cross\_arch\_benchmark.py, deliberately tests whether the no-backprop hypothesis survives contact with real benchmark data. It uses torchvision MNIST and FashionMNIST, keeps all feature extractors fixed after random or analytic wiring, and learns only closed-form ridge readouts. The profiles are architectural analogues rather than trained networks: DCT operator coefficients, fixed local convolutional banks, liquid row/column scans, fixed patch attention, grid-graph diffusion, class-balanced dendritic RBF memory cells, and a fused ridge readout. The comparison baselines are small MLP/CNN models trained by ordinary backprop for one epoch under the same CPU-only runner. This is not a SOTA protocol; it is a fast falsification test for whether local operator learning has benchmark-scale signal beyond the controlled dendritic-cell setting.
llrrrr@
Dataset/split
Profile
Accuracy
Train ms
Est. MB
Backprop
MNIST 10k/10k
CNN local ridge
97.44%
1082.0
289.8
no
MNIST 10k/10k
CNN local ensemble ×4
97.85%
4065.0
289.8
no
MNIST 10k/10k
DOS-NC fused ridge
97.58%
7102.1
515.1
no
MNIST 10k/10k
CNN, one epoch
92.79%
2083.6
4.4
yes
MNIST full
CNN local ridge
97.97%
4247.8
912.5
no
MNIST full
CNN local ensemble ×4
98.03%
16325.2
912.5
no
MNIST full
DOS-NC fused ridge
98.07%
26602.2
1531.9
no
MNIST full
CNN, one epoch
98.25%
12278.1
4.4
yes
Fashion 10k/10k
CNN local ridge
88.06%
1094.2
289.8
no
Fashion 10k/10k
CNN local ensemble ×4
89.10%
4072.5
289.8
no
Fashion 10k/10k
DOS-NC fused ridge
88.30%
7104.4
515.1
no
Fashion 10k/10k
CNN, one epoch
77.60%
2079.6
4.4
yes
Fashion full
CNN local ridge
89.48%
4264.5
912.5
no
Fashion full
CNN local ensemble ×4
89.53%
16477.9
912.5
no
Fashion full
DOS-NC fused ridge
89.56%
26816.8
1531.9
no
Fashion full
CNN, one epoch
86.36%
12664.3
4.4
yes
Real-data no-backprop cross-architecture stress test. The no-backprop rows use fixed local operator features and closed-form ridge readouts; the CNN baseline is a small gradient-trained model, not a tuned SOTA model. Full MNIST/Fashion use the standard 60,000/10,000 train/test split.
Table [tab:dendritic-cross-arch-realdata] is encouraging but not a moonshot. On the low-data MNIST split, the local CNN ensemble reaches 97.85% and beats the one-epoch backprop CNN by 5.06 percentage points. On FashionMNIST, the no-backprop local ensemble reaches 89.10% with 10,000 training examples and 89.53% on the full split, beating the bounded one-epoch CNN by 11.50 and 3.17 percentage points respectively. However, full MNIST remains below the same one-epoch CNN (98.07% fused no-backprop versus 98.25%), the fixed-attention and grid-GNN views are weak, and the memory estimate for the closed-form full-feature solves is much larger than the small CNN. The correct interpretation is therefore not ``SOTA without backprop.'' The result is a fast, CPU-only sample-efficiency signal for local operator features plus algebraic readouts, and a clear boundary: generic CNN/RNN/attention/GNN replacement will require true stacked local learning and streaming/local normal equations, not merely wider fixed feature banks.
No-backprop optimization: neuromodulated local control
The next runner, neuromodulated\_local\_learning\_benchmark.py, moves from fixed features toward a biologically motivated optimization loop. The intended replacement for global reverse-mode differentiation is not ``no loss.'' It is a different decomposition of the loss. Each cell or local branch receives a local state, a local eligibility trace, and a low-dimensional modulatory innovation. For a branch state vib(s,t), ∂tvibτiV˙i=Dib∂ssvib−λibvib+j∑θijbrj(t)δ(s−sijb),=−Vi+b∑aibvib(0,t), the branch-level update should be a three-factor control rule, Δθijb=ηmi(t)eijb(t)−ηh∂θijbHijb,eijb(t)=∫rj(τ)Gib(t−τ)χib(τ)dτ. Here eijb is the local eligibility trace induced by the dendritic Green function, mi is a reward, dopamine, prediction-error, or observation-innovation field available to the cell or region, and H is a homeostatic stability cost. This is closer to feedback control than to backpropagation. The global task loss is allowed to create a modulatory signal, but it is not differentiated through every downstream operation to produce an exact adjoint for every upstream synapse.
This leads to a different architecture search space. A layer should be a population of multicompartment cells with lateral competition and residual identity highways, xℓ+1=xℓ+PℓΓℓ({Vℓi}i), where Γℓ can include soma thresholds, local winner-take-all inhibition, liquid leak filters, and branch-local sparse solves. The residual path is not a stylistic copy of ResNets; it is a stability/control channel that prevents local cell updates from having to preserve the whole signal while they learn a correction. Attention should also be reinterpreted. Instead of backpropagating through dense learned QKV matrices, a biological attention analogue stores local keys or prototypes, routes by kernel similarity and competition, and updates the keys only when a modulatory innovation indicates that a region was informative or surprising. Transformer-like selectivity is still needed, but the learning mechanism must be local memory deposition and residual routing rather than exact gradient transport.
The benchmark implements a first small version of this idea on MNIST and FashionMNIST. It learns patch filters by local competitive quantization, learns class-gated prototype keys by per-class local clustering, solves the readout by ridge normal equations, then performs a residual-memory correction: R0=Y−Y0,Cc=topK{xi:yi=c,∥R0,i∥2},A⋆=argAmin∥ΦCA−R0∥F2+λ∥A∥F2, and predicts by Y=Y0+ΦCA⋆. The residual centers Cc are a simple dopamine analogue: high-innovation examples deposit class-local memory, and the correction controller is fitted algebraically. No representation layer is trained by reverse-mode AD.
llrrrr@
Dataset/split
Profile
Accuracy
Train ms
Est. MB
Backprop
MNIST 10k/10k
Fixed local CNN ridge
96.76%
2052.7
660.7
no
MNIST 10k/10k
Hebbian CNN ridge
86.76%
3206.0
1113.0
no
MNIST 10k/10k
Dopamine attention ridge
93.56%
3970.2
52.1
no
MNIST 10k/10k
NML fused local ridge
98.04%
2175.3
1193.9
no
MNIST 10k/10k
Dopamine residual memory
98.12%
2518.2
1294.8
no
MNIST 10k/10k
CNN, one epoch
92.79%
2076.1
4.4
yes
Fashion 10k/10k
Fixed local CNN ridge
87.90%
2084.2
660.7
no
Fashion 10k/10k
Hebbian CNN ridge
88.28%
3166.6
1113.0
no
Fashion 10k/10k
Dopamine attention ridge
80.75%
4008.6
52.1
no
Fashion 10k/10k
NML fused local ridge
88.31%
1967.0
1193.9
no
Fashion 10k/10k
Dopamine residual memory
88.53%
2827.5
1611.4
no
Fashion 10k/10k
CNN, one epoch
77.60%
2075.5
4.4
yes
Neuromodulated local-learning stress test. Patch filters and prototype keys are learned by local clustering, readouts are closed-form ridge solves, and the residual-memory row uses high-innovation examples as class-local memory centers. The Fashion residual-memory row uses 256 centers per class; the MNIST row uses 64 centers per class.
Table [tab:neuromodulated-local-learning] gives a precise result rather than the desired universal breakthrough. On MNIST, the residual-memory controller improves the fused no-backprop model from 98.04% to 98.12%, beating the one-epoch CNN control by 5.33 points on the same split. On FashionMNIST, increasing residual memory from 64 to 128 and 256 centers per class improves the residual row from 88.39% to 88.51% and 88.53%, but it still remains below the earlier four-bank fixed local ensemble at 89.10%. The attention-like prototype branch is also weak as a stand-alone model. The conclusion is important: a scalar/vector modulatory innovation can improve no-backprop local learning, but the current single-stage memory deposition is still too shallow and too memory-heavy to replace stacked backprop-trained architectures at SOTA scale. The next serious architecture must stack these local residual controllers, expose intermediate local targets or predictive residuals at each layer, and update the residual memory by streaming/local normal equations rather than by one global dense solve.
Cell-operator mismatch audit: poles, conductance, and activation
The criticism of the previous liquid/cellular experiments is correct: a hand-chosen linear cable basis with a sigmoid release transform is not yet the correct neuron model. Hasani's LTC formulation and the exact multi-synapse extension instead make the synapse the nonlinear operator [hasani2020ltc,hasani2022cfc,cantini2025exact]. In scalar form, x˙(t)=−ωx(t)+s=1∑Sfs(gs(t);θs)(As−x(t)), so the instantaneous pole is not fixed. It is p(t)=−(ω+s∑fs(gs(t);θs)), and the driving equilibrium is the conductance-weighted reversal potential. This means that the ``activation function'' is not a pointwise ReLU/SIREN-style nonlinearity after a linear map. It is a synaptic conductance field that simultaneously controls gain, sign, equilibrium, and time constant. A linear exponential-pole dictionary can approximate its traces, but it is structurally mismatched because it does not include the multiplicative feedback term (As−x).
The diagnostic runner cellular\_operator\_model\_audit.py isolates this issue. It generates a teacher from the exact zero-order-hold multi-synapse LTC recurrence xk+1=γkxk+(1−γk)ω+∑sfs(gs,k;θs)∑sfs(gs,k;θs)As,γk=exp[−Δtk(ω+s∑fs(gs,k;θs))], then compares four no-backprop identification families: fixed linear poles, conductance with wrong gates, conductance with oracle gates, and sparse search over an overcomplete conductance-gate dictionary. The conductance learners use the locally observed voltage and solve x˙+ωx=s∑wsfs(gs;θs)(As−x) by ridge or OMP/ridge, followed by exact rollout. This is a local operator-identification rule, not reverse-mode training through a network.
lrrrrr@
Profile
Test RMSE
PSNR
Terms
Train ms
Est. MB
Linear raw multi-pole ridge
2.8644e−02
16.947
49
157.6
0.58
Linear sigmoid multi-pole ridge
2.5585e−02
17.928
49
161.7
0.58
Wrong-gate conductance ID
8.2627e−03
27.745
8
12.0
0.09
Oracle-gate conductance ID
3.0990e−04
56.263
8
11.8
0.09
Dense grid-gate conductance ID
6.3447e−02
10.040
384
476.5
5.04
Sparse grid-gate OMP ID
2.1460e−03
39.455
24
39.6
4.48
Cell-operator mismatch audit on a synthetic exact multi-synapse LTC teacher with 16 training sequences, 64 test sequences, 192 time steps, 8 synapses, and observation noise 10−3. Linear pole dictionaries are the wrong operator family. Matched conductance identification is both more accurate and leaner. Blind dense gate expansion is ill-conditioned; sparse gate selection is the viable unknown-operator path.
Table [tab:cellular-operator-mismatch-audit] identifies the current mistake sharply. The old fixed-pole view is not merely under-tuned; it is the wrong operator for an LTC-style cell. Matching the conductance law reduces test RMSE by roughly 83× versus the best linear-pole row while using only 8 terms and about 0.09 MB in this audit. Even a wrong fixed gate improves substantially over linear poles, proving that the multiplicative reversal-potential structure matters. Conversely, a dense overcomplete gate dictionary fails, while sparse OMP over gate candidates recovers much of the gap. The next no-backprop architecture should therefore start with constrained local identification of ω, As, θs, and active synapses, under positivity/stability bounds, before any CNN/RNN/transformer-scale benchmark. Architecture comes after the cell operator is right.
Grown-topology operator networks and sample-efficient closed-form identification
The cell-operator mismatch audit establishes that an LTC-style cell is governed by a conductance operator, not a fixed-pole linear filter. This subsection develops the learning-time consequence. Mainstream artificial networks fix the architecture in advance and brute-force a generic function approximator by backpropagation. Biological networks instead grow: capacity is added developmentally while the system is learning, and the learning rules and topology themselves were shaped by evolution [stanley2002neat]. A single biological neuron is correspondingly far richer than a weighted-sum-plus-activation unit; a layer-five pyramidal cell requires a five-to-eight layer temporal network to reproduce [beniaguev2021single]. These two observations motivate a different training regime, which we state as a falsifiable thesis.
Thesis.
For systems whose structure is specifiable as a known operator family — a connectome-shaped dynamical system, a governing differential operator — one should not learn a generic function. One should parameterize the operator and identify its few free parameters: solve everything that is linear in its coefficients by an OSNR closed-form solve, and reserve a small gradient-free evolutionary search for the nonlinear and structural parameters, growing the topology with warm starts. The claim is that such a model matches or beats a backpropagation network of equal budget on sample-efficiency, parameter count, and out-of-distribution robustness, not on raw task score.
Two-timescale decomposition.
The regime separates exactly along the linear/nonlinear boundary already used throughout OSNR.
Inner (fast), linear. Given a frozen nonlinearity shape and topology, the conductance/readout coefficients enter the operator linearly. They are recovered by a single scale-invariant ridge solve over an operator-matched feature dictionary, with derivatives supplied by the autograd-free spline ladder of Section [sec:autograd-free] rather than by automatic differentiation. No iteration, no backward graph.
Outer (slow), nonlinear and structural. Time constants, synaptic gains, gate shapes, and the topology itself are searched by an evolutionary method (CMA-ES/PGPE [sehnke2010pgpe,salimans2017es]) with NEAT-style complexification [stanley2002neat]. A neuron is added only when capacity saturates, and its outgoing weights are initialized near zero so the addition is a function-preserving no-op [chen2016net2net]; the inner solve then re-fits in milliseconds and the new unit is retained only if held-out fitness improves.
The mutual dependency is the point: NEAT-style growth is normally bottlenecked because scoring each candidate topology needs a full training run, while a closed-form inner solve makes candidate evaluation nearly free. Cheap convex identification and evolutionary growth each enable the other. This is the operator-spline counterpart of sparse governing-equation discovery [brunton2016sindy] and universal differential equations [rackauckas2020universal], specialized to grown conductance networks.
Rung 0: verifying the identification engine.
Before any closed-loop control study, the premise must be checked in isolation: when the operator form is exactly known, is the closed-form solve genuinely more sample-efficient and cheaper than backpropagation on the same model class? The runner bio\_growth/rung0\_osnr\_id\_verification.py fixes a known LTC teacher with N=8 neurons, M=2 inputs, a known sigmoidal synaptic feature bank, and unknown (τi,wij,vik,Ai). The expanded right-hand side lies exactly in the matched dictionary Φi(x,I)=[xi,{σj(x)}j,{xiσj(x)}j,{Ik}k,{xiIk}k], so identification reduces to per-neuron ridge regression of x˙i onto Φi. Four methods of the identical model class are compared on 24 held-out clean trajectories (rollout normalized RMSE): OSNR, closed-form ridge on a Tikhonov/curvature-smoothed derivative (the operator-spline derivative); FD, the same closed-form ridge on a raw finite-difference derivative — i.e. the exact minimizer of the one-step linear least-squares objective; 1-step SGD, that same linear objective optimized by Adam for 4000 epochs; and rollout BPTT, the naive recurrent fit by backpropagation through a 300-step integrator (300 epochs). Observation noise is 10% of per-state standard deviation.
trajectories
OSNR nRMSE
OSNR time
FD nRMSE
1-step SGD nRMSE
rollout BPTT nRMSE
BPTT time
1
9.146
0.00 s
10.136
4.139
0.134
40.5 s
2
1.010
0.01 s
2.267
1.603
0.088
45.5 s
4
0.0815
0.01 s
2.793
1.314
0.103
45.3 s
8
0.0387
0.03 s
0.0771
0.734
0.109
45.2 s
16
0.0318
0.06 s
0.0585
0.842
0.0494
46.0 s
32
0.0276
0.11 s
0.0215
0.189
0.0481
47.3 s
Rung 0 LTC operator identification (held-out rollout nRMSE) versus number of
training trajectories at 10% observation noise. The closed-form OSNR solve runs in
0.01–0.11 s versus ∼45 s for backpropagation-through-time — a
400–4000× wall-clock reduction — and from four trajectories upward it is
also more accurate than the trained recurrent BPTT fit. Finite-difference closed-form is
the exact one-step least-squares optimum; one-step SGD on the identical objective has
not reached it after 4000 epochs, illustrating that the direct solve dominates
iterative optimization even on the linear sub-problem at fixed budget.
observation noise
OSNR-spline nRMSE
finite-difference nRMSE
0%
0.0357
0.0225
5%
0.0361
0.0303
10%
0.0387
0.0771
20%
0.0626
1.509
40%
0.674
5.662
Rung 0 noise robustness at eight training trajectories. The operator-spline
(Tikhonov-curvature) derivative is the active ingredient: at low noise it is unnecessary
(finite difference is marginally better, since the smoother tends to the identity), but
finite-difference identification collapses as noise grows while spline-OSNR degrades
gracefully — a 24× advantage at 20% noise.
Interpretation and honest scope.
Three conclusions hold robustly. First, the wall-clock advantage is unconditional: a deterministic closed-form solve in tens of milliseconds replaces tens of seconds of backpropagation, which is precisely the property that makes evolutionary topology growth affordable. Second, the operator-spline derivative, not merely the closed form, is what buys sample-efficiency and noise robustness: the finite-difference control collapses at four trajectories (Table [tab:rung0-sample-efficiency]) and under noise (Table [tab:rung0-noise]), whereas the spline-smoothed solve remains accurate. Third, from four trajectories upward the closed-form identification is at least as accurate as a fully trained recurrent backpropagation fit. The honest caveats are equally explicit. In the data-starved regime (≤2 trajectories) the closed-form solve is unstable while rollout BPTT, which is implicitly regularized by having to produce a stable trajectory, is more robust; a low-data identification therefore needs stronger rank-revealing regularization. And the accuracy comparison is against a recurrent BPTT fit whose difficulty is partly the long-horizon credit-assignment problem the closed form sidesteps — so the unconditional claim is wall-clock and compute, with the accuracy advantage holding once a minimal data threshold is met. Rung 0 thus validates the inner-solve premise and clears the path to the closed-loop control study (Rung 1), where the evolutionary outer loop and warm-started growth are exercised directly.
Rung 1: closed-loop control, and an honest negative.
We exercised the full regime on a 2D chemotaxis control task (a noisy gradient-climbing agent, the canonical C.\ elegans behaviour), with the policy a small liquid reservoir whose readout is the closed-form solve and whose dynamics and topology are grown by evolution. Two findings, one methodological and one sobering. First, behaviour cloning from a privileged teacher fails for both the structured network and a backpropagation baseline, because teacher-forced training drifts off-distribution in closed loop; reframing the task as direct reward optimisation (the evolutionary outer loop) fixes this and the structured policy solves the task with ∼50 parameters. Second, and honestly, the architectural advantage on control is modest: at convergence a generic recurrent network nearly matches the structured liquid network on task score and robustness, and the structured model's remaining edge is roughly a factor of four in trained-parameter count, not a decisive win. The clean, decisive advantages of the operator-matched approach are therefore in identification, not control, which the next sections quantify against the standard identification baselines. On the related continual-learning axis—catastrophic forgetting, often cited as a place where biological learning outperforms backpropagation—the operator-matched approach admits a more ambitious construction that brings together the three ingredients of the biological thesis: grow structure on demand, learn each piece by a local closed-form solve rather than global backpropagation, and do not overwrite what was already learned. We test it in the hardest fair setting: a stream of dynamical regimes arrives in blocks with no regime labels and no replay, and the learner must itself detect when the dynamics have changed. The grown learner maintains a bank of closed-form matched experts; each incoming window is routed to the expert that best explains it, a new expert is grown whenever the best residual exceeds a novelty threshold, and a final consolidation pass merges experts that turn out to capture the same regime. Crucially, the baselines are given the same matched feature library, so the comparison isolates the mechanism (grow-and-consolidate with closed-form solves) rather than the feature prior: a single shared matched model updated online by stochastic gradient descent, a black-box neural vector field trained online, the same field regularized by elastic weight consolidation (EWC, the canonical deep continual-learning method [kirkpatrick2017overcoming], given the task boundaries and its best regularization strength—an advantage the grown learner is not given), and—as an upper bound—the same black-box field trained jointly on all regimes with full replay.
The result is decisive and robust across three independent system sets (five regimes each; Figure [fig:grown-continual]). The grown learner reaches mean forecast nRMSE 0.007–0.024 with essentially no forgetting of the first regime (0.006–0.032), while every backpropagation baseline exhibits textbook catastrophic forgetting (mean 0.9–3.1, retaining only the most recent regime). Critically, this includes EWC—the method deep learning built specifically to prevent forgetting: even with task boundaries handed to it and its regularization strength tuned to its best, EWC (0.9–2.7) is essentially no better than naive online training, because the five regimes are genuinely distinct vector fields and no single network can hold them all. The grown learner is roughly two orders of magnitude better than EWC. The replay upper bound, despite seeing every regime jointly, never falls below mean ≈1.2—a single context-free vector field cannot represent five distinct dynamics at once—so the grown bank is roughly two orders of magnitude better than even the strongest backpropagation control. The growth mechanism also recovers the latent structure: from eleven to thirteen experts grown online, consolidation returns exactly the five true regimes on every run. The memory is moreover persistent with instant recognition: when the entire five-regime sequence is presented a second time, the learner grows zero new experts on the revisit—it routes each returning regime straight back to its existing expert, with no relearning—whereas the eleven experts were all grown on the first pass. The single hyperparameter—the novelty threshold for growing—does not need tuning: sweeping it over a 3.5× range (0.20 to 0.70) leaves the forecast accuracy unchanged (0.0089–0.0092) and the consolidated regime count exactly five, even though the number of raw experts grown varies from twenty down to seven; the consolidation pass absorbs the difference. Unlike the naive version, this is not merely an argument from modularity handed regime labels for free; it is a label-free, online, self-structuring learner whose only prior is the matched feature family, and on this home turf of biological learning it decisively outperforms backpropagation.
The advantage is also a sample-efficiency advantage, biology's other reputed strength. As the data per regime is reduced (Figure [fig:sample-efficiency]), the grown learner degrades gracefully—mean nRMSE 0.013→0.08→0.21→0.70 at eight, four, two, and one trajectory per regime—and already beats every backpropagation variant, including full replay, at a single trajectory per regime (by roughly 4×, widening to ∼170× at eight). The backpropagation baselines never improve with more per-regime data because forgetting, not data, is their binding constraint. The grown learner thus wins on both axes at once: it needs little data per regime and it does not overwrite.
The separation widens as the stream lengthens (Figure [fig:continual-scaling]). Scaling the number of sequential regimes from eight to twenty, the grown learner's mean forecast nRMSE stays flat at ≈0.02 and its forgetting of the first regime is constant at 0.009—adding twenty regimes after the first degrades the first not at all—while every backpropagation baseline remains pinned near nRMSE 1, retaining essentially only the most recent regime. The structure discovery also stays exact at scale: consolidation returns precisely the true number of regimes at every K tested (recovering eight, twelve, sixteen, and twenty experts from fourteen, twenty-six, thirty-four, and forty-two grown online). The grown learner therefore scales the way biological memory is supposed to: capacity is added as needed, old skills are untouched, and the cost of a new regime is one new expert rather than interference with all the others.
Label-free, online, growing continual learning across five sequential dynamical regimes
(no labels, no replay), averaged over three independent system sets; bars show the mean and
whiskers the min–max across sets, on a log scale. Left: mean forecast nRMSE over all regimes after
the full stream. Right: forgetting, measured as the error on the first regime once the stream has
ended. The grown learner (grow a closed-form matched expert on novelty, then consolidate
duplicates) is roughly two orders of magnitude better than a single shared matched model updated
online, a black-box neural field trained online, EWC (the canonical deep continual-learning method,
given task boundaries and its best regularization strength), and even a black-box field trained
jointly with full replay—all of which see the same matched features. Growth recovers exactly the
five true regimes on every run. This is the biological recipe—grow, learn locally in closed form, do not
forget—winning on its home turf.
Sample efficiency on the same five-regime continual stream: mean forecast nRMSE versus the
number of trajectories seen per regime (log–log). The grown learner improves steeply with data and
sits below both the online and full-replay backpropagation baselines at every data budget, beating
them even at a single trajectory per regime. The backpropagation curves stay flat because their
limiting factor is catastrophic forgetting, not the amount of per-regime data.
Scaling the continual stream from eight to twenty sequential regimes. Left: mean forecast
nRMSE stays flat at ≈0.02 for the grown learner while the online and full-replay
backpropagation baselines remain near 1 regardless of stream length. Right: the consolidated
expert count (solid) tracks the true number of regimes (dashed) exactly at every scale, while the
raw number of experts grown online (dotted) runs ahead before consolidation collapses it. Capacity
grows with the task; old regimes are not disturbed.
Generalization beyond dynamical systems: Permuted- and Rotated-MNIST.
The grow-and-consolidate mechanism is not specific to differential equations. To test it on a benchmark the continual-learning community actually tracks, we apply the identical recipe to Permuted-MNIST (ten tasks, each a fixed random permutation of the 784 pixels, presented sequentially without replay). The only change is the per-task solver: a frozen bank of random ReLU features feeds a per-task closed-form ridge classifier, grown on each task. Against the same EWC, online, and joint-replay deep baselines (an MLP with a shared head), the grown bank attains 96.1% average accuracy across the ten tasks when the task is known at test time—exceeding EWC at its best (87.6%) and even the joint upper bound (95.8%), with no backpropagation and no forgetting (the online MLP collapses to 68.5%, retaining 38% on the first task). When the task is not given it must be inferred from the input. A confidence router (send each image to the most confident expert) recovers 81.6%; a simple generative router—a per-task diagonal Gaussian over the same random features, choosing the task of highest likelihood—routes perfectly (100%), so task-free accuracy equals the task-known accuracy at 96.1%, again well above EWC. The earlier task-free shortfall was thus a weak router, not a limitation of the modular learner: the per-task feature distributions are cleanly separable. The same recipe on Rotated-MNIST (ten tasks, each a fixed rotation of the digits) is also decisive: 96.1% task-known and 86.9% task-free, both far above EWC's ∼64–71% on this harder shift. Interestingly the routers swap roles here—adjacent rotation angles overlap distributionally, so the Gaussian router degrades, but the confidence router is rescued by cross-generalization (routing to a neighboring-angle expert still classifies the digit). In every case at least one simple router beats EWC. Across both standard benchmarks, then, the grown learner beats the dedicated deep continual-learning method in all four settings (task-known and task-free, permuted and rotated), with no backpropagation and no forgetting—and where the tasks are distributionally distinct, task inference is essentially exact, so task-free operation is free.
The most stringent test is class-incremental learning, where deep methods are weakest: on Split-MNIST (five tasks of two digits each, no task label at test, classification over all ten classes), EWC and naive online training both collapse to about 19–20%, the well-documented class-IL failure of regularization-based methods, since a single shared head cannot keep ten once-seen classes separable. The grown bank instead reaches 91.1% with a shared-covariance Mahalanobis router (and 99.6% with an oracle task label, confirming the per-task experts are near-perfect and the only loss is in routing), against a joint upper bound of 97.5%. That is roughly a four-and-a-half-fold improvement over EWC in precisely the setting deep continual learning finds hardest—again with no backpropagation, no replay, and no forgetting. Figure [fig:continual-vision] summarizes the three vision settings.
Continual learning on standard vision benchmarks with no task label at test. Across
Permuted-MNIST and Rotated-MNIST (domain-incremental, task-free) and Split-MNIST
(class-incremental), the closed-form grown bank (blue) stays close to the joint upper bound (green)
and well above EWC (red), the deep continual-learning baseline—most dramatically in the
class-incremental setting, where EWC collapses to near chance. No backpropagation, no replay, no
forgetting.
Scaling up, and a fair fight against strong replay.
Two caveats must be met for these results to mean anything to the continual-learning community: EWC is by now a weak baseline, and MNIST is a toy. We therefore move to a frozen ImageNet-pretrained ResNet-18 backbone (the parameter-efficient protocol in which all methods share identical features, so only the continual mechanism differs) and to Split-CIFAR, scored against dark experience replay (DER++ [buzzega2020dark]), a strong rehearsal baseline rather than EWC. On Split-CIFAR-10 class-incremental, the closed-form bank reaches 77% versus 49% for EWC. On the standard Split-CIFAR-100 (ten tasks of ten classes), the per-task-routed bank is held back by the harder ten-way task inference (48% routing), but the class-incremental instance of the same closed-form idea—a per-class prototype classifier with a shared-covariance (Mahalanobis) metric, which is simply a prototype grown for each class as it is seen—reaches 57.0%, ahead of DER++ at 47.6% and essentially at the joint upper bound of 58.7%, while EWC and online training collapse to 9% (Figure [fig:cifar100-classil]). The point is not that the prototype classifier is novel—nearest-class-mean on frozen features is a known strong rehearsal-free baseline—but that the entire family is closed-form: it carries no replay buffer, takes no gradient step, and fits in under a second, whereas DER++ requires backpropagation and a two-thousand-example buffer for a lower score. This is exactly the efficiency axis the field has turned to.
The same picture holds, and sharpens, on the backbone the prompt-based literature actually uses. With a frozen ImageNet-21k ViT-B/16 (features extracted once on a laptop GPU), Split-CIFAR-100 class-incremental accuracy rises to 88.0% for the Mahalanobis prototype classifier and 89.8% for a random-projection Gram-ridge variant (the closed-form core of RanPAC [mcdonnell2023ranpac]), versus 86.1% for DER++ and a collapse to 16% for EWC; the published numbers for prompt-tuning on the identical backbone are roughly 83–84% (L2P [wang2022l2p]) and 84–86% (DualPrompt [wang2022dualprompt]). Our closed-form classifiers thus exceed DER++ and the prompt methods while training in seconds with no buffer and no gradient step. We are explicit that the prototype and random-projection classifiers are not our invention—they are the established strong rehearsal-free baselines of this regime—and that we did not re-run the prompt methods; the contribution is the unification (the same closed-form, operator-/structure-matched principle that identifies dynamical systems and PDEs also drives a competitive continual-vision learner) together with the efficiency demonstration: in the frozen-backbone regime, closed-form modular learning matches or beats strong replay and prompt-tuning at one to two orders of magnitude less compute and no stored data.
We also report the boundary honestly. On Split-ImageNet-R—a deliberately harder benchmark whose renditions shift away from the pretraining distribution—the same closed-form classifiers reach 67% (random-projection) and 66% (Mahalanobis), still above DER++ (59%) and the published L2P (∼61–65%), and on par with DualPrompt (∼66–69%), but now below CODA-Prompt (∼73–75%) and the full RanPAC (∼74–78%). The gap is attributable to first-session backbone adaptation, a one-time gradient pass those methods include and our purely closed-form variant omits; restoring it is possible but would forfeit the zero-backpropagation property that is the point here. The honest summary across both benchmarks is that closed-form modular learning is at or near the accuracy frontier while being categorically cheaper—decisively so when the frozen features already suit the data, competitively so when they do not.
Backbone quality, not the classifier, turns out to be the lever. Swapping the supervised ViT-B for a self-supervised DINOv2 ViT-L/14 (still frozen, features extracted once on a laptop GPU) raises the closed-form Split-ImageNet-R class-incremental accuracy to 89.6% for the random-projection variant and 84.1% for the Mahalanobis prototype, above DER++ (86.7%) on the same features and well above the published prompt-tuning and RanPAC numbers reported on ViT-B (∼61–78%). We are careful about what this does and does not show: it is a stronger-backbone result, not a classifier-versus-classifier victory over RanPAC, whose closed-form core our random-projection variant essentially is. The honest takeaway is that the entire closed-form family scales with backbone quality at no training cost—a better frozen representation is free to adopt and turns a covariance solve into a state-of-the-art-rivaling continual learner—so the practical frontier in this regime is set by representation quality and a closed-form read-out, not by the expensive prompt- or replay-based adaptation machinery.
Table [tab:ptmcil] places our method on the canonical rehearsal-free pretrained-model CIL leaderboard [mcdonnell2023ranpac] (ten tasks, final average accuracy, no rehearsal buffer). On the identical ViT-B/16-in21k backbone our closed-form learner uses no gradient step at all, yet it surpasses every prompt-based method—L2P, DualPrompt, CODA-Prompt—and the adapter method ADaM, and it reproduces RanPAC's own no-adaptation ablation (89.9% measured versus their 89.0% on CIFAR-100). The one method above us is the full RanPAC, whose advantage is precisely its first-session gradient adaptation of the backbone; matching that with a purely closed-form routine is left open, and we instead claim the strongest zero-backpropagation position on the board. With a stronger frozen backbone (DINOv2 ViT-L/14) the same closed-form method reaches 92.3% on CIFAR-100 and 89.6% on ImageNet-R, exceeding the best published numbers—though we note two caveats in fairness: full RanPAC would also benefit from the stronger backbone, and our ImageNet-R figure uses an 80/20 split rather than the reference split.
Method (ViT-B/16-in21k, final acc.)
CIFAR-100
ImageNet-R
gradient?
L2P [wang2022l2p]
84.6
72.5
yes (prompts)
DualPrompt [wang2022dualprompt]
81.3
71.0
yes (prompts)
CODA-Prompt
86.3
75.5
yes (prompts)
ADaM
87.6
72.3
yes (adapter)
RanPAC (full, with FSA) [mcdonnell2023ranpac]
92.2
78.1
yes (FSA)
RanPAC, no FSA (= ours)
89.0
71.8
no
NCM only
83.4
61.2
no
Ours, closed-form (ViT-B/16-in21k)
89.9
–
no
Ours, closed-form (DINOv2 ViT-L/14)
92.3
89.6
no
Rehearsal-free pretrained-model class-incremental learning (final average accuracy, ten
tasks; published numbers from [mcdonnell2023ranpac]). Among methods that take no gradient
step, ours is the strongest; the only entry above it on the shared backbone is full RanPAC, whose
edge is a gradient first-session adaptation. A stronger frozen backbone lifts the closed-form method
past the best published numbers (caveats in text).
Split-CIFAR-100 class-incremental on a frozen ResNet-18, against a strong replay baseline.
The closed-form per-class prototype classifier (blue) edges out DER++ (orange) and reaches the joint
upper bound (green), while EWC (red) collapses—and it does so with no replay buffer and no
gradient steps, in under a second versus several seconds plus a two-thousand-example buffer for
DER++.
The sparse-stochastic-process view: when does matched beat random?
The continual-vision results motivate a conditional modeling question: when does a physically matched representation improve the task beyond a generic feature map? The SSP framework [unser2014sparse1,unser2014sparse2] supplies the innovation model Ls=w, with an admissible inverse and stated boundary conditions. The generalized white noise w is defined through test functions, not pointwise samples; L whitens and L−1 colors. For a suitable localization filter Ld, its useful spline bridge is Lds=βL∗w, where βL=LdL−1δ. Overlapping increment kernels can retain dependence. Non-Gaussian compressibility is not necessarily finite-rate innovation; a Green-atom expansion for atomic forcing is not a universal cardinal expansion of white noise.
Correction of an earlier overstatement.
A random feature map is not itself the driving white noise, and a frozen vision embedding has not thereby been shown to be white or Gaussian. We withdraw the claimed if-and-only-if theorem that every non-scalar L gives a strict matched-ridge risk advantage. For orthonormal design, ridge has effective degrees of freedom d/(1+λ) and risk (λ2∥c∥2+σ2d)/(1+λ)2 under independent zero-mean variance-σ2 noise, irrespective of the unknown coefficient support. An orthogonal feature rotation preserves all ridge predictions. A non-scalar orthogonal L can leave Gaussian covariance isotropic; a nonorthogonal inverse can concentrate it despite dense Gaussian innovations. Five executable checks in tests/test\_ssp\_claim\_boundaries.py establish these elementary counterexamples. They are not a new SSP theorem.
The historical feature-swap results remain: PCA whitening 0.881, linear discriminant 0.849, Nystr\"om-RBF 0.846, covariance-shaped random weights 0.894, Student-t0.894, Laplace 0.895, and Gaussian random projection 0.894. In the reported ten-examples-per-class comparison, the ℓ1 readout obtains 0.827 versus ridge 0.832. These outcomes establish no advantage for those tested alternatives, not Gaussian optimality, absence of all exploitable structure, or universal optimality of random projections. An ℓ1 estimator is not generically optimal for every sparse stochastic process.
The defensible interpretation is conditional: appropriate operator priors, observation models, regularization and computational structure can help. The cited dynamical-system/PDE gains are measured under their particular information and baseline protocols. Improvement in those experiments is not a universal ordering of representations, and changing coordinates within one function space is distinct from changing that space or its prior.
Closed-form meta-adaptation as a stability primitive for recursive self-improvement
Fixed-feature pooled estimation is useful in a recursive self-improvement loop because it retains earlier objective contributions without raw-data replay. It is not structurally immune to catastrophic forgetting or self-label error amplification. With conflicting labels the pooled optimum can worsen an old task; repeated erroneous pseudo-labels can bias the sufficient statistics. A frozen backbone and exact readout W=(G+λI)−1C prevent backbone updates and avoid iterative readout-solver error, but neither fact guarantees correct self-generated supervision. The following low-drift comparisons are empirical outcomes of their stated protocols, not consequences of universal immunity.
We test this directly with a ``telephone-game'' self-training loop on a frozen backbone. Starting from a tiny labelled seed (one or two examples per class on CIFAR-100 features), the agent repeatedly pseudo-labels a fresh batch of unlabelled data with its own current model, folds it into its experience, and updates—for many generations, with no further ground truth. The only thing that differs between the two learners is the update rule: an online gradient step on each self-labelled batch (the dense-weight route), versus accumulation into the closed-form Gram memory. The outcome is unambiguous. The closed-form learner bootstraps and stays stable: from a single label per class it climbs from 49% to 62% and holds there, its self-generated labels improving across generations (0.42→0.62). The gradient learner collapses: it drifts downward from 40% to 34% as its pseudo-label accuracy decays generation over generation—the textbook autoregressive failure—ending roughly 28 points below the closed-form learner (the same pattern holds at two labels per class, 73% stable versus 44% and falling). The closed-form update is thus not merely a cheaper continual learner; it is a mechanism for self-improvement that accumulates capability without forgetting or drift where the gradient loop degenerates. The pattern is not an artifact of one backbone: across a supervised ViT-B, a self-supervised DINOv2 ViT-L, and a weaker ResNet-18, and on both CIFAR-100 and ImageNet-R features, the closed-form learner stays stable or improves across generations while the gradient learner drifts downward in every case. This connects the operator-matched, closed-form philosophy of this paper to the stability of open-ended, self-improving systems (Figure [fig:rsi]).
Recursive self-training (``telephone game'') from a one-label-per-class seed on a frozen
backbone: each generation the agent pseudo-labels fresh data with its own current model and updates.
The closed-form Gram learner (blue) bootstraps and stays stable; the gradient learner (red) drifts
downward as its self-generated errors compound—the autoregressive collapse. Only the update rule
differs.
Evolving architecture plus stable self-improvement: the full loop.
The same closed-form memory composes with architecture growth to give the open-ended picture in full. We stream a hundred skills (CIFAR-100 classes) at ten new skills per generation; at each generation the agent grows fresh closed-form capacity for the new skills and must retain all earlier ones, against a gradient agent that expands its head and trains by stochastic gradient descent. The contrast is categorical (Figure [fig:grow-rsi]): growing from a small labelled seed, the closed-form agent holds 83% accuracy over all hundred accumulated skills and answers the first generation's skills at 89%, while the gradient agent collapses to 1% overall with zero retention of the first skills. Capacity grows on demand and nothing already learned is disturbed.
Closing the loop requires self-improvement to add capability rather than corrupt it, and this is where the design matters. A naive self-labelling loop does corrupt even the closed-form learner—a new skill's unlabelled data is confidently mislabelled as old skills before its prototype is established, and a plain confidence gate only softens this (52%). The fix uses the temporal structure that is genuinely available without labels: a freshly arrived unlabelled batch belongs to the new skills, so its pseudo-labels are confined to the current generation's classes and admitted under a confidence gate, letting the new prototypes bootstrap from unlabelled data instead of being absorbed by old ones. With this, self-improvement exceeds the labelled-seed-only model (83.4%→88.5% over all hundred skills) while retaining the first skills at 94%, approaching the fully-supervised ceiling of 89.8%—all with growth on demand, no forgetting, and the gradient agent still at 1%. The open-ended loop is therefore complete: grow new capacity, improve it from a few labels plus unlabelled experience, and never forget or drift—a closed-form realisation of the stability a self-improving system requires.
Evolving architecture with stable self-improvement: a hundred skills streamed ten at a time,
the agent growing capacity for each and self-improving from a few labels plus unlabelled data
(temporally-restricted, confidence-gated). The closed-form agent (blue/green) reaches 88.5% over
all skills and 94% on the first generation after ten generations—above its own labelled-seed-only
model (83.4%) and approaching the fully-supervised ceiling (89.8%); the gradient agent
(red/orange) collapses to chance and forgets the first skills entirely. Same frozen backbone; only the
learning machinery differs.
Operator-matched identification versus generic polynomial discovery
The grown-topology study (S[sec:grown-topology]) and the cell-operator mismatch audit establish that the basis must match the operator. This subsection makes the consequence quantitative against the standard equation-discovery baseline, sparse identification of nonlinear dynamics (SINDy) [brunton2016sindy]. SINDy regresses estimated state derivatives onto a fixed feature library, typically polynomial, with sequential thresholded least squares. The OSNR position is identical in algorithm but insists that the library be the operator-matched dictionary rather than a generic polynomial one, and that derivatives come from the autograd-free operator-spline ladder of S[sec:autograd-free].
Polynomial systems: parity.
On the Lorenz system, which lies in a degree-two polynomial span, the library is matched for every method and only the derivative estimator differs. Across a noise sweep (five seeds) the operator-spline derivative ties a well-tuned smoothed-SINDy and clearly beats the naive finite-difference SINDy default (coefficient error at five percent noise: finite difference 0.046, smoothed 0.015, operator-spline 0.012); at the highest noise a hand-tuned Gaussian smoother is marginally better. The honest conclusion is parity: when the operator is unknown or polynomial, OSNR is never worse than SINDy but does not dominate it.
Non-polynomial systems: decisive separation.
The picture changes when the operator is non-polynomial, which is the physically relevant regime for saturating, conductance-like, or biological dynamics. Consider the coupled saturating system x˙i=j∑Aijtanh(xj)−dixi, which is linear in its coefficients under the matched library [xi,tanh(xj)]. Training uses multiple trajectories spanning the saturating regime ∣x∣≤3 so the nonlinearity is identifiable, and models are tested both in-distribution and on extrapolation to ∣x∣≤6. Table [tab:nonpoly-sysid] reports normalized RMSE over six seeds.
method
in-distribution nRMSE
extrapolation nRMSE
terms
SINDy, degree-3 polynomial (standard)
0.31
12.5 (diverges)
12
SINDy, polynomial +tanh features
0.0066
0.026
11
OSNR, operator-matched [x,tanhx]
0.0050
0.0065
11
Identification of a non-polynomial saturating system. The standard polynomial
SINDy library is 40–60× worse in-distribution and roughly 400–2000×
worse on extrapolation, where the polynomial approximation of tanh yields an unstable
identified model that diverges outside the training range. The matched basis extrapolates
exactly. The result is robust to observation noise (extrapolation nRMSE for the polynomial
library stays near 12.5 at 0, 5, and 10 percent noise, versus 0.006, 0.007,
0.029 for the matched basis).
Combined with the polynomial-system parity above, OSNR is never worse than generic sparse discovery and is decisively better when the governing operator is known and non-polynomial, the physics-informed and biological regime this work targets.
Generality across operator families.
The separation is not specific to the saturating tanh nonlinearity. Repeating the experiment with an oscillatory family (g=sin, matched library [x,sinx,cosx]) and a rational family (g=1/(1+x2), matched library [x,1/(1+x2)]) gives the same outcome: the standard polynomial library diverges on extrapolation (nRMSE 13.3 and 7.2 respectively) while the matched basis remains accurate (0.0076 and 0.0034), a three-order-of-magnitude gap in every case.
Scaling with dimension.
The separation is not a small-system artifact; it widens with dimension. For N-unit saturating networks x˙i=∑jAijtanh(xj)−dixi, the matched library has 2N features (linear in N), whereas a polynomial library grows combinatorially. Across N∈{3,6,12,20} the matched basis stays accurate (forecast nRMSE 0.006 to 0.028) and fast, while a degree-two polynomial library degrades catastrophically (0.10 at N=3 to 65 at N=20, already diverging by N=6): the matched representation is roughly three orders of magnitude better at N=20 while using 40 features against the polynomial's 231.
Extrapolation error versus system dimension. The operator-matched basis stays
accurate while a polynomial library degrades and diverges; the gap widens with dimension.
Physical example: the large-angle pendulum.
The effect is not an artifact of synthetic systems. For the damped pendulum θ¨=−(g/L)sinθ−γθ˙, physics supplies the matched feature sinθ, whereas the small-angle polynomial approximation sinθ≈θ−θ3/6 is famously wrong at large amplitude. Training at moderate amplitude and forecasting a near-inverted swing (θ0≈2.8, five seeds), the matched library reaches extrapolation nRMSE 0.30, while a degree-five polynomial SINDy model diverges (7.2) and a black-box neural ODE—which fits in-distribution best—fails to extrapolate the physics (1.3).
Biological capstone: a conductance neuron.
The motivating case for this entire program is a neuron, whose dynamics are conductance ODEs with sigmoidal gating. For a Morris–Lecar-type model, V˙=I−gL(V−EL)−gKw(V−EK)−gNam∞(V)(V−ENa),w˙=τw∞(V)−w, with m∞,w∞ sigmoidal (tanh) gating, the model is linear in its conductances under the matched library [1,V,w,wV,m∞,m∞V,w∞] once the gating midpoints and slopes are taken from biophysics. Identifying the model from noisy voltage traces (five seeds), the matched basis reaches forecast nRMSE 0.0019 in-distribution and 0.0023 extrapolated to a large voltage excursion, versus 0.061/0.080 for a degree-five polynomial SINDy model and 0.018/0.156 for a neural ODE—a 30–70× advantage, largest in extrapolation. This is the thesis in its native setting: when the operator is the biophysics, encoding it is decisively better than approximating it.
Comparison with neural ordinary differential equations.
The other deep-learning approach to learning dynamics from data is the neural ODE, a black-box multilayer-perceptron vector field fθ trained by backpropagation through an ODE solver [chen2018neuralode]. Table [tab:osnr-vs-node] compares OSNR's closed-form matched identification against a well-trained neural ODE (sixty-four hidden units, roughly nine thousand parameters, fifteen hundred epochs of RK4 backpropagation) on the saturating system, as a function of the number of training trajectories.
trajectories
OSNR in/extrap
OSNR time
neural ODE in/extrap
NODE time
speedup
2
0.021/0.076
0.002 s
0.317/0.487
11.1 s
4621×
8
0.010/0.008
0.007 s
0.118/0.451
11.6 s
1608×
32
0.010/0.007
0.027 s
0.037/0.114
11.7 s
430×
OSNR closed-form matched identification versus a well-trained neural ODE on the
saturating system (forecast nRMSE, four seeds). OSNR is more accurate at every data budget
(its two-trajectory model already beats the neural ODE trained on thirty-two), extrapolates
far better, and is 430–4600× faster to fit. The neural ODE is converged, not a
strawman; it loses because it is a structure-free black box. This is the physics-informed
regime: OSNR exploits the known operator family while the neural ODE learns it from scratch.
Interpretation.
Two controls make the claim honest. First, the SINDy advantage is not a derivative trick: both methods use the same derivative and sparse regression. Second, a SINDy variant given the matched features also succeeds, so the separation is entirely about exploiting known operator structure—precisely the OSNR premise. The matched library is additionally leaner and more noise-robust than the augmented one. Together with the closed-form speed and data-efficiency over neural ODEs and the autograd-free identification speed over backpropagation (S[sec:grown-topology], Rung 0), the operator-matched representation is Pareto-dominant for identifying known-structure dynamical systems.
Scope and a negative control.
The advantage is specific to the regime where the operator family is known. When the basis must instead be discovered from a large overcomplete library deliberately populated with features collinear to the truth, sparse thresholded regression is already strong: on such a coherent library SINDy attains extrapolation nRMSE 0.015, whereas an OSNR rank-revealing column selection followed by sparse refitting reaches only 0.23, despite yielding a leaner (13 versus 33 terms), far better-conditioned (108 versus 1013), and more seed-stable model. We therefore do not claim that operator-matched machinery improves blind equation discovery; the claim is narrower and the experiments support it: when the governing operator family is known—the physics-informed setting—encoding it in the basis decisively beats both generic sparse discovery and black-box neural learning, especially in extrapolation.
Scaling to PDEs: operator-matched identification versus neural operators
The dimension-scaling result of S[sec:matched-vs-sindy] predicts that the operator-matched advantage should be largest for high-dimensional spatial operators, i.e. partial differential equations. We test this against the neural-operator state of the art, the Fourier neural operator (FNO) [li2021fno], on the one-dimensional viscous Burgers equation ut=−uux+νuxx on a periodic domain. The OSNR method estimates a small matched spatial-operator library [ux,uxx,uux] from trajectory data, solves the coefficients in closed form, and forecasts by integrating the identified equation (method of lines, spectral spatial derivatives). The FNO is trained autoregressively.
Data-efficiency, extrapolation, speed.
Table [tab:pde-fno] forecasts fresh initial conditions over a horizon twice as long as any seen in training, as a function of the number of training trajectories. A well-trained FNO (74k parameters, GPU) is the baseline.
trajectories
OSNR in/extrap
OSNR time
FNO in/extrap
FNO time
2
0.043/0.062
0.002 s
0.451/0.720
8.1 s
8
0.029/0.041
0.007 s
0.171/0.433
8.2 s
32
0.014/0.020
0.029 s
0.041/0.073
8.3 s
Burgers forecasting, OSNR operator-matched identification versus a Fourier neural
operator (nRMSE, in-horizon / twice-horizon extrapolation; twelve test trajectories). OSNR is
more accurate at every data budget—its two-trajectory model beats the FNO trained on
thirty-two—roughly three to four times better on long-horizon extrapolation because it
integrates the identified equation rather than rolling out a learned autoregressive map, and
is two-to-four orders of magnitude faster to fit.
The advantage is not specific to the FNO. Against the other neural-operator family, DeepONet [lu2021deeponet], in its native operator-map mode (branch encoding the initial condition, trunk the space-time query) and fully trained, the forecast nRMSE on the same problem is 1.14, 0.67, and 0.70 at two, sixteen, and thirty-two training trajectories—worse than both the FNO and OSNR, partly because the forecast horizon extends beyond the training window and the trunk does not extrapolate in time. OSNR therefore outperforms both standard neural-operator baselines, with the FNO the stronger of the two.
Cross-regime adaptation.
The black-box weakness is generalization across physical regimes. An FNO trained at one viscosity cannot forecast another; OSNR re-identifies the viscosity from a short snippet of the new regime in closed form and forecasts any of them. Training at ν0=0.06 and testing across ν∈[0.045,0.12] (Table [tab:pde-crossreg]), the FNO degrades by up to an order of magnitude away from ν0, while OSNR re-identifies the viscosity to within one percent from thirty frames and forecasts at sub-percent error throughout. A frozen-coefficient OSNR control degrades just as the FNO does, confirming that the advantage is the closed-form re-identification, not the representation alone.
test ν
FNO (trained at 0.06)
OSNR frozen
OSNR adapt
ν recovered
0.045
0.109
0.046
0.0059
0.0447
0.060
0.085
0.006
0.0017
0.0598
0.100
0.174
0.143
0.0013
0.0998
0.120
0.231
0.210
0.0012
0.1196
Cross-regime forecasting (nRMSE). OSNR adapts to an unseen viscosity by closed-form
re-identification from a thirty-frame snippet and forecasts at sub-percent error across the
range; the FNO, trained at a single viscosity, cannot adapt and degrades away from it. The
frozen-coefficient OSNR control degrades like the FNO, isolating re-identification as the
mechanism.
Cross-regime forecasting versus test viscosity. The FNO (trained at one viscosity)
and frozen-coefficient OSNR both degrade away from the training regime; OSNR with closed-form
re-identification stays at sub-percent error across all regimes.
The same adaptation holds in two dimensions: training the FNO at ν0=0.07 on a 64×64 grid and testing across ν∈[0.05,0.09], OSNR re-identifies the viscosity near-exactly and forecasts at ∼10−3 at every regime, while the FNO sits at 0.17–0.28 throughout (and a frozen-coefficient OSNR control again degrades off-regime).
Generality across PDEs.
The advantage is not specific to Burgers. On a bistable reaction-diffusion (Allen–Cahn) equation ut=νuxx+u−u3 with matched library [uxx,u,u3], OSNR recovers the exact coefficients from two trajectories and forecasts at nRMSE ≈10−4 at every data budget (the dynamics are shock-free, so integrating the identified equation is near machine precision), whereas the same well-trained FNO ranges from 0.40 at two trajectories to 0.045 at thirty-two—a three-to-four-order-of-magnitude gap.
Noise robustness.
Real data is noisy, and under noise both derivatives are fragile: spectral spatial derivatives amplify high-wavenumber noise by k2, and finite-difference time derivatives amplify white noise. The OSNR remedy is operator-spline smoothing before differentiating—a data-adaptive spatial low-pass (the cutoff set from the high-wavenumber noise floor) plus per-point Tikhonov temporal smoothing. Identifying Burgers from noisy trajectories and forecasting from a clean state (Table [tab:pde-noise]), the smoothed solver degrades gracefully and beats the FNO at every noise level, while the unsmoothed solver collapses—confirming that the smoothing, not merely the closed form, is the essential ingredient—and the FNO's autoregressive rollout diverges at ten percent noise.
observation noise
OSNR raw
OSNR spline
FNO
0%
0.007
0.017
0.082
2%
0.495
0.090
0.192
5%
0.748
0.173
0.281
10%
0.913
0.245
diverged
Burgers identification under observation noise (forecast nRMSE from a clean state).
Operator-spline smoothing (adaptive spatial low-pass + Tikhonov temporal) degrades gracefully
and beats the FNO at every level; the unsmoothed spectral solver collapses; the FNO rollout
diverges at 10%. At high noise the smoothed solver trades mild coefficient bias for
stability.
Identification under observation noise. The operator-spline solve degrades
gracefully; the unsmoothed solve collapses and the FNO rollout diverges at high noise.
Chaotic fourth-order operators via the weak form.
The hardest case is a chaotic, fourth-order operator, the Kuramoto–Sivashinsky equation ut=−uux−uxx−uxxxx. Here the strong form fails: the uxxxx feature amplifies high-wavenumber content by k4, so differentiating chaotic data directly gives unstable, trajectory-dependent coefficients (recovered values scattered over −0.5 to −1.0, a relative error of 0.32). The remedy is the operator-spline weak form: integrate the equation against smooth compactly-supported test functions and move every derivative analytically onto the test function, so the data enters only as u and u2 and is never differentiated. With this, the coefficients are recovered as (−1.0002,−0.998,−0.998) against a true (−1,−1,−1)—a relative error of 0.002 with negligible seed-to-seed variance, two orders of magnitude better than the strong form. The matched-operator approach thus extends, with the weak form, even to chaotic high-order PDEs.
Two-dimensional capstone.
Neural operators are used above all in two and three spatial dimensions, and the dimension-scaling argument predicts the matched-operator advantage should be largest there. On a 64×64 periodic 2D reaction-diffusion ut=ν(uxx+uyy)+u−u3 with matched library [∇2u,u,u3], against a well-trained 2D FNO (4.6×105 parameters, fifteen hundred epochs, 34 s per fit), Table [tab:pde-2d] shows the largest separation of all: OSNR recovers near-exact coefficients from two trajectories, forecasts at ∼10−3 at every budget, and is about sixty times more accurate than the FNO even at the FNO's best data budget, two to three orders of magnitude better in extrapolation, and three to four orders of magnitude faster to fit.
trajectories
OSNR in/extrap
OSNR time
FNO in/extrap
FNO time
2
0.0010/0.0012
0.01 s
0.524/0.926
35 s
8
0.0009/0.0010
0.05 s
0.116/0.243
34 s
16
0.0009/0.0011
0.11 s
0.060/0.105
34 s
2D reaction-diffusion forecasting (nRMSE, in-horizon / twice-horizon). In the
two-dimensional setting where neural operators are normally deployed, OSNR's two-trajectory
model is roughly sixty times more accurate than the FNO trained on sixteen, with two-to-three
orders of magnitude better extrapolation and three-to-four orders faster fitting. The gap is
not a single-seed artifact: over five FNO training seeds the in-horizon error is
0.27±0.09 at eight trajectories and 0.13±0.04 at sixteen, far larger than OSNR's
∼10−3.
OSNR versus FNO forecast error against the number of training trajectories, in 2D
(left) and 3D (right). OSNR is flat near machine precision; the FNO needs far more data and
never closes the gap.
The same separation holds in three dimensions, the most demanding neural-operator setting: on a 323 grid OSNR recovers near-exact coefficients from two trajectories and forecasts at ∼10−3, while a 3.7×105-parameter 3D FNO reaches only 0.25–0.32 in-horizon and 0.61–0.71 in extrapolation—roughly a 300× and 500× gap—at two orders of magnitude more compute. The matched-operator advantage thus grows monotonically across one, two, and three dimensions, exactly as the dimension-scaling argument predicts.
Partial knowledge: a hybrid of known core and learned residual.
The scope caveat of the whole approach is that it presumes the operator family is known. Real physics is usually only partially known. The operator-matched representation extends naturally to this case by the sparse-plus-smooth construction of S[sec:matched-vs-sindy]: keep the known operator terms as a closed-form core and add a small learned residual dictionary for the unmodeled part. On Burgers with an unknown non-polynomial reaction added, ut=−uux+νuxx+0.7sin(2.5u), where the method is told only the advection-diffusion core, Table [tab:pde-hybrid] compares a misspecified pure core, the hybrid (core plus a fourteen-function radial-basis residual in u, still one closed-form ridge solve), and an FNO.
trajectories
OSNR core only (misspecified)
OSNR hybrid
FNO
2
0.446
0.027
0.560
8
0.453
0.024
0.264
32
0.464
0.022
0.060
Partial-knowledge regime (Burgers plus an unknown reaction). The misspecified core
cannot improve with data (a model-form error), and the black-box FNO needs many trajectories;
the hybrid—known core plus a small learned residual, solved in closed form—captures the
unknown reaction and is data-efficient (its two-trajectory model beats the FNO trained on
thirty-two). This extends the approach from a fully known operator to physics-plus-discrepancy.
Partial-knowledge regime. The misspecified core is flat (model-form error); the FNO
improves slowly with data; the hybrid (known core plus a small learned residual) is accurate
and data-efficient.
The PDE results mirror the ODE ones (S[sec:matched-vs-sindy]) at higher dimension and against a stronger black-box baseline: in the regime where the operator family is known, encoding it yields a representation that is more data-efficient, more accurate in extrapolation, far faster, and—unlike a learned operator—instantly adaptable to a new physical regime. The same scope caveat applies: this is the physics-informed regime, and a neural operator remains the tool of choice when the governing equations are unknown.
Real measured data
All results so far use synthetic or simulated systems. We close the loop with three real measured benchmarks from the nonlinear system-identification literature, scored by free-run simulation error against published results, and summarized on a common axis in Figure [fig:real-data].
Silverbox (a real Duffing oscillator).
The Silverbox is a measured electronic circuit implementing a Duffing oscillator, my¨+cy˙+ky+k3y3=u. We identify the discrete-time matched form (a NARX whose nonlinear term is the known cubic) by a closed-form least-squares solve, then refine the coefficients by output-error (free-run) minimization from that initialization. On the three official test sets, the linear model gives 9.2, 14.9, 8.3 mV free-run RMSE; the closed-form matched model 8.3, 9.5, 7.1 mV; and the refined model 2.0, 3.4, 1.8 mV. The matched cubic is necessary (it separates from the linear model on the high-amplitude multisine), and the refined result is competitive with strong published nonlinear-identification methods (which report roughly 0.2–1 mV for the very best and 1–7 mV more typically); it is not the absolute state of the art on this much-studied benchmark, but it validates the matched-operator-plus-closed-form-plus-refinement pipeline on genuinely measured data.
Cascaded Tanks (an honest negative).
The Cascaded Tanks benchmark is a real two-tank fluid system with only 1024 training samples, a hidden upper-tank state, and an unknown overflow saturation. Here the matched-physics advantage does not materialize: the closed-form matched and hybrid models (0.87 and 0.98 V) do not beat a linear model (0.84 V), and only output-error refinement reaches the edge of the competitive range (0.75 V versus a published 0.3–0.7 V). The cause is structural and worth stating: the governing physics lives partly in the unobserved upper tank, so it cannot be expressed in lags of the measured lower-tank level alone, and the data is scarce. This sharpens the scope of the whole approach: the operator-matched advantage requires the relevant dynamics to be observable (as in Silverbox), and degrades to parity with generic models when a dominant state is hidden.
The remedy the diagnosis prescribes (latent-augmented matched model).
If the failure on Cascaded Tanks is caused by a hidden state, the principled fix is to restore that state explicitly: a grey-box model in which the unobserved upper-tank level is a learned latent variable evolved by its own known-form dynamics (Bernoulli square-root outflow, linear pump inflow, an overflow spill into the lower tank), with the lower tank as the observed output, fit end-to-end by output-error (back-propagation through the two-state free-run rollout, all physical parameters positive). This is the matched-operator principle carried into the partially-observed regime: a known-physics core paired with the minimal latent state the system requires. It works. The latent-augmented model reaches 0.55 V free-run RMSE on the test set, down from 0.75 V for the observed-only model and now inside the published competitive range of 0.3–0.7 V. Restoring the hidden state turns the honest negative into a competitive result, which is the strongest possible confirmation of the diagnosis: the obstacle was observability, not the matched-operator idea, and the same latent-augmentation recipe is what a real conductance-neuron recording (with its hidden gating variables) would require. The result is robust: across ten random initializations the fit converges to the same input-output behavior (0.55 V on every restart), with only mild non-identifiability in the absolute scale of the latent state (which is itself unobservable)—it is a stable basin, not a lucky seed.
EMPS (a friction-dominated positioning system).
The EMPS benchmark is a real electro-mechanical positioning system, a double integrator dominated by friction: Mq¨=u−Fvq˙−Fcsign(q˙)−τ0. The known nonlinearity is the Coulomb friction term sign(q˙), which we encode in the discrete-time matched NARX (with a small Stribeck residual of velocity radial basis functions for the hybrid model). The result isolates the value of the matched nonlinearity cleanly: the linear (no-friction) model diverges in free-run (it is numerically unstable on this marginally-stable plant), whereas adding the known Coulomb term makes the simulation stable at 22 mm RMSE, and the hybrid Stribeck residual reaches 14 mm (about 17% of the output standard deviation), inside the published range of roughly 3–15 mm. Two points are worth noting. First, the matched friction term is decisive for stability, not merely accuracy: without it the free-run model has no usable prediction at all. Second, output-error refinement yields no improvement here—the closed-form one-step fit is already at a free-run optimum—in contrast to Silverbox, where refinement was essential. The closed-form solution is thus sometimes already output-error-optimal, and sometimes only a good initialization; which case obtains depends on the conditioning of the simulated rollout. We also tried, for completeness, the latent-state continuous grey-box that succeeds on Cascaded Tanks (below)—treating velocity as an explicit hidden state and integrating the friction ODE—but on EMPS it is markedly worse (148 mm) because the plant is a pure double integrator: open-loop integration of a slightly imperfect acceleration accumulates unbounded position drift over the long free-run, whereas the position-feedback NARX form is anchored and stable. The right matched form therefore depends on the stability character of the operator (dissipative versus integrating), not only on observability—latent augmentation helps the bounded, dissipative tank dynamics and hurts the marginally-stable integrator.
Real measured data, three benchmarks, on a common axis (free-run RMSE as a percentage of
the test-output standard deviation; log scale). Where the governing dynamics are observable in the
measured output (Silverbox, EMPS), the matched/hybrid OSNR model improves by a large factor over a
linear baseline and approaches the strong end of the published range; for EMPS the linear
no-friction model is numerically unstable in free-run (hatched bar, capped). When a dominant state
is hidden (Cascaded Tanks, unobserved upper tank), the matched model collapses to parity with the
linear baseline and stays far from the state of the art. Observability—not the presence of a
nonlinearity per se—governs whether the matched-operator advantage materializes. The black
diamond on the Cascaded Tanks group is the remedy: a latent-augmented matched model (the hidden
upper tank restored as a learned state) drops back into the published competitive band.
Bio-conductance vision: retina, V1, and predictive residual fields
The next runner, apps\_industrial\_breakthrough/bio\_conductance\_vision\_benchmark.py, is the first real-data attempt to move away from a rigid fixed-feature interpretation of the liquid thesis. The architecture is deliberately cellular rather than MLP-like. A retinal front end performs local contrast normalization. A V1-like bank computes oriented even/odd quadrature energy, divisive normalization, and lateral inhibition. The inhibited visual field is pooled into 7×7 token maps and scanned in row, reverse-row, column, and center-out orders by exact conductance cells.
For a token zk, each cell uses excitatory and inhibitory conductances gi,k+=σ(γi+(wi+⊤zk+bi+)),gi,k−=σ(γi−(wi−⊤zk+bi−)), and the exact zero-order-hold update xi,k+1=ρi,kxi,k+(1−ρi,k)λi+gi,k++gi,k−gi,k+Ai++gi,k−Ai−,ρi,k=exp[−Kλi+gi,k++gi,k−]. A lateral competition step xi←tanh(xi−ηxˉ) follows each token update. This keeps the activation mechanism aligned with the LTC/conductance audit: the nonlinearity changes the pole and reversal equilibrium, not merely a pointwise post-activation.
The readout is still algebraic. The base controller fuses four streams: random/analytic local convolution responses, low-frequency DCT identity coefficients, V1 inhibited energy, and conductance-cell scan states. A dopamine-style residual memory then treats the output innovation as a local control signal. For stage q, Rq=Y−Yq,Cc,q=topK{si:yi=c,∥Rq,i∥2}, where si is a compact retinal state plus a fixed random feedback projection of the V1/conductance state. Gaussian class-local memory features ΦCq solve Bq⋆=argBmin∥ΦCqB−Rq∥F2+λ∥B∥F2,Yq+1=Yq+αΦCqBq⋆. This is a small predictive-coding stack: each stage reselects high-innovation examples and applies a damped local correction. There is no reverse-mode differentiation through the visual front end, the conductance scan, or the residual stack.
llrrrr@
Dataset/split
Profile
Accuracy
Train ms
Est. MB
Backprop
Fashion 10k/10k
Prior local ensemble ×4
89.10%
4072.5
289.8
no
Fashion 10k/10k
Fixed local CNN ridge
89.19%
2118.3
660.7
no
Fashion 10k/10k
Retina/V1 inhibited ridge
88.46%
3964.6
134.8
no
Fashion 10k/10k
Conductance scan ridge
84.54%
5.9†
60.9
no
Fashion 10k/10k
Bio residual fused ridge
89.10%
62.2†
223.9
no
Fashion 10k/10k
Bio+local fused ridge
90.35%
1053.8
1013.8
no
Fashion 10k/10k
Bio dopamine residual
90.41%
1139.7
1217.1
no
Fashion 10k/10k
Predictive dopamine stack, 2 stages
90.44%
1177.8
1420.4
no
Fashion 10k/10k
Predictive dopamine stack, 3 stages
90.36%
1274.1
1623.6
no
MNIST 10k/10k
Bio+local fused ridge
98.14%
1012.7
1013.8
no
MNIST 10k/10k
Bio dopamine residual
98.23%
1139.0
1217.1
no
MNIST 10k/10k
Predictive dopamine stack, 2 stages
98.22%
1198.0
1420.4
no
Fashion 10k/10k
CNN, one epoch
75.83%
2145.8
4.4
yes
MNIST 10k/10k
CNN, one epoch
91.43%
2108.8
4.4
yes
Bio-conductance vision benchmark. The daggered rows reuse previously computed shared V1/conductance or DCT features, so their train time should not be read as a standalone full pipeline cost. All no-backprop rows use closed-form ridge or local residual solves.
Table [tab:bio-conductance-vision] is a real improvement over the earlier no-backprop vision frontier, but not a SOTA claim. On FashionMNIST 10k/10k, the best profile improves the previous local-ensemble mark from 89.10% to 90.44%, a +1.34 point gain. The improvement does not come from the conductance scan alone; by itself that scan reaches only 84.54%. The useful mechanism is the combination of residual identity channels, local convolutional evidence, conductance/V1 state, and shallow predictive residual correction. The 3-stage row is also important: more local memory is not automatically better, and undamped repeated correction can overfit or destabilize the class field. On MNIST, the same family transfers, with the single dopamine residual reaching 98.23% and the second stage slightly lowering accuracy to 98.22%.
This result answers part of the architectural criticism. The model is no longer just a rigid MLP/CNN/RNN/attention analogy with fixed random features; it contains retina-like normalization, V1-like competition, conductance-pole cellular dynamics, modulatory residual memory, and a predictive-coding correction loop. The boundary is equally clear. The best Fashion row still consumes about 1.42 GB in dense feature/readout memory, and it remains far below heavily tuned backprop SOTA on MNIST/FashionMNIST. The next liquid architecture must therefore make the residual stack local and streaming: solve many small region/cell normal equations, sparse-select conductance gates and feedback projections, and expose intermediate predictive targets, rather than fitting one dense global readout over all cellular features.
Bio-plasticity continual learning without replayed gradients
The dense-readout limitation suggests a different validation regime. Biological learning is not an offline i.i.d. fit over a stationary dataset; it is sequential plasticity under interference. The runner apps\_industrial\_breakthrough/bio\_plasticity\_continual\_benchmark.py therefore tests class-incremental MNIST/FashionMNIST. The learner receives five tasks with two classes per task and is evaluated after each task on all classes seen so far, then on the full ten-class test set. Backprop controls are small MLP/CNN models trained sequentially with AdamW. Two control regimes are reported: no replay, which exposes catastrophic forgetting, and equal-exemplar replay, which stores the same number of old images per class as the bio learner stores local center states.
The bio learner uses the same fixed retina/V1/conductance/local-operator front end as Table [tab:bio-conductance-vision], but projects the temporary feature stack into a compact state zi=∥tanh(P[di,vi,ci,ℓi])∥2tanh(P[di,vi,ci,ℓi]), where di are DCT coefficients, vi are V1 inhibited-energy features, ci are conductance scan states, and ℓi are fixed local-convolution responses. The stored biological memory is not the raw image and not the full feature stack. When class c arrives, the learner deposits a small diverse center set Cc,t=FPSK{zi:yi=c,i∈Tt}, using farthest-point selection in the compact state space. After each task, only the accumulated centers solve a local normal equation Wt⋆=argWmin∥[1,ZCt]W−YCt∥F2+λ∥W1:∥F2. Thus old classes are retained by deposited center states and a small algebraic readout, not by replaying old images through backpropagation. The stronger variant replaces the center buffer by a streaming covariance/eligibility field. For each task it updates only local sufficient statistics Gt=Gt−1+i∈Tt∑[1,zi]⊤[1,zi],Bt=Bt−1+i∈Tt∑[1,zi]⊤yi, and then solves the modulatory controller Wt⋆=argWminτ≤t∑∥[1,ZTτ]W−YTτ∥F2+λ∥W1:∥F2. This is recursive least squares written as a local co-activity field: Gt is an eligibility covariance accumulated from presynaptic cellular states, Bt is the dopamine/label-modulated cross-covariance, and the solve minimizes a quadratic control energy without reverse-mode gradients or raw-image replay. We also include ablations: max/mean RBF prototype voting, diagonal Gaussian local statistics, a class-subspace attractor energy, and a multi-head attention-fusion controller. The last two are useful negative results under this budget; splitting the state into weak heads or class subspaces did not beat the single stable covariance field.
llrrrr@
Dataset/order
Profile
Final acc.
Train ms
Est. MB
Backprop
Fashion canonical
MLP no replay, 3 ep/task
19.84%
212.8
0.65
yes
Fashion canonical
CNN no replay, 3 ep/task
19.91%
10222.1
0.91
yes
Fashion canonical
MLP equal replay, 3 ep/task
82.89%
345.0
16.00
yes
Fashion canonical
CNN equal replay, 3 ep/task
82.93%
16772.5
16.26
yes
Fashion canonical
Bio center linear readout, D=768
86.38%
2091.9
17.29
no
Fashion canonical
Bio covariance field, D=2048
89.37%
125.3
16.09
no
Fashion canonical
Bio covariance field, D=4096
89.98%
673.8
64.19
no
Fashion canonical
Bio dendritic covariance, 3×3072
90.13%
927.9
108.42
no
Fashion shuffled
MLP equal replay, 3 ep/task
80.88%
346.2
16.00
yes
Fashion shuffled
CNN equal replay, 3 ep/task
80.90%
16737.9
16.26
yes
Fashion shuffled
Bio center linear readout, D=768
86.39%
1457.2
17.29
no
Fashion shuffled
Bio covariance field, D=4096
89.98%
681.3
64.19
no
Fashion shuffled
Bio dendritic covariance, 3×3072
90.13%
926.8
108.42
no
Fashion full canonical
Bio MPS covariance, D=4096
91.54%
983.8
64.19
no
Fashion full canonical
Bio MPS covariance, D=8192
92.05%
4498.9
256.38
no
Fashion full shuffled
Bio MPS covariance, D=8192
92.06%
4511.5
256.38
no
Fashion full canonical
Bio streaming MPS covariance, D=12288
92.32%
40241.0
576.63
no
Fashion full shuffled
Bio streaming MPS covariance, D=12288
92.29%
40272.4
576.63
no
Fashion full canonical
Bio streaming MPS covariance, D=16384
92.24%
73964.3
1024.8
no
MNIST canonical
MLP equal replay, 3 ep/task
93.37%
347.6
16.00
yes
MNIST canonical
CNN equal replay, 3 ep/task
96.46%
16751.5
16.26
yes
MNIST canonical
Bio center linear readout, D=768
97.72%
1482.1
17.29
no
MNIST canonical
Bio covariance field, D=4096
98.33%
674.8
64.19
no
MNIST shuffled
MLP equal replay, 3 ep/task
93.18%
349.6
16.00
yes
MNIST shuffled
CNN equal replay, 3 ep/task
95.86%
16661.4
16.26
yes
MNIST shuffled
Bio center linear readout, D=768
97.73%
1505.4
17.29
no
MNIST shuffled
Bio covariance field, D=4096
98.31%
675.1
64.19
no
Bio-plasticity class-incremental learning on real MNIST/FashionMNIST subsets. Unless marked ``full'', each run uses 10,000 training and 10,000 test examples, five two-class tasks, and 3 backprop epochs per task for the replay controls. Full Fashion rows use all 60,000 training images and the same 10,000 test images. Center rows store 512 compact states per class. Covariance-field rows store only sufficient statistics Gt,Bt over the fixed cellular state and no raw exemplars. The dendritic covariance rows were selected by a separate 250-examples-per-class validation split over degree, projection depth, dendrite count, neuron count, ridge, and probability-fusion temperature, then refit on the full 10,000 training examples. MPS capacity rows use a larger 14,992-dimensional sensory stack and torch.mps for projection/covariance solves. Streaming MPS rows keep the sensory operators, image-to-state projection, covariance accumulation, solve, and evaluation on MPS and avoid the earlier host-side feature matrix; the memory column reports the covariance field, while projection/cache footprints are listed in the artifacts. Train time for the bio rows is the plasticity/readout update after fixed feature extraction or state caching; feature extraction/cache and projection time are recorded separately in the artifacts.
Table [tab:bio-plasticity-continual] is the strongest no-backprop learning result in this branch so far. It is not an offline SOTA classifier claim. It is a continual-learning claim under a specific replay/memory protocol: local cellular evidence plus algebraic plasticity resists class-incremental interference better than the bounded backprop controls, including equal-exemplar replay. On FashionMNIST canonical, the memory-matched covariance field reaches 89.37% with 16.09 MB of sufficient-statistic state, while the equal-replay CNN reaches 82.93% with 16.26 MB. The larger D=4096 covariance field reaches 89.98%. A validation-selected dendritic ensemble of three independent 3072-neuron covariance fields with probability fusion reaches 90.13% on both canonical and shuffled FashionMNIST, crossing the previous single-field ceiling. The larger MPS capacity runner apps\_industrial\_breakthrough/bio\_plasticity\_mps\_capacity\_benchmark.py uses 20 DCT modes, 12 V1 orientations, 4 scales, 384 conductance cells, 96 local convolution filters, and an 8192-neuron covariance field. On the full 60,000/10,000 FashionMNIST protocol it reaches 92.05% canonical and 92.06% shuffled, with 256.38 MB of covariance state; the MPS projection and solve take about 7.5 s and 4.5 s respectively after about 38.6 s of feature extraction.
The first implementation still underused the GPU because it materialized a multi-GB host feature matrix and then used MPS mostly for projection and dense solves. The streaming correction, apps\_industrial\_breakthrough/bio\_plasticity\_streaming\_mps\_benchmark.py, caches the DCT, V1/Gabor, local-convolution, and conductance operators on MPS, maps image batches directly to projected cellular state on MPS, accumulates covariance fields on MPS, and caches only the projected state. This reduces the full Fashion image-to-state cache pass to about 6.6–7.2 s for the D=8192–12288 rows, then exposes the true bottleneck: the covariance solve. The D=8192 streaming row reaches 92.10%, the D=12288 row reaches 92.29% with a float16 state cache and 92.32% with a float32 state cache, and shuffled order reaches 92.29%. Pushing to D=16384 lowers final accuracy to 92.24% while increasing the solve to about 74 s, so raw covariance width is now hitting a conditioning/credit-allocation wall. The next architectural step should not be another global dense field; it should use local/block covariance fields, low-rank Woodbury updates, gated dendritic subfields, or residual-modulated cell groups that preserve GPU residency without an O(D3) global solve. On MNIST, covariance-field plasticity reaches 98.33% canonical and 98.31% shuffled, versus 96.46% and 95.86% for equal-replay CNN. No-replay backprop collapses to about 18–20% final accuracy on both datasets, confirming that the task is measuring interference rather than ordinary stationary classification.
The next MPS experiment moved from algebraic readouts over fixed cellular states to a genuinely trainable neural network without reverse-mode differentiation. The runner apps\_industrial\_breakthrough/bio\_feedback\_alignment\_mps\_benchmark.py trains a two-hidden-layer local-feedback MLP on real MNIST/FashionMNIST images. The input is a fixed retinal/operator sensory stack: normalized pixels; 2× pooled identity channels; a low-frequency DCT block; and a local convolution bank with signed rectified responses and pooled/statistical summaries. For the strongest FashionMNIST row this gives 7832 sensory channels. The trainable network is x∈R7832tanh(W1x+b1)h1∈R4096tanh(W2h1+b2)h2∈R2048W3h2+b3y^∈R10. noindentArchitecture diagram.
No-backprop local-feedback MLP used in Table [tab:bio-local-feedback-mps-training]. The forward path is an ordinary two-hidden-layer neural classifier, but the backward path is not reverse-mode differentiation. The output innovation is broadcast through fixed random feedback matrices B1,B2; each layer updates only from its presynaptic activity and its local postsynaptic/modulatory signal.
Figure [fig:bio-local-feedback-mps-architecture] makes the key architectural distinction explicit. The forward pass is conventional enough to compare against backprop-trained MLP/CNN controls, but the credit path is a broadcast-modulatory path rather than a reverse-mode computational graph. No autograd graph is built for the OSNR/bio rows. The output innovation is e=softmax(y^)−y, and each hidden layer receives a fixed random feedback projection rather than the transpose of downstream weights: δ2=(eB2)⊙(1−h22),δ1=(eB1)⊙(1−h12). The local plasticity updates are the three-factor eligibility rules ΔW3=−ηh2⊤e,ΔW2=−ηh1⊤δ2,ΔW1=−ηx⊤δ1, with optional weight decay, no reverse-mode chain rule, no stored computation graph, and epochwise plasticity decay ηt=η0γt. The strongest full FashionMNIST run uses all 60,000 training images, 10,000 test images, batch size 512, η0=0.012, γ=0.94, feedback scale 1.0, no momentum, rank-free online updates, and MPS tensors. It reaches 92.34% online test accuracy, while the same trained representation with a final closed-form ridge readout reaches 92.41%. Under the same runner, the bounded 12-epoch backprop controls reach 88.74% for the small MLP and 91.14% for the shallow CNN. The exact metrics are in apps\_industrial\_breakthrough/bio\_feedback\_alignment\_mps\_outputs\_fashion60k\_h4096\_decay094/bio\_feedback\_alignment\_mps\_metrics.json.
Dataset/protocol
Profile
Accuracy
Train ms
Est. memory
Backprop
Fashion full
Local-feedback MLP, online
92.34%
54395.9
717.95 MB
no
Fashion full
Local-feedback MLP, ridge readout
92.41%
55630.7
717.95 MB
no
Fashion full
3-member local-feedback ensemble, online
92.44%
199815.0
2153.8 MB
no
Fashion full
3-member local-feedback ensemble, ridge
92.64%
199815.0
2153.8 MB
no
Fashion full
Prior cross-run OSNR/local-feedback fusion
93.23%
853738.1
2919.9 MB
no
Fashion full
Heterogeneous residual OSNR/local-feedback fusion
93.27%
557494.8
4464.3 MB
no
Fashion full
Small MLP control, 12 epochs
88.74%
983.3
2.60 MB
yes
Fashion full
Small CNN control, 12 epochs
91.14%
9118.6
3.63 MB
yes
MNIST full
Local-feedback MLP, online
98.71%
32142.5
529.50 MB
no
MNIST full
Local-feedback ensemble, ridge, 5 members
98.89%
493270.3
3864.4 MB
no
MNIST full
Streaming OSNR covariance, 2×8192
98.90%
62464.5
1976.1 MB
no
MNIST full
Margin-weighted OSNR/local-feedback fusion
99.14%
644565.9
4464.3 MB
no
MNIST full
Small CNN control, 12 epochs
99.09%
10745.3
3.63 MB
yes
MNIST full
Local-feedback trainable conv sheet
98.00%
42217.1
98.65 MB
no
Real neural training on MPS without reverse-mode differentiation. The local-feedback MLP rows train hidden weights by fixed-feedback three-factor plasticity, not by backpropagation. The FashionMNIST result is the first full-data stationary-vision run in this project where an online no-backprop neural network beats the bounded shallow CNN backprop control. The best Fashion fusion row uses a wider 6144/2048 local-feedback branch, a damped class-local residual-memory readout over trained h2 states, and two streaming OSNR covariance sources; the saved-logit postprocess applies fixed label-free margin weighting to centered logits. No labels, learned fusion weights, or reverse-mode graph are used in that postprocess. The MNIST fusion row is the first full-data run here to cross the same internal shallow-CNN control: it fuses the five-member local-feedback logits with streaming OSNR covariance-state logits, with no learned fusion weights and no reverse-mode graph. This is still not a public MNIST SOTA claim.
The follow-up ensemble runner apps\_industrial\_breakthrough/bio\_feedback\_alignment\_ensemble\_mps\_benchmark.py tests whether the remaining error is primarily local-credit variance. Three independently seeded 7832→4096→2048→10 local-feedback learners improve the online result to 92.44%, and averaging their closed-form ridge logits gives the current best stationary FashionMNIST no-backprop row, 92.64%. Pushing to five members raises online averaging only to 92.47% and lowers ridge averaging to 92.43%, so naive ensembling is not a route to a large jump. It reduces some variance, but the shared architecture still makes correlated errors.
The stronger MNIST experiment is apps\_industrial\_breakthrough/bio\_osnr\_no\_backprop\_fusion\_mps\_benchmark.py. It deliberately fuses mechanisms instead of merely widening one model. The first source is a five-member local-feedback ensemble with 11348-D sensory states, 4096/2048 hidden cells, 24 epochs, η0=0.01, γ=0.96, and no autograd; its ridge-logit average reaches 98.89%. The second source is a streaming OSNR/V1/conductance covariance field with a 20032-D operator feature stack, a two-head 8192-cell dendritic projection, and an online covariance solve; it reaches 98.90%. A single 16384-cell streaming source reaches 98.87%. Probability averaging is not enough—the all-source probability fusions reach only 98.83% and 98.80%—but centered-logit fusion exposes complementary evidence. Centering each source logit vector per example and averaging the four sources \local online, local ridge, 2×8192 streaming, 16384 streaming\ reaches 99.13% on the full 60,000/10,000 MNIST protocol. The deterministic postprocess runner apps\_industrial\_breakthrough/bio\_osnr\_fusion\_postprocess.py then evaluates label-free confidence rules on the saved logits; margin-weighted centered fusion reaches 99.14%. This crosses the bounded 12-epoch shallow-CNN control at 99.09% without reverse-mode differentiation. The result is important because it says the bottleneck is not only local-feedback seed variance: the OSNR covariance field makes different errors from the trainable local-feedback network.
The same transfer now holds on FashionMNIST. A direct full-data fusion run with the previous 4096/2048 three-member local-feedback ensemble and streaming sources single12288,dendrites2\_8192 reaches 93.10% when the local ridge logits are centered and averaged with the prior-best single12288 OSNR source; the all-source centered row reaches 92.90%. A more ambitious 6144/3072 local-feedback branch is not better alone—its ridge ensemble is 92.60% and its online ensemble is 92.22%—but it is more complementary to the OSNR stream. Fixed label-free margin-weighted fusion of \local online, local ridge, single12288\ reaches 93.19%. A subsequent saved-logit cross-fusion, generated by apps\_industrial\_breakthrough/bio\_osnr\_cross\_fusion\_postprocess.py with --keep\_duplicates and stored in apps\_industrial\_breakthrough/bio\_osnr\_no\_backprop\_cross\_fusion\_outputs\_fashion\_valid\_best.json, fixed-centers and equally averages the seven saved source logits from those two runs and reaches 93.23%, moving the Fashion no-backprop boundary by +0.59 points over the previous 92.64% ensemble row.
The next MPS sprint tested whether that boundary was caused by too little architecture diversity, weak residual control, or weak activation modeling. The new runner apps\_industrial\_breakthrough/bio\_spline\_residual\_feedback\_mps\_benchmark.py compares tanh, conductance-softsign, normalized RBF-spline, sinusoidal-pole, mixed-cell, and fixed sensory-residual variants under the same local-feedback rule. On the 10,000/3,000 FashionMNIST sweep, softsign plus a fixed sensory residual was the best small-split ridge row (88.47%), but under the stronger full-data no-normalization recipe it fell to 91.98% ridge versus 92.39% for the homogeneous tanh control. Thus the apparent activation/residual-skip gain was not robust. A readout-mirror feedback variant, where h2 receives current classifier weights as top-down apical feedback without autograd, also underperformed on the small split (88.60% ridge). Finally, an EGGROLL-inspired antithetic low-rank refinement over the trained W2 matrix reduced reward-batch cross-entropy but did not improve held-out accuracy, indicating that naive weight-space evolution is optimizing the wrong local objective.
The useful improvement came from a damped residual-memory controller over trained h2 states. With the original 4096/2048 local-feedback network, 64 high-residual centers per class, residual ridge 1.0, γ=0.1, and scale 0.1, the single-run FashionMNIST row improves from 92.41% ridge to 92.52% residual memory; 128 centers worsens to 92.48%, and a wider/lower-amplitude setting gives 92.51%. We then extended bio\_osnr\_no\_backprop\_fusion\_mps\_benchmark.py so residual-memory logits become first-class fusion sources. The 4096/2048 residual fusion with three local members and two OSNR streams reaches 93.21% after deterministic label-free postprocessing. A wider 6144/2048 local-feedback run with the same residual controller and two OSNR streams reaches 93.22% in-run and 93.27% after fixed margin-weighted centered-logit postprocessing, recorded in apps\_industrial\_breakthrough/bio\_osnr\_no\_backprop\_fusion\_mps\_outputs\_fashion60k\_lf3\_h6144\_residual64\_stream2/. Cross-run all-source fusion does not improve further (93.21%); a label-selected diagnostic subset reaches 93.35% but is explicitly not a benchmark claim. Thus the latest positive result is small but mechanistic: local residual memory adds a complementary error mode, while unstructured same-architecture columns, naive activation swaps, readout mirroring, and raw EGGROLL weight perturbations do not break the ceiling. The next version should make complementarity endogenous, by adding local predictive targets and topographic residual pathways inside the cellular network rather than fusing two finished systems after the fact.
noindentFusion architecture diagram.
No-backprop OSNR/local-feedback fusion architecture for the full MNIST 99.13% result. The top branch trains neural hidden weights by fixed-feedback local plasticity. The bottom branch builds operator-state evidence and solves local covariance fields. Fusion is fixed centered-logit averaging, not an extra trained classifier, so the positive result measures complementarity between two no-backprop evidence streams.
We also implemented a more explicitly spatial no-backprop CNN in apps\_industrial\_breakthrough/bio\_local\_feedback\_conv\_mps\_benchmark.py. Its visual sheet is image -> tanh(conv5x5) -> 2x average pool -> tanh(hidden) -> logits. The convolutional weights are trainable without loss.backward(): each update forms local image patches with torch.nn.functional.unfold, multiplies them by a neuromodulatory membrane delta projected from the output innovation, and applies the resulting local eligibility tensor to the 5×5 filters. On a 64-channel, 2048-hidden MNIST full run, this reaches 98.00%. That is a useful proof that trainable spatial filters can be updated by local tensor rules on MPS, but it is not yet the winning architecture. We then added residual readout variants and synaptic-homeostasis/validation-restoration controls to the local-feedback MLP. The residual readout variants helped the 2000-example smoke test but did not improve the full FashionMNIST result: x\_h1\_h2 finished at 92.22% and h1\_h2 at 92.16%, below the 92.34%h2-only online row. Homeostatic normalization with a 5% validation split finished at 92.30% online. The newer residual-memory fusion results above supersede the earlier conclusion that residual control only matches ridge; the corrected statement is that a carefully damped residual controller helps, but only by about 0.1 point as a single-network readout and about 0.04 point at the fusion frontier.
The first heterogeneous operator-cell moonshot, apps\_industrial\_breakthrough/bio\_heterogeneous\_operator\_dopamine\_mps\_benchmark.py, explicitly mixes leaky tanh cells, conductance-like softsign cells, oscillatory pole cells, and sparse event cells. It also adds regional predictive heads that can provide dopamine-like local class-prediction errors to each hidden layer. The small-split tuning showed that injecting regional heads into the final logits hurt, and that regional modulators did not yet improve over global feedback. The best full FashionMNIST heterogeneous run therefore used heterogeneous cells with global feedback only. It reached 92.23% online and 92.37% with a ridge readout, below the homogeneous local-feedback ensemble. This is a useful negative result: heterogeneity is likely necessary at scale, but the current mixture of cell operators and regional signals is not sufficient. We also tested a class-local high-residual RBF controller over trained hidden states as a dopamine residual memory. The undamped residual controller overcorrected badly on the small split, while a damped variant matched but did not improve the ridge readout. This repeats the earlier warning from dense memory kernels: residual controllers need local structure and validation gates, not just high-error centers in one global state space.
The current interpretation is sharper than before: no-backprop training can beat bounded backprop controls on full FashionMNIST when the sensory operator stack is rich and the local-feedback dynamics are stabilized, and a fixed fusion of local-feedback and OSNR covariance-state evidence now beats the bounded shallow CNN control on full MNIST. It still does not crush public SOTA. The next serious step is a deeper topographic local-feedback stack with residual identity paths, normalization/homeostatic targets, validation-selected plasticity schedules, and local predictive losses for intermediate layers, rather than another global covariance solve, a single random feedback projection, a wider final readout, a larger same-architecture ensemble, or unstructured heterogeneous cell mixing.
The first full-color vision stress is apps\_industrial\_breakthrough/bio\_physical\_resnet\_cifar\_mps\_benchmark.py. Unlike the older cross-architecture runners, it keeps CIFAR-10 as RGB 32×32 images. The physical model is a residual topographic reservoir x→C1(5×5)→pool→C2(3×3)→pool→C3(3×3)→s(x)→h→ℓ, where each convolutional sheet uses heterogeneous cell channels: tanh membrane cells, conductance-softsign cells, Gaussian spline-like cells, and rational-pole cells. When local training is enabled, each sheet receives a fixed direct neuromodulatory projection of the output innovation, and the convolutional update is the local product of unfolded presynaptic patches with the broadcast membrane delta. No reverse-mode graph is built for the physical rows. The final readout can either use online logits or a closed-form ridge solve over the physical state summaries.
Full RGB CIFAR-10 physical-ResNet stress on MPS. The physical rows use fixed heterogeneous ODE-like cell nonlinearities and closed-form readouts or local eligibility updates, not loss.backward(). The result is not a SOTA win: the best no-backprop physical reservoir reaches 57.33%, far below the small backprop TinyResNet at 82.29%. The useful signal is that a fixed physical reservoir plus ridge readout already extracts meaningful CIFAR evidence quickly, while the current direct-feedback local convolutional plasticity does not improve the reservoir and spatial random projections actually hurt.
This experiment changes the biological-learning diagnosis. The failure is not that physical states are useless; the 57.33% full-CIFAR row is far above chance and comes from a fixed heterogeneous physical reservoir. The failure is credit assignment inside the reservoir. The current local convolutional dopamine rule optimizes online logits weakly and does not make the final physical state more linearly separable than the untrained reservoir. The next no-backprop architecture therefore needs local predictive targets, contrastive/target-propagation-like regional objectives, or layerwise self-supervised physical fields before the supervised dopamine signal, rather than only a fixed output-error broadcast to every sheet.
The first response to that diagnosis is apps\_industrial\_breakthrough/bio\_scattering\_patch\_cifar\_mps\_benchmark.py. This runner keeps the same full RGB CIFAR-10 protocol, but replaces global-error-driven convolutional plasticity by a sensory growth model. The state contains color grid statistics, RGB DCT coefficients, color-opponent V1-like complex Gabor energy, class-balanced Hebbian patch filters, signed feature hashing into a compact cortical field, and the stronger random physical branch from Table [tab:cifar-physical-resnet]. The learned non-readout objects are local image patches sampled in a class-balanced way, or optionally from low-margin hard examples. The classification readout is still closed-form ridge. No reverse-mode graph is built for the physical rows.
First full-color CIFAR-10 architecture lift after the negative physical-ResNet stress. The fused physical+scattering reservoir improves the best full no-backprop CIFAR-10 result from 57.33% to 65.60%, an 8.27 point absolute gain, while remaining below the bounded TinyResNet backprop control at 82.29%. The negative rows are equally important: patch/Gabor scattering without the physical reservoir overfits badly, hard-example patch growth does not help, and a larger hash field or too many patch filters can reduce test accuracy.
The interpretation is more constructive than the previous negative result. The strong row is not a trained deep CNN in disguise; it is a fixed physical branch plus local sensory fields and a ridge readout. It therefore validates that OSNR-style operator states, color-opponent scattering, and local patch growth can add substantial linearly decodable evidence without backpropagation. It does not validate the full replacement thesis yet. Accuracy remains 16.69 points below the small TinyResNet control, and the best row still relies on a global algebraic readout rather than a fully local multilayer credit mechanism. A follow-up nonlinear mixed-cell readout expansion is negative on the 10k/3k protocol: concatenating a 4096-cell expansion drops the tuned fused state to 55.13%, and expansion-only drops to 53.63%. The next step should not be a larger patch bank or generic random nonlinear readout. It should use the 65.60% fused state as the sensory substrate, then add local predictive targets between regions so the hidden physical branch itself is shaped by non-terminal, reward-gated objectives.
We then pushed the architecture in two directions motivated by recent no-backprop and biologically plausible learning work: local target fields and population codes. The local target runner, apps\_industrial\_breakthrough/bio\_forward\_target\_cifar\_mps\_benchmark.py, learns layerwise class-prototype fields by local covariance/ridge solves and mixed ODE-like cells. It is a negative result: the target layers learn their own prototypes on the training set but do not improve test accuracy. The stronger direction is apps\_industrial\_breakthrough/bio\_population\_columns\_cifar\_mps\_benchmark.py, which trains independent physical/scattering columns and fuses their logits by a simple population mean or a small validation ridge head.
CIFAR-10 protocol
Model
Accuracy
Recorded time
Backprop
10k/3k
forward-only target stack, two 2048-cell target layers
57.13%
7.43 s
no
10k/3k
Fisher-selective patch synaptogenesis, single column
58.27%
4.75 s
no
50k/10k
Fisher-selective patch synaptogenesis, single column
65.35%
13.77 s
no
10k/3k
4 independent physical/scattering columns, mean logits
62.40%
19.01 s
no
10k/3k
4 columns, 2 deterministic train/test views, no validation holdout
65.73%
29.70 s
no
10k/3k
4 columns, 4 deterministic train/test views, no validation holdout
66.83%
152.73 s
no
10k/3k
8 heterogeneous hard-margin columns, 4 test views
67.77%
117.62 s
no
10k/3k
heterogeneous columns plus ES-CNN logits, confidence fusion audit
68.83%
504.52 s
no
10k/3k
local-feedback MLP over retinal/DCT/conv features
55.80%
11.03 s
no
10k/3k
auxiliary local-error RGB CNN, local sheet heads
32.63%
95.94 s
no
10k/3k
normalized residual CNN, binary direct feedback, ridge over best state
55.73%
207.70 s
no
10k/3k
normalized residual CNN, readout-aligned head feedback, online best
46.60%
275.08 s
no
10k/3k
normalized residual CNN, readout-aligned head feedback, ridge over best state
Post-failure CIFAR-10 architecture search. The target-field stack, Fisher-selective patch growth, trainable direct-feedback MLP, and auxiliary local-error CNN are useful negative controls. The first large positive jump after the 65.60% single-column result is a population code: independent no-backprop physical/scattering columns reach 69.79% on full CIFAR-10 when full-train fitting and two deterministic train/test views are used. The newer heterogeneous hard-margin population route improves the clean full-data no-backprop result to 72.06%; a rerun with saved core logits plus a prechosen hard-state spline controller reaches 72.37%. A validation-selected post-training controller over the population columns plus self-supervised residual heads reaches 72.66%. A later all-in population/augmentation audit reaches 73.69% with physical-only inference and no reverse-mode training. The 72.99% cross-run logit fusion remains a diagnostic audit until source sets are pre-registered or selected on validation. This is still far from SOTA and below the bounded TinyResNet control, but it is the strongest evidence so far that structured diversity of locally learned physical columns matters more than simply widening one column or attaching a naive local-error CNN.
The population result changes the next experiment. Generic target-field readouts, hard-example patch growth, Fisher-selected patches, random nonlinear readout depth, and a direct-feedback MLP are insufficient. The new auxiliary local-error RGB CNN in apps\_industrial\_breakthrough/bio\_cifar\_auxiliary\_local\_feedback\_mps\_benchmark.py is also negative: a 10k/3k run with four trainable convolutional sheets, local class heads, and local-head feedback ends at 32.63% after 30 epochs, with an unstable transient peak near 41.77%. We then tested a literature-inspired normalized residual CNN in apps\_industrial\_breakthrough/bio\_cifar\_cbdfa\_reservoir\_mps\_benchmark.py, adding batch-style homeostatic normalization, residual same-shape blocks, crop/flip/cutout augmentation, leaky/spline cell variants, direct binary feedback, rec-LRA-style reachable targets, and readout-aligned top-down head feedback. The best pure local-feedback CNN result is still only 46.60% online and 54.77% with a closed-form ridge readout over the best checkpointed state; the strongest ridge variant with random binary feedback reaches 55.73%.
We therefore added the missing evolutionary test to the same runner. After local no-backprop training, the script now starts from the best checkpoint and applies forward-only antithetic low-rank EGGROLL-style perturbations to either the readout plus deepest convolutional sheets or all convolutional sheets. Each ES pair is scored only by reward-subset cross-entropy; no reverse-mode graph, layerwise backpropagated gradient, or test-set selection is used. The smoke test verifies that sign-scored antithetic ES moves the actual CNN state: on 2k/1k, online accuracy rises from 36.10% to 37.90%. On the harder 10k/3k protocol, head/deep-sheet ES improves the online checkpoint from 41.93% to 44.40% and lifts the ridge readout over the evolved state to 59.13%. Perturbing all convolutional sheets is stronger: online accuracy reaches 46.20% and the ridge readout reaches 60.73%. This is the first positive CIFAR evidence in this branch that evolutionary, no-backprop weight-space refinement can improve a trainable convolutional physical network, but it is not a SOTA result and it still trails the fixed physical/scattering population code at 69.79%.
The failure is now more specific. Simply adding normalized direct feedback, reachable local targets, spline-like activations, readout-aligned feedback, or low-rank evolutionary perturbations does not reproduce the strong feature geometry of the fixed physical/scattering reservoir. Independent columns help because they preserve different local patch samples, random physical poles, and feature hash collisions. ES helps when it can perturb the convolutional sheets, but the present reward signal is still a shallow terminal classifier objective over a weak state.
We then changed the population route itself. The updated population runner can save logits, vary member architecture, and let later columns sample class-balanced patches from low-margin training examples discovered by the first columns. The heterogeneous recipe cycles filters per class, patch sizes, DCT ranks, hash dimensions, physical pole-bank widths, hidden physical dimensions, and ridge values; after three columns, subsequent patch banks are grown from the lowest-margin 35% of core examples. This is a different mechanism from simply adding more identical columns. On the 10k/3k split, eight heterogeneous hard-margin columns reach 67.77%, and a label-free confidence fusion audit with the ES-CNN logits reaches 68.83%. On full CIFAR-10, the same eight-column heterogeneous/hard run reaches 71.33% by mean logits and 72.06% by a core ridge fusion over column logits. A second independent four-column offset run reaches 71.40% by core fusion; fixed normalized averaging of the two runs' core and pairwise-fusion logits reaches 72.99%. The latter is recorded as a post-hoc audit, not yet as a clean benchmark claim, because the source set should be pre-registered or selected on validation before being treated as a final headline.
We then cross-pollinated this route with post-training model fusion, optimal-control language, and spline theory in apps\_industrial\_breakthrough/bio\_cifar\_control\_spline\_fusion\_audit.py. The saved member logits are treated as a columnar state: normalized class voltages, margins, entropies, votes, and disagreement terms. A confidence-bin reliability gate estimates source weights from core examples. A closed-form ridge controller then maps this state to class drives, analogous to an algebraic LQR-style terminal controller over a fixed dynamical state. Finally, a hard-state RBF spline residual appends kernels centered on core examples that are wrong or low margin, so that the controller can correct local residual geometry without reverse-mode gradients. On the rerun full-data sources, the eight-member population with saved core logits reaches 71.35% by mean logits and 72.11% by core ridge; the four-member offset run reaches 70.62% by mean logits and 71.48% by core ridge. Combining all twelve members, normalized mean logits reach 71.75%, reliability gating reaches 71.77%, the closed-form control feature ridge reaches 71.98%, and the prechosen 512-center hard-state spline controller reaches 72.37%. A diagnostic sweep with 1024 centers and a tighter scale reaches 72.57%, while 2048 centers does not improve it (72.54%). The diagnostic sweep is not a clean benchmark claim because the hyperparameter was chosen after seeing the test result.
The next cross-pollinated audit, apps\_industrial\_breakthrough/bio\_cifar\_ssl\_jepa\_attention\_fusion\_audit.py, explicitly tests whether pretraining and post-training ideas can add residual information without backpropagation. It constructs four closed-form self-supervised heads on full CIFAR: a spline/pole random-feature head, a JEPA-style deterministic-view prediction head, a denoising/diffusion-style latent recovery head, and a low-margin landmark-attention head. These heads are weak as standalone classifiers, reaching only 46.17%, 45.19%, 45.56%, and 46.52% respectively for the 512-latent/384-attention run. Naively appending them as equal experts hurts: mean fusion falls to 69.22%, reliability fusion to 69.82%, and the hybrid spline controller to 72.11%. The useful result appears only when the post-training controller treats them as candidate residual sources and chooses source set plus spline hyperparameters on held-out core examples. With a 5000-example core-validation split, the selected source is all twelve population columns plus all four self-supervised heads, with a 1024-center RBF spline controller, scale 0.5, and ridge 100. Refit on all 50k core examples, this reaches 72.66% on the CIFAR-10 test set. A larger 768-latent/512-attention profile improves the standalone attention head to 47.81% but falls to 72.52% after validation-selected fusion; a 10k validation split selects only the attention head as residual source and reaches 72.63%. Thus the clean gain is real but small, and capacity scaling over these shallow SSL heads overfits rather than compounding.
The next all-in push asked whether the gap to a frozen AlexNet sensory prior is mainly a missing post-training controller or a missing representation. We first added eight more heterogeneous hard-margin physical columns with a new seed offset. They reach 71.33% by mean logits and 72.34% by core ridge fusion, so identical population scaling is already saturated. A second four-column block with four deterministic train and test views reaches only 71.49% by core ridge, showing that simple view augmentation is also insufficient. However, fusing the original twelve columns with this augmented four-column block and a narrower 4096-center hard-state spline controller gives the strongest clean physical-only CIFAR row so far: 73.69%. Adding all twenty columns is worse: the core ridge remains only 72.27% and the large RBF controller collapses to 70.53%, which is a conditioning/overfit failure rather than a capacity gain.
We then made the AlexNet-geometry bridge explicit in apps\_industrial\_breakthrough/bio\_cifar\_alexnet\_logit\_bridge\_audit.py. The audit uses the frozen AlexNet CIFAR logits only as train targets, then evaluates a student whose inference path contains only physical-column logits. This distinction matters: teacher-guided rows are physical-only at inference, but they are not fully independent from scratch because their targets came from an external ImageNet-pretrained model. Over the best sixteen-column physical state, the label-trained hard-state spline reaches 73.69%; replacing the label target by AlexNet logits is worse at 73.42%; mixing AlexNet logits with the label target improves only to 73.98%. Thus the current gap to the 86.22% frozen-AlexNet readout is not a missing final controller. The physical columns do not yet contain enough of the AlexNet-class sensory geometry, and soft teacher logits can only add a fractional correction.
The all-in pretraining stress uses a different protocol and must not be mixed with the from-scratch physical-column claim. In apps\_industrial\_breakthrough/bio\_cifar\_frozen\_alexnet\_prior\_audit.py, the only locally cached external model was an ImageNet-pretrained AlexNet checkpoint. We froze it, extracted CIFAR features on MPS, and fitted only closed-form ridge readouts or the same spline/control fusion heads. There is no CIFAR backpropagation, but the sensory prior was trained externally and is therefore marked as an external-pretrained-prior result in Table [tab:cifar-external-alexnet]. The result crosses the requested 80% line easily: the frozen AlexNet multi-layer ridge readout reaches 84.78% with one deterministic view, 85.79% with two views, 86.02% with four views, and 86.22% with eight views. Adding the weaker physical/scattering population logits to the AlexNet logits hurts the best readout, although the fused spline controller still reaches 83.23% at eight views and the fused control ridge reaches 80.57%. The interpretation is precise: current post-training spline control can exploit a strong pretrained sensory cortex, but it does not yet make the weaker from-scratch physical columns competitive with that external prior.
CIFAR-10 protocol
External-pretrained sensory prior
Post-training readout
Views
Accuracy
CIFAR backprop
50k/10k
frozen ImageNet AlexNet
ridge over multi-layer features
1
84.78%
no
50k/10k
frozen ImageNet AlexNet
ridge over multi-layer features
2
85.79%
no
50k/10k
frozen ImageNet AlexNet
ridge over multi-layer features
4
86.02%
no
50k/10k
frozen ImageNet AlexNet
ridge over multi-layer features
8
86.22%
no
50k/10k
frozen ImageNet AlexNet plus 12 physical columns
control ridge over logits
8
80.57%
no
50k/10k
frozen ImageNet AlexNet plus 12 physical columns
hard-state spline controller
8
83.23%
no
50k/10k
AlexNet logits as training target only
16-column physical spline student
–
73.98%
no
External-pretrained-prior CIFAR-10 stress. The AlexNet weights are a locally cached ImageNet-pretrained prior, frozen during all CIFAR experiments. The readouts/controllers are closed-form and do not use CIFAR backpropagation. These rows show that the post-training OSNR/spline controller can cross 80% when given a strong pretrained sensory cortex. The final row removes AlexNet from the inference path but still uses its logits as a training target, so it is teacher-guided physical-only inference rather than a fully independent from-scratch no-backprop result.
For reproducibility, all CIFAR rows above use torchvision CIFAR-10 stored under artifacts/torchvision\_data, source-order seed 20260601, the full 50,000/10,000 train/test split, PyTorch 2.12.0, torchvision 0.27.0, and no CIFAR reverse-mode training in the reported readouts/controllers. The physical population logits used by the 80%+ AlexNet fusion are generated by the following two MPS runs; the runpy wrapper is intentional because direct script execution can hide the MPS backend in this environment:
Each population member builds a local physical/scattering state from RGB color statistics, RGB DCT coefficients, color-opponent Gabor energy, class-balanced Hebbian image patches, a signed hash field, and a random physical branch. The heterogeneous recipe cycles filters/class, patch sizes, DCT ranks, hash dimensions, physical widths, hidden dimensions, and ridge values; after member three, hard-margin synaptogenesis samples patches from the lowest-margin 35% of the core examples. The saved arrays member\_core\_logits, member\_test\_logits, y\_core, and y\_test are the only population inputs used by the downstream fusion audits.
The 86.22% frozen-AlexNet row and the 80.57%/83.23% AlexNet-plus-physical fusion rows are then reproduced by:
The only external prior in this command is the locally cached ImageNet AlexNet checkpoint \textasciitilde/.cache/torch/hub/checkpoints/alexnet-owt-7be5be79.pth. The script constructs torchvision.models.alexnet(weights=AlexNet\_Weights.IMAGENET1K\_V1), freezes every parameter, resizes CIFAR images to 224×224, applies ImageNet normalization, averages eight deterministic views, and concatenates fc6, fc7, and ImageNet logits into a 9192-dimensional feature vector. The CIFAR readout is a closed-form ridge solve with ridge 300 and targets 2onehot(y)−1. For the fusion rows, the AlexNet CIFAR logits are appended as one additional source to the twelve physical-column logit sources; normalized source voltages, margins, entropies, votes, and disagreement features feed either a closed-form control ridge or a 1024-center hard-state RBF spline controller with scale 0.5 and ridge 100. No CIFAR loss is backpropagated through AlexNet or through the fusion controller.
The clean 73.69% physical-only row and the 73.98% teacher-guided bridge row require one more four-column augmented population block and then two algebraic audits:
The bridge audit reports both label-only and teacher-guided rows from the same physical state. The label-only hard-state spline target gives 73.69%. AlexNet-logit targets alone give 73.42%. The mixed target AlexNet logits + 1.0*(2*onehot-1) gives 73.98% while using only physical-column logits at inference. This is why the bridge row is reported separately: it removes AlexNet from the inference path, but it does not remove the external teacher from training.
The immediate biological/predictive follow-up was to ask whether the missing AlexNet-like sensory geometry can be built internally by local physical pretraining rather than by another final controller. Two new MPS runners test this directly. The stricter runner, apps\_industrial\_breakthrough/bio\_cifar\_physical\_predictive\_pretrain\_mps.py, splits CIFAR into a 4×4 cortical sheet. Each local column receives RGB patch descriptors and an unsupervised Hebbian patch bank, then passes them through four fixed heterogeneous cell branches: a tanh leak cell, conductance-softsign cell, damped oscillatory pole cell, and signed Gaussian event cell. The column states are laterally diffused and inhibited on the sheet. Learning before labels is a closed-form local predictive map: for each region, north/south/east/west/global neighboring states and coordinates predict a deterministic target sensory view by a ridge normal equation. Labels enter only in the final ridge or spline readout. The strongest 10k/3k run used 96 unsupervised filters per kernel and 256 hidden units per column:
It reaches only 55.17% from the source state, 55.37% from the predicted target state, and 54.03% after predictive-logit spline control. The predictive residual state collapses to 45.43%, so the residual channel is not a useful class geometry. That early 10k/3k result was not the end of the path, however. After the pair-graph audit below showed that late logit-level residuals were saturating, we returned to this runner and scaled the representation itself on the full 50k/10k CIFAR protocol with guarded readout profiles, using the new --profiles option to omit the oversized residual feature solve. With 64 unsupervised patch filters/kernel and 128 hidden units per cortical region, the full run reaches 61.90% from the source state and 62.46% after the predictive-logit spline controller. Scaling to 96 filters/kernel and 256 hidden units per region raises the controller to 65.29%. Scaling once more to 128 filters/kernel and 384 hidden units per region gives the strongest from-scratch predictive-column source so far: source-state ridge 65.98%, target-state ridge 65.73%, predicted-target-state ridge 66.01%, and predictive-logit spline controller 66.20%. The key observation is that, at full scale, the predicted target state slightly exceeds the raw source state; the local view-prediction operator is no longer only smoothing away class geometry.
The second runner, apps\_industrial\_breakthrough/bio\_cifar\_scattering\_predictive\_geometry\_mps.py, starts from the stronger existing OSNR/V1/scattering substrate: color/DCT statistics, color-opponent Gabor energy, class-balanced Hebbian patch filters, signed hashing, and the fixed physical branch. It then fits a closed-form mixed-cell view-prediction map in a latent state and adds a hard-state feature-space RBF spline readout. The two registered 10k/3k profiles were:
The 12-filter profile reaches 59.83% with the source scattering state, 60.40% after feature-space hard-state RBF splines, and 60.50% after predictive-logit spline control. The predictive latent itself reaches only 51.33%, and concatenating source plus predictive geometry falls to 55.77%. The 24-filter profile is similar: source ridge 59.03%, feature RBF 59.53%, source-plus-predictive geometry 56.67%, and predictive-logit spline control 60.73%. The conclusion is not that predictive pretraining is useless in general. It is that one-shot deterministic-view prediction in a global latent space does not build the missing sensory hierarchy. It smooths or compresses away class geometry while the RBF spline recovers only a small local correction. The next internal-pretraining attempt must therefore use interacting columns with local target selection, contrastive negative states, or distance-forward residual objectives that preserve discriminative patch identity, rather than a single view-prediction ridge map attached after a fixed scattering encoder.
The clean conclusion at this point was that population diversity plus residual-style hard synaptogenesis moves the full no-backprop CIFAR frontier from 69.79% to 73.69% when fixed post-training spline control is allowed, while an external-teacher bridge reaches 73.98% with physical-only inference. The later full-scale predictive-column runs sharpen this story further. Appending the three logits from the 4×4, 384-hidden/region predictive physical column to the 14-source physical population and fitting the same closed-form hard-state spline controller reaches 75.51% on the full CIFAR-10 test set, with no external pretrained model and no reverse-mode CIFAR training. This beats the prospective pair-graph frontier below at 74.75%. The controller sweep is also informative: RBF scale 0.125 gives 75.36%, scale 0.20 gives 75.38%, ridge 5 gives 74.54%, ridge 20 gives 75.28%, 2048 centers gives 75.03%, and 8192 centers overfits to 74.74%; the best remains 4096 centers, scale 0.15, ridge 10. Adding the old top-six pair specialists on top of the predictive source drops to 74.99%, so the new gain is not a late residual-pair effect. It is a complementary early physical representation effect. This is still not SOTA and still below the small TinyResNet control at 82.29%, but it is the first full-CIFAR result in this branch where internal no-backprop predictive column formation gives a clear jump beyond the hand-built residual-controller frontier. The next serious architecture should therefore deepen this early route: multiple predictive sensory views, local contrastive negatives, recurrent column settling, and validation-selected predictive targets should be learned inside the column state before logit compression, rather than only voting at the end, training isolated sheet-local heads, or broadcasting one-step readout feedback. This also gives a concrete bridge to the broader biological thesis: pretraining can build sensory columns, post-training controllers can act as fast neuromodulatory adaptation, and spline/control residuals can target the hard state manifold without storing a reverse-mode computation graph.
The next MPS run tested whether the useful augmented block was a one-off or a reproducible high-information physical column. A second full 50k/10k four-member population with train views 4, test views 8, and seed offset 700000 again produced a strong member-2 column (71.61% test), close to the previous offset-500000 member-2 column (71.75%). However, adding both view-rich member-2 columns to the controller dropped the spline row to 73.84%, and replacing the old member by the new one reached only 73.91%. The conclusion is that this member type is reproducible but highly correlated across seeds; source diversity, not raw ensemble count, controls the residual spline gain.
We then pushed the same mechanism harder by increasing the sensory orbit inside the standout member rather than adding more columns. This required two implementation changes. First, bio\_population\_columns\_cifar\_mps\_benchmark.py now accepts --member\_ids, so a targeted run can instantiate only the heavy heterogeneous member-2 architecture. Second, bio\_scattering\_patch\_cifar\_mps\_benchmark.py now detects oversized MPS ridge systems and accumulates the normal equations from CPU-held feature chunks streamed through MPS. The original single train.T @ train path hits an MPSGraph tensor-dimension limit for the 400,000-row train-view-8 system; the streamed path preserves the same closed-form ridge objective without backpropagation.
The selected member uses the heterogeneous recipe's third architecture: 16 filters/class, patch sizes 3 and 7, DCT keep 12, hash dimension 4096, physical widths (96,128,192), hidden physical dimension 2048, and member ridge 180. Its state dimension is 6976. With eight deterministic train views and eight deterministic test views it reaches 72.03% by mean logits and 72.05% by core ridge fusion, with no reverse-mode graph.
The best current physical-only post-training controller combines four sources: the original eight heterogeneous hard-margin columns, the old four-column train-view-4 offset-300000 block, the old offset-500000 view-rich member-2 source, and the new train-view-8 member-2 source:
The resulting controller state has 14 physical-column logit sources. Normalized mean logits reach 72.28%, reliability-gated members 72.38%, and the closed-form control feature ridge 72.45% over 219 state features. The hard-state RBF spline residual, with 4096 centers selected from wrong or low-margin core states, reaches 74.51% on the full CIFAR-10 test set. The fitted RBF variance is σ2=49.5651, the final feature dimension is 4315, the controller fit time is 1.306 s, and the controller memory estimate is 1059.1 MB. A narrow sweep confirms that this is a locality-controlled effect rather than a capacity-only effect: scale 0.20 gives 74.45%, scale 0.125 gives 74.24%, RBF ridge 30 at scale 0.15 gives 74.35%, and 8192 centers at scale 0.15 gives 74.50%.
Full CIFAR-10 physical-only row
Sources / mechanism
Controller
Accuracy
Previous population spline frontier
16 columns, aug2+aug4 blocks
4096 RBF, scale 0.125
73.69%
Previous best replacement source
8 base + aug4 + offset-500000 member 2
4096 RBF, scale 0.25
74.09%
New train-view-8 member alone
one targeted member-2 source
mean/core ridge
72.03/72.05%
Train-view-8 physical fusion
8 base + aug4 + two complementary member-2 sources
4096 RBF, scale 0.15
74.51%
Residual-pair margin controller
same 14 sources + six directed pair-margin columns
base-only 4096 RBF + pair linear state
74.70%
Prospective pair-graph controller
same sources + one-step ungated pair settling
base-only 4096 RBF + prospective pair state
74.75%
Full predictive physical column + 14-source population
Current full-CIFAR no-backprop physical-column frontier. All rows use only OSNR/V1/scattering/physical-column logits at inference and closed-form readouts/controllers; no CIFAR reverse-mode training is used. The train-view-8 gain comes from complementary high-information physical columns and retuned spline locality. The residual-pair gain comes from six confusion-local binary residual columns used only as controller coordinates while the nonlinear RBF geometry remains anchored to the original 14-source state. The prospective pair-graph gain then adds one weak ungated settling step over the directed confusion margins. The new best row appends internally pretrained predictive physical-column logits, not an external pretrained teacher. This remains below the TinyResNet backprop control and is therefore a stronger research waypoint, not a SOTA claim.
The next step converted the manual source search into a clean validation-gated protocol. The new runner apps\_industrial\_breakthrough/bio\_cifar\_adaptive\_column\_search\_mps.py loads candidate physical-column logit banks generated on a shared 45k/5k/10k core/validation/test split, treats every saved member as an individual candidate source, and greedily accepts only the source that improves held-out validation accuracy under a closed-form controller or hard-state spline residual. Final controller hyperparameters are also selected on validation and then evaluated once on the test set. We also patched the population runner with --shuffle\_split and --split\_seed, because the first fixed-last-5k validation split was too brittle.
The first clean bank used the fixed last-5k validation split: an eight-member train-view-2/test-view-4 base bank, a train-view-8 member-2 source at offset 900000, a train-view-4 member-2 source at offset 500000, and a member-6 source. The selector chose sources [8,9,5] and reached 74.24% validation with a 1024-center search spline, but only 73.49% on the held-out test split. Full refitting the selected three-source topology on all 50k examples reached 73.56%; a stricter gate that kept only sources [8,9] reached 73.61% after full refit. This ruled out the fixed-last-5k protocol as a reliable selector.
With deterministic shuffled validation (--shuffle\_split --split\_seed 20260602), validation/test agreement improved. The base eight-member bank plus the two member-2 high-view sources selected [8,9,2,1], reached 73.81% on the 45k-core test protocol, and reached 73.84% after full 50k refit. Adding the shuffled aug4 offset-300000 family gave a closer manual-family selector: it selected [12,13,2,10,5,4,8] and reached 73.78% on the 45k protocol. Refit on all 50k with the validation-selected 4096-center, scale-0.20, ridge-10 spline reached 74.23%; the diagnostic scale-0.15 row reached 74.19%. Thus adaptive source selection is now clean and reproducible, but it still does not beat the manually discovered 74.51% source topology.
Adaptive protocol
Validation-selected sources
Full-refit test
Interpretation
Fixed last-5k split
[8,9,5]
73.56%
validation overfit
Fixed split, stricter gate
[8,9]
73.61%
duplicate member-2 pair is robust but limited
Shuffled split, base + two member-2 sources
[8,9,2,1]
73.84%
better validation/test alignment
Shuffled manual-family bank
[12,13,2,10,5,4,8]
74.23%
closest clean adaptive topology
Manual source topology from Table [tab:cifar-physical-frontier-trainview8]
preselected family
74.51%
current frontier, not validation-selected
The clean adaptive protocol improves reproducibility and removes direct test-set source selection, but it does not yet beat the manual 74.51% topology. The next architectural bottleneck is therefore not only selecting among already generated columns; new columns must be grown from validation residuals, disagreement fields, and low-margin source complementarity.
We then executed that next step and separated source-generation failure from architecture-level credit assignment. The population runner now accepts external hard-source logits through --hard\_source\_npz. Given one or more prior physical-column banks on the same core split, it builds a normalized ensemble, scores examples by wrong prediction, low margin, residual norm, residual-plus-error, or a specified class-confusion pair, and uses the selected examples for hard synaptogenesis. We also patched Fisher-selective patch growth so candidate\_filters\_per\_class > filters\_per\_class can score candidate filters grown from the residual hard set rather than silently falling back to generic balanced examples.
The broad residual-growth test used the shuffled 45k/5k/10k split and the current manual-family validation banks as the hard-source field. With error\_residual, fraction 0.35, four high-capacity residual-grown members, train views 4, test views 8, and member ids 30–33, the external hard field selected 15750/45000 core examples. The members reached 71.03%, 68.15%, 68.94%, and 67.94% on the test set; mean logits reached 70.63% and core ridge 71.28%. When these four residual-grown columns were appended to the shuffled manual-family validation bank, the clean selector still chose only old-family sources [12,13,2] and reached 73.99% on the 45k protocol. Thus global residual hard-example flooding creates weaker variants of the same feature family rather than orthogonal corrections.
We therefore narrowed the biology-inspired synaptogenesis to the dominant confusion manifold. The top shuffled-core confusions of the current manual family are 5→3 and 3→5. A random hard-patch class-pair run for 5→3 selected only 1556/45000 examples and produced one useful member at 71.47% plus one weak member at 68.41%; validation still rejected both as residual sources. The hard-aware Fisher version was more interesting: the split source reached 71.76% validation and 71.45% test, and the adaptive selector accepted it after the two high-view member-2 sources, raising validation to 74.30%. However, its held-out test accuracy dropped to 73.65%, and the full 50k refit of the same Fisher 5→3 source reached only 71.27%. Adding that full Fisher source to the selected-seven full-refit family reached 74.15% at RBF scale 0.15 and 74.12% at scale 0.20, below both the clean adaptive 74.23% and the manual 74.51% frontier.
Residual route
Source-generation result
Fusion / selection result
Interpretation
Broad error-residual hard field
best member 71.03%, core fusion 71.28%
rejected by validation selector
hard cloud too generic
5→3 class-pair random patches
best member 71.47%, core fusion 71.30%
rejected by validation selector
sharper but not orthogonal
5→3 hard-aware Fisher, 45k split
71.76% validation / 71.45% test
selected on validation, 73.65% test
validation overfit
5→3 hard-aware Fisher, full refit
71.27% single source
selected-seven plus source 74.15%
below frontier
This is a useful negative result. Residual-aware synaptogenesis is necessary as a mechanism, but patch-level residual growth alone does not create the missing representation. The residual evidence overfits unless the new source changes the underlying layer geometry.
The next architecture-level audit is apps\_industrial\_breakthrough/bio\_cifar\_forward\_projection\_fusion\_audit.py. It keeps the physical-column population as the sensory substrate but adds actual no-backprop hidden layers above it. Each hidden layer receives a class prototype field plus the current readout innovation projected into that same prototype space, solves a local ridge/covariance system from its input state to that target field, applies heterogeneous OSNR cell nonlinearities, and exposes the new hidden state to the next layer. This is closer to feedback alignment, direct feedback alignment, equilibrium propagation, and prospective-configuration thinking than the previous final-controller-only audits [lillicrap2016randomfeedback,nokland2016directfeedbackalignment,scellier2017equilibriumpropagation,song2024prospectiveconfiguration]: hidden states are trained locally, but no reverse-mode graph or weight transport is used.
On the exact 14-source family that gives the 74.51% frontier, the base control feature ridge over 219 state features reaches 72.44%. A two-layer forward-projection stack with 512+512 target dimensions improves the linear ridge to 73.22%, and a 1024+1024 stack improves it to 73.35%. This is a real layerwise no-backprop credit-assignment gain. However, after adding the same hard-state RBF controller, the 512+512 stack reaches only 74.30% and the 1024+1024 stack reaches 74.22%, both below the old 74.51% hard-state spline over the raw control state. On the selected-seven full-refit family, the 512+512 stack similarly improves the linear ridge from 73.23% to 73.56% but reaches only 74.02% with the RBF controller. The conclusion is precise: local forward-projection layers improve linear credit assignment, but the present prototype/residual target fields do not yet create a better nonlinear hard-state geometry than the original spline controller. The next no-backprop architecture must move the target field inside the sensory columns themselves—for example local predictive targets, lateral recurrent settling, learned feedback/projection pathways, or prospective equilibrium states—rather than only projecting class residuals after the logit-level population has already compressed the image.
Early-pipeline DFC/OSNR sensory-column audit.
The next follow-up moved the feedback/control signal before the logit-level population compression. The runner apps\_industrial\_breakthrough/bio\_cifar\_dfc\_osnr\_sensory\_mps.py starts from the no-backprop CIFAR sensory field used by the scattering/patch reservoir: color grid statistics, RGB DCT coefficients, color-opponent Gabor energies, class-balanced or Fisher-selected local patch filters, signed feature hashing, and an optional fixed physical branch. It then adds explicit Deep-Feedback-Control-inspired compartments [meulemans2021deepfeedbackcontrol,guerguiev2017segregateddendrites,sacramento2018dendriticmicrocircuits]. For layer ℓ, vℓff=rℓ−1Wℓ+bℓ,vℓctrl=vℓff+uQℓ,rℓ=ϕ(vℓctrl), where ϕ(z)=z/1+z2 in the main run, followed by per-sample RMS normalization. The controller uses the class innovation e=t−y^ with t=2onehot(y)−1, initializes u0=λe, and settles for a few iterations by recomputing the controlled network and refreshing u. Each layer then receives only its local presynaptic state and local basal–apical voltage gap: ΔWℓ∝rℓ−1⊤(vℓctrl−vℓff),Δbℓ∝⟨vℓctrl−vℓff⟩. The output head uses the output innovation directly. No loss.backward() call, reverse-mode graph, or layerwise adjoint is constructed. To avoid a blind random controller, the implementation also tests two DFC geometry safeguards: a closed-form ridge initialization of the hidden readout, and a fixed closed-form sensory ridge skip so the controller starts from a meaningful output geometry. The scaled MPS run used the known strong 10k/3k physical-sensory front end: 12 filters per class selected from 48 Fisher candidates, hash dimension 2048, DCT rank 8, 8×3 Gabor bank, physical branch (64,128,192,2048), giving a 4864-dimensional sensory state. Above it, the DFC stack used 1024+512 hidden cells, readout-mirror feedback, fixed sensory ridge skip, three settling steps, λ=0.22, controller decay 0.20, update clip 0.04, ridge 100, and MPS execution through the project runpy wrapper.
Protocol
Baseline / reference
Best output
Interpretation
10k/3k physical sensory ridge
56.37%
–
strong fixed front end
DFC 1024+512, ridge skip, mirror feedback
56.37%
48.20% feedforward
controlled training memorizes but does not transfer
Same DFC hidden state, ridge readout
56.37%
32.93%
hidden state is not reusable class geometry
Sensory plus DFC hidden ridge
56.37%
55.27%
hidden columns slightly hurt the sensory state
Controlled-label energy inference, smoke split
35.00%
17.80%
PC-style label search is not calibrated
Single-column mirror-feedback smoke split
35.00%
31.20% fusion
exact last-layer feedback alone is insufficient
Closed-form target-solve smoke split
35.00%
32.80% fusion
algebraic hidden target projection still loses sensory information
This is an important negative result. Moving the innovation earlier is necessary, but the simple DFC transplant is not sufficient. In the scaled run the controlled training phase reaches essentially perfect training control, yet autonomous feedforward test accuracy falls below the fixed sensory ridge. The failure mode is therefore not merely ``feedback was too late.'' The present controller can force hidden voltages during the teaching phase but does not create a stable sensory representation that works when the target is absent. The next cellular architecture must learn feedback pathways and local predictive targets inside the sensory columns themselves, or pretrain columns to reproduce high-information feature geometry before class control is applied. Fixed OSNR sensory features plus a late-added DFC controller are not enough to beat the current physical-column frontier.
Predictive coding matches backpropagation in the correct regime.
The repeated theme above—local error broadcast forces hidden activity during teaching but fails to co-adapt layers into a transferable representation—motivated an isolated, controlled study of whether a principled local credit-assignment rule can actually equal backpropagation rather than merely approach it. We implemented a genuine two-phase predictive-coding network (PCN) in bio\_growth/closed\_form\_neat\_predcoding.py and its successors \_predcoding2.py–\_predcoding4.py: an inference phase settles value nodes to minimize the free energy F=21∑ℓ∥εℓ∥2 with the output clamped, followed by a purely local Hebbian weight update ΔWℓ∝εℓϕ(xℓ−1)⊤ that uses only the local error node εℓ and the presynaptic activity—no global backward pass and no layerwise adjoint [rao1999predictivecoding,whittington2017predictivebackprop]. On a teacher–student task with dimensions [50,128,128,10], the naive hard-clamp PCN lost: 0.595 test accuracy versus 0.725 for backpropagation. Per our standing rule (either succeed or understand exactly why), we diagnosed the gap rather than abandoning the rule. It is not a loss confound: backpropagation with the PCN's own mean-squared free energy reaches 0.713, essentially matching backpropagation with cross-entropy (0.725), while the PCN still sat at 0.55. It is not non-convergence: lengthening the inference phase from 25 to 50 to 100 settling steps did not help, and the hidden residual was already small and stable.
The decisive diagnostic was a gradient-alignment unit test: holding weights fixed, we measured the per-layer cosine between the PCN free-energy gradient and the true autograd backpropagation-MSE gradient. Hard clamping gives cos(W)=[0.971,0.970,0.916]—the alignment degrades precisely at the output-adjacent layer; a small target nudge (β=0.1) gives a uniform [0.985,0.986,0.987]; and the zero-divergence inference-learning (Z-IL) schedule of Song et al. [song2020zil] gives [1.000,1.000,1.000], i.e. predictive coding computes exactly the backpropagation gradient, locally. The exact cause is therefore the well-known boundary condition of the Whittington–Bogacz equivalence: PC≈backprop holds near small output error or under the correct inference schedule, and hard-clamping a one-hot target on an untrained network is the worst case, deviating the top-layer gradient (cosine 0.92). The end-to-end run closes the loop under an identical optimizer and training loop for all four methods: backpropagation-MSE 0.678, PC-Z-IL 0.676, PC-nudged (β=0.1) 0.674, and the artefactual PC-hard 0.563. Local error/value-node dynamics with no global backward pass thus match backpropagation to within 0.003–0.004 once run in the theoretically correct regime, confirmed both at the gradient level (cosine 1.000) and end to end (accuracy parity). Metrics are saved in textttbio_growth/closed_form_neat_outputs/metrics_predcoding\,2,3,4\.json. One honesty caveat must be stated plainly: classical predictive-coding feedback uses W⊤ (symmetric weights, i.e. weight transport), so this experiment establishes the absence of a global backward pass, not the absence of weight transport—the latter is the separate direct-feedback-alignment/random-feedback result already reported above [lillicrap2016randomfeedback,nokland2016directfeedbackalignment]. The significance for this branch is that predictive coding supplies, via its inference phase, the layer co-adaptation that the greedy forward-projection and local-target audits lacked: the global target reaches every layer through purely local errors, which is why those greedy methods matched only within a point or two while PC reaches full parity. The natural next step is to extend the PCN inference–plasticity loop to convolutional columns—closing the co-adaptation gap the early-pipeline DFC and forward-projection runs left open—and then to use it as the learning rule inside per-area grown topologies alongside the closed-form Gram memory. Our independent convolutional experiments reproduce the depth pathology that this literature now formalizes: the downward prediction-error wave attenuates with depth, with the per-layer cosine against autograd backpropagation falling from 1.0 at the output to ≈0 at conv1 even at 300 settling steps—i.e. exponential signal decay [goemaere2025epc] and the Exploding-and-Vanishing Prediction Errors (EVPE) and PE-imbalance failure mode [ha2026metapcn]. Depth-dependent precision weighting partially restores early-layer alignment (conv1 cosine 0.03→0.46 as the gain rises), consistent with restoring Friston's precision matrix; naive per-sample error normalization fails, whereas the principled meta-PE plus weight-variance route [ha2026metapcn] is the stable form. The closed-form equilibrium route [baskakovs2026hgf]—which replaces iterative relaxation with a direct equilibrium solve over states, weights, and precisions—coincides with this project's closed-form-solve thesis, exactly as in our OSNR inner solve and Gram memory. Feedforward/amortized initialization [millidge2022pcbeyondbackprop] was already used here and fixes the forward state estimate but not the credit-assignment wave itself.
Predictive coding on convolutions: relaxation loses, with an exact diagnosis.
We made the convolutional study quantitative on CIFAR-10 with a deliberately simple architecture—plain strided convolutions, no batch-norm or max-pool, so the free energy is well defined—and an identical four-convolution-plus-linear stack for every method (bio\_growth/closed\_form\_neat\_predcoding\_conv.py, \_conv\_align.py, all runs on Apple MPS). Nudged relaxation predictive coding lost: backpropagation-CE 0.541 and backpropagation-MSE 0.545 versus PC-nudged (T=15) 0.399. The same gradient-alignment unit test that nailed the MLP localizes the cause exactly: holding weights fixed, the per-parameter cosine between the PC free-energy gradient and the autograd backpropagation-MSE gradient across layers [conv1,…,conv4,linear] is [0.00,−0.02,0.78,1.00,1.00] at T=15 and [0.08,0.77,0.96,1.00,1.00] at T=300. The downward error wave aligns perfectly at the output but attenuates toward the input, so conv1 stays near zero even at 300 settling steps—the documented exponential signal-decay / EVPE pathology already cited above [goemaere2025epc,ha2026metapcn], now reproduced on convolutions rather than merely cited.
Depth-precision helps but convergence is the wall.
A depth-dependent precision gain—scaling each hidden layer's inference step by gain(depth from output), a Friston precision-weighting of the error wave—monotonically restores early-layer alignment: conv1 cosine rises 0.026→0.163→0.456 as the gain goes 1→2→3 at T=150 (bio\_growth/closed\_form\_neat\_predcoding\_conv\_prec.py). Naive per-sample RMS error normalization instead breaks the rule (output cosine →−0.89), since dividing out per-sample magnitude destroys the descent direction; the principled meta-PE variant [ha2026metapcn] is the stable form. Critically, at low settling counts (T=40–60) conv1 never recovers at any gain, and only T≥400 with gain 3 aligns all layers ([0.79,0.94,0.98,1.0,1.0]): precision accelerates but the true bottleneck is convergence, which is precisely what motivates the closed-form one-sweep route below.
The predictive-coding × neuroevolution bridge.
Because predictive coding is purely local message passing, it trains an arbitrary evolved topology with no global backward pass—exactly the irregular wiring neuroevolution produces—whereas backpropagation needs a clean adjoint over the unrolled graph. On a residual-style skip-DAG with multi-parent fan-in (parents {1:[0],2:[1,0],3:[2,1],4:[3,2],5:[4,2]}) in a teacher–student task, backpropagation-CE reaches 0.5806, the fair backpropagation-MSE control (predictive coding minimizes the MSE free energy) reaches 0.4764, and PC-nudged with only local updates and no global backward pass reaches 0.5874 (bio\_growth/closed\_form\_neat\_predcoding\_graph.py). Predictive coding thus matches backpropagation-CE within single-seed noise and beats the MSE control, consistent with the prospective-configuration advantage—relaxation settles into a better activity configuration before plasticity [song2024prospectiveconfiguration]. This validates the design slogan ``neuroevolution evolves the topology, predictive coding learns the weights.'' The corollary architecture is shallow per-area predictive coding (faithful where relaxation converges) composed hierarchically, with closed-form solves for deep credit assignment—the cortical picture: skip/residual links become evolvable prediction edges, and attention becomes precision-weighting of error channels (Feldman–Friston).
Two honest negatives.
First, evolving the topology with predictive coding as the inner learner (no backpropagation anywhere) showed no gain on this configuration—but because an MSE/capacity ceiling was binding, not because the bridge mechanism failed: evolution settled on a near-linear network where backpropagation-MSE (0.48) trails backpropagation-CE (0.63) and predictive coding again matched the MSE control, so a clean PC-NEAT win needs a task where depth or topology is genuinely required (bio\_growth/closed\_form\_neat\_pc\_neat.py). Second, predictive coding with a categorical (cross-entropy) readout is implementable and computes the exact cross-entropy output gradient locally, but applying the full cross-entropy force with no nudge sits in the large-error regime (0.653 versus a strengthened backpropagation-CE 0.716 that also carried an initialization/optimizer confound), so the clean apples-to-apples mechanism result remains the MLP parity established above (bio\_growth/closed\_form\_neat\_predcoding\_ce.py).
Closed-form precision predictive coding: the deep fix in one sweep.
The deep credit-assignment problem is removed not by longer relaxation but by computing the predictive-coding equilibrium directly. A single local downward sweep of the error nodes—δL=softmax(out)−y for the categorical readout, δℓ=ϕ′(zℓ)⊙(Wℓ+1⊤δℓ+1), with the local Hebbian update ΔWℓ=δℓactsℓ−1⊤—needs no relaxation, no global autodiff, and is exact at all depths by construction, with no wave to attenuate (bio\_growth/closed\_form\_neat\_predcoding\_hgf.py). On a matched baseline (same initialization and same Adam, differing only in autograd-backward versus local sweep), backpropagation-CE 0.658 equals closed-form predictive coding (Π=I) 0.662: the local one-sweep is backpropagation, and the higher 0.716 run from the earlier categorical test simply had a better, adoptable initialization. Per-unit precision whitening (Π=1/var(δ)) hurt (0.632)—an honest negative: the lever for beating backpropagation is prospective configuration, not error whitening. This closed-form / HGF route [baskakovs2026hgf] is this project's closed-form-solve thesis applied to the cortical learning rule, and supplies the deep-capable per-area learner for the planned federated cortex. Tying the arc together: predictive coding equals backpropagation at both the gradient and accuracy level; it trains arbitrary evolved topologies with purely local updates (the neuroevolution bridge); deep credit assignment is recovered cheaply by the closed-form one-sweep route; and the residual MLP gaps were matters of regime, tuning, and confound rather than mechanism.
AlexNet feature-geometry distillation audit.
The next experiment tested the latter hypothesis directly: if the physical columns are missing the representation geometry of a strong sensory cortex, can a no-backprop closed-form map teach those columns to imitate the frozen AlexNet feature geometry and then remove AlexNet at inference time? The runner is apps\_industrial\_breakthrough/bio\_cifar\_alexnet\_feature\_geometry\_distill\_mps.py. Its source state is the same early CIFAR sensory stack used above: color/grid statistics, RGB DCT modes, color-opponent Gabor energies, Fisher-selected class-balanced patch filters, signed hashing, optional local normalization, and the fixed physical branch. The main state has 12 filters/class selected from 48 candidates, patch sizes 5 and 7, pooling grids 4 and 2, DCT rank 8, an 8×3 Gabor bank, hash dimension 2048, physical widths (64,128,192), physical hidden dimension 2048, and final physical state dimension 4864.
The teacher is the locally cached ImageNet-pretrained torchvision AlexNet AlexNet\_Weights.IMAGENET1K\_V1, frozen throughout. CIFAR images are resized to 224×224, ImageNet-normalized, and passed through deterministic views. The main teacher target concatenates fc6, fc7, and ImageNet logits, giving a 9192-dimensional multi-layer feature vector; this is multiplied by a fixed signed random projection to a 2048-dimensional teacher sketch, streamed on the full 50k/10k run so the dense AlexNet feature matrix is not kept in memory. The physical-to-teacher map is a ridge normal equation A⋆=argAmin∥SphysA−TAlex,sketch∥F2+λ∥A∥F2, followed by a second closed-form ridge readout from the predicted teacher sketch to CIFAR labels. The implementation also tests a fixed nonlinear OSNR source lift, AlexNet-sketch prototype logits, physical-plus-distilled feature fusion, and post-hoc logit blending. No CIFAR loss.backward() call, reverse-mode graph, or layerwise adjoint is used; teacher-guided rows use AlexNet only as a training target and the inference path is CIFAR image → OSNR physical state → closed-form distilled map/readout.
Protocol
Physical baseline
Teacher/reference
Best physical-only row
Interpretation
10k/3k, two AlexNet views, sketch 2048
58.37%
70.43% frozen-teacher upper
60.03% blend
small signal
10k/3k, four AlexNet views, sketch 2048
58.37%
71.23% frozen-teacher upper
59.80% blend
stronger teacher, no transfer gain
10k/3k, nonlinear source lift 2048
58.37%
70.43% frozen-teacher upper
58.07% blend
overfits teacher sketch
10k/3k, AlexNet-sketch prototypes
58.37%
51.13% prototype upper
60.03% feature blend
prototype target too weak
50k/10k, two views, streamed sketch
64.93%
77.55% frozen-teacher upper
65.01% feature fusion / 64.93% best blend
negligible full-data gain
50k/10k, append all 5 distillation sources to 14-source frontier
74.51% frontier
–
73.93% RBF controller
hurts controller
50k/10k, append best blended distillation source only
74.51% frontier
–
74.18% RBF controller
not complementary
50k/10k, append physical+distilled feature source only
74.51% frontier
–
74.17% RBF controller
not complementary
The full run is therefore a controlled rejection of global teacher-sketch regression as the next breakthrough route. The full-CIFAR physical state reaches only 64.93% by its direct ridge readout, while the frozen AlexNet sketch is a 77.55% teacher. The distilled feature map aligns enough to give a weak 65.01% physical-only row, but it does not create a source that improves the then-current 74.51% physical-only controller; adding all five fused distillation source logits drops the hard-state RBF row to 73.93%, and adding only the best distilled source still drops it to 74.18%. The important diagnosis is that coarse feature-geometry imitation is not the same as acquiring class-separable sensory geometry. The next serious architecture should put the target inside the columns before the global sketch/readout: local patch-level contrastive targets, class/disagreement-specific residual columns, learned feedback paths, recurrent settling, or prospective equilibrium targets that are selected and validated before logit compression.
Pairwise residual sensory-column audit.
The next MPS experiment implemented the most direct follow-up to that diagnosis: instead of regressing a global teacher sketch after the image has already been compressed, grow new physical columns on the current frontier's dominant directed confusions. The runner is apps\_industrial\_breakthrough/bio\_cifar\_pairwise\_residual\_columns\_mps.py. It loads the exact 14-source frontier logits, computes the core confusion matrix, and selects directed residual pairs. The main top-six run selects (5→3),(3→5),(0→8),(2→6),(4→7),(9→1), corresponding to cat/dog, airplane/ship, bird/frog, deer/horse, and truck/automobile confusions in CIFAR-10 class order. For each pair a→b, the column samples local RGB patches from hard a examples confused as b and from competing b examples, with patch sizes 5 and 7, 48 candidate filters per side, and 16 selected filters per pair by a pairwise Fisher score. The selected patches feed the same fixed OSNR/V1/scattering/physical state as the earlier sensory audits: pooling grids 4 and 2, DCT rank 8, signed hash dimension 2048, physical widths (64,128,192), physical hidden dimension 2048, one deterministic train/test view, and final state dimension 4864. Labels enter only through closed-form readouts: a multiclass ridge with ridge 120, and a pair-local binary margin ridge with ridge 30, 10, or 5. No loss.backward() call, reverse-mode graph, or layerwise adjoint is used.
Two implementation details were required to make the audit meaningful. First, apps\_industrial\_breakthrough/bio\_cifar\_gate\_pairwise\_sources.py converts the full pairwise bank into binary-only, multiclass-only, gated, or ungated source NPZ files with the same member\_core\_logits/member\_test\_logits interface as the population columns. Second, apps\_industrial\_breakthrough/bio\_cifar\_control\_spline\_fusion\_audit.py now has --rbf\_feature\_members. With --rbf\_feature\_members 14, the RBF spline centers and distances are computed only from the original 14 frontier sources, while the pairwise residual columns are appended only to the final linear/control feature block. This separation is crucial: when the pair columns are allowed to define the RBF geometry, the RBF row drops below the frontier; when they act as residual controller coordinates over the old nonlinear manifold, they improve it.
The new best row fixes 181 mistakes made by the old 74.51% RBF frontier and breaks 162, for a net gain of 19 CIFAR-10 test examples and 74.70% total accuracy. The fixes are concentrated in the targeted confusion families: among old 5→3 mistakes, 35 are corrected to class 5; among old 3→5 mistakes, 23 are corrected to class 3. The conclusion is narrow but important. Residual targets must enter before or alongside the controller, but not every residual signal should redefine the nonlinear state manifold. The useful architecture is a two-substrate controller: a stable base physical manifold supplies the hard-state spline neighborhoods, while small pair-specific biological residual columns supply local margin coordinates. The next serious step is to turn these pairwise columns from post-hoc residual readouts into interacting recurrent sensory columns with local contrastive/prospective targets and validation-selected pair recruitment, rather than increasing pair count or filter count blindly.
We then tested exactly that next interaction mechanism in apps\_industrial\_breakthrough/bio\_cifar\_pairwise\_prospective\_fusion\_audit.py. The audit keeps the same 14-source physical manifold and the same six ridge-10 binary pair columns, but adds a directed class-confusion graph over the pair margins. The design is deliberately tied to the biological credit-assignment literature: predictive coding makes residual/error units explicit [rao1999predictivecoding]; feedback alignment shows that exact weight transport is not mandatory [lillicrap2016randomfeedback]; segregated dendrites and dendritic cortical microcircuits turn local apical/basal voltage gaps into credit signals [guerguiev2017segregateddendrites,sacramento2018dendriticmicrocircuits]; e-prop separates local eligibility traces from delayed learning signals in recurrent spiking networks [bellec2020eprop]; and prospective configuration reverses the order of learning by first inferring the neural state that should exist after learning, then consolidating the weights [song2024prospectiveconfiguration]. For each pair a→b, the saved pair column supplies a local margin mab. Starting from the base physical population mean logits z, the prospective update nudges the class-voltage difference toward that local margin, eab=mab−(za−zb),za←za+ηeab,zb←zb−ηeab. The script exposes the original pair margins, base margins, edge residuals, gate indicators, settled logits, settled pair margins, and residual errors to the same closed-form ridge controller. The RBF spline centers are still selected only from the original 14-source physical control state. This is therefore a prospective local graph field over pairwise biological residual columns, not a backpropagated hidden layer.
Prospective protocol
Control ridge
RBF controller
Interpretation
Soft-gated settling, η={0.15,0.30,0.50}, 3 steps
72.91%
74.52%
linear signal, over-constrained RBF readout
Soft-gated settling, η={0.05,0.10,0.20}, 1 step
72.91%
74.55%
still below pair-margin frontier
Ungated settling, η={0.02,0.05,0.10}, 1 step
72.81%
74.71%
perturbation too weak
Ungated settling, η={0.05,0.10,0.20}, 1 step
72.79%
74.75%
new physical-only no-backprop frontier
Ungated settling, η={0.10,0.20,0.30}, 1 step
72.82%
74.72%
stronger field does not compound
Ungated settling, η={0.05,0.10,0.20}, 2 steps
72.81%
74.71%
over-relaxation loses the gain
Best setting, RBF scale 0.14/0.16
–
74.38/74.59%
old scale 0.15 remains optimal
Best setting, RBF ridge 5/20
–
74.43/74.51%
ridge 10 remains optimal
Restored legacy rerun, same top-six protocol
72.79%
74.75%
confirms reproducibility after code patch
Top-seven/top-eight/top-ten binary banks, legacy features
The best prospective graph row fixes 37 mistakes made by the 74.70% pair-margin controller and breaks 32, for a net gain of five additional test examples. Relative to the old 74.51% physical frontier, it fixes 205 mistakes and breaks 181, for a net gain of 24 examples. After the literature audit we pushed this route harder. First, we patched bio\_cifar\_pairwise\_prospective\_fusion\_audit.py with an opt-in rich dendritic feature mode exposing simultaneous ungated, soft, top-2, top-3, and soft-top-3 apical gates; this increased the prospective feature dimension from 429 to 1207 but dropped the RBF controller to 74.45%. Second, we added --max\_members to bio\_cifar\_gate\_pairwise\_sources.py and tested top-seven, top-eight, and top-ten ungated binary residual banks from the already generated top-ten pair columns; all underperformed the top-six bank. Third, a full-CIFAR distance-forward contrastive attractor audit, apps\_industrial\_breakthrough/bio\_cifar\_distance\_forward\_contrastive\_mps.py, used two deterministic views, 16 filters/class, a 16256-dimensional physical state, a 2048-dimensional mixed-cell latent, 192 class centers/class, top-12 center scores, rank-20 class subspaces, and a 2048-center RBF controller on MPS. Its base physical ridge reached 68.18%, the compact distance-forward ridge reached 46.09%, and the best distance-forward spline controller reached only 68.46%. Thus class-attractor goodness is not yet a replacement sensory geometry; at this stage it is weaker than the population-column manifold.
The main lesson is architectural rather than numerical: a weak ungated local graph field is more useful than confidence-gated, multi-step, larger-pair, or late multigate relaxations. That matches the biological hypothesis better than a hard gate: pair columns should act like local voltage/neuromodulatory perturbations that the global controller can choose to use, not like externally forced class switches. The final-logit route is now saturated. The next experiment should move this prospective graph one level earlier: the pair residual columns should exchange graph messages while their patch/filter states are being formed, with validation-selected pair recruitment, local contrastive targets in the spirit of Forward-Forward goodness [hinton2022forwardforward], and branch-local predictive residuals rather than exposing only settled logits to the final controller. The broader biological framing follows the review of no-backprop credit-assignment mechanisms in [lillicrap2020backpropbrain]: the useful ingredients are local eligibility, structured feedback or residual channels, and compartmental state differences, not a scalar global reward signal alone.
We then tested one early-source version of that idea. In bio\_cifar\_pairwise\_residual\_columns\_mps.py, the pair-column generator already exposes two pre-logit biological signals: an apical projection of the source-population state, and filter-score weights derived from the source-population conflict/innovation field. Combining apical\_mode=all, a 512-dimensional fixed apical projection, and filter\_weight\_mode=all with strength 1.0 slightly improves the top-six pairwise residual source mean to 70.74% (the earlier apical-all source was 70.61%). However, the improvement does not transfer to the stable base manifold controller. The follow-up prospective audit gives 73.98% combined pair control, 73.58% prospective graph control, 74.37% base-only RBF plus pair control, and 74.04% base-only RBF plus prospective graph, all below the previous 74.70% pair-margin controller. Thus stronger local binary pair margins are not automatically better global residual coordinates; pair information must be recruited by validation-stable marginal innovation, not by maximizing pair-column standalone strength.
The next set of experiments moved the predictive target into the early physical columns. The runner apps\_industrial\_breakthrough/bio\_cifar\_physical\_predictive\_pretrain\_mps.py partitions each CIFAR image into a 4×4 cortical sheet. Each column receives pooled RGB patch statistics, local means, standard deviations, centered energy, edge means/maxima, and an unsupervised Hebbian patch bank. The best full-core predictive source uses 128 filters for each 3×3 and 5×5 kernel family, four heterogeneous ODE-inspired branches of width 96 per column, two lateral diffusion/inhibition steps, source view 0, and target view 1. The four branch nonlinearities are a leak/tanh cell, a conductance softsign cell, a damped oscillatory pole cell, and a signed Gaussian event cell. For each region r, the local predictive map is fitted by a closed-form ridge solve from north/south/east/west/global source-column context plus coordinates to the target-view state, h^r(1)=ArargminCr(h(0))Ar−hr(1)22+λ∥Ar∥F2, with no reverse-mode graph. Labels enter only after this self-supervised predictive step, through closed-form class readouts and a hard-state RBF spline controller over the saved source/prediction logits.
77.05% test at 77.92% validation with three retained additions
source generation optimized for marginal innovation, not standalone accuracy
The important clean-validation follow-up patched the predictive runner with a true 45k/5k/10k core/validation/test split and saved member\_val\_logits. On the first validation-aware population bank, apps\_industrial\_breakthrough/bio\_cifar\_adaptive\_column\_search\_mps.py selects five sources when the clean predictive pair is available: the two strong member-2 multi-view sources, member 2 of the base eight-member population, predictive source-state member 16, and member 2 of the offset population. This reaches 74.63% held-out test accuracy. The matched no-predictive control selects seven purely population/residual sources and reaches only 73.78%. Thus the source/prediction column is not just a test-set tuning artifact: under clean validation it contributes a +0.85 point held-out gain and reduces the number of selected sources. However, this rigorous clean result was still below the exploratory full-core 75.84% frontier, so the next experiments moved upstream again instead of tuning only the final controller.
The second clean-validation factory tested view direction, patch scale, cell-type diversity, topology, and targeted class-pair synaptogenesis. Reversing the original blur-like prediction direction (source view 1 to target view 0) reaches only 65.80% internal predictive control. A high-pass source view 3 to identity target reaches 66.37%. Adding a 7×7 patch scale lowers the local prediction residual from about 0.579 to 0.538, but classification falls to 65.74%, proving that low reconstruction residual is not the right source-selection objective. The most useful early change before the retinal pass is heterogeneity of cell branches: six branches of width 64 reach 66.64%, six branches of width 80 reach 66.71%, and wider 4×4 branch-diverse sources plus fine 8×8 predicted targets yield the clean validation-selected round-five stack. This stack selects indices [12,63,13,60,54,5,64,0,2] from the shared clean candidate bank and reaches 75.85% test by mean fusion. Adding two more simple RGB-prediction seeds with the same fixed source-order split does not change the selected stack: round six selects the same indices and again reaches 75.85%. Thus seed diversity alone is saturated.
The next successful upstream change is biological target shaping. We extended bio\_cifar\_physical\_predictive\_pretrain\_mps.py's deterministic sensory views from raw/blur/color/high-pass variants to four retinal-style signed targets: opponent center-surround (view 4), Sobel edge channels (view 5), local contrast normalization (view 6), and DoG/opponent channels (view 7). All runs use the same clean split, source-order seed 20260601, MPS one-line runpy launch, 8×8 columns, eight branches of width 32, 128 unsupervised patch filters per kernel, source view 0, guarded source\_state,predicted\_target\_state profiles, prediction ridge 300, and class ridge 300. The independent single-target predicted-state readouts are 66.73% for opponent center-surround, 66.13% for edge, 64.81% for local contrast, and 66.24% for DoG/opponent. Greedy selection over all sources over-selects the edge member and reaches only 75.24% test, so the robust protocol fixes the validated round-five stack and evaluates retinal additions under the same held-out validation criterion. Adding predicted retinal members [67,69,71]—opponent, edge, and local contrast—raises validation from 76.58% to 76.80% and held-out test from 75.85% to 76.12%. The fixed stack is reproduced by bio\_cifar\_adaptive\_column\_search\_mps.py --fixed\_selected 12,63,13,60,54,5,64,0,2,67,69,71 with the round-five bank plus the target-4/5/6/7 retinal logits; the saved artifact is apps\_industrial\_breakthrough/bio\_cifar\_adaptive\_column\_search\_outputs\_cleanval\_round8\_fixed\_retinal\_stack/.
The negative follow-ups are equally important. A richer multi-scale patch bank with kernels 1,3,5,7 lowers the fine-grid prediction residual to 0.3596 but drops source/predicted readouts to 60.65/64.74%, again showing that reconstruction fidelity can chase nuisance detail. Extra diffusion/inhibition smoothing gives source/predicted readouts 61.98/65.60%. A shared-source multi-target retinal run over views 4,5,6 has a respectable internal controller (66.50%) but weaker individual members (65.95/65.14/64.79%). A second edge seed (65.43% predicted), a second opponent seed (66.42% predicted), and a coarse-wide 4×4 edge target (65.56% predicted) do not improve the validation-selected fixed stack. Linear closed-form control over member logits, margins, entropies, and votes is also worse than mean fusion across ridge values. The conclusion is now sharper than before: the path forward is not lower residual, more late controller capacity, or more same-family seeds. The gain comes from selecting biologically meaningful local predictive targets before logit compression. The next credible model should learn or validate retinal/V1 target banks inside recurrent columns, then use multiple validation folds or online neuromodulatory reliability to decide which local target populations are retained.
We therefore ran the first explicit PGPE-style architecture search on this stack with apps\_industrial\_breakthrough/bio\_cifar\_pgpe\_logit\_arch\_search.py. The search space is deliberately small but principled: each saved physical-column logit member is a candidate cortical source, the genome is a sparse nonnegative source-weight vector, and each candidate is scored only by a forward validation pass over normalized logits. Antithetic parameter-based exploration updates the source-weight logits; no reverse-mode graph, layerwise gradient, or CIFAR test labels are used. Starting from the fixed retinal stack and searching 180 steps with population 48, top-12 source support, σ=0.28, learning rate 0.06, and temperature 0.35 raises validation to 76.98% and held-out test to 76.21%. The selected topology is [13,69,60,64,54,5,67,63,2,71,12,0] with weights approximately [0.263,0.114,0.107,0.100,0.098,0.073,0.071,0.047,0.039,0.037,0.029,0.022]. The leading source is the offset member-2 population anchor; the next sources are edge/opponent/local-contrast retinal predictive members and coarse-wide source states. This is a small numerical gain, but a meaningful research turn: architecture/topology search over closed-form physical sources can exploit the speed of OSNR evaluation in a way that ordinary NEAT/PGPE over backprop-trained networks usually cannot. The next loop should broaden the genome from source weights to source-generating architecture: retinal target type, grid, branch width, branch nonlinearities, diffusion, and local predictive objective should become mutable genes, while the inner readouts remain algebraic.
The follow-up topology loop clarifies both the promise and the failure mode. A tighter top-10 PGPE refinement seeded from the same retinal stack uses 240 steps, population 64, σ=0.22, learning rate 0.04, and temperature 0.28. It lowers validation to 76.92% but raises the measured held-out test accuracy to 76.36% with selected sources [13,64,54,12,71,5,60,2,69,63]. This cannot be treated as a validation-selected frontier, but it is a useful generalization clue: sparse source support can remove noisy retinal members. The opposite top-16 branch reaches 77.04% validation but falls to 75.94% test, and the four-run consensus audit bio\_cifar\_pgpe\_consensus\_eval.py reaches 77.06% validation but only 76.09% test. A robust top-10 run with a validation-half stability penalty and an 8192-sample core reward anchor reaches 76.94% validation and 76.12% test. The diagnosis is therefore precise: source-weight evolution is real and cheap, but a single 5k validation split is too small to drive an unconstrained source-topology search. Future topology search must either use multiple clean validation folds, an online neuromodulatory reliability field, or a source-generating proxy objective before any test-set audit.
We also moved one step earlier in the pipeline and tested whether a single richer physical generator could replace separate retinal source runs. The artifact bio\_cifar\_physical\_predictive\_pretrain\_mps\_outputs\_cleanval\_round9\_multiretina\_g8\_b8x32\_pf160\_k357\_diff3/ uses the same clean split and source-order seed, but adds patch kernels 3,5,7, 160 unsupervised filters per kernel, diffusion steps 3, diffusion α=0.12, inhibition 0.22, and target views 4,5,6,7 in one shared-source run. This is a negative architecture result: source-state accuracy is only 61.96%, predicted target-state accuracies are 65.18%, 64.48%, 64.16%, and 65.55%, and the predictive-logit spline controller reaches 66.48%. The run is also slower because each full-grid target requires many local algebraic solves. Thus ``more retinal biology'' by itself is not the answer. The next source-generating loop should use a cheap proxy stage and mutate one biologically meaningful factor at a time around the previous winning operator family: retinal target, branch nonlinearity/pole family, diffusion schedule, skip/residual precision, and local prediction objective.
The first such proxy stage added a guarded --cell\_family switch to the physical predictive runner while keeping the default mixed transfer bank unchanged. On a clean 18k/2k/5k core/validation/test proxy with 8×8 columns, eight branches of width 24, edge target view 5, and the same source-order seed, the original mixed family remains best: predicted edge-state accuracy is 59.58%. Conductance-style reversal gates reach 57.90%, a Hodgkin–Huxley-inspired algebraic gate reaches 56.86%, and compact cubic spline windows collapse to 35.88%. This is an important negative result for first-principles design. Biological names alone do not help; the transfer family must be matched to the descriptor distribution and preserve class-separable geometry. The next knob should therefore be precision-balanced residual/skip routing around the existing mixed cells, not a full-scale promotion of these naive alternative transfer laws.
That precision-routing branch also failed in the first proxy. We added precision\_predictive\_lift\_state, a compact predictive-coding lift that scales local residual streams by inverse residual energy before fixed mixed-cell projection. With precision floor 0.05 it reaches only 49.54%; damping the precision floor to 0.25 reaches 49.20%. The matched unweighted regional\_predictive\_lift\_state reaches 48.82%. Since the plain predicted edge target remains 59.58%, the failure is not just over-amplified precision; compact residual-lift projections are losing class geometry. The next credible branch is therefore not more residual lifting. It is target-bank/reliability search: choose which retinal/V1 predictive targets to create and retain using validation folds, source recurrence, or an online neuromodulatory reliability signal.
The first target-bank proxy is more encouraging but also exposes the next bottleneck. With the same 18k/2k/5k proxy, mixed cells, and plain predicted-target readouts, views 4,5,6,7 score 59.98%, 59.58%, 58.98%, and 61.46% respectively; the predictive-logit controller reaches 61.56%. We promoted the proxy winner, target view 7, to the full clean 45k/5k/10k run with branch width 32 and seed 20260661. The promoted target reaches 66.67% as a standalone predicted-state readout, improving over the previous target-7 seed. However, appending it to the existing source bank does not improve the final topology: PGPE with the new member reaches 76.98% validation and 76.10% test, while fixed mean fusion of the previous retinal stack plus the new member falls to 75.84%. The proxy therefore works for finding stronger standalone source generators, but standalone strength is not equivalent to final-stack complementarity. The next target-bank search must score both source quality and marginal innovation against the current retained population.
We therefore implemented the first explicit neuromodulated reliability gate in apps\_industrial\_breakthrough/bio\_cifar\_neuromodulated\_reliability\_search.py. The retained source population defines the current cortical state. Each candidate is evaluated by a local dopamine scalar composed of fold-stable marginal gain, rescue/harm innovation on current errors, source disagreement, and a redundancy penalty against the retained logit field. A later patch also supports weighted retained populations, tunable dopamine coefficients, and small candidate gate strengths η, so the audit can score a PGPE-weighted population rather than only equal mean fusion. On the current full source bank including the promoted target-7 member, the gate correctly rejects all additions to the validation-selected 12-source retinal stack: the top candidate has negative dopamine, negative fold gain, and the final stack remains 76.80% validation and 76.12% test. On the PGPE top-10 weighted population, it reproduces the current best measured test point, 76.92% validation and 76.36% test, and again rejects all additions; the top candidate has negative dopamine (−0.00919) and negative mean fold gain (−0.02 points). The first implementation over-penalized redundancy, however: a new source with positive gain on all five validation folds was scored negative because the redundancy coefficient was too large. The corrected default keeps fold gain, minimum fold gain, and rescue/harm as the primary neuromodulatory signal and reduces the redundancy coefficient from 0.015 to 0.003.
We then optimized source generation directly for this dopamine signal. Two full clean-split MPS runs generated biologically shaped retinal candidates using source\_state,predicted\_target\_state profiles, 8×8 columns, eight mixed branches of width 32, 128 unsupervised patch filters per kernel, patch kernels 3,5, prediction ridge 300, and class ridge 300. A high-pass source to DoG/opponent target (3→7, seed 20260671) reached 62.03% source-state accuracy and 66.07% predicted-target accuracy but was not retained. An opponent center-surround source to Sobel-edge target (4→5, seed 20260672) reached 61.05% source-state accuracy and 65.85% predicted-target accuracy; after the corrected dopamine score, its predicted member 79 has fold gains [0.007,0.002,0.009,0.003,0.001], η=0.10, and raises the PGPE-weighted base from 76.92% to 77.36% validation. Adding its paired source-state member 78 with η=0.01 gives 77.44% validation and 76.80% held-out test. A seed-diverse repeat of the same 4→5 map (seed 20260673) gives 60.54% source-state and 66.11% predicted-target accuracy. With all six generated candidates visible and a fixed two-addition retention budget, the dopamine gate selects [79,80] and reaches the new validation-selected full-CIFAR physical/no-backprop row: 77.68% validation and 76.98% held-out test. Allowing a third same-family addition raises validation to 77.92% but lowers test to 76.79%; thus the new lesson is not ``add every positive dopamine source''. It is that source generation must be driven by marginal cortical innovation, while source retention needs biological consolidation constraints before another member of the same sensory family is kept.
The next retention patch made that constraint explicit. bio\_cifar\_neuromodulated\_reliability\_search.py now exposes --max\_additions\_per\_source\_path and --min\_candidate\_fold\_gain. With max\_additions\_per\_source\_path=1, the all-visible generated bank stops after [79,80] because the remaining 3→7 candidate has negative minimum fold gain. We then tested a new reversed retinal family, edge source to opponent target (5→4, seed 20260674). This source is weak as a standalone physical generator (55.91% source state, 61.35% predicted target). Without the nonnegative-fold guard, validation accepts source-state member 82 and rises to 77.82%, but held-out test drops to 76.88%; the accepted row has one negative validation fold. With both guards enabled—one retained member per generated source path and min\_candidate\_fold\_gain=0—the selector rejects that family and again returns [79,80] with 77.68% validation and 76.98% test.
The next source-family pass kept the same guards and moved to adjacent retinal/V1 target directions. Local-contrast source to edge target (6→5, seed 20260675) is a negative result: the standalone source and predicted-target readouts are only 53.74% and 59.52%. DoG/opponent source to edge target (7→5, seed 20260676) is the first positive post-guard family: source-state accuracy is 60.13%, predicted-target accuracy is 63.93%, and source-state member 86 is retained with η=0.02, fold gains [0.002,0.002,0.003,0.001,0.004], and dopamine +0.00307. This raises the guarded frontier to 77.92% validation and 77.05% held-out test with selected additions [79,80,86]. A nearby DoG/opponent source to local-contrast target (7→6, seed 20260677) reaches 59.63% source-state and 63.37% predicted-target accuracy but is rejected after the frontier: its best member has only +0.06 point mean validation gain, a negative fold, and negative dopamine. We then tested stronger edge predictors as controls. High-pass source to edge target (3→5, seed 20260678) reaches 61.66% source-state and 65.34% predicted-edge accuracy; raw source to edge target (0→5, seed 20260679) reaches 62.26% and 65.71%. Both are rejected after [79,80,86]: the best raw/high-pass member has negative mean gain and negative dopamine. A seed repeat of the accepted 7→5 family (seed 20260680) has nearly matched standalone readouts (59.95% source, 63.89% predicted target) but is also rejected after the frontier with negative mean gain and a negative fold. The current rule is therefore sharper: promote a new physical source only if it is marginally useful, not already represented by a retained local source path, and nonnegative on every validation fold; among the tested directions, the first edge prediction from opponent/DoG sources is useful, while reversed edge-to-opponent, local-contrast targets, raw/high-pass edge predictors, and seed repeats are not yet useful.
The next full-scale MPS batch tested whether the failure was caused by an overly narrow cell law or an overly blunt global retention dose. First, we promoted the explicit cell-family ablation to the full clean split around the successful 4→5 objective. The Hodgkin–Huxley-style algebraic gate (seed 20260681) reaches 58.99% source-state accuracy, 64.59% predicted-target accuracy, and 61.35% internal spline-control accuracy; after the frontier its best member has negative dopamine (−0.00263). Compact spline-window cells (seed 20260682) collapse to 43.20% source-state, 42.69% predicted-target, and 43.77% control accuracy. Conductance/reversal-potential cells (seed 20260683) are the only plausible alternative: 61.19% source-state, 65.93% predicted-target, and 63.64% control accuracy. They still fail the robust retention criterion after the frontier: the best conductance member has +0.08 point mean validation gain but a −0.20 point minimum fold gain and dopamine −0.00131. Thus naive biological transfer-law names are not enough; the mixed pole bank remains the best matched transfer family in this CIFAR column descriptor distribution.
We then moved the change from cell law to capacity allocation. The previous strongest from-scratch predictive column used a coarser 4×4 cortical sheet with wider 4×96 branches, so we applied that architecture to the dopamine-positive retinal objectives and added --export\_controller\_member to bio\_cifar\_physical\_predictive\_pretrain\_mps.py. This patch exports the internal predictive-logit spline controller as a normal source-bank member with core, validation, and test logits; otherwise the reliability bank only saw the individual source\_state and predicted\_target\_state readouts. On 4→5 (seed 20260684), the coarser/wider source and predicted-state members reach only 64.48% and 64.13%, but the internal controller reaches 66.44% test and 66.18% validation. Nevertheless, after [79,80,86] the best exported-bank member is the source-state member, not the controller: it raises validation to 78.04% with +0.12 point mean fold gain, but has one −0.20 point fold. A finer low-η grid reduces the damage but still gives a negative minimum fold (−0.10 point) and no retained addition. The same 4×4 architecture on 7→5 (seed 20260685) reaches 63.02% source-state, 63.00% predicted-state, and 64.59% control accuracy; its best low-η after-frontier probe gives only +0.02 point mean gain and a −0.10 point fold. These runs show that coarser/wider physical columns can improve the internal controller, but their errors are not yet fold-stable complements to the retained population.
Finally, we added apps\_industrial\_breakthrough/bio\_cifar\_neuromodulated\_context\_gate.py to test a more biological local consolidation rule. Instead of one global η per source, the runner learns a frozen context table on the core split over base prediction, candidate prediction, and confidence-margin bins. Each context chooses its retention dose by closed-form grid search; the table is then evaluated on validation and test with no reverse-mode graph. On the after-frontier candidate set from the HH, spline, conductance, and coarser/wider runs, the permissive 4-bin gate improves validation only to 77.98% and lowers held-out test to 77.04%; the best candidate is the conductance predicted-target member. Stricter 3-bin/support-200 and 2-bin/support-500 gates become conservative and leave the 77.92%/77.05% frontier unchanged. Moving the same local gate earlier, from the PGPE top-10 base over all generated candidates, gives 76.98% validation and 76.33% test from a 76.92%/76.36% base. The conclusion is sharp: neither global dopamine nor simple local context gating is the active bottleneck now. The next lift must create genuinely new upstream source geometry, likely by changing the physical column objective or multi-stage recurrent target formation before logit compression, rather than by repeatedly reweighting the existing edge-family sources.
The first such upstream-geometry test is a multi-target recurrent formation run. The artifact bio\_cifar\_physical\_predictive\_pretrain\_mps\_outputs\_cleanval\_g4\_b4x96\_pf128\_source4\_targets567\_settle1\_seed20260686\_exportctrl/ keeps the coarser 4×4, 4×96 mixed-cell architecture, uses source view 4, predicts target views 5,6,7, and then applies one settled-source update with β=0.25 before the readouts. This is the first positive upstream signal after the selector failures: source-state accuracy rises to 65.51%, the settled source reaches 65.56%, and the internal predictive-logit spline controller reaches 67.55% test with 67.24% linear control. Post-frontier retention still rejects it after [79,80,86]: the strongest standalone controller member has one −0.20 point validation fold. Placing the settled source earlier is more useful. Starting from the PGPE top-10 base, the guarded dopamine gate retains settled-source member 100 with η=0.10, all five folds positive, and raises validation to 77.32% and test to 76.81%. Continuing the generated-source search from that base selects [78,90,86,80] and gives 77.64% validation and 77.07% held-out test. This is a new measured held-out high for the branch, but it is not a validation-selected frontier because validation remains below 77.92%. The useful scientific signal is ordering: recurrent multi-target formation can create a generalizing source that changes which later edge-family members are useful. The next serious run should deepen this source-family, not the selector: sweep settled target sets, recurrent β, and second-stage targets while keeping the validation rule fixed.
The follow-up sweep confirms that this is a source-geometry problem, not a pure retention problem. All runs used the same clean 45k/5k/10k split, source\_order\_seed=20260601, Apple MPS, no reverse-mode graph, 4×4 columns, four mixed branches of width 96, 128 unsupervised patch filters per kernel, patch kernels 3,5, prediction ridge 300, class ridge 300, control ridge 30, 512 RBF centers, and exported controller logits. Removing target view 6 (source4\_targets57\_settle1\_seed20260687) improves some raw readouts but lowers the controller to 67.13% and is not retained after the frontier; from the earlier PGPE base it gives only 77.18% validation and 76.84% test. Increasing recurrence to two settled steps (source4\_targets567\_settle2\_seed20260689) lowers the source/predicted readouts and controller to 67.06%; the selector can extract a tiny validation-only after-frontier gain (77.96% validation) but held-out test falls to 77.02%. Stronger settling with β=0.40 is rejected (77.92%/77.03%), and weaker settling with β=0.15 reaches 78.00% validation but drops test to 76.98%, showing that residual RMS improvements are not sufficient when the induced class geometry is wrong. The reciprocal directed graph, source view 5 predicting 4,6,7 (seed 20260691), is a hard negative: source/predicted readouts are about 60% and the controller reaches only 61.79%. Thus the useful column is directed: view 4 is a good source for the 5,6,7 target bank, but the reverse source is not.
The successful push is seed-diverse cortical population formation around the same directed source graph. A repeat of the source-4, targets-5,6,7, one-step β=0.25 architecture with seed 20260692 produces source/predicted/settled readouts 65.79%, 65.41%, 64.94%, 65.32%, and 66.01%, plus a 67.47% predictive-logit controller. After the existing validation-selected frontier [79,80,86], a one-member gate retains member 97 (predicted target state v5) with η=0.06 and raises validation/test to 78.10%/77.06%. Allowing the same dopamine rule to add a second member from this source retains member 100 (settled source state) with η=0.018 and reaches the new full-CIFAR physical/no-backprop frontier: $78.18\%$ validation and $77.11\%$ held-out test. The final retained source indices are [13,64,54,12,71,5,60,2,69,63,79,80,86,97,100] with weights approximately [0.215,0.084,0.072,0.071,0.059,0.052,0.049,0.046,0.046,0.040,0.081,0.090,0.018,0.059,0.018]. This is not a SOTA CIFAR result, but it is a clean no-backprop improvement over the previous validation-selected 77.92%/77.05% frontier and over the previous measured 77.07% held-out high. Importantly, continuing to add old generated-source candidates after this base raises validation to 78.28% but lowers test to 77.04–77.08%, and further source-4 seed repeats are quality-gated out: seed 20260693 has a 67.03% controller and contributes nothing after the new base, while seed 20260694 drops to a 66.55% controller. The working rule is therefore specific: generate a population of directed multi-target physical columns, retain only class-aligned seed members that improve every validation fold at small η, and reject validation-only additions even when their local residuals improve.
We then hardened the consolidation rule itself. The new runner apps\_industrial\_breakthrough/bio\_cifar\_neuromodulated\_resample\_consolidation.py keeps the same no-backprop source logits, but accepts a candidate only if its small-η gain survives many random validation resamples. A candidate must have enough full-validation support, a sufficiently high resample win rate, and a non-catastrophic lower-tail gain before it is consolidated. Re-auditing the seed-20260692 source under this rule still selects members 97 and 100, now with weights ending in 0.0588 and 0.0200, and raises the reproducible frontier to 78.22% validation and 77.13% held-out test. The first selected member has full validation gain +0.18 points, resample lower-tail gain +0.036 points, 90.1% resample win rate, and support +9 examples. This is a better rule than the deterministic five-fold gate because it rejects candidates whose apparent gain is carried by a few validation examples.
Finally, we implemented an explicit two-stage predictive hierarchy in bio\_cifar\_physical\_predictive\_pretrain\_mps.py. The new --second\_stage\_target\_views option first fits the usual source-4 maps to targets 5,6,7, settles the source state, and then fits a second set of local ridge maps from the settled source to the same targets before readout. A two-stage-only seed 20260695 improves physical residuals for all targets (v5:0.773→0.756, v6:0.833→0.813, v7:0.697→0.688) but is rejected after the consolidated base; its best marginal row has only +0.04 validation points, support +2, and a negative resample lower tail. Exporting both first-stage and second-stage profiles from the same seed lifts the internal controller to 67.85%, but it is still redundant after the consolidated source-20260692 base. A second hybrid seed, 20260696, has weaker standalone readouts and controller (67.35%), yet its first-stage predicted target-6 member is marginally complementary. The resampled gate retains this member with η=0.015, full validation gain +0.10 points, lower-tail gain 0.00, 85.9% resample win rate, and support +5 examples, producing the new full-CIFAR physical/no-backprop frontier: $78.32\%$ validation and $77.14\%$ held-out test. The final retained source indices are [13,64,54,12,71,5,60,2,69,63,79,80,86,97,100,113]. A targeted v6-only hierarchy (seed 20260697) gives the best v6 residual in the batch (0.828→0.799) but weak class readouts and is rejected. The scientific conclusion is sharper than the numerical gain: physically better target reconstruction is not enough; useful no-backprop source formation requires class-aligned multi-target context plus resampled neuromodulatory consolidation.
The next push made that conclusion explicit. We added two supervised-but-still-local training signals to bio\_cifar\_physical\_predictive\_pretrain\_mps.py. The option --class\_align\_strength injects a training-only class-centroid dopamine signal into the target states used by the local predictive ridge maps; validation/test features still use only the image-derived source state and the learned maps. A full source-4, targets-5,6,7 run with strength 0.20 gives source/predicted/settled readouts 64.97%, 65.08%, 64.73%, 65.36%, 65.18% and a 67.02% controller. The resampled gate rejects all exported members after the 78.32%/77.14% base; the best candidate has full validation gain −0.06 points, lower-tail gain −0.109 points, 2.6% win rate, and support −3. Thus naively adding class centroids to reconstruction targets does not create useful marginal geometry.
The stronger variant is a local dopamine classifier field. For each cortical region, the script now fits a closed-form map from the region's neighboring source context to the global class signal, producing local\_class\_context\_state and settled\_local\_class\_context\_state profiles. This is closer to a biological three-factor rule: source activity supplies the eligibility field, labels provide a broadcast neuromodulator on the core split, and inference uses only the learned local maps. With ridge 300, seed 20260703 reaches 66.37% local-class readout, 66.32% settled local-class readout, and a 67.86% controller. Seed 20260704 improves to 67.00%, 66.84%, and a 67.96% controller. A ridge sweep shows the regularization boundary: ridge 30 overfits tiny-split regional train accuracy to about 99.8% and hurts readout; ridge 1000 improves the tiny-split smoke but at full scale slips to 66.93%, 66.66%, and 67.88%. These local dopamine fields are therefore a real standalone architectural improvement over raw source/predictive readouts, but they still do not beat the current source bank after consolidation. The resampled additive gate over seeds 20260703/20260704 rejects all members; the best row is candidate 131 with η=0.008, zero full-validation gain, lower-tail gain −0.073 points, 38.3% win rate, and support 0. A fixed closed-form spline controller over the retained base plus all local-class members drops to 75.10% test, and a context-gated dopamine table over the strongest local-class candidates leaves validation flat at 78.32% while test slips to 77.13%. The current frontier therefore remains 78.32%/77.14%, and the next necessary change is not another late gate; it is to make the local dopamine classifier field participate earlier in source formation, for example by feeding its regional error/context back into patch selection, target selection, or multi-stage recurrent state formation before logits are compressed.
We implemented that early-feedback test in the same runner. The new --early\_class\_context\_gain, --early\_class\_context\_temperature, and --early\_class\_feedback\_steps options first fit the local class-context field on the raw source columns, convert the resulting regional logits into centered class-probability mixtures over training-set class-centroid displacements, inject that dopamine-like displacement into the source state, renormalize, and diffuse/inhibit the state before any predictive maps or class readouts are fitted. Labels enter only through the core-set local class maps and centroid table; validation/test source shaping uses the learned regional logits. On a 3000/500/1200 smoke split, moving the field earlier raises the source-state readout from 45.58% to 51.08% at gain 0.50, temperature 0.45, confirming that the feedback changes the representation rather than merely adding a late logit source. On the full clean split with the previously strong seed 20260704, raw source is 65.79%, early-shaped source is 66.37%, early class-context is 67.00%, shaped local context is 67.42%, and predicted targets v5/v6/v7 reach 66.56%/66.06%/65.96%. The shaped local field's mean regional train accuracy rises from 63.88% to 72.08% before settling and 72.30% after settling. Thus the upstream geometry hypothesis is validated. However, a global resampled additive gate still rejects all early-feedback members after the 78.32%/77.14% base; the best raw additive candidate has full validation gain −0.04 points and lower-tail gain −0.073 points. We therefore added bio\_cifar\_predictive\_recontroller.py for closed-form subset controllers and upgraded bio\_cifar\_neuromodulated\_context\_gate.py so it can inherit prior source banks and export chainable member logits. Subset controllers improve some held-out test rows but overfit validation. The only robust consolidation lift is an ultra-conservative context replacement gate with 2 confidence bins, minimum bin support 300, and support shrink 500: it selects early-feedback candidate 129, activates only two contexts with mean eta 0.0129, and improves the current base from 78.32%/77.14% to $78.36\%$ validation and $77.15\%$ held-out test, with nonnegative fold minimum. This is numerically tiny, not a SOTA claim, but it is the first evidence that early physical dopamine plus context-local retention can improve both validation and held-out test beyond the resampled frontier. Chaining the exported context member through the generic source-bank normalizer changes its calibration, so the next implementation task is a calibration-aware context-source loader or a replacement-base consolidation protocol, not more blind global addition.
The calibration-aware follow-up resolves that artifact. Both bio\_cifar\_neuromodulated\_context\_gate.py and bio\_cifar\_neuromodulated\_resample\_consolidation.py now accept --calibrated\_logits\_npz; such members are appended to the source bank but bypass the per-member core mean/std/RMS normalization, because they are already fused logits in the ensemble's calibrated decision space. Reloading the round-60 context member this way exactly preserves its replacement-base accuracy, 78.36%/77.15%. Narrow context gating over the early-dopamine members is validation-flat and lowers test to 77.14%; the strict resampled gate rejects the same candidates, with the best row having full validation gain −0.08 points, lower-tail gain −0.109 points, 0% resample win rate, and support −4 examples. A broad all-nonbase context gate over the whole available bank is also exactly flat at 78.36%/77.15%. We then tested a recurrent cellular version of early dopamine: --early\_class\_feedback\_refit\_rounds refits the local class-context field on the shaped source and applies another closed-form class-centroid displacement, while --early\_class\_feedback\_gate can modulate the displacement by local entropy, confidence, or margin. On the smoke split, ungated refit improves source/local readouts to 51.92%/51.75%; entropy gating improves settled source and predicted-target readouts (52.00% settled source, v5/v6/v7=51.17%/50.83%/50.17%); confidence gating collapses the source to 45.17% and the controller to 36.42%. Full clean seed 20260704 is more decisive. Ungated refit slightly improves source and settled-source readouts (66.45%, 66.65%) but lowers local context and controller (67.25%, 67.17%) relative to the one-pass early-dopamine run (67.42%, 67.59%). Entropy gating improves the physical residuals (v5/v6/v7=0.770/0.831/0.693) but hurts every class readout (source 66.11%, local context 66.90%, controller 66.74%). Context and resampled consolidation reject the recurrent-refit members after the calibrated 78.36%/77.15% base. The useful conclusion is not merely negative: residual reconstruction, uncertainty-gated dopamine, and discriminative class geometry are empirically different objectives. The next source architecture should therefore optimize local discriminative predictive targets directly—for example class-conditional residual fields, local contrastive target formation, or target selection by validation-stable class innovation—rather than adding more blind residual reconstruction or late logit gates.
We then made that target explicit with --discriminative\_target\_mode and --discriminative\_target\_strength. For each target view, the runner computes the class-conditional target-state center cy(r) for every cortical region and hidden channel, the global center cˉ(r), and the residual xtarget(r)−cy(r). The tested modes are class\_delta, which uses cy−cˉ; class\_center, which uses cy; and class\_residual, which uses (xtarget−cy)+(cy−cˉ). The surrogate field is RMS-balanced to the raw target field, then either blended with the raw target for strengths in [0,1] or added for strengths above 1. This is a training-only local target transform: validation and test states are image-derived, and the learned source-to-target maps receive no validation/test labels. On a matched 2500/500/1200 MPS smoke split with source view 4, targets 5,6,7, early dopamine gain 0.50, temperature 0.45, and one settled step, the no-discriminative-target controller is 42.92%. class\_delta at strength 1.0 raises individual predicted-target readouts to 51.42%/52.17%/53.25% but leaves the controller at 45.00%. class\_center at strength 1.0 is better: predicted-target readouts become 52.50%/53.08%/53.25% and the controller reaches 47.67%. Adding regional and precision predictive lifts with --regional\_lift\_dim=16 does not improve the individual lift readouts beyond the best predicted-target row, but it improves fusion diversity and raises the smoke controller to 50.92%.
The full clean result shows both the promise and the current limitation. The artifact bio\_cifar\_physical\_predictive\_pretrain\_mps\_outputs\_cleanval\_disctarget\_center100\_lifts16\_earlydop\_g050\_t045\_seed20260704/ uses the same 45k/5k/10k split, Apple MPS, source view 4, targets 5,6,7, 4×4 columns, four mixed branches of width 96, 128 patch filters, prediction ridge 300, class/context ridge 300, class\_center strength 1.0, lift dimension 16, and exported controller logits. The raw source, shaped source, early class-context, and local class-context rows reproduce the previous clean one-pass geometry (65.79%, 66.37%, 67.00%, 67.42%). Pure predicted-target readouts do not improve (66.01%/66.26%/65.74%), confirming that class-center synthesis alone is not the missing mechanism. The useful channel is lifted local predictive geometry: regional predictive lifts reach 68.77%/68.69%/68.14%, precision lifts reach 68.33%/68.45%/67.88%, and the internal predictive-logit spline controller reaches 69.31% (69.12% linear control). This is a real upstream improvement over the previous 67–68% predictive-column family. However, it is redundant with the calibrated frontier. Context gating over these new candidates after the calibrated round-60 base keeps validation flat at 78.36% and lowers test to 77.14%; strict resampled consolidation rejects all members, with the top candidate having full validation gain −0.08 points, lower-tail gain −0.109 points, 0% win rate, and support −4 examples. An orthogonal source view 0 to targets 1,2,3 run (seed 20260711) is weaker internally: regional/precision lifts top out at 67.40% and the controller reaches 68.01%. It is also rejected after the calibrated base, and the combined source-4 plus source-0 candidate pool remains exactly flat at 78.36%/77.15%.
The immediate follow-up tested whether the lifted predictive geometry could be moved earlier by feeding it back into the source state before readout. The new --predictive\_feedback\_* options fit local class-context maps from source, target, prediction, residual, and multiplicative agreement streams. The compact lift context projects those streams to a small random branch basis per region; the resulting regional class logits are averaged over selected target views and passed through the existing class-centroid dopamine displacement. This is still a closed-form no-backprop update, but it is not the missing mechanism. On the smoke split, all-view feedback with gain 0.25 gives feedback source/local rows 51.83%/52.25% and a 50.58% controller; gain 0.50 gives 52.08%/52.58% and a 50.42% controller; view-5-only routing gives 51.67%/52.25% and a 49.42% controller. The full clean artifact bio\_cifar\_physical\_predictive\_pretrain\_mps\_outputs\_cleanval\_predfeedback\_g050\_center100\_lifts16\_earlydop\_g050\_t045\_seed20260704/ confirms the negative: the original regional and precision lift rows reproduce exactly, but the feedback source reaches only 66.72%, the feedback local-context row reaches 67.48%, and the controller drops to 69.09%. The conclusion is therefore precise: discriminative lifted predictive targets are the strongest new upstream single-family result, but post-predictive class-centroid source displacement is too blunt and still too late. The next architecture must move the lifted predictive geometry earlier into source formation, for example by letting regional lift errors select patches, target views, source-view routing, or recurrent class-conditional target fields before the first physical column state and first class readout are formed.
We then moved the signal all the way into patch formation. The new --predictive\_patch\_growth\_* options implement a pilot synaptogenesis pass: the runner first builds the ordinary unsupervised patch bank, collects pilot source/target columns on the training core, solves the same closed-form local predictive maps, scores each image region by raw residual, discriminative mapped residual, or entropy-weighted innovation, samples new fixed patch filters from the high-score image cells, rebuilds the final physical columns with the augmented bank, and only then fits the final readouts. This remains forward-only; the pilot uses training-core residual fields to choose fixed filters, and validation/test images only pass through the resulting filter bank. On the source-4 target-5,6,7 smoke split, mapped-residual growth with 16 filters per kernel/view improves several target rows but overfits the small RBF controller (49.67%, linear control 53.33%), while entropy-weighted innovation is weaker (50.08%, control 51.83%). The best smoke is mapped-residual growth with 32 filters per kernel/view, top-25% residual sampling, and power 1.5: predicted target rows reach 54.50%/53.92%/55.08%, the spline controller reaches 51.08%, and the linear control row reaches 53.50%. A stricter top-15%, power-2.0 sampler regresses to a 50.50% spline controller, so lower residual alone is again not the objective.
The full clean run is a useful split decision rather than a victory. The artifact bio\_cifar\_physical\_predictive\_pretrain\_mps\_outputs\_cleanval\_predpatch\_mapped\_f32\_top025\_center100\_lifts16\_earlydop\_g050\_t045\_seed20260704/ uses the exact previous 45k/5k/10k split and source/target configuration, but appends the residual-grown filters to the bank before final column formation. Early source geometry improves: raw source rises from 65.79% to 66.21%, shaped source from 66.37% to 66.48%, and settled source from 66.30% to 66.93%. However, the strongest lifted predictive geometry weakens: regional lifts are 68.40%/68.17%/67.98% instead of 68.77%/68.69%/68.14%, precision lifts are 68.19%/67.81%/67.80% instead of 68.33%/68.45%/67.88%, and the internal controller drops to 69.07% (linear control 68.69%) instead of 69.31%. A validation-aware logit fusion audit shows that the grown-bank columns are nevertheless orthogonal within the physical predictive family: the original discriminative-target logits alone give 68.58%/69.42% validation/test under the same RBF fusion audit, while original plus grown-patch logits give 71.88%/71.95%. But this does not survive the global calibrated source bank. Feeding the four fused members into the round-60/round-43 calibrated context gate, either normalized or marked as already calibrated, remains exactly flat at 78.36%/77.15%, with no active contexts for the strongest fused control members. The conclusion is architectural: predictive residual patch growth creates useful new source evidence, but appending those filters into the same bank perturbs the best lifted target geometry and remains redundant after the larger calibrated ensemble. The next serious version should keep baseline and residual-grown patches as parallel cortical populations with source-local routing before predictive maps, rather than replacing the baseline population by concatenating filters.
That follow-up is now a negative result. We extended bio\_cifar\_physical\_predictive\_pretrain\_mps.py with a parallel patch-growth mode, a separate grown or augmented physical population, and a no-backprop confidence route fitted from local class-context solves. The grown-only route mostly rejects the grown population: on the smoke split its mean gate is 0.341 and only 1.7% of image-regions prefer the grown branch. Routed predictive residuals worsen and the controller reaches only 47.75%. The augmented base+grown population is no better: mean gate 0.345, grown-preferred fraction 1.4%, routed predictive rows about 49–51%, and controller 48.92%. A relaxed resample consolidation of the full physical source bank after the calibrated 78.36%/77.15% base also selects no physical additions; the best candidate has η=0.002, full validation gain −0.040 points, and zero resample win rate. Thus the 71.95% family-fusion signal is late-logit diversity, not evidence for state interpolation by a class-confidence gate.
We then pushed the same smoke protocol across source-generation knobs. Seed diversity and width-only scaling do not help: seed 20260711 and branch width 128 give best predicted-target rows 52.83% and 52.75%, with controllers 47.50% and 49.92%. Lower early dopamine gain improves target residuals but destroys class geometry, while stronger or weaker class-center targets also regress. A class\_residual target law gives very low physical residuals (0.907/0.944/0.802 for target views 5/6/7) and a 52.67% local-context row, but predicted-target class rows collapse to 47–48%; adding it as an auxiliary neuromodulator in the main class-center run still yields only a 46.33% controller. The biological lesson is concrete: physically easy target prediction is not the same as class-aligned representation formation.
Finally, we repeated the cell-law ablation in this residual-growth setting. Conductance/reversal-potential cells are the only plausible single-family alternative, with 52.42% source accuracy and 54.00% best predicted-target accuracy, but their controller remains 48.58%. Hodgkin–Huxley-style algebraic gates overfit the local class context and trail at 48.08% controller. Compact spline-window cells produce smoother target residuals but collapse discriminative geometry to about 32% and a 20.50% controller. The mixed pole bank remains the best matched transfer law for the current CIFAR descriptors. The next architecture should therefore not promote naive biological naming, width, confidence state routing, or residual reconstruction. It should generate stronger mixed-cell physical sources with pre-registered diversity and retain them by resampled logit-level consolidation, or replace the confidence gate by a marginal predictive-innovation route that is selected before class-logit compression.
We next made the synaptogenesis reward explicitly discriminative rather than reconstructive. The same runner now supports --predictive\_patch\_growth\_score values class\_error, class\_margin, class\_error\_mapped\_residual, and class\_margin\_mapped\_residual. In the pilot pass, a local class-context map is fitted from the source columns to the training labels by the same regional ridge solves used for early dopamine. For each image and cortical region, the class\_error score is 1−py, where py is the local probability assigned to the correct class; the margin score uses the best competing class against py. These scores sample new fixed filters from regions where the source representation is locally class-hard, before the final source columns, target maps, readouts, and controllers are fitted. This is a closer three-factor biological signal than raw residual reconstruction: presynaptic image patches define candidate synapses, the local class-context solve defines a postsynaptic error field, and the sampled patch bank changes the future source representation without reverse-mode differentiation.
The smoke results identify the correct objective. Repeating the source-4 target-5,6,7 split with 24 filters per kernel/view, top-20% sampling, power 1.5, early dopamine gain 0.50, temperature 0.45, class\_center target strength 1.0, and lift dimension 16, pure class\_error growth reaches a 54.75% regional-lift row and a 52.33% spline controller on seed 20260720. The same setting on seed 20260710 reaches a 54.25% precision-lift row and the same 52.33% controller. Combining class hardness with mapped residual is worse: class\_error\_mapped\_residual falls to a 46.42% controller, and class\_margin\_mapped\_residual reaches only 48.08%. A wider class\_error bank with 32 filters and top-25% sampling also regresses to a 51.67% controller. Thus the local reward field itself is useful, but multiplying it by target residual geometry reintroduces the wrong objective.
The full clean result is the new strongest upstream single-family CIFAR result in this branch. The artifact bio\_cifar\_physical\_predictive\_pretrain\_mps\_outputs\_cleanval\_classhard\_classerror\_f24\_top020\_center100\_lifts16\_earlydop\_g050\_t045\_seed20260704\_exportctrl/ uses the canonical 45k/5k/10k split, Apple MPS, source view 4, targets 5,6,7, 4×4 columns, four branches of width 96, 128 base patch filters, 24 class-error-grown filters per kernel/view, prediction and class ridges 300, class\_center strength 1.0, regional lift dimension 16, and an exported predictive-logit spline controller. It raises the raw/source/local rows to 66.67%, 67.24%, and 67.91%, versus 65.79%, 66.37%, and 67.42% for the earlier discriminative-lift baseline. The best regional lift reaches 69.19%, precision lifts reach 68.81%/68.76%/68.40%, and the exported internal controller reaches 69.99% with a 69.82% linear control row. This beats the previous 69.31% discriminative-lift controller and the 69.07% residual-growth full run, while preserving the no-backprop protocol.
A stricter promotion of the predictive-feedback variant closes that branch. The full clean artifact bio\_cifar\_physical\_predictive\_pretrain\_mps\_outputs\_cleanval\_classhard\_predfeedback\_g025\_f24\_top020\_seed20260704\_exportctrl/ uses the same 45k/5k/10k split, source view 4, targets 5,6,7, class-error growth with 24 filters and top-20% sampling, plus a conservative predictive-feedback gain 0.25. It preserves the upstream class-hard rows but does not improve the controller: source/local rows are 67.24%/67.91%, best regional lift is 69.19%, and the predictive-logit spline controller reaches 69.85% with 69.33% linear control, below the 69.99% class-hard baseline. Strict resampled consolidation after the calibrated base rejects all exported predictive-feedback members; the best candidate gives only a +0.020 point full-validation bump, a −0.073 point lower-tail gain, 45.3% resample win rate, and is not retained. Thus predictive feedback is currently redundant once the class-hard local reward field and calibrated source bank are present.
The calibrated frontier audit remains negative. Appending these 16 exported class-error members after the round-43 inherited source bank and using the round-60 context member as a calibrated base gives base\_selected=136 and candidates 120,…,135. The context gate finds only a tiny validation bump, 78.36%→78.40%, while held-out test falls from 77.15% to 77.14%. Strict resampled consolidation rejects every addition; the best candidate has η=0.020, full validation gain −0.020 points, lower-tail gain −0.145 points, 32.8% resample win rate, and support −1, so the final calibrated frontier remains 78.36%/77.15%. The scientific conclusion is sharper than the score: local class-hardness is the first patch-growth objective that improves all upstream source and controller geometry on full CIFAR, but the high-70s global source bank is now saturated by similar errors. The next step should not be another residual selector; it should create source diversity around the class-hardness signal itself, for example distinct local reward heads, class-pair-specific hard-region filters, or validation-stable source families whose errors differ from the round-60 calibrated base.
We also reran the self-supervised JEPA/diffusion/attention audit with the same clean split and the current source bank. The patched bio\_cifar\_ssl\_jepa\_attention\_fusion\_audit.py now shuffles and validates the candidate source split consistently before checking source labels. On the 17-member current-frontier bank with latent\_dim=1024 and 768 attention centers, the standalone closed-form heads remain weak: spline/pole random features 47.06%, JEPA view prediction 47.52%, diffusion denoising 47.58%, and low-margin attention 49.34%. The source-bank baselines are much stronger, with mean normalized logits 69.61%, control ridge 72.20%, and hard-state RBF 73.99%. Appending the SSL heads is flat or worse: hybrid mean 69.86%, reliability-gated hybrid 70.15%, hybrid control ridge 72.31%, and hybrid hard-state RBF 73.42%. Validation selects the baseline hard-state RBF, not an SSL-augmented source set. This is a direct negative control for late JEPA/diffusion/attention attachments. If these ideas help OSNR, they must become source-forming local objectives inside the physical columns, not shallow heads appended after class logits.
The first diversity attempt was deliberately class- and competitor-local. The new --predictive\_patch\_growth\_class\_focus, --predictive\_patch\_growth\_competitor\_focus, and --predictive\_patch\_growth\_focus\_leak options restrict the pilot class-hardness field to selected true labels or best-competing labels while retaining a small off-focus sampling leak. Narrow cat/dog focus is not a win. In append mode, a true-label focus on classes 3,5 or a competitor focus on 3,5 collapses the fair smoke split to about 49–50% best rows before or at the controller, far below the global class-error smoke. Keeping the focused filters as a separate parallel grown population avoids replacing the base bank, but it still fails: with the fair 4×96 branch width, shuffled split, prediction ridge 300, and settle step, the focused grown-only row reaches 45.75%, concatenating it with the base source reaches 52.83%, the best predicted-target row is 54.92%, and the controller falls to 50.42%. A broader animal-family focus on classes 2,…,7 with leak 0.20 is also negative: best row 54.50%, controller 49.58%. The interpretation is that class-hardness should remain a dense regional reward field; hard class subsets are too sparse and perturb the random pole projection or add weak auxiliary populations.
We also tested two more biologically plausible follow-ups on top of the class-hard growth. First, predictive-geometry dopamine feedback was added after the class-hard maps. A conservative all-view feedback gain 0.25 preserves the old best row (54.75%) and nudges the smoke controller from 52.33% to 52.42%, but this is only marginal. Entropy-gated feedback improves the feedback-local member to 54.17% but lowers the controller to 52.00%, so the signal is member-level diversity rather than a stable source-family improvement. Second, an extra early class-feedback refit round increases local training confidence but overfits the field: with class-hard growth it gives a 53.83% source row, only 54.17% best regional lift, and a 50.92% controller. These smokes close an important loop. The next serious route is not narrower hard-class masks, post-predictive global displacement, or more refit confidence. The higher-leverage target is a validation-selected source generator: mutate the source view, target views, pole family, branch width, growth objective, feedback gate, and consolidation rule together, then keep only source families that improve the calibrated frontier under resampled consolidation.
The next source-generator mini-matrix fixed the class-hard objective and mutated the retinal source/target graph. Local-contrast source 6 to targets 4,5,7 is a clear failure: despite nearly perfect local training context, held-out rows are only 42–44.5% and the controller is 39.75%. DoG/opponent source 7 to targets 4,5,6 reaches only a 52.25% target row and a 48.50% controller. High-pass source 3 to targets 4,5,7 is closer, with a 54.67% predicted-target row and a 52.00% controller, but it remains below the source-4 baseline. Raw RGB source 0 to targets 4,5,7 is the only promoted candidate: smoke seed 20260733 reaches a 55.08% predicted-target-4 row and a 53.25% controller, and repeat seed 20260734 gives a 54.50% regional row and a 53.17% controller. The full clean promotion is competitive but not better than source 4: raw/source/local rows are 66.47%/66.64%/67.61%, best regional lift 69.11%, and the exported controller 69.87% with 69.59% linear control, versus 69.99% for the source-4 class-hard run. The calibrated round-60 audit is flat and strict resampling rejects every addition; top candidate 121 has η=0.004, full validation gain −0.060 points, lower-tail gain −0.109 points, 1.0% win rate, and support −3. Thus source 0 is a real upstream variant but not a frontier-complementary family. The working source graph remains opponent/center-surround source 4 with edge/contrast/DoG targets 5,6,7; future generation must mutate more than the source view, for example local reward heads and pole/branch families jointly.
We then ran that joint direction as small controlled smokes, still on source 4→5,6,7. Conductance/reversal-potential cells under class-hard growth are worse than the mixed pole bank: best row 51.33%, controller 49.58%, despite perfect local training context. Increasing mixed capacity from 4×96 to 6×80 also regresses: best row 53.33%, controller 50.08%. Pure margin-based synaptogenesis is not the missing reward head: best row 53.67%, controller 50.08%. Changing class-error selectivity confirms the top-20% sampler. A sharper top-10% sampler reaches 54.58% best row but only a 49.25% controller, while top-30% falls to a 52.50% best row and 49.67% controller. These negative ablations are useful because they narrow the mechanism: the current winner is not simply more biological cell naming, more width, margin-only reward, or arbitrary hard-region sparsity; it is the specific combination of mixed poles, opponent source geometry, class-error regional reward, and moderate top-20% patch growth.
As a final consolidation check, we made the PGPE logit-architecture search calibrated-aware. Without this correction, the PGPE loader normalized the already calibrated round-60 base and artificially lowered it from 78.36% to 78.12% validation, producing a misleading 77.29% test result at lower validation. The updated runner accepts --calibrated\_logits\_npz and bypasses per-member normalization for those sources, matching the context-gate protocol. Running PGPE over the calibrated round-60 base plus the source-4 and source-0 class-hard banks with 180 steps, population 48, top-12, validation-half stability penalty, and an 8192-sample core reward anchor selects only the base: final validation/test remain 78.36%/77.15%. Thus even population-level nonnegative source weighting does not rescue these candidates once calibration is handled correctly.
The next MPS smokes close the current class-hard patch-growth family. Keeping the grown class-error filters as a separate parallel physical population does not solve the overwrite problem: grown-only reaches 47.33%, base-plus-grown concatenation reaches 50.92%, the best predicted-target row is 54.92%, and the controller is 49.92%. A no-backprop confidence route fitted from local class-context solves mostly rejects the grown population; only 1.7% of regions prefer it, and the routed controller falls to 47.75%. The sampling-sharpness sweep is also negative. Flattening the class-error sampling power to 1.0 gives a 54.25% best row but only a 44.08% controller; sharpening to 2.25 gives 53.08% best and a 45.08% controller. A broad source-4 predictive stack to all non-source retinal views reaches only 54.17% best and a 52.25% controller, below the selected target-5,6,7 graph. Finally, we tested an explicit forward-only contrastive goodness profile: for each target map, the local state compares the predicted target against the true target and several rolled negative targets. The compact contrastive score is stable but not frontier-moving (54.25% best contrastive score), while the high-dimensional contrastive residual state is weaker (53.58% best). The best overall row in that run remains the old regional lift at 54.75%, with controller 52.25% and linear control 55.00%. The conclusion is that class-hardness top-20% on source 4→5,6,7 is a local optimum for this substrate; the next attempt must create a different source-forming mechanism, not another patch-growth or late routing variant.
We then tested a more explicit cellular architecture in apps\_industrial\_breakthrough/bio\_columnar\_predictive\_control\_benchmark.py. The model has a retinal/V1 front end, 49 L1 cortical columns with 64 cable cells each, 16 L2 association columns with 128 cells each, a 4096-cell global field, optional thalamic sensory skip cells, four dendritic branches, and four cable modes per branch. Each branch has stable leak/synapse/diffusion poles, conductance gates, reversal potentials, lateral inhibition, block-local covariance readouts, fusion, and a dopamine-like residual controller. This is closer to the proposed biological architecture than the single global projection, but the first full FashionMNIST run is a negative result: the full-field columnar profile reaches 90.74% on the 60,000/10,000 protocol, below the simpler streaming covariance field at 92.32%. Matching the streaming ridge and fp32 cache lowers it further to 90.04%. The diagnosis is useful. Explicit dendritic geometry alone does not solve credit assignment; the current columnar stack discards or overcompresses class-separable sensory evidence before the closed-form controller. A serious next cellular model must learn or select intermediate predictive targets locally, not merely route fixed cable states into a larger final covariance solve.
The ablations matter. Prototype voting retains useful memory but is weaker than the linear center-state solve. Diagonal Gaussian statistics are too crude. The dense kernel memory is pathological at high capacity, collapsing to about 10–12% final accuracy despite excellent early-task performance; the failure is a conditioning/credit-allocation warning against treating every stored center as a dense global kernel atom. The class-subspace attractor reaches only 75.93% and fixed-budget multi-head attention fusion reaches 79.32% on canonical FashionMNIST, so attention-style splitting is not automatically useful without a stable local credit field. The architecture search in apps\_industrial\_breakthrough/bio\_plasticity\_architecture\_search.py adds the knobs the biological thesis actually needs—state degree, local projection depth, neuron count, dendrite count, ridge, and fusion temperature. Degree-2 lifts and extra random projection depth increase memory without improving Fashion test accuracy; five-dendrite variants tie validation but cost substantially more memory. The selected three-dendrite row is therefore the current best tradeoff. The working mechanism is more specific: project cellular evidence into independent pole-rich dendritic states, accumulate local eligibility covariances and dopamine cross-covariances, and fuse the resulting quadratic controllers. This is closer to biological plasticity than replayed global-gradient updates, because old information persists as local co-activity statistics rather than as raw examples or weights repeatedly overwritten by backpropagation.
The first application-level validation script, apps/01\_fluid\_dynamics/run\_vortex\_street.py, extends the Hermite trunk from separate 1D passes to a true 2D tensor-product coefficient system. Each grid node stores nine Hermite streams corresponding to value, first derivatives, second derivatives, and mixed derivative channels. The continuous stream function az defines the incompressible velocity field by the analytical curl vx=∂yaz,vy=−∂xaz. Consequently, incompressibility is structural rather than imposed by a penalty. Solid cylinder and wall masks overwrite all nine coefficient streams to zero at masked vertices, enforcing no-slip and zero-flux constraints by coefficient assignment.
The 2D Hermite Gram is assembled as a tensor product of the 1D cross-correlation filters. Applying torch.fft.fft2 diagonalizes the spatial part of the block-circulant system, reducing the global solve to independent 9×9 complex systems at each frequency coordinate (νy,νx). In the current synthetic unrolled vortex-street validation, the script runs 36 frames on a 32×48 grid with a mean processing duration of 1.2850 ms per frame. Boundary leakage remains 0.000000e+00 and the coefficient clamp residual remains 0.000000e+00 through the unroll. The final momentum residual is 1.876066 in the script's synthetic nondimensional units.
Visual validation: multi-obstacle CFD cinema
The breakthrough visual script, apps\_breakthrough/fluid\_vortex\_cinema.py, scales the same tensor-product Hermite fluid engine to a 64×256 canvas with a multi-obstacle mask inspired by the ``Smiley Face / HI!'' geometry used in the Spline-PINN visual demonstrations. The mask combines disk, capsule, and rectangular primitives to produce a dense nonconvex obstacle field. The solid set includes the obstacle geometry and the domain walls. At every time step, all nine Hermite coefficient streams are overwritten to zero on this set, so no-slip and zero-flux constraints enter as direct coefficient assignments rather than differentiable penalties.
The state variable remains a scalar stream function az. Velocities are recovered by the analytical curl (vx,vy)=(∂yaz,−∂xaz), which structurally removes the need for an incompressibility loss. The unrolled update evaluates a synthetic transport-diffusion step azn+1=azn+Δt(νΔazn−η(vn⋅∇)azn−κωn+fn), where ω=∂xvy−∂yvx is the vorticity field and fn is a time-dependent wake forcing. The updated scalar field is mapped into the nine Hermite streams, clamped on the solid set, passed through the block-circulant Hermite Gram, and recovered by the parallel 9×9 Fourier solver. A conservative amplitude limiter is applied to keep the synthetic visualization stable over the full 100-frame unroll; this limiter is a numerical display stabilizer, not a replacement for calibrated Navier–Stokes time integration.
The run exports every fifth step as a high-contrast PNG visualization of speed and signed vorticity. In the verified local run, it produced 20 frames, maintained boundary leakage 0.000000e+00 and coefficient clamp residual 0.000000e+00, and completed with mean latency 13.6396 ms per frame. The final synthetic momentum residual was 8.212812 in the script's nondimensional units. The entire run executes under torch.no\_grad() with 0.00 B autograd graph allocation.
Measured execution ledger for the multi-obstacle CFD cinema validation.
Full exported speed-vorticity rendering sequence from the OSNR multi-obstacle CFD cinema run. Dark regions are hard-clamped obstacle and wall coefficients; color encodes velocity magnitude modulated by signed vorticity.
The graphics super-resolution script, apps\_industrial\_breakthrough/graphics\_superres\_engine.py, applies the Tier 2 adaptive sparse core to a high-density geometric rendering problem. The target is a synthetic industrial graphics asset with high-frequency directional contours. Each horizontal scanline is modeled as a finite-rate-of-innovation signal with six discontinuity locations, corresponding to three filled geometric bands. The raw comparison image is produced by evaluating the same asset on a coarse uniform grid and expanding it to the display canvas, which exposes block aliasing at the sub-pixel boundaries.
The OSNR path passes the scanline moments into the TLS matrix-pencil tracker, recovers the fractional transition coordinates, and snaps the sparse step dictionary to those coordinates before reconstruction. The cross-Gram-shielded ADMM sieve suppresses the empty background and uniform interior atoms, while the scale-invariant ridge debiasing pass stabilizes the active discontinuity support. The final continuous field is evaluated on a 512×512 canvas and exported as a side-by-side PNG: coarse block-aliased rendering on the left, FRI-snapped OSNR reconstruction on the right.
The aerodynamic wind-tunnel script, apps\_industrial\_breakthrough/aerodynamic\_wind\_tunnel.py, extends the tensor-product Hermite fluid path to a high-Reynolds engineering surrogate. The obstacle is a multi-element NACA 0012-style body composed of a main airfoil, slat, and deflected flap. The geometry is rasterized into a curved solid mask on a 64×192 wind-tunnel grid. The simulated regime uses Re=50,000 with reference velocity U=0.24, chord c=0.72, and kinematic viscosity ν=3.45600000e−06.
As in the CFD cinema experiment, the state variable is a scalar stream function az and the velocity field is recovered through vx=∂yaz,vy=−∂xaz. This curl parameterization enforces incompressibility structurally. The wall and airfoil masks overwrite all nine tensor-product Hermite coefficient channels to zero at each step, imposing no-slip and zero-flux constraints by assignment. The update evaluates advection, viscous diffusion, and a high-frequency wake forcing through forward finite-difference ladders, then applies the 2D block-circulant Hermite Fourier solve. Every tenth time step is exported as a raw state matrix containing velocity magnitude, vorticity, the solid mask, Reynolds number, and boundary leakage.
Measured execution ledger for the high-Reynolds multi-element airfoil wind-tunnel validation. The exported state matrices are stored under apps\_industrial\_breakthrough/wind\_tunnel\_states/.
Industrial benchmark ingestion: Hugging Face video challenger
The Hugging Face challenger script, apps\_industrial\_breakthrough/huggingface\_sota\_challenger.py, is the first repository path that ingests an external hosted video asset rather than a manufactured field. The script uses the official datasets library and Hugging Face Hub APIs to inspect video metadata, resolves local or HTTPS media references when they are available, and decodes real video containers with a prioritized backend chain: decord, then PyAV, then imageio-ffmpeg. The production execution profile targets T=30 frames at 256×256 RGB resolution. The public APRIL-AIGC/UltraVideo rows currently expose metadata and YouTube identifiers rather than direct .mp4 payloads, so the script requires OSNR\_HF\_VIDEO\_FILE for a local or HTTPS UltraVideo media export. If no decodable media file is provided, it records this condition explicitly and falls back to a real Hugging Face video fixture so the decoding, algebraic compression, and metric path remains executable.
Each RGB scanline is processed as a composite sparse-plus-smooth color track. Before moments are formed, the decoded tensor is passed through a localized separable cubic B-spline prefilter with kernel [1,4,6,4,1]/16 along both image axes. This shift-invariant smoothing step suppresses quantization and compression perturbations that otherwise dominate the algebraic roots. Gradient-selected transitions are then de-duplicated by non-maximum suppression, converted into moments, and routed into the packaged AdaptiveSparseSolver. For this noisy-video path, the solver's FRI tracker is replaced by a ridge-regularized TLS matrix pencil: the denoised Hankel coordinate equation is solved as (H0⊤H0+γI)Z=H0⊤H1,γ=10−6mean(diag(H0⊤H0)). This prevents near-null Hankel directions from snapping knots to compression artifacts. The remaining sparse recovery uses the cross-Gram-shielded ADMM sieve with ridge debiasing to suppress inactive background atoms. Because natural video is not a pure step-edge signal, the repaired pipeline adds a pruned smooth residual tier after sparse recovery. The residual is projected onto an orthonormal DCT row dictionary, ridge solved, and hard-pruned to retain only the largest coefficients. This implements the continuous-domain composite model f(x)=fsparse(x;τk)+fsmooth(x), with FRI atoms representing geometry and low-frequency DCT atoms representing illumination, texture, and compression residuals.
Implemented video algorithm.
The current video path is not a neural training loop. It is a deterministic composite inversion pipeline whose components are tied to the spline theory above. For a decoded video tensor Y∈[0,1]T×H×W×3,T=30,H=W=256, the implementation proceeds as follows.
Decode and normalize. Load consecutive frames through the prioritized decoder chain decord/PyAV/imageio-ffmpeg; resize to 256×256 and normalize RGB values to [0,1].
Spline prefilter. Apply the separable cubic B-spline smoothing kernel b=161[1,4,6,4,1] along x and y. This produces a denoised tensor Y used only for edge moment estimation, not for final metric evaluation.
Scanline FRI moments. For every time, row, and color channel, flatten the horizontal trace into a one-dimensional signal yt,h,c(x). Select K=32 non-maximum-suppressed gradient transitions and convert them into innovation moments mℓ=k=1∑Kakτkℓ,ℓ=0,…,2K+1.
Ridge TLS matrix pencil. Build Hankel pairs (H0,H1) from the moments, project the concatenated pencil to rank K, and solve the stabilized shift equation (H0⊤H0+γI)Z=H0⊤H1,γ=10−6mean(diag(H0⊤H0)). The eigenvalues of Z give the snapped sub-pixel step coordinates τk.
Cross-Gram sparse solve. Form a smooth sinusoidal dictionary As and a snapped step dictionary Ax(τk). The ADMM block updates use the cross-Gram shield As⊤Ax exactly as in the hybrid normal equations: (As⊤As)cs=As⊤y−As⊤Axz,(Ax⊤Ax+ρI)cx=Ax⊤y−Ax⊤Ascs+ρz−u. This is the oblique projection step that prevents smooth illumination atoms and sparse step atoms from absorbing each other's energy.
Scale-invariant debiasing. Debias the active FRI atoms with (A⊤A+ϵdI)c=A⊤y,d=mean(diag(A⊤A)),ϵ=10−6. The sparse tensor is stored in a 1024-slot accounting dictionary; only the FRI-snapped active atoms are solved, while the remaining slots are explicit hard zeros.
Frame-wise 2D-DCT residual. Convert the sparse scanline prediction back to a video tensor Ysparse. For each frame and color channel, represent the residual Rt,c=Yt,:,:,c−Ysparse,t,:,:,c in an orthonormal two-dimensional DCT basis Rt,c(i,j)≈p=0∑Py−1q=0∑Px−1dt,c,p,qψp(i)ψq(j),Py=Px=256. Coefficients are computed by the separable projection Dt,c=Ψy⊤Rt,cΨx. The practical profile keeps the 32768 largest coefficients per frame/channel; the ceiling profile keeps all 65536 coefficients.
Composite synthesis and export. The final reconstruction is Y=Ysparse+ΨyDΨx⊤, clipped to [0,1]. The script exports target and OSNR frames for both Pareto profiles and evaluates PSNR, SSIM, LPIPS, sparsity, latency, and memory.
This algorithm explains the main empirical observation. The row-DCT variant had no vertical basis functions and therefore generated visible scanline ripple. The 2D-DCT tier restores a true image-plane smooth residual space, eliminating that artifact when enough coefficients are retained. The price is that the ceiling profile becomes a dense transform-codec upper bound rather than a sparse representation claim.
The experimental record is cumulative. We keep the earlier small high-quality run because it provides a reconstructable baseline for the quality ceiling of the current sparse-plus-DCT path: T=2 frames at 32×32 RGB resolution, K=8 transitions per scanline, one B-spline smoothing pass, matrix-pencil ridge scale 10−6, and a 32-term DCT residual tier pruned to 30 coefficients per row. That configuration improved PSNR from the original sparse-only 18.5182 dB and the ridge-prefiltered 21.1817 dB result to 60.2304 dB, with SSIM 0.999686, LPIPS 0.000004, 80.21% combined hard-zero parameters, and 95.00% sparse-tier hard-zero parameters.
The overhauled production-profile run used the UltraVideo metadata row as the ingestion target; because that row exposed the non-decodable identifier BJRpaBau\_QI, the script decoded the Hugging Face video fixture while retaining the requested T=30, 256×256 RGB tensor layout and K=32 transitions per scanline. It used two B-spline smoothing passes, matrix-pencil ridge scale 10−6, a 1024-slot sparse accounting dictionary, and a 128-term DCT residual candidate dictionary. The live ADMM solve uses only the FRI-snapped active knot bank; the remaining sparse slots are retained as explicit hard-zero background atoms, avoiding a wasteful B×1024×1024 Gram expansion. With the strict sparsity setting of 22 retained DCT coefficients per scanline, the representation reaches 29.1223 dB PSNR, SSIM 0.805561, and LPIPS 0.243320 with 95.31% combined hard-zero parameters. With a quality-prioritized setting of 120 retained DCT coefficients per scanline, the same decoded tensor reaches 35.0321 dB PSNR, SSIM 0.958564, and LPIPS 0.033300 while retaining 86.81% combined hard-zero parameters and 96.88% sparse-tier hard-zero parameters. The latter is the better production direction because image fidelity is the decisive benchmark; the strict sparsity profile is retained as an ablation, not as the preferred operating point.
Visual inspection of the row-DCT reconstructions revealed coherent vertical ripple artifacts. This is a structural artifact of treating each row independently: the smooth residual tier has no vertical coupling, so natural two-dimensional texture is forced into separable scanline corrections. We therefore added a frame-wise two-dimensional DCT residual tier and unrolled it across the complete T=30 decoded sequence. The script exports both the target and OSNR reconstruction at every time step to apps\_industrial\_breakthrough/ultravideo\_cinema/, with separate practical\_sparse and quality\_ceiling directories. The multi-frame run produced 120 PNG frames: target and reconstruction pairs for both profiles over t=0,…,29.
With 32768 retained 2D coefficients per frame/channel, the full cinema profile reaches 41.3874 dB and LPIPS 0.021302, and the visible ripple is largely suppressed. With the full 256×256 2D residual basis retained, the current algebraic framework reaches its quality ceiling across the full sequence: 142.7187 dB PSNR, SSIM 1.000000, and LPIPS 0.000000. This ceiling run is not presented as a compression result; it is an upper-bound diagnostic proving that the sparse FRI geometry plus 2D smooth residual path can reproduce the decoded video exactly when quality is unconstrained. The next meaningful engineering target is therefore the intermediate regime between 32768 and 65536 retained 2D residual coefficients, or a more structured perceptual residual dictionary that concentrates the same visual quality into fewer active parameters.
Cumulative Hugging Face video ingestion benchmark ledger using the B-spline-prefiltered, ridge-regularized, sparse-plus-DCT OSNR codec path. LPIPS is computed with the official torchmetrics AlexNet-backed implementation. Rows are retained as experiment memory rather than overwritten by later ablations.
First-frame visual comparisons for the cumulative Hugging Face video challenger ledger. The earlier small high-quality profile is preserved as a reconstructable historical result, while the full 2562 2D-DCT ceiling profile records the current maximum image quality of the algebraic sparse-plus-smooth path.
clearpage
Frame-by-frame UltraVideo cinema comparison
Table [tab:ultravideo-cinema-frames] gives the direct visual audit requested for the multi-frame cinema export. Each row shows the decoded target frame, the full-quality OSNR 2D-DCT ceiling reconstruction, and the practical sparse reconstruction with the same FRI geometry tier but a pruned 2D residual budget. The comparison is intentionally image-first: the ceiling column records the maximum quality attainable by the current algebraic sparse-plus-smooth representation, while the practical column records the visible cost of residual pruning.
c c c c@
caption(Frame-by-frame visual comparison for the T=30 UltraVideo cinema export. The three image columns are, respectively, the decoded target, the quality-ceiling OSNR reconstruction, and the practical sparse OSNR reconstruction.)
Frame
endfirsthead
Frame
endhead
000
001
002
003
004
005
006
007
008
009
010
011
012
013
014
015
016
017
018
019
020
021
022
023
024
025
026
027
028
029
Frame-by-frame visual comparison for the T=30 UltraVideo cinema export. The three image columns are, respectively, the decoded target, the quality-ceiling OSNR reconstruction, and the practical sparse OSNR reconstruction.
endgroup
Multi-view DL3DV scene reconstruction runner
The next industrial runner, apps\_industrial\_breakthrough/dl3dv\_sota\_challenger.py, lifts the video pipeline from a regular x,y,t tensor to a multi-view scene tensor governed by camera rays. The runner is designed for the gated DL3DV/DL3DV-Benchmark repository. It deliberately refuses blind full-dataset downloads and requires either a local scene directory via OSNR\_DL3DV\_SCENE\_DIR or an authenticated single-scene Hugging Face prefix via OSNR\_DL3DV\_SCENE\_PREFIX. This is necessary because the public benchmark repository is multi-terabyte scale and requires acceptance of dataset access conditions.
For a selected scene, the runner parses transforms.json, resolves the first 30 frame images, downsamples them to 256×256, and constructs a target tensor Y∈[0,1]30×256×256×3. For each view k, pixel coordinates are mapped to continuous camera rays by the usual NeRF/COLMAP transformation xk(s;i,j)=ok+sdk(i,j), where ok is the camera center from the camera-to-world matrix and dk is obtained by applying the camera rotation to the normalized intrinsic-coordinate direction ((i−cx)/fx,−(j−cy)/fy,1). The current algebraic solve then uses the same composite partition as the UltraVideo runner: a ridge-regularized TLS matrix pencil estimates scanline discontinuity coordinates, the FRI step dictionary is snapped to those coordinates, and the cross-Gram ADMM shield prevents the smooth sinusoidal tier from absorbing discontinuity energy. The residual is lifted from a frame-wise 2D DCT to a view-volume 3D DCT, R(v,y,x,c)≈p,q,r∑dc,p,q,rψp(v)ψq(y)ψr(x), with an optional 3D FFT ridge conditioner Rcond(ω)=1+γR(ω). This FFT division is the implemented block-circulant identity/ridge solve for the current prototype; a full physically coupled 3D radiance operator remains future work.
The gated live run was executed against the locally cached official scene prefix 0853979305f7ecb80bd8fc2c8df916410d471ef04ed5f1a64e9651baa41d7695. The selected path contains the DL3DV nerfstudio layout with transforms.json and RGB source frames. The runner resolved the first image as nerfstudio/images/frame\_00001.png, parsed the camera intrinsics and camera-to-world matrices, constructed per-pixel ray origins and directions of shape 30×256×256×3, and reconstructed the observed multi-view image stack. This is not yet a novel-view renderer and should not be read as a physical x,y,z radiance-volume solve; the current 3D residual axes are view index, image row, and image column. The experiment is therefore a memory-bounded multi-view stack reconstruction with camera-ray metadata, establishing the ingestion and algebraic reconstruction path before the future physically coupled radiance-field step.
The initial DL3DV ledger measured two profiles. The first retained only 56 three-dimensional DCT residual coefficients per color channel. It produced 20.4800 dB PSNR, SSIM 0.448128, LPIPS 0.718625, 96.92% combined hard-zero parameters, 96.88% sparse-tier hard-zero parameters, 193,302,032 measured peak bytes, and 140.1238 ms/view. This sparse run is useful as a stress test, but not as the preferred visual-quality profile. The second profile prioritized image quality by retaining the complete 30×256×256 orthonormal DCT support per color channel. It produced 117.2378 dB PSNR, SSIM 0.99999988, LPIPS 1.0539e−10, 77.50% combined hard-zero parameters, 96.88% sparse-tier hard-zero parameters, 215,813,648 measured peak bytes, and 135–136 ms/view across repeated runs. Dense NumPy tensor export remained disabled; only metrics and PNG frames were written.
The first failed quality-profile attempt exposed a real implementation bottleneck: the four-operand 3D-DCT projection einsum was killed externally during contraction planning/execution despite the preflight estimate. The corrected implementation now computes the separable DCT projection and synthesis axis-by-axis: x projection, y projection, view projection, followed by view, y, and x synthesis. This keeps the DCT stage inside the same memory envelope and makes the quality-profile run reproducible on the local CPU. The TLS matrix-pencil amplitude solve also gained an absolute ridge floor for degenerate rows, so blank or nearly flat scanlines no longer produce singular Gram failures.
DL3DV multi-view stack reconstruction ledger for scene 0853979305f7ecb80bd8fc2c8df916410d471ef04ed5f1a64e9651baa41d7695. Both profiles run under torch.no\_grad() with a 0 byte autograd graph and dense NPZ tensor export disabled. The quality ceiling is an image-fidelity upper bound, not a sparsity claim.
After establishing the quality ceiling, the follow-up Pareto sweep in apps\_industrial\_breakthrough/dl3dv\_pareto\_sweep.py re-ran the same scene while monotonically pruning the 3D-DCT residual support. Each profile used the same 30×256×256 input tensor, K=32 sparse knots, 5 LPIPS views, torch.no\_grad(), disabled dense NPZ export, and the same 2 GiB preflight guardrail. The sweep deliberately keeps the sparse tier fixed, so the measured curve isolates the visual effect of residual support pruning rather than conflating it with a new boundary locator.
The measured curve is more conservative than the optimistic pre-run hypothesis. The 50% DCT-support profile remains high fidelity at 42.7562 dB, SSIM 0.981283, LPIPS 0.002584, and 87.50% combined hard-zero parameters. The 25% profile reaches 92.50% hard-zero parameters but falls to 35.5073 dB, and the 10% profile reaches 95.50% hard-zero parameters but falls to 30.4278 dB. The edge-localization error remains at 0.958244 for every sweep point, confirming that the present sparse tier is not yet carrying enough of the geometric boundary load; additional gains should come from improving the TLS/Hankel conditioning and boundary model rather than from further blind DCT pruning.
Compression sanity check.
The same DL3DV quality-ceiling result also motivates a direct compression audit, because an exact orthonormal residual expansion is not automatically a competitive codec. The script apps\_industrial\_breakthrough/osnr\_compression\_audit.py therefore takes the same 30×256×256 RGB target stack, whose raw unsigned-byte footprint is 5,898,240 bytes, and measures payload size after scalar quantization and np.savez\_compressed entropy compression. Two OSNR-style transform payloads are tested: a global separable 3D DCT over view, row, and column axes, and an independent per-frame 2D DCT. Both use low-frequency support masks and quantized integer coefficients. The audit then decodes the stored coefficients and evaluates PSNR/SSIM against JPEG and WebP encodings at quality 90 using the same source images.
The result is intentionally conservative and negative. At similar bits per pixel, the naive OSNR transform payloads are far below mature image codecs: the 5% 3D-DCT profile reaches only 25.2894 dB at 1.7118 bpp, and the 5% per-frame 2D-DCT profile reaches 26.4132 dB at 1.8927 bpp, while WebP reaches 39.4283 dB at 1.8132 bpp. Increasing OSNR support restores quality but destroys payload efficiency: the 50% per-frame 2D-DCT profile reaches 37.4358 dB, but costs 12.8751 bpp. This confirms that the 117 dB observed-stack ceiling is a completeness result, not a compression claim. A serious OSNR codec would need at least perceptual quantization, coefficient ordering, block or geometry-conditioned prediction, motion/view compensation, and a real entropy coder before it should be compared against JPEG, WebP, AV1, or neural codecs. For the present manuscript, compression is therefore recorded as a promising but unfinished direction rather than the next flagship validation target.
A more appropriate compression target is neural-scene distillation rather than still-image coding. The audit apps\_industrial\_breakthrough/neural\_scene\_distillation\_audit.py tests this narrower claim using the frozen nerfstudio even/odd split. It serializes only the fifteen training key views into quantized OSNR 2D-DCT coefficient packages, decodes those key views, and predicts the held-out views by the same deterministic adjacent-view interpolation rule. This is a lightweight scene-streaming proxy: the payload is a compact mathematical scene package rather than a trained radiance field, and the metric is held-out view quality per transmitted byte.
The first result is a foothold, not a SOTA win. At 5% support and q=0.004, the OSNR keyview package is 234,555 bytes and reaches 22.3882 dB held-out PSNR, slightly above the WebP-keyview stream at 255,070 bytes and 21.9900 dB. An even smaller 2% OSNR package is only 92,228 bytes and still reaches 21.7984 dB. However, the locally available nerfacto CPU pilot checkpoint is 242,859,619 bytes and reaches 27.0830 dB on the same odd-view protocol. Therefore the current OSNR package is dramatically smaller, but not yet quality-competitive with even a reduced NeRF pilot. The next scene-compression experiment would need a true geometry-aware residual package–for example plane-sweep depth support, sparse COLMAP anchors, and view-dependent residual coefficients–before claiming neural-field model compression.
The follow-up audit apps\_industrial\_breakthrough/neural\_scene\_geometry\_package\_audit.py performs exactly this geometry-aware test. For each package profile, it first decodes the transmitted key views and then runs the deterministic COLMAP-pose plane-sweep renderer from those decoded key views into the held-out cameras. The package byte count includes the serialized keyview payload plus a conservative 17,394,552 byte COLMAP camera metadata budget from cameras.bin and images.bin. With uncompressed key views, the geometry package reaches 29.3178 dB in 20.34 MB. More importantly, compressed keyview packages still beat the local nerfacto CPU pilot: WebP key views plus geometry reach 29.0865 dB in 17.62 MB, and the OSNR 50% keyview package reaches 28.9853 dB in 18.98 MB. Compared with the 242.86 MB nerfacto CPU checkpoint at 27.0830 dB, this is a concrete local model-compression win: better held-out PSNR with roughly 12–14× smaller serialized scene state. The limitation is equally clear. The current OSNR keyview transform is not yet the best keyview codec inside the geometry package–WebP remains slightly better at lower payload–so the next OSNR-specific compression gain must come from geometry-conditioned residual coefficients or a more mature entropy-coded spline payload rather than from naive per-frame DCT pruning alone.
The residual-package audit apps\_industrial\_breakthrough/neural\_scene\_residual\_package\_audit.py tests that next hypothesis directly. It computes leave-one-out plane-sweep residuals on the even training views, projects those residuals onto quantized spatial OSNR DCT packets, interpolates the decoded residual packets to the odd held-out views, and adds them to the held-out geometry render. This is deliberately quality-first: it tests 50% and 25% residual support and does not force extreme sparsity. The result is negative. For raw key views, the geometry-only package remains best at 29.3178 dB; adding the best residual packet falls to 29.1727 dB. For WebP key views, geometry-only reaches 29.0865 dB, while the best residual packet falls to 28.9533 dB. The residual stream therefore encodes view-specific plane-sweep errors that do not transfer cleanly from even leave-one-out views to odd held-out views. The practical conclusion is that the current quality bottleneck is not residual coefficient capacity; it is visibility/depth correctness. The scene package should next improve geometry–depth maps, occlusion masks, or multi-source visibility confidence–before adding larger residual payloads.
The spline inverse-problem bridge apps\_industrial\_breakthrough/dl3dv\_tomographic\_radiance\_bridge.py then tests whether the McCann–Donati H⊤H convolution idea can already help the real DL3DV held-out split. The script recolors the 81,120 sparse COLMAP points from even training views, deposits them into a 643 compact spline voxel grid, applies a cubic-B-spline FFT normal solve with ridge 0.005, and renders the regularized radiance grid into the odd held-out cameras. The result is a small but measurable PSNR foothold rather than a finished renderer. Adjacent-view interpolation reaches 22.7803 dB, SSIM 0.611987, and LPIPS 0.150600. The best FFT-tomographic blend uses only 2% of the regularized grid prediction and reaches 22.7955 dB, but SSIM falls to 0.609815 and LPIPS rises to 0.155779. Larger blends degrade quickly. This confirms that the inverse grid contains some held-out radiance signal, while the dominant problem remains visibility-aware measurement construction and occlusion reasoning rather than the speed of the FFT normal solve itself.
Compression sanity check on the same 30-view DL3DV target stack. The audit measures actual serialized payload bytes after quantization and compression. The observed-stack OSNR reconstruction remains a completeness result; these naive transform payloads are not yet competitive with mature codecs.
p0.36linewidthp0.13linewidthp0.13linewidthp0.13linewidthp0.15linewidth@
Scene package profile
Bytes
bpp/eval
Held-out PSNR
Held-out SSIM
Uncompressed key views + interpolation
2,949,120
24.0000
21.9836 dB
0.889469
JPEG key views, quality 90
322,203
2.6221
21.9584 dB
0.888935
WebP key views, quality 90
255,070
2.0758
21.9900 dB
0.889673
OSNR keyview DCT, 5%, q=0.004
234,555
1.9088
22.3882 dB
0.900308
OSNR keyview DCT, 2%, q=0.008
92,228
0.7506
21.7984 dB
0.884918
nerfacto CPU pilot checkpoint
242,859,619
n/a
27.0830 dB
0.818526
First neural-scene distillation audit on the frozen DL3DV even/odd split. The OSNR keyview package is smaller than JPEG/WebP keyview streams at comparable held-out interpolation quality, but it does not yet match the trained nerfacto pilot's PSNR. This supports scene-streaming potential, not a completed neural-field compression result.
Geometry-aware neural-scene package audit. Each compressed-keyview row is decoded before rendering; the deterministic plane-sweep renderer then predicts the odd held-out views from the decoded even views. Under this local CPU-pilot comparison, compact geometry packages are both smaller and higher-PSNR than the available nerfacto checkpoint, while the OSNR-specific keyview transform still trails WebP inside the package.
Geometry-conditioned residual package audit. Residual coefficients are fitted only from even-view leave-one-out geometry errors and then evaluated on odd held-out views. The negative result is informative: residual capacity does not solve the current error mode; visibility and depth correctness dominate.
First real-scene spline-tomographic radiance bridge on the frozen DL3DV even/odd split. The 643 FFT-normal grid solve takes 17.2450 ms; rendering the unoptimized voxel splats takes 848.5788 ms for the held-out stack. The slight PSNR gain at tiny blend weights is useful evidence, but not yet a NeRF-quality renderer.
The next implementation step adds an explicit world-edge proposal layer without modifying the packaged AdaptiveSparseSolver. The module apps\_industrial\_breakthrough/dl3dv\_world\_edge\_atoms.py computes cubic-B-spline derivative magnitudes in each view, selects non-maximum-suppressed gradient pixels, backprojects them through the parsed camera intrinsics and camera-to-world matrices, triangulates adjacent-view ray pairs by closest point of approach, clusters the resulting candidate points, and reprojects those world atoms back into each camera. The DL3DV runner exposes this through --edge\_atom\_mode scanline|world|hybrid; the hybrid mode merges projected world-edge knots with the original scanline knots while preserving the same no-autograd, RAM-guarded execution path.
This first world-edge ablation is intentionally diagnostic rather than presented as an improvement. With 128 edge rays per view, 1024 clustered world atoms, CPA threshold 0.025, cluster radius 0.015, and ridge 10−6, the hybrid projection uses 24.03% projected knots and lowers the scanline-referenced edge delta from 0.958244 to 0.624169. However, the multi-view reprojection error is still 13.5410 pixels, so the injected atoms are not yet selective enough: at 25% DCT support, PSNR falls from 35.5073 dB to 33.6019 dB; at 10% support, PSNR falls from 30.4278 dB to 29.2703 dB. A stricter CPA run with threshold 0.005 and cluster radius 0.005 gives similar quality (33.6076 dB at 25% support) and worse reprojection error (14.3163 pixels). The conclusion is precise: the world-space atom path is now executable and measurable, but adjacent-view CPA alone must be augmented with epipolar-consistency scoring, depth/COLMAP support, or multi-view consensus pruning before it can replace the DCT cushion.
To separate observed-view reconstruction from genuine view generalization, apps\_industrial\_breakthrough/dl3dv\_heldout\_challenger.py implements an interleaved held-out light-field challenge on the same scene. Even-indexed views {0,2,…,28} are the only training/input images; odd-indexed views {1,3,…,29} are held out for metrics. The runner compares three deterministic, no-autograd profiles: adjacent-view linear interpolation, DCT interpolation along the camera sequence, and a ray-kernel OSNR model that fits RGB from sampled training rays (o,d) and evaluates the held-out camera rays directly. This experiment is quality-first and does not impose sparsity pruning on the ray model.
The held-out result is a useful boundary marker rather than a new SOTA claim. Linear neighbor interpolation reaches 22.7803 dB PSNR, SSIM 0.611987, and LPIPS 0.150600 on the fifteen held-out views. DCT view interpolation reaches 21.2918 dB, SSIM 0.534591, and LPIPS 0.141890. The first ray-kernel OSNR profile, using 131,072 sampled training rays and 1024 RBF centers, reaches only 18.0553 dB, SSIM 0.466200, and LPIPS 0.904455; increasing to 262,144 samples and 4096 centers with a broader kernel worsens PSNR to 14.0241 dB. This confirms that the earlier 117 dB DL3DV ceiling is an observed-stack completeness result, not yet a NeRF-style novel-view synthesis result. The next graphics step therefore needs depth-aware or epipolar-consensus geometry rather than a larger ray-only kernel.
The asset audit then found a usable sparse COLMAP reconstruction: nerfstudio/colmap/sparse/0/points3D.bin contains 81,120 points, with matching images.bin and cameras.bin. No depth maps, NumPy geometry arrays, or Gaussian-splat .ply file are cached. The follow-up renderer apps\_industrial\_breakthrough/dl3dv\_colmap\_heldout\_renderer.py therefore tests a deterministic geometry-backed held-out baseline: parse the COLMAP points and image poses, recolor visible points from even training views, z-buffer splat them into the odd held-out cameras, and blend the sparse render with the linear-neighbor fallback. This is still not a trained NeRF or dense 3DGS renderer, but it is the first held-out result in this section that uses actual scene geometry.
The geometry-backed profile beats the interpolation-only baseline. The first refined run used COLMAP poses, training-view recolored points, radius-1 splats, and a 45% geometry blend, reaching 23.6510 dB PSNR. Pushing quality further showed that the limiting artifact is high-frequency splat noise: applying a 3×3 smoothing kernel to the sparse geometry before blending raises the held-out score to 23.9984 dB PSNR and SSIM 0.684186, compared with 22.7803 dB and SSIM 0.611987 for linear neighbor interpolation. The next visibility-aware recoloring pass resolves each training view with a deterministic nearest-depth test before sampling point colors; this removes occluded color assignments and raises the peak to 24.1063 dB PSNR and SSIM 0.686748 at a 67% geometry blend. Adaptive surfel splats increased projected coverage from 65.82% to 75.87%–82.73%, but did not improve PSNR because the additional coverage carried too much color and visibility noise. Finally, apps\_industrial\_breakthrough/dl3dv\_osnr\_residual\_renderer.py fits residuals only on even training views after subtracting the visibility-clean geometry anchor, then evaluates the residual correction on odd held-out views. A small linear residual correction raises the peak to 24.2145 dB PSNR, while higher-capacity DCT residuals underperform because they begin to inject view-dependent residual noise.
The stronger Track-A result is the dense, deterministic plane-sweep renderer apps\_industrial\_breakthrough/dl3dv\_plane\_sweep\_renderer.py. It uses the same COLMAP camera convention as the successful sparse renderer and warps the nearest even training views into each odd held-out camera over a bounded depth lattice. A diagnostic depth pass over the COLMAP points gives average visible-depth quantiles q50≈6.31, q90≈12.06, and q95≈13.68, explaining why the 0.4–12.0 range outperforms the initial 0.4–8.0 sweep. The first dense pass, with 64 depth planes, two source views, and a 5% residual correction, reached 26.3525 dB PSNR. The v2 pass then replaces hard winner-take-all depth selection with soft plane aggregation, introduces patch-averaged photometric costs, and removes residual correction once it becomes detrimental. This staged refinement raises held-out quality to 29.1731 dB PSNR at 192 depths. The v3 ablation shows that confidence-fused multi-pair sources, bilateral edge-aware cost aggregation, and local depth refinement do not improve PSNR on this scene; the best path is still nearest two-view soft aggregation with a uniform patch cost and more depth support. A final high-depth push raises the deterministic ceiling to 29.4187 dB PSNR, SSIM 0.898331, and LPIPS 0.099257 using 512 linear depth planes, a 13×13 patch cost, two source views, and no neural training or autograd. The gain over 320 planes is measurable but small relative to the added compute, so the 320-plane profile remains the practical operating point while the 512-plane profile records the quality ceiling of this deterministic renderer.
The bridge from the controlled multi-ray FRI studies to real COLMAP imagery is apps\_industrial\_breakthrough/dl3dv\_multiview\_fri\_geometry\_bridge.py. The runner keeps the same even/odd held-out split, but augments the depth-lattice score with cubic-B-spline derivative edge coherence from the even training views. At each candidate depth, RGB disagreement between warped source views is combined with a source-edge disagreement term and a small coherent-edge reward; this is a real-image analogue of the controlled multi-ray residual selection loop, but still avoids target-view leakage. A diagnostic no-edge four-source setting reaches only 27.5992 dB at 1602 resolution, while the nearest-two-source edge-coherent setting reaches 30.1786 dB at 1602. At the standard 2562 DL3DV size, the 256-plane edge bridge reaches 29.8085 dB PSNR, SSIM 0.913584, and LPIPS 0.070140 in 21.481 s for the fifteen held-out views. This exceeds the earlier deterministic 512-plane photometric ceiling in all three perceptual metrics while using half the number of depth planes, and it also improves over the trained liquid-residual polishing row. The edge bridge is therefore the strongest current geometry result: it is autograd-free, uses only observed training views and camera geometry, and demonstrates that sparse edge evidence transfers from the controlled Haouchat setting into real multi-view reconstruction.
The liquid-neural-network follow-up deliberately tests a different question: whether a very small continuous-time residual corrector can remove systematic plane-sweep artifacts without replacing the deterministic geometry engine. The runner apps\_industrial\_breakthrough/dl3dv\_liquid\_residual\_challenger.py first computes leave-one-out plane-sweep renders on the even training views, fits only the residual image with a tiny exact liquid cell, and evaluates the learned residual on the odd held-out views. Each hidden unit follows the closed-form multi-synapse liquid update xk+1=γk(xk−ω+∑sfs(ck)∑sAsfs(ck))+ω+∑sfs(ck)∑sAsfs(ck),γk=exp(−ω−s∑fs(ck)), where the conditioning vector ck contains the plane-sweep RGB estimate, the linear-view fallback, confidence, normalized pixel coordinates, normalized view index, and the plane-sweep/fallback discrepancy. This hybrid no longer has the zero-autograd property during fitting, but the learned component is intentionally small and residual-only. At 320 depth planes it improves the held-out result from 29.3178 to 29.3994 dB and reduces LPIPS from 0.102406 to 0.096436. At the 512-plane quality setting, a 64-state liquid residual improves the plane-sweep ceiling from 29.4187 to 29.5291 dB, SSIM from 0.898331 to 0.901776, and LPIPS from 0.099257 to 0.093192. The gain is consistent but modest; it supports liquid residuals as a polishing layer, not as a substitute for denser visibility-aware geometry.
To make the SOTA comparison falsifiable rather than rhetorical, apps\_industrial\_breakthrough/dl3dv\_external\_baseline\_protocol.py freezes an external baseline protocol for the same scene, view count, resolution, and even/odd split. The exporter writes a nerfstudio\_even\_odd\_256 dataset with 30 resized frames whose basenames explicitly contain train or eval, transforms.json, transforms\_train.json, transforms\_eval.json, and split.json, plus explicit nerfacto, splatfacto, and ns-eval command lines. The regenerated protocol report also audits the local device stack. The machine itself supports Apple Metal: direct shell testing shows that the project venv can allocate device='mps' tensors. However, the current .venv\_nerfstudio Python 3.10 environment reports mps=False, and fresh Homebrew Python 3.11/3.13/3.14 torch-2.12 test environments failed the same MPS runtime gate in this execution context. A targeted follow-up repinned the Python 3.10 nerfstudio environment to torch-2.5.1/torchvision-0.20.1 and then to torch-2.3.1/torchvision-0.18.1; both variants still failed torch.ones(1, device='mps') with the same PyTorch OS-version gate. Python 3.14 cannot be used directly for nerfstudio because Open3D has no compatible wheel. A local CPU-only nerfacto pilot with 1000 iterations and 1024 rays per batch reaches 27.0830 dB PSNR, SSIM 0.818526, and LPIPS 0.176857 on the odd held-out views. This remains only a protocol-validation point. The current comparison target for the external run is no longer the older plane-sweep row but the multi-view FRI edge bridge: 29.8085 dB PSNR, SSIM 0.913584, and LPIPS 0.070140. A final trained-NeRF/3DGS comparison therefore requires either a CUDA-capable nerfacto/splatfacto run or a local nerfstudio environment whose PyTorch installation is first verified to allocate MPS tensors.
The consolidation script apps\_industrial\_breakthrough/dl3dv\_heldout\_sota\_ledger.py collects the scattered held-out outputs into one ranked ledger. This makes the current competitive status unambiguous. The naive ray-kernel OSNR row is not competitive, reaching only 18.0553 dB, so the project cannot claim that coordinate-ray regression alone beats NeRF. The geometry-aware rows are different: the deterministic edge-consistent bridge reaches 29.8085 dB without neural training, the deterministic 512-plane sweep reaches 29.4187 dB, and the tiny residual liquid polishing layer reaches 29.5291 dB. These exceed the available nerfacto CPU pilot at 27.0830 dB on the identical split, but the comparison remains a local pilot until a full GPU nerfacto/splatfacto run is executed. The ledger therefore defines the next hard target: retain the held-out quality advantage while cutting the plane-sweep latency and replacing the external CPU pilot with a complete CUDA baseline.
The adaptive-depth follow-up apps\_industrial\_breakthrough/dl3dv\_adaptive\_depth\_bridge.py tests whether this latency can be reduced by replacing the uniform depth lattice with a two-stage proposal scheme. The renderer first runs a coarse edge-aware bridge, then evaluates local per-pixel depth offsets around the selected depth and optional sparse COLMAP point-depth proposals. The first single-depth adaptive profile reduces the held-out-stack latency but loses quality: a 64+17+9 proposal profile reaches 28.7523 dB in 6830.4 ms, while a denser 128+17 profile reaches 28.8875 dB in 10598.5 ms. A top-K proposal-recall upgrade is stronger. Keeping the best three coarse depth hypotheses per pixel and refining each with seven local offsets reaches 29.0585 dB at 64 coarse depths, 29.3291 dB at 128 coarse depths, and 29.5615 dB at 192 coarse depths. The 192-depth top-K profile is faster than the uniform bridge (15882.1 ms versus 21481.0 ms) and improves perceptual metrics (SSIM 0.918645, LPIPS 0.064826), but it still trails the uniform bridge in PSNR. We also tested a sparse COLMAP visibility-consistency penalty that rejects candidates landing behind the source view's nearest sparse point depth. Even a weak penalty (visibility\_weight=0.02, visibility\_eps=0.12) reduces the top-K profile to 29.3382 dB and increases latency to 21878.8 ms. Thus sparse COLMAP depth is useful as a proposal hint but too noisy as a direct occlusion veto. Since the adaptive residual variant again reduces PSNR, the dominant error is not a missing smooth residual; it is missed depth/visibility proposal quality. The next DL3DV improvement must therefore use stronger epipolar source-edge intersections or a dense source-depth confidence field before enforcing bidirectional consistency.
The final graphics push in this cycle tests whether the EGGROLL low-rank evolution-strategy idea can serve as a minimal neural visibility selector without turning the method into a full radiance MLP. The script apps\_industrial\_breakthrough/dl3dv\_eggro\_visibility\_scorer.py renders three candidate stacks—linear interpolation, the top-K adaptive bridge, and the uniform FRI edge bridge—then trains a tiny low-rank antithetic ES scorer on even-view leave-one-out pixels. The scorer sees only candidate confidence, candidate disagreement, and local edge features; it blends candidate RGB values but does not synthesize new color. With 100 ES iterations, population 40, and rank 8, the learned blend reaches 29.6504 dB, SSIM 0.914176, and LPIPS 0.082802 on the odd held-out views. This improves over the candidate bridge rendered inside the same joint run, but it still does not beat the frozen deterministic FRI edge bridge at 29.8085 dB or the top-K profile's LPIPS. A subsequent affine color calibration overfits the leave-one-out training views and falls to 26.4261 dB. We then expanded the scorer with an OSNR feature lift: local pooled spline-style neighborhoods, candidate-rank channels, luminance/chroma terms, and Fourier coordinate features. The low-resolution smoke improves, but the full 2562 held-out run reaches only 29.6202 dB, below the simpler ES scorer. The conclusion is narrow but useful: a tiny ES visibility scorer is plausible, but raw scorer capacity over finished RGB candidates is not yet a SOTA-grade substitute for better geometry proposals, epipolar edge evidence, or dense visibility/depth reasoning.
To isolate the precise difference from NeRF-style training, apps\_industrial\_breakthrough/dl3dv\_osnr\_volume\_renderer.py implements a shared OSNR volume with actual alpha compositing. Sparse COLMAP points are recolored from even training views, deposited into a compact 643 radiance/density grid, regularized by the FFT B-spline/Laplacian normal solve, and rendered into odd cameras by ray-marching 64 samples per ray. This borrows NeRF's volumetric visibility equation but not its MLP. The result exposes the missing ingredient. The pure alpha-composited volume reaches only 11.4149 dB, SSIM 0.376490, and LPIPS 0.911937. Blending 1% of the volume render with 99% interpolation gives a tiny PSNR foothold at 22.7898 dB, but no meaningful view-synthesis gain. Thus the NeRF advantage is not merely alpha compositing; it is direct optimization of a dense occupancy/transmittance field from multi-view ray losses. Sparse COLMAP deposition plus smooth FFT regularization does not provide enough empty-space or surface evidence to create that field.
We then tested the obvious next bridge, apps\_industrial\_breakthrough/dl3dv\_osnr\_pinn\_density\_renderer.py: an OSNR-PINN density volume whose color and density spline-grid coefficients are initialized from the same COLMAP/FFT volume but fitted against even-view rays with a differentiable volume-rendering loss, total-variation regularity, sparsity, an initialization anchor, and an eikonal-style surface-gradient proxy. At 1282 resolution with a 483 grid, 48 samples per ray, and 120 training iterations, the training photometric loss drops from 0.118027 to 0.025061. However, the held-out pure density render reaches only 15.7504 dB, SSIM 0.327521, and LPIPS 0.850917; the best 5% blend with interpolation reaches 24.7899 dB, slightly below the interpolation baseline at the same resolution (24.8264 dB). This negative result is useful: weak physics-style regularity is insufficient. A NeRF-competitive OSNR geometry model needs either dense depth/occupancy supervision, stronger epipolar surface constraints, or an optimizer that directly solves the nonlinear visibility ambiguity rather than only smoothing sparse COLMAP evidence.
The denser initializer apps\_industrial\_breakthrough/dl3dv\_dense\_depth\_volume\_renderer.py then replaces sparse COLMAP deposition with leave-one-out FRI depth estimates for all even training views. At 1602 resolution, the runner backprojects 257,276 confidence-weighted dense depth samples into a shared 723 spline/FFT radiance-density volume and renders held-out odd views with 72 alpha samples per ray. This is a stronger geometry initializer, but the single smoothed volume still fails: the pure dense-depth alpha volume reaches 12.9803 dB, SSIM 0.325629, and LPIPS 0.652960; the best 2% blend reaches 24.1111 dB, essentially tied with but not better than the same-resolution interpolation baseline (24.1113 dB). The failure mode is now clear. The deterministic depth bridge succeeds because it keeps view-conditioned depth hypotheses and local source evidence alive until rendering. Collapsing those hypotheses into one global smoothed density grid discards too much visibility structure. The next NeRF-facing OSNR attempt should therefore preserve surface/depth hypotheses explicitly (for example as layered splines or surfel sheets) or learn opacity with direct multi-view transmittance constraints, not by smoothing depth maps into a volumetric average.
The layered follow-up apps\_industrial\_breakthrough/dl3dv\_dense\_surfel\_renderer.py confirms this diagnosis. Instead of averaging the dense FRI depths into a volume, it keeps the 361,290 backprojected training-view samples as explicit weighted surfels and z-buffers them into held-out views. At 1922 resolution with 160 depth planes, the same-resolution interpolation baseline is 23.5474 dB. The best layered surfel profile, radius 1 with a 3×3 smoothed 10% blend, reaches 23.6577 dB and SSIM 0.659134. Pure surfels remain poor (18.6729 dB) because coverage, view-dependent color, and visibility ordering are still imperfect, but this is the first volumetric/surface variant in this sequence to improve over its same-resolution interpolation baseline. The conclusion is narrow: preserving layered surface hypotheses is directionally correct, while collapsing geometry into a single smoothed voxel field is not.
We also tested whether the EGGROLL-style low-rank selector could turn the surfel signal into a stronger visibility model. The scorer in apps\_industrial\_breakthrough/dl3dv\_surfel\_visibility\_scorer.py is trained on even-view leave-one-out pixels over five candidates: interpolation, raw surfels, smoothed surfels, confidence-blended surfels, and the fixed 10% smoothed surfel blend. This did not improve the frontier. The fixed blend remains best at 23.6577 dB, while the learned ES selector reaches only 23.0595 dB and the affine-calibrated selector falls to 22.2272 dB. The current candidate scorer therefore overfits or selects the wrong corrections.
The next quality-first probe replaces the shallow ES selector with a spline-native network in apps\_industrial\_breakthrough/dl3dv\_deep\_spline\_surfel\_network.py. The model uses two dense hidden layers, but replaces standard pointwise activations by learnable per-channel compact spline activations. It receives the same explicit surfel candidate stack, local OSNR feature lift, confidence fields, candidate disagreement, coordinates, and candidate RGB values, and is fitted only on even-view leave-one-out pixels. At 1922 resolution, 160 depth planes, 64 hidden channels, 25 activation knots, and 60 epochs over 131,072 samples, the deep spline selector reaches 23.7305 dB and SSIM 0.663583, improving over the same-resolution linear baseline (23.5474 dB) and fixed surfel blend (23.6577 dB). However, LPIPS worsens to 0.265901, and the total diagnostic latency rises to 52.433 s. This is a useful but bounded result: additional spline-network capacity can extract a little more PSNR from the preserved surfel hypotheses, but it does not repair missing visibility, coverage, or view-dependent color evidence. A stronger NeRF/SIREN competitor needs explicit source-view agreement, occlusion ordering, normal/facing estimates, epipolar edge intersections, or dense confidence geometry before adding still larger learned layers.
We therefore added those first-order evidence channels directly in apps\_industrial\_breakthrough/dl3dv\_evidence\_spline\_surfel\_network.py. The renderer records per-pixel surfel confidence, color-consistency variance, depth-coherence variance, projected depth-edge strength, and a crude facing proxy derived from the rendered depth gradient. These measurements define conservative and stronger evidence-gated surfel blends, and the spline network is constrained to choose among these candidates with no free RGB residual by default. This improves the safety of raw surfel corrections but does not change the frontier. At the same 1922 diagnostic setting, the fixed 10% smoothed-surface blend remains best at 23.6576 dB and LPIPS 0.148888. The evidence-conservative blend reaches 23.6183 dB but slightly improves LPIPS to 0.146167, while the evidence-aware spline selector falls to 23.5142 dB despite a higher SSIM of 0.670806. This is the strongest negative constraint so far: simple confidence, variance, and facing features are not enough to infer NeRF-grade visibility. The next real push must create better geometry hypotheses themselves, specifically multi-view epipolar edge intersections, source-view agreement at candidate depths before surfel projection, and occlusion-ordered layered surfaces.
To remove ambiguity about benchmark quality scales, the next runner, apps\_industrial\_breakthrough/nerf\_synthetic\_spline\_benchmark.py, switches to the canonical NeRF Synthetic/Blender data layout: transforms\_train.json, transforms\_test.json, RGBA images composited over white, and OpenGL camera-to-world matrices. The script supports real scenes such as lego or chair, but also creates a tiny generated Blender-style sphere scene for camera-convention and memory smoke testing. This benchmark records three profiles: nearest-pose image transfer, deterministic OSNR-style plane sweep, and a small deep spline ray network with Fourier ray features and learnable compact spline activations. On the generated sphere diagnostic (24 train views, 8 held-out views, 962 resolution, 96 depth planes), nearest-pose transfer reaches 24.0446 dB, SSIM 0.869139, and LPIPS 0.016128. The deterministic plane sweep reaches only 19.1317 dB but a higher SSIM of 0.887576, exposing a cost/depth ambiguity despite correct Blender projection. The deep spline ray network fits the training rays down to MSE 2.5344×10−3, but collapses on held-out views at 6.2504 dB and LPIPS 0.694466. This reproduces the DL3DV lesson under a cleaner benchmark convention: a ray-only coordinate network, even with spline activations, is not a NeRF replacement. The next canonical run should therefore use the actual NeRF volume-rendering transmittance equation with spline-parameterized density/radiance, then evaluate on real lego/chair metrics against published NeRF-family scores.
That volume-rendering step is implemented in apps\_industrial\_breakthrough/nerf\_synthetic\_spline\_volume\_renderer.py. The model keeps NeRF's alpha-compositing equation but replaces the MLP with a compact trilinear spline grid storing density and color coefficients. Training is still by ray loss, so this is not an autograd-free solver; it is a controlled test of whether the missing ingredient is the transmittance model rather than the spline representation. On the same generated sphere benchmark at 962 resolution, a 643 spline volume with 80 samples per ray and 800 AdamW iterations drives the training loss to 9.064×10−4 and reaches 28.6898 dB PSNR, SSIM 0.963119, and LPIPS 0.021464 on held-out views. This beats nearest-pose transfer by 4.6452 dB and improves substantially over both plane sweep and the ray-only spline network. The result is the first positive canonical-NeRF-path evidence: OSNR should compete through spline-parameterized density/radiance under the correct volume-rendering operator, not through direct ray-to-RGB regression.
We then moved from the generated smoke scene to real canonical NeRF Synthetic Blender scenes. To keep the diagnostic memory-bounded and comparable to the smoke run, we downloaded only the per-scene lego and chair archives from the NerfBaselines data mirror, evaluated 24 training views and 8 held-out test views at 962 resolution, and kept the same 643 trilinear spline grid, 80 samples per ray, 800 AdamW iterations, and 87.60 MiB estimated active footprint. On lego, nearest-pose transfer reaches only 14.6577 dB, SSIM 0.650337, and LPIPS 0.180211, while the spline volume reaches 22.4265 dB, SSIM 0.891044, and LPIPS 0.068127. On chair, nearest-pose transfer reaches 22.6878 dB, SSIM 0.883803, and LPIPS 0.128070, while the spline volume reaches 29.4390 dB, SSIM 0.954310, and LPIPS 0.062394. These real-scene results show that the volume-rendering mechanism transfers beyond the generated sphere and produces large held-out gains over image transfer. They are not yet SOTA: the current model is still a first-order trilinear grid without view-dependent radiance, hierarchical sampling, cubic/exponential spline interpolation, sparse occupancy priors, or closed-form color updates. The next technical bottleneck is therefore not whether to use the NeRF operator, but how to replace the primitive trilinear lattice by a higher-order operator-spline volume with better density localization and view-dependent color.
We next replaced the primitive trilinear sampler by an explicit tensor-product cubic B-spline interpolation path. Each query point now accumulates over a compact 4×4×4 support stencil using the cardinal cubic weights, rather than the 2×2×2 trilinear hat stencil. This tests whether higher-order spline regularity alone improves the NeRF-style volume without changing the loss, grid size, camera model, or radiance parameterization. Because the cubic support is eight times larger, we used a matched reduced ray budget (64 samples per ray and 1024 rays per batch) and ran both cubic and linear controls. On chair, cubic interpolation improves the matched-budget PSNR from 27.4948 dB to 28.2934 dB and SSIM from 0.934462 to 0.947217, but worsens LPIPS from 0.090736 to 0.131908 and is roughly an order of magnitude slower. On lego, cubic improves the matched-budget PSNR from 21.8256 dB to 22.0616 dB and SSIM from 0.876625 to 0.882167, but again worsens LPIPS from 0.095020 to 0.148289. The conclusion is therefore nuanced: higher-order tensor-product splines do help global least-squares-style reconstruction at fixed stochastic ray budget, but regularity alone is not the missing NeRF/SIREN ingredient. The model still needs sharper density localization, hierarchical occupancy sampling, and view-dependent radiance before it can approach published NeRF-family quality.
The strongest geometry result comes from replacing diffuse volumetric density by explicit spline-surface intersections. Inspired by the closed-form convolution/Gram acceleration used in spline snake resampling, we implemented apps\_industrial\_breakthrough/nerf\_synthetic\_spline\_surface\_renderer.py. The diagnostic represents the sphere geometry as a compact tensor-product cubic spline surface
s(u,v)=i,j∑cijβ3(Mu−i)β3(Nv−j)
and solves ray intersections by local Newton updates on the three unknowns (u,v,t):
s(u,v)−o−td=0.
More generally, let βα denote a compact exponential or polynomial spline generator with pole vector α and support length equal to the number of poles. A tensor-product spline surface is
Because the support of βα is compact, only a small stencil of coefficients contributes to each (u,v) evaluation. For cubic polynomial splines this stencil is 4×4 for a surface, while for an order-Pu by order-Pv exponential spline it is Pu×Pv. This is the surface analogue of the convolution/Gram trick used in spline resampling: all repeated products between basis functions and derivative basis functions can be pretabulated as local functions of fractional coordinates, and candidate patches can be culled by compact support before solving eqref(eq:spline-surface-newton). The expensive global scene query is therefore reduced to a small number of local 3×3 systems.
The boundary conditions are determined by the topology of the parameter domain:
Rectangular open patches. For a surface patch over [0,1]×[0,1], both parameters are non-periodic. Newton updates are clamped or damped to keep (u,v) inside the valid domain, and only basis functions whose support overlaps the rectangle are evaluated. This covers trimmed sheets, local surface charts, and open spline patches used for piecewise object shells.
Cylindrical topology. For a cylinder-like surface, one parameter is periodic and the other is open. Typically u∈R/Z wraps around the circumference, while v∈[0,1] remains clamped along the height. The coefficient index i is evaluated modulo M, but j uses boundary-aware open support. Newton updates wrap u←umod1 and clamp v.
Toroidal topology. For a torus-like surface, both parameters are periodic: (u,v)∈(R/Z)2. Both coefficient indices are circular, the Gram matrices are block-circulant, and the support search can be diagonalized or accelerated by FFT-style periodic convolution. Newton updates wrap both parameters.
Spherical topology. A sphere is periodic in longitude but singular at the poles. The practical implementation used here treats u as periodic and v∈[0,1] as a clamped latitude coordinate, with duplicated/regularized polar control rows. A more invariant construction can use multiple overlapping charts, such as two or six rectangular charts, to avoid polar degeneracy. In either case, the intersection equation remains eqref(eq:spline-surface-root); only the index wrapping and chart transition rules change.
For closed surfaces the first positive root along the ray is selected, while for multi-layer or self-occluding surfaces the renderer evaluates all candidate local roots and keeps the smallest valid t after residual and normal-facing checks. This gives an explicit alternative to NeRF's volumetric opacity integral: geometry is stored as a low-dimensional spline manifold, visibility is resolved by root ordering, and radiance can be attached to the surface as a second tensor-product field ρ(u,v,d) rather than diffused through a dense 3D volume. This is not a full unknown-scene method yet: the surface family is known and the initialization uses the sphere's analytic support. It is nevertheless a critical density-localization experiment because the renderer evaluates only compact surface support instead of fitting an opaque 3D density field. On the same generated NeRF Synthetic sphere benchmark, a coarse 32×17 surface reaches 24.9397 dB, SSIM 0.965720, and LPIPS 0.008263; a 64×33 surface reaches 31.0490 dB, SSIM 0.991639, and LPIPS 0.002731; and a 96×49 surface reaches 50.8061 dB, SSIM 0.999790, and LPIPS 0.000023 in 530.8 ms at only 16.29 MiB estimated active footprint. This confirms the central geometric hypothesis: when the shape class can be represented explicitly, tensor-product spline surfaces plus direct ray intersection can outperform diffuse volumetric fitting by a large margin. The next research problem is to infer such surfaces from multi-view data, using edge/epipolar evidence and occupancy fields, rather than assuming them.
The first unknown-scene bridge experiment is apps\_industrial\_breakthrough/nerf\_synthetic\_occupancy\_shell\_renderer.py. It trains the same spline density/radiance volume on chair, then attempts to collapse the learned opacity field into a single explicit shell sample per held-out ray. We tested three extraction rules: maximum transmittance weight, first alpha-threshold crossing, and expected-depth projection. This is the direct test of whether a NeRF-style learned density can be turned into an OSNR-style surface renderer without first improving the geometry prior. The result is negative but informative. The alpha-composited spline volume repeats the previous 29.4390 dB, SSIM 0.954310, LPIPS 0.062394 result. The best shell collapse, maximum weight, falls to 24.2233 dB, SSIM 0.880552, and LPIPS 0.103145; first-alpha reaches only 21.5798 dB, and expected-depth reaches 22.4439 dB. The density active ratio remains 0.998177, showing that the learned field is still a diffuse opacity cushion rather than a localized surface shell. This explains why the explicit sphere-surface experiment succeeds while direct shell extraction from the primitive volume fails: explicit spline geometry is powerful, but the current volume training objective does not yet produce extractable geometry. The next step must add an occupancy/surface regularizer, multi-view depth agreement, or an edge-driven shell proposal before collapsing to tensor-product patches.
We therefore recast the Blender held-out problem as a spline-tomographic inverse problem rather than as pure coordinate-network fitting. The implementation is apps\_industrial\_breakthrough/nerf\_synthetic\_silhouette\_volume\_solver.py. It uses the NeRF Synthetic RGBA alpha channel as an explicit silhouette measurement and solves the geometry stage before the radiance stage. Let σc(x)≥0 be the compact-support spline density volume with coefficients c, and let ri(t)=oi+tdi be a camera ray. The opacity forward operator is
H(c)i=1−exp(−∫tmintmaxσc(ri(t))dt),
which is the nonlinear analogue of the tomographic projector Hc used in the spline CT and cryo-EM papers. The first stage estimates occupancy by minimizing a balanced foreground/background silhouette data term with positivity built into σc=softplus(c~):
This is the NeRF-facing counterpart of the constrained regularized weighted-norm reconstructions of Nilchian and Donati: the unknown is a spline coefficient volume, the data term is a ray projection model, and the priors enforce support, positivity, sparsity, and bounded variation. After the density stage, the density is frozen and a separate compact spline color volume is fitted through the standard alpha compositing integral. This cleanly separates geometry recovery from radiance fitting and prevents the color loss from using diffuse density as an unrestricted numerical cushion.
The result is a useful diagnostic. On lego, 50 training views, 8 held-out views, 962 resolution, a 643 grid, 80 samples per ray, and 500+500 density/color iterations produce held-out alpha IoU 0.944604 and reduce the density active ratio from the RGB-only volume's roughly 0.997 to 0.418766. The RGB score is 22.6041 dB, SSIM 0.896264, and LPIPS 0.076846. On chair, the same protocol yields alpha IoU 0.937057, active density 0.341656, and 28.6128 dB, SSIM 0.956229, LPIPS 0.064510. A quality-relaxed Chair pass with weaker TV/sparsity and a longer color stage reaches 28.9733 dB, SSIM 0.957624, LPIPS 0.058525, alpha IoU 0.937829, and active density 0.307980. The interpretation is precise: silhouette tomography fixes the diffuse-geometry failure and creates a compact, visible occupancy field, but it does not yet beat the RGB-only spline volume in PSNR. The remaining bottleneck is surface-aware radiance assignment and visibility, not silhouette geometry. The next NeRF-facing solver should therefore use the recovered occupancy field as an initialization/preconditioner for an adjoint or variable-projection radiance solve, or convert the high-confidence occupancy boundary into explicit tensor-product spline patches.
Silhouette-constrained spline-volume inverse reconstruction on NeRF Synthetic held-out views. The alpha reconstructions show that the occupancy inverse problem localizes the object support accurately (Lego IoU 0.944604, Chair IoU 0.937829), while the RGB images expose the remaining surface-radiance and visibility bottleneck.
The next NeRF-facing bridge is apps\_industrial\_breakthrough/osnr\_nerf\_spline\_mlp.py, which keeps NeRF's alpha-compositing and hierarchical coarse/fine ray sampling but replaces the plain coordinate MLP by a compact multiresolution spline-grid feature field. The important engineering correction is that compact support must be used as a local-control mechanism, not as a dense expanded positional feature vector. Early variants with dense spline encodings and spline activations improved the matched Fourier/ReLU baseline but were prohibitively slow. The current quality-first configuration samples one shared compact cubic spline grid, feeds the local spline features to split density/color heads, disables dense spline encodings by default, and runs on Apple Metal through torch.device='mps'. On the generated sphere diagnostic, a 4-level grid with base resolution 6 and 6 features per level reaches 22.6470 dB, SSIM 0.859884, and LPIPS 0.112952 at 482 resolution, while the matched Fourier/ReLU profile reaches only 18.9988 dB. On real NeRF Synthetic lego, the same compact-grid mechanism transfers: at 642 resolution, 24+16 samples per ray, 24 training views, 8 held-out views, and 2000 MPS iterations, the shared-grid OSNR profile reaches 21.2960 dB, SSIM 0.858446, and LPIPS 0.069849, versus 20.4836 dB, SSIM 0.830932, and LPIPS 0.120454 for the matched Fourier/ReLU model. Increasing MLP width from 64 to 128 hidden channels does not help, and raw grid scaling to base resolution 8 or 5 levels trades PSNR for perceptual metrics rather than producing a clean improvement. A light ray-geometry concentration prior is more useful: with entropy weight 10−4 and depth-variance weight 10−5, the 3000-step Lego run reaches 21.6605 dB, SSIM 0.873037, and LPIPS 0.062367, while the matched Fourier/ReLU baseline in the same run reaches 21.0730 dB, SSIM 0.846366, and LPIPS 0.099627. Scaling to 962 shows that the 4-level grid underfits perceptual detail (21.2913 dB, SSIM 0.847601, LPIPS 0.149864), but adding a fifth compact grid level recovers the high-resolution frontier. With 5000 MPS iterations, the 962 five-level OSNR profile reaches 21.9754 dB, SSIM 0.869955, and LPIPS 0.083635 without reintroducing dense spline features; extending the same run to 8000 iterations lowers held-out PSNR to 21.8090 despite lower training loss, indicating overfitting or stochastic ray-sampling mismatch. Increasing view Fourier frequencies to 10 worsens the same setting to 21.0993 dB, doubling the ray batch to 1536 reaches only 21.5024 dB, lowering the learning rate to 3×10−4 reaches only 21.4863 dB, and re-enabling compact spline activations reaches only 21.3165 dB while increasing training time beyond 1000 s. The stronger quality lever is camera coverage: increasing the Lego training set from 24 to 50 views at the same 962 five-level setting raises the OSNR profile to 22.5796 dB, SSIM 0.883554, and LPIPS 0.088734 at 5000 iterations, 23.0810 dB, SSIM 0.896646, and LPIPS 0.077989 at 8000 iterations, and 23.7099 dB, SSIM 0.909780, and LPIPS 0.066458 at 12000 iterations. Using all 100 training views with only 5000 iterations improves SSIM to 0.886317 but lowers PSNR to 22.4800 and LPIPS to 0.094478, suggesting that the fixed update budget is then spread too thinly across cameras. A direct 1282 scaling run with the 50-view, five-level configuration reaches 23.1842 dB, SSIM 0.887382, and LPIPS 0.132941. Adding a sixth grid level improves the 1282 perceptual score to LPIPS 0.109930 and SSIM to 0.889882, but leaves PSNR essentially unchanged at 23.1901 dB while increasing training time to 834.1 s. At 962, the six-level Lego profile is a perceptual/detail tradeoff rather than a universal improvement: PSNR drops from 23.7099 to 23.4193 and SSIM from 0.909780 to 0.903694, but LPIPS improves from 0.066458 to 0.058955. A stronger Lego improvement comes from decoupling opacity and radiance support: enabling a separate compact color grid while keeping the five-level density grid raises the 962, 50-view, 12000-step result to 24.0426 dB, SSIM 0.916105, and LPIPS 0.055512. Combining separate color support with a sixth level gives the best Lego LPIPS so far, 0.046647, but drops PSNR to 23.7521 dB and takes 1629.4 s to train; this makes it a quality-ceiling/perceptual point, not the efficient frontier. We also implemented edge-weighted ray sampling, mixing CPU-selected high-gradient rays with uniform MPS batches to avoid a large-vector MPS multinomial failure. A 50% edge-biased mixture on the five-level separate-color Lego model improves LPIPS slightly to 0.053491 but lowers PSNR/SSIM to 23.8910 dB and 0.910451, while a gentler 25% mixture falls further to 23.5742 dB, SSIM 0.909236, and LPIPS 0.056093. Static image-gradient sampling is therefore not the correct hard-ray policy; it prioritizes apparent silhouettes before the model has estimated which rays are actually underfit. The successful sampler is residual-driven: after a 3000-step uniform warmup, replacing 25% of each batch with the highest-error rays from a 2× no-gradient candidate pool raises Lego to a new quality frontier of 24.4613 dB, SSIM 0.916640, and LPIPS 0.041176. A cheaper 1.25× candidate pool reaches 24.2211 dB, SSIM 0.916619, and LPIPS 0.044676 in 1596.4 s, preserving most of the perceptual gain while reducing runtime. The same residual-mined policy transfers to chair, raising the separate-color frontier from 29.8809 dB to 30.8454 dB, SSIM 0.961629, and LPIPS 0.047816, which also beats the previous six-level shared-grid Chair LPIPS. These two scenes confirm that model-aware hard-ray selection is a general OSNR-NeRF mechanism; its cost shows that the next engineering problem is a cheaper cached or amortized residual map rather than more model capacity. Chair also benefits from both earlier capacity mechanisms, but in different ways: a sixth shared grid level improves perceptual quality to LPIPS 0.048553 and raises PSNR to 29.6390 dB, while a separate five-level color grid gives 29.8809 dB and 0.961321 SSIM with LPIPS 0.053649. The efficient frontier is therefore no longer a single grid-depth setting: decoupled radiance support and residual-driven evidence selection are the strongest general quality levers, while extra local spline scale is a perceptual/detail lever whose value depends on scene content. The limiting factor is not angular encoding bandwidth, stochastic batch noise, learning-rate instability, or pointwise activation expressivity; it is the amount and organization of geometric/radiance evidence available to the compact field. The current lesson is specific: the spline advantage is real when it is implemented as compact local grid control with split radiance/opacity heads, separate radiance support, and model-aware hard-ray selection; larger dense MLPs, dense spline feature expansion, and simply adding more depth samples are not the path forward.
Reproducible residual-mined OSNR-NeRF protocol.
The residual-mined experiments use the same held-out NeRF Synthetic split throughout: 50 training views, 8 test views, 962 render resolution, 24 coarse samples, 16 fine hierarchical samples, 12000 MPS training iterations, 768 rays per training batch, a 5-level compact cubic spline grid with base resolution 6 and 6 features per level, separate compact radiance support via --separate\_color\_grid, entropy regularization 10−4, and depth-variance regularization 10−5. Training starts with a uniform-ray warmup of 3000 iterations. After warmup, for each step a candidate ray set of size κB is sampled uniformly, rendered under torch.no\_grad(), scored by per-ray RGB MSE, and the top ρB candidates replace part of the training batch:
ei=31∥c^(ri)−ci∥22,Ht=TopKi∈Ct(ei,ρB),
where B=768, ρ=0.25, and κ∈{1.25,2.0}. The final batch is
Bt=Ut∪Ht,∣Ut∣=(1−ρ)B,∣Ht∣=ρB.
This differs from fixed image-gradient sampling: the hard rays are selected by the current model's residual after a warmup, not by an a-priori edge detector. The exact frontier commands are:
Consolidated held-out DL3DV SOTA ledger generated by dl3dv\_heldout\_sota\_ledger.py. The current win condition is geometry-backed rendering, not naive ray-coordinate regression. The nerfacto row is a local CPU pilot and must be replaced by a full CUDA baseline before making final SOTA claims.
Ranked DL3DV held-out ledger across interpolation, pure OSNR ray kernels, spline-tomographic radiance, deterministic geometry, hybrid residual correction, and the available nerfacto CPU pilot.
p0.15linewidthp0.16linewidthp0.1linewidthp0.1linewidthp0.1linewidthp0.12linewidthp0.12linewidthp0.11linewidth@
DCT support
Kept/channel
PSNR
SSIM
LPIPS
Sparsity
Max edge error
Time
100%
1,966,080
117.2378 dB
1.000000
0.000000
77.50%
0.958244
134.5901 ms/view
50%
983,040
42.7562 dB
0.981283
0.002584
87.50%
0.958244
136.9464 ms/view
25%
491,520
35.5073 dB
0.927023
0.049832
92.50%
0.958244
136.9298 ms/view
10%
196,608
30.4278 dB
0.824789
0.230503
95.50%
0.958244
135.8118 ms/view
5%
98,304
28.0242 dB
0.746550
0.391810
96.50%
0.958244
136.1801 ms/view
1%
19,661
24.6923 dB
0.604307
0.591024
97.30%
0.958244
136.6344 ms/view
DL3DV 3D-DCT residual Pareto sweep for the same locally cached scene. The preflight estimator reports 552,895,644 bytes and measured peak memory stays at 215,813,648 bytes for all sweep points. The constant edge-error column is a diagnostic: this sweep changes residual capacity only, not the sparse-tier Hankel locator.
First DL3DV world-edge atom ablation. The hybrid mode activates ray-consistent projected knots and improves the scanline-referenced edge-delta diagnostic, but the reprojection error remains too high for a visual-quality gain. This table is included to document the structural progression and the current bottleneck, not as a final compression result.
Held-out DL3DV light-field generalization challenge. Metrics are computed only on odd-indexed views that were excluded from the fitting/input set. The ray-kernel rows evaluate held-out camera rays directly from (o,d) coordinates, but do not yet include depth-aware visibility or epipolar-consensus geometry.
Geometry-backed held-out renderers on the same DL3DV split. Sparse COLMAP splats improve PSNR over interpolation but remain coverage-limited; the dense COLMAP-pose plane sweep is the first Track-A renderer to deliver a large held-out quality gain without neural training or autograd. The multi-view FRI edge bridge adds real-image derivative coherence to the deterministic depth score and becomes the strongest autograd-free geometry row. The liquid rows add a tiny trained residual corrector on top of the deterministic renderer and are therefore reported as hybrid quality-polishing experiments.
Held-out DL3DV view-one comparison for the multi-view FRI edge bridge. Left: target odd view excluded from fitting. Middle: adjacent even-view interpolation. Right: deterministic edge-consistent geometry render using only even training views and COLMAP camera poses.
Held-out DL3DV view-one comparison for the quality-first liquid residual run. Left: target odd view excluded from fitting. Middle: deterministic 512-plane sweep. Right: plane sweep plus exact liquid residual correction trained only from even-view leave-one-out residuals.
Representative view-zero reconstructions from the DL3DV Pareto sweep, ordered left-to-right and top-to-bottom by retained 3D-DCT support: 100%, 50%, 25%, 10%, 5%, and 1%. The corresponding target view is identical to the target in Table [tab:dl3dv-stack-frames].
c c c@
caption(Frame-by-frame visual comparison for the 30-view DL3DV stack quality-ceiling export. The columns are the observed target view and the OSNR reconstruction of the same view.)
View
endfirsthead
View
endhead
000
001
002
003
004
005
006
007
008
009
010
011
012
013
014
015
016
017
018
019
020
021
022
023
024
025
026
027
028
029
Frame-by-frame visual comparison for the 30-view DL3DV stack quality-ceiling export. The columns are the observed target view and the OSNR reconstruction of the same view.
The structural-shell validation script, apps/03\_structural\_shells/biharmonic\_plate.py, uses the same 2D tensor-product Hermite machinery for a fourth-order clamped-plate operator Δ2Φ=∂xxxxΦ+2∂xxyyΦ+∂yyyyΦ=f(x,y). The validation constructs a manufactured clamped deflection field, computes the load by an autograd-free fourth-order finite-difference ladder, maps the deflection into nine Hermite streams, applies the 2D block-circulant Gram, and recovers the coefficient tensor through parallel 9×9 Fourier-domain solves. Rigid plate edges are enforced by overwriting all boundary coefficient streams to zero.
On a 48×48 grid, the current run completes in 42.9853 ms, reports boundary clamping residual 0.000000e+00, interior deflection RMS error 4.697917e−08, and relative biharmonic operator residual 1.065298e−04. The high maximum frequency-system condition number, approximately 1.8757e09, identifies the expected low-frequency stiffness of fourth-order tensor-product Gram systems and motivates more specialized biharmonic preconditioning before external structural-mechanics comparisons.
SOTA comparison: SIREN versus adaptive sparse OSNR
The first comparative benchmark, benchmarks\_sota/compare\_siren\_sdf.py, evaluates a standard sinusoidal representation network against the Tier 2 adaptive sparse OSNR solver on an identical non-bandlimited geometric target. The target is a two-dimensional silhouette with high-frequency wavy boundaries and step discontinuities along each scanline. It is intentionally hostile to smooth coordinate MLPs because the field is not bandlimited and its boundary locations fall between grid samples.
The SIREN baseline follows the Sitzmann et al. implicit representation pattern: a fully parameterized multilayer perceptron maps coordinates (x,y) to occupancy values through sinusoidal hidden layers. The benchmark uses a 64×192 coordinate grid, a hidden width of 64, three hidden sine layers, ω0=30, and Adam optimization for 1000 full-batch epochs. This produces a dense model with 12,737 trainable weights. Its final prediction is thresholded to estimate boundary locations, yielding 24.32 dB PSNR and edge blurring error 4.803569e−03 after 5,622.11 ms of optimization in the current rerun.
The OSNR path uses the same target samples but does not optimize a coordinate network. Each scanline is encoded as a finite-rate-of-innovation signal with two step horizons. The TLS matrix-pencil pre-filter recovers those continuous edge coordinates from moments, the sparse knot frame is snapped to the recovered horizons, and the cross-Gram-shielded sparse solver debiases the active shock atoms with scale-invariant Tikhonov stabilization. The OSNR pass runs under torch.no\_grad(), uses 128 active sparse knots across the batch, hard-zeros 97.9% of the sparse parameter tensor, and achieves 111.89 dB PSNR with edge localization error 1.443290e−15 in 19.23 ms.
A third path, apps\_industrial\_breakthrough/osnr\_operator\_atlas\_rank\_probe.py, tests the broader no-backprop idea inspired by local random-feature FBPINNs and rank-revealing feature filtering. It does not use the exact FRI edge moments. Instead, it covers the same field with 12×12 and 24×24 partition-of-unity charts, evaluates frozen local polynomial, DCT, and SIREN probe features, greedily keeps locally independent rank directions, orthogonalizes the retained chart directions, and solves one global ridge system. The best fixed-grid pilot keeps 10,080 of 32,400 local candidate directions, drops 68.9% of the atlas, and reaches 35.318555 dB PSNR with edge error 2.642476e−03 in 49,629.10 ms for the unoptimized Python prototype. This row is not a substitute for the exact FRI solver; it is evidence for the more general operator-atlas thesis: local frozen features plus rank-revealed algebra can beat a trained global SIREN even when the exact sparse innovation coordinates are not supplied.
The next adaptive variant, apps\_industrial\_breakthrough/osnr\_operator\_atlas\_adaptive\_refine.py, removes the hand-picked dense fine grid. A coarse 8×8 pilot atlas scores a 40×40 candidate chart lattice by residual energy, target-gradient energy, and rank density. The solver then keeps only 641 refined charts after one-cell dilation, evaluates frozen polynomial, DCT, SIREN, and curved local edge-step atoms, applies target-aware local rank filtering, and solves one global ridge system. The condition-diagnostic run keeps 8,820 of 131,130 local candidate atoms, drops 93.27% of the candidate atlas, and reaches 86.360935 dB PSNR with edge error 2.615928e−03. Two seed repeats reach 89.188964 dB and 84.082762 dB, respectively. Thus the result is not a one-seed random-feature accident. It also improves the fixed rank-revealed atlas by 51.04 dB while retaining fewer directions. The remaining gap to the exact FRI row is expected: the FRI row is given the exact sparse innovation model, whereas the adaptive atlas only receives samples and a frozen local operator dictionary.
The same adaptive runner now includes a global eigentruncated right-preconditioned solve. Instead of trusting the full ridge normal system, it diagonalizes the global atlas Gram matrix, removes directions below a relative eigenthreshold, and solves in the retained eigenspace. On the silhouette target, the publication-friendly threshold 3×10−11 keeps 86.093140 dB while reducing the effective condition from 5.389×1012 to 3.235×1010, a 166× reduction, with retained eigenspace rank 6,778. A more aggressive 10−8 threshold still reaches 84.601181 dB while reducing the effective condition to 9.841×107, about 5.48×104 lower than the raw system. Thus the adaptive atlas result survives explicit right-preconditioned rank truncation rather than depending on hidden nearly-null directions.
The dense eigentruncation is now cross-checked by a randomized projected eigensolve. With Gaussian sketching, one subspace iteration, and projected dimension 6,714, the randomized solve reaches 84.601150 dB with the same edge error 2.615928e−03 and retained Ritz rank 6,669, matching the dense aggressive row to within 3.1×10−5 dB. A projected dimension of 6,592 still reaches 82.817051 dB. More aggressive compression is not free: projected dimensions 6,464, 6,208, and 4,160 reach 75.836651, 63.725327, and 20.082664 dB, respectively. This negative boundary is useful: the atlas win is broad-rank and rank-revealed, not a tiny hidden low-rank shortcut.
The global solve can also avoid dense normal-matrix formation. A row-block Jacobi-preconditioned CG mode applies the field/operator design only through matrix-vector products, with optional weighted row sampling. On the silhouette target, full-row PCG reaches 84.500732 dB with edge error 2.615928e−03 after 6,400 iterations, within 0.10 dB of the dense aggressive eigentruncated row while never forming the dense Gram matrix. Fixed-policy replay separates solver effects from adaptive chart drift: replayed row-norm sampling at 10,000 of 12,288 rows keeps 80.457415 dB and edge error 2.615923e−03, whereas coarse spatial block and stratified schedules collapse on this discontinuity target. Thus row scheduling must respect the operator/objective geometry; it is not a generic block-dropping problem.
First SOTA-style comparison on a non-bandlimited multi-edge silhouette. The SIREN row reports trained coordinate-network performance after 1000 Adam epochs. The rank-revealed atlas rows are no-backprop local-feature prototypes that do not use the exact FRI edge moments; the adaptive row selects charts by pilot residual/rank maps and uses local discontinuity atoms. The OSNR Tier 2 row reports the matched FRI-snapped sparse solver, which remains the exact-structure oracle for this target.
As a first non-silhouette cross-check, the same adaptive runner now supports a mixed SPDE target generated by a smooth, Gaussian, and sparse-event innovation passed through a periodic advection-diffusion-reaction inverse. In this setting edge atoms are actively harmful, which is the expected guardrail for a smooth operator field. A field-only adaptive atlas with polynomial/DCT/SIREN atoms reaches 56.786560 dB versus a 512-feature global DCT baseline at 46.317344 dB, but its operator relative RMSE remains 0.982192. Adding operator rows to the algebraic normal equation, cmin∥Ac−u∥22+λ2∥LAc−Lu∥22+γ∥c∥22, improves the best mixed-SPDE point to 57.140339 dB with operator relative RMSE 0.209998, compared with global DCT operator relative RMSE 2.773887. This is a positive second validation of the adaptive atlas idea outside the silhouette benchmark, while also exposing the next numerical issue: the operator-augmented normal system is ill-conditioned, with diagnostic condition estimate 4.805×1013, so RRQR/right-preconditioned block solves are the next required improvement before making external PDE benchmark claims.
The eigentruncated solve materially improves that numerical story. With relative eigenthreshold 10−8, the mixed-SPDE atlas keeps 56.927949 dB and operator relative RMSE 0.212386 while reducing the effective condition from 4.805×1013 to 9.860×107. Only 2.24% of global eigendirections are removed, indicating that the SPDE instability is concentrated in a small global null-like subspace. This is still a dense diagnostic solve, not yet a scalable PDE production method, but it validates the intended RRQR/right-preconditioning direction.
The randomized projected solve also preserves the operator-aware SPDE result. At projected dimension 7,404, it reaches 56.911218 dB with operator relative RMSE 0.212675, essentially matching the dense eigentruncated row. At projected dimension 7,164, it still reaches 56.343715 dB and operator relative RMSE 0.217228, while projected dimension 6,208 collapses to 13.261593 dB and operator relative RMSE 3.790479. Thus the next scaling target is not smaller global rank alone; it is matrix-free or block-randomized least squares that avoids full Gram formation while preserving the broad well-conditioned Ritz subspace.
The no-dense-normal PCG path gives the same conclusion on the SPDE target. With all 24,576 augmented field/operator rows, PCG reaches 56.984770 dB and operator relative RMSE 0.210942 after 6,400 iterations, slightly stronger field PSNR than the dense eigentruncated row. Replayed row schedules then identify the correct compression geometry. Independent row-norm sampling at 20,000 rows preserves field PSNR, 56.967060 dB, but degrades operator relative RMSE to 0.709043; uniform, spatial-stratified, equal field/operator quota, and field-full/operator-sampled controls also fail to preserve both objectives. The positive schedule keeps all operator rows exactly and samples only the field rows by row norm. At 20,000 of 24,576 rows it reaches 57.600055 dB and operator relative RMSE 0.209357, slightly beating the full-row PCG anchor while using 18.6% fewer augmented rows. The same operator-shell schedule remains strong at 18,000 rows (57.153551 dB, 0.209805), 16,000 rows (56.930231 dB, 0.215080), 14,000 rows (55.953736 dB, 0.210489), and 13,000 rows (55.315051 dB, 0.210262). The 12,288-row operator-only cliff collapses to 7.165436 dB and operator relative RMSE 1.656706, showing that the sampled field equations are nullspace anchors for the operator shell rather than expendable data rows.
The same result now survives streamed atlas-column construction. In streamed\_row\_sketch\_pcg mode, retained atlas directions are stored as chart-local support blocks, and the solver evaluates Av, A⊤v, and A⊤L∗LAv without materializing either the dense field design A or the dense operator design LA. A smoke run matches materialized PCG to within 3.91×10−8 dB PSNR and 7.56×10−9 operator relative RMSE. On the replayed SPDE policy, streamed full-row PCG reaches 56.991617 dB and operator relative RMSE 0.210990 in 17.109 s, compared with 56.984770 dB and 0.210942 in 77.045 s for the materialized full-row run. Streamed operator-shell PCG keeps the compressed frontier: 20,000 rows reaches 57.598052 dB and 0.209398 in 16.830 s, while 13,000 rows reaches 55.321231 dB and 0.210299 in 17.831 s; the streamed 12,288-row operator-only cliff still collapses to 7.165883 dB and 1.656653. Thus the SPDE atlas result is now no-dense-Gram and no-dense-design in the global solve. The next scaling step is larger streamed operator-atlas validation, not blind row dropping.
External PDEBench Darcy sparse OSNR assimilation
The first external Darcy result is deliberately framed as sparse-observation assimilation rather than a blind coefficient-to-solution solver claim. On the real PDEBench\_2D\_DarcyFlow\_beta0.01 shard, a hand-coded finite-volume elliptic bridge maps each coefficient field to a solution shape, but high low-conductivity inclusion regimes expose a hidden amplitude/interface convention. The previous scalar sparse-sensor audit showed that a few target observations can calibrate the dominant amplitude. The new runner, pdebench\_darcy\_osnr\_sparse\_residual\_ladder.py, asks a stronger question: can sparse observations also identify a compact OSNR residual dictionary on top of the PDE-shaped solution, u(x)=αuCG(x)+β+k=1∑Kckϕk(x), where ϕk are low-frequency DCT residual modes. All ranks, ridges, and sensor policies are selected on the training split and reported once on 300 held-out Darcy fields.
The first ladder uses 1000 samples, 700 for training, 240 of those for hyperparameter selection, sensor budgets from 1 to 128, residual ranks 0,2,4,8,16,32, and random, grid, training-residual-variance, and energy sensor policies. The blind finite-volume CG bridge has held-out mean nRMSE 0.3262416; the full-field scalar oracle, which uses all target pixels only to choose one amplitude, has mean nRMSE 0.0585205. The first deployable sparse row uses only 64 fixed grid sensors out of 1282 pixels (0.390625% of the field), selects the PDE shape plus a rank-32 residual dictionary with ridge 10−4, and reaches mean nRMSE 0.0153122, median 0.0122784, p90 0.0325075, and max 0.0501816.
The active-design follow-up keeps the same train/held-out split but chooses additional sensor locations by the leverage geometry of the PDE+DCT feature system. Policies such as grid\_dopt128 first allocate a coarse grid prefix, then greedily add points by target-independent D-optimal posterior leverage. They use the coefficient field, the CG solution, and the frozen residual dictionary, but not unobserved target residuals. With residual ranks 96 and 128, the fixed grid already breaks the 0.01 barrier at 256 sensors, reaching mean nRMSE 0.0097030. The best active row, grid\_dopt128, reaches 0.0088075 at 256 sensors and 0.0086891 at 384 sensors. Thus the external sparse-assimilation frontier moves from ``below the scalar oracle'' to a sub-10−2 held-out PDEBench error with no neural retraining.
The neural follow-up then performs the comparison that this result demands. The runner pdebench\_darcy\_neural\_sparse\_assimilation\_baseline.py trains a U-Net sparse assimilator on the same 700 training fields. Its inputs are the coefficient field, CG base, same-sensor scalar-calibrated base, sparse observed target values, sparse residual values, the binary observation mask, and coordinate channels. After the neural prediction, the same OSNR residual adapter is fitted from the same sparse observations, but now around the neural field rather than the CG field: u(x)=αuneural(x)+β+k=1∑Kckϕk(x). This makes the claim harder: OSNR must improve an already trained sparse neural assimilator rather than only beat a scalar or DCT control.
External PDEBench Darcy sparse-observation assimilation. The PDE-shaped residual model is selected on the training split. Plain DCT-only interpolation is a negative mechanism control; it is much worse than the PDE-shaped residual row, showing that the elliptic operator bridge supplies the dominant field prior. Grid-D-opt policies are target-independent active measurement designs based on PDE+DCT feature leverage. The U-Net rows are trained sparse assimilators under the same train/held-out split; the OSNR adapter is then fitted from the same sparse observations at test time. The final budgeted row uses a separately budgeted target-independent 256-sensor design rather than the nested 256-prefix from the joint 256,384 sweep.
The improvement is strongest exactly where the blind bridge was weakest. In the high-inclusion bin, low_fraction≥0.75 (44 held-out fields), blind CG has mean nRMSE 0.8641549 and the full-field scalar oracle has 0.1287972. The 64-sensor rank-32 PDE+DCT residual row reduces this to 0.0274448, and the 384-sensor grid-D-opt rank-128 row reduces it further to 0.0149138. The neural adapter pushes the same hard bin to 0.0107779 at 384 sensors and 0.0101835 in the focused 256-sensor run. Mid/high bins show the same pattern: for low-fraction 0.50–0.75, the rank-128 CG+OSNR row improves 0.4736753 blind and 0.0900559 scalar-oracle nRMSE to 0.0120798, while the focused neural+OSNR row reaches 0.0076549; for 0.25–0.50, it reaches 0.0037485.
This result is the first external PDEBench row in the manuscript where OSNR beats the scalar-oracle ceiling rather than only calibrating amplitude, and the neural follow-up changes the status of the claim. OSNR is no longer only a standalone no-retraining sparse assimilator; it is also a test-time correction layer that improves a trained neural sparse assimilator. In the main MPS run, the 256-sensor sparse U-Net reaches mean nRMSE 0.0080866, while U-Net+OSNR reaches 0.0063245. The focused 256-sensor run reaches 0.0059962 mean nRMSE, median 0.0047928, p90 0.0119379, and max 0.0161753, using only 1.5625% of pixels. The scientific claim remains sparse-observation assimilation rather than blind coefficient-to-solution neural-operator SOTA, but the mechanism is now more general: OSNR can operate both as the primary PDE-shaped residual solver and as a plug-in residual adapter on top of a learned neural prior.
SOTA comparison: Spline-PINN regime versus tensor-product OSNR CFD
The second comparative benchmark, benchmarks\_sota/compare\_spline\_pinn\_cfd.py, targets the fluid-surrogate regime studied by Wandel et al. for Spline-PINN. The script uses a DFG-style cylinder domain with a 41×220 spatial layout and compares the reported Spline-PINN training regime against the measured Tier 3 OSNR tensor-product Hermite path. The Spline-PINN row is therefore not a rerun of the authors' training code; it is an explicit reported-regime reference capturing the relevant structural cost: a spline-interpolated U-Net update model trained with physics-informed losses and data recycling over one to two days, with real-time inference reported at approximately 30 updates per second.
The OSNR path uses the same grid and obstacle geometry but does not train a time-step network. A stream function az generates the velocity field by the hard incompressible curl map vx=∂yaz,vy=−∂xaz, and the cylinder no-slip boundary is enforced by overwriting the nine tensor-product Hermite coefficient channels at masked vertices. The nonlinear advection and viscous diffusion terms are evaluated through the forward Hermite derivative ladder, and the coefficient update is resolved through the 2D block-circulant Fourier solver. The measured OSNR rows run under torch.no\_grad() with zero autograd graph allocation.
For Reynolds-number calibration, the benchmark uses Re=μρUD,ρ=1,U=0.20,D=10grid cells. This gives μ=0.100000 for Re=20 and μ=0.020000 for Re=100. Boundary leakage is measured directly on the obstacle mask after coefficient overwriting. The divergence residual is the root-mean-square divergence of the reconstructed velocity field on the discrete validation grid.
SOTA-style CFD comparison on a DFG-style cylinder grid. The Spline-PINN rows summarize the reported training and inference regime; the OSNR rows are measured package outputs from the tensor-product Hermite CFD benchmark.
SOTA comparison: classic PINN versus operator-spline boundary solver
The third comparative benchmark, benchmarks\_sota/compare\_classic\_pinn.py, isolates the cost of high-order automatic differentiation in the classical physics-informed neural-network formulation. The validation problem is the fourth-order boundary-value system u(4)(x)=(2π)4sin(2πx),x∈[0,1], with strict Dirichlet and curvature constraints u(0)=u(1)=0,u′′(0)=u′′(1)=0. The exact interior solution is u(x)=sin(2πx), which satisfies both the boundary values and the curvature clamps.
The PINN baseline follows the Raissi et al. pattern: a deep fully connected tanh network is trained with Adam on a joint physics-plus-boundary objective. The physics loss is evaluated by repeated backward-mode automatic differentiation through the network to obtain u(4)(x), while the boundary loss separately differentiates the boundary predictions to obtain u′′(0) and u′′(1). The benchmark uses 2000 optimization epochs. On the CPU validation run, where PyTorch does not expose a global peak autograd allocator analogous to CUDA peak memory, the script reports a conservative graph-footprint estimate built from the derivative tapes, layer activations, parameters, gradients, and Adam state tensors.
The OSNR path uses the same spatial dimension but replaces the learned function with a calibrated knot grid and a Fourier biharmonic symbol inversion. In the periodized operator basis, the fourth derivative is diagonalized by the Fourier symbol (2πν)4, so the coefficient recovery is a single element-wise division in the frequency domain. Boundary constraints are then represented as hard coefficient-layer constraints rather than soft penalties. The complete OSNR segment runs under torch.no\_grad() and allocates no autograd graph.
Classic PINN comparison on a fourth-order boundary-value problem. The PINN row measures iterative tanh-network training with repeated fourth-derivative autograd; the OSNR row measures the calibrated operator-spline Fourier inversion.
The absolute PSNR values are extremely high because these are controlled algebraic verification problems with exact synthetic data and matched model assumptions. Future external comparisons must include noisy measurements, non-exact operators, multidimensional fields, and standardized PINN/SIREN baselines.
Operator-compiled constitutive KANs
The edge-function idea of a Kolmogorov–Arnold network becomes more useful for operator learning when the learnable function is placed at the constitutive uncertainty, rather than at the entire PDE right-hand side. Consider ut=νuxx−∂xF(u),Fc(u)=j=1∑Jcjϕj(u). The chain rule compiles the unknown flux into a linear coefficient problem, ut−νuxx=−Fc′(u)ux=j=1∑Jcj[−ϕj′(u)ux]. Consequently, coefficient identification is one regularized least-squares solve even though the resulting PDE is nonlinear in u. At rollout we evaluate Fc(u) and apply the discrete spectral derivative to the complete flux, rather than separately sampling the chain-rule factors. On a periodic grid this gives dtdn∑un=0 up to floating-point roundoff for every learned coefficient vector. The resolution-sensitive differential operator is never approximated by the network.
Uniform cubic cardinal B-splines provide local adaptation and a matrix–vector evaluation path, but compact support creates an unavoidable amplitude- extrapolation ambiguity. If training states occupy only an interval Itr, coefficients whose supports lie outside Itr are not identified; minimum-norm fitting makes the represented flux flatten outside the observed interval. A global carrier is therefore not an implementation detail but an identifiability requirement. We test polynomial carriers and a sparse exponential-polynomial atlas A={u,u2,u3,sin(ωu),cos(ωu):ω=1,…,6}. These atoms are the null-space functions associated with repeated zero poles and conjugate imaginary poles. Cardinal exponential splines reproduce the same spaces; the present experiment operates directly in the reproduction space and is therefore a pole-discovery/compiler test, not yet a compact E-spline implementation.
The matrix–vector qualification is operational, not merely asymptotic. If z=(u−u0)/h=j+t with t∈[0,1), a centered cardinal cubic edge is evaluated from only four adjacent coefficients by Fc(u)=[1tt2t3]611−33−140−63133−30001[cj−1cjcj+1cj+2]⊤. Thus a layer is a batched gather followed by a fixed small matrix contraction; neither Cox–de Boor recursion nor a dense all-knot basis tensor is needed. The same compilation extends to objectives. Badoual, Schmitter, and Unser's periodic inner-product calculus gives ⟨Fc,Fd⟩L2=c⊤Ad,Akℓ=⟨ϕ(⋅−k),ϕ(⋅−ℓ)⟩, and derivative energies replace A by precomputed derivative cross-Grams. For an equal periodic cardinal grid these matrices are circulant. A compact cubic mass-matrix application is therefore a seven-tap O(J) stencil, while regularized inversion and broader composite operators are diagonalized by the DFT. The Hermite construction of Appendix [app:hermite-block-gram] is the multichannel version of exactly this identity. Consequently cardinality provides two separate accelerators: a fixed local evaluation kernel for neural edges and exact coefficient-space calculus for training losses and operator solves.
A dedicated CPU benchmark tests both claims against independent paths. For 8,192–32,768 edge outputs and 16, 32, and 64 knots, the local cubic matrix kernel agrees with generic vectorized Cox–de Boor evaluation to at worst 1.58×10−16 and is approximately 4.8×–12.2× faster across the sweep. Its gathered payload remains four values per edge, whereas the recursive dense basis workspace grows with the knot count. For 64–4096 periodic coefficients, compact-stencil and FFT Gram products agree to at worst 1.09×10−15, the coefficient bilinear form agrees with 32-point-per-cell continuous quadrature to at worst 2.67×10−8, and FFT ridge solves have relative residual at most 1.03×10−15. At J=4096, a dense Gram alone is 128 MiB, compared with approximately 0.063 MiB for its real kernel and complex half-spectrum. The timing boundary is useful: the seven-tap direct stencil is faster than FFT at the largest tested compact case, so FFT should be reserved for inversion or noncompact/composite circulant operators rather than applied dogmatically. These are NumPy CPU prototype timings. An unfused PyTorch CPU control is negative: although the dense explicit-cardinal and local paths agree to 2.33×10−6 in float32, composing gather, mask, and reduction primitives makes the local path 1.95×–4.31× slower in the forward pass and 3.44×–7.68× slower for forward–backward. The direct KAN therefore keeps the dense explicit-cardinal contraction as its CPU default and exposes the local matrix path as an opt-in reference. Apple MPS exposes the predicted resolution crossover even without a custom kernel. In a synchronized sweep with batch 1024, 32 inputs, 64 outputs, and 16–512 knots per edge, a memory-bounded implementation streams the four taps instead of materializing a batch–input–tap–output tensor. Forward evaluation first beats the dense explicit-cardinal path at 128 knots; forward–backward first wins at 256 knots. At 512 knots, streamed evaluation is 4.23× faster forward and 2.95× faster forward–backward, while its explicit workspace is 8 MiB rather than the dense basis's 64 MiB. The two paths agree to relative 1.44×10−5 in float32 at this finest grid. The direct KAN now selects the streamed path automatically on MPS from 256 knots and retains the dense path below the crossover. A truly fused Metal/Triton kernel remains valuable because it should lower the crossover into the small-grid regime; the present result already establishes a measured high-resolution GPU training-throughput advantage.
The same inner-product compiler removes another KAN-specific failure mode: changing grid resolution need not change the represented edge. For source basis Φ and target basis Φ, the continuous L2 projection is the precomputable map c=A−1Ac,A=⟨Φ,Φ⟩,A=⟨Φ,Φ⟩. This is the resampling identity in Badoual et al. and in the accompanying E-snake derivation, which further gives a short characteristic-polynomial description of a special exponential-spline circulant inverse. We test the projection identity for periodic cubic cardinal edges using four-point Gauss integration on the union knot partition, which is exact for the piecewise degree-six cross-products. Nested refinements 24→48 and 24→96 preserve the continuous edge to worst relative L2 error 1.71×10−15, whereas periodic coefficient interpolation and using new-knot values as coefficients incur 0.95%–2.66% error. The 24→96→24 coefficient round trip has relative error 1.18×10−15. A nonnested 48→72 projection has 0.060% median error versus 0.53%/0.77% for the two heuristics. For 64→32,24,16 coarsening, the exact projection reduces median continuous error by 1.8×–2.6× and preserves the integral to 1.39×10−16, while heuristic integral drift reaches 2.76%. After matrix precomputation, 32-edge projection solves take 15–67 μs on CPU. Thus cardinal grid extension can be exactly function preserving when the spaces are nested and L2 optimal otherwise. Implementing the E-spline-specific reciprocal-root inverse as an IIR/parallel-scan kernel is a separate gate, which we close next.
The E-snake Gram has the more specific real symmetric circulant form A=pI+q(S+S−1)+r(S2+S−2). Consequently A−1 must itself be symmetric circulant: if g is its first row, gk=gM−k. Let z1,z2 be the two roots inside the unit disk of rz4+qz3+pz2+qz+r, and set γ=r/(z1z2). Reciprocal pairing gives A=γi=1∏2(I−ziS)(I−ziS−1). Partial fractions therefore produce the explicitly symmetric periodic Green kernels Hz[k]=(1−z2)(1−zM)zk+zM−k,gk=a1Hz1[k]+a2Hz2[k], where a1=γ(z1−z2)(1−z1z2)z1,a2=γ(z2−z1)(1−z1z2)z2. This exposes a consequential correction. Under the circulant convention stated in the E-snake note, its one-sided equations (33)–(34), transcribed as printed, have inverse residual 0.474–0.492 and symmetry defect 0.659–0.666 for M=8–128. The expression above has exactly zero measured symmetry defect and worst inverse residual 1.28×10−15. The equivalent four cyclic first-order filters agree with dense inversion to worst residual 1.18×10−15 in both sequential and logarithmic-depth parallel implementations.
The factorization is also a stable learnable parameterization. For this real-root family, zi=−sigmoid(θi) and γ=exp(η) put the roots strictly inside the unit disk and induce the positive spectrum A(ω)=γi=1∏2(1−2zicosω+zi2)>0. Autograd root/scale derivatives agree with centered differences to at worst 1.52×10−10. Both FFT and associative-scan realizations propagate finite MPS gradients through the right-hand side, roots, and scale. The systems control is negative for an unfused scan: across two synchronized batch-256 sweeps over M=64–4096, the root-spectrum FFT is 2.35–9.08× faster forward and 2.20–7.66× faster forward–backward, with scan/FFT float32 disagreement at most 7.95×10−7. At M=4096, FFT takes 0.305–0.307 ms forward versus 2.77 ms for the scan and avoids a 64 MiB dense float32 inverse. Thus symmetric-circulant Fourier diagonalization is the preferred compiled path on FFT-capable hardware; the exact bidirectional filters are retained for streaming or no-FFT targets rather than claimed as an MPS speed improvement.
We then insert this transfer into a trainable six-edge additive KAN. Every arm starts from the same 24-knot checkpoint, moves to 96 knots, resets the Adam optimizer to isolate parameter transport, and continues training. On a target with deliberately unresolved frequency-15 and localized detail, exact transfer has median/worst immediate function drift 1.97×10−15/1.98×10−15 across five seeds, loss-jump ratio one, and integral drift at roundoff. Coefficient interpolation changes the function by 2.63% and raises loss by 5.67%; a zero restart loses the entire function and raises loss by 70.98×. At 1% training noise the coarse median test MSE is 2.0122×10−2. Fine training after exact transport reaches 2.2615×10−5 without regularization. The exact curvature Gram Rkℓ=⟨ϕk′′,ϕℓ′′⟩,λe∑ce⊤Rce, with λ=10−9 selected on development seeds, improves the median to 2.0896×10−5 on five untouched seeds. It wins all five paired comparisons with a 1.100× geometric improvement factor, giving roughly 963× improvement over the coarse unresolved model. In the clean control, exact initialization reaches 3.37×10−8 after 250 steps, where a zero restart remains at 6.20×10−6.
Refinement is not universally beneficial. On the earlier smooth target, the 24-knot model already has median test MSE 5.02×10−6 at 1% noise; unregularized refinement worsens it to 2.24×10−5, and the best tested curvature setting only returns to approximately 5.94×10−6. Thus the complete rule is guarded: project exactly so a proposed refinement cannot damage the current function, train new fine modes under the continuous Sobolev Gram, and accept the enlarged grid only if held-out evidence improves. This separates stability of grid transport from necessity of grid growth.
The controlled benchmark uses three periodic viscous conservation laws (quadratic, quadratic plus sinusoidal, and saturating flux), centered temporal- difference targets, six training trajectories at 64 points, and held-out long-horizon rollouts at unseen amplitude and at 128 points. Across three seeds the pole atlas obtains median/worst amplitude-OOD nRMSE 9.287×10−7/8.475×10−6 on the oscillatory flux. Median errors for the compact cardinal edge, degree-five polynomial/SINDy, a 345-parameter direct cardinal KAN, and a 337-parameter MLP are respectively 0.1870, 0.06477, 0.3300, and 0.2977. The selected reference law is 0.499993u2+0.079997sin(4u), recovering the true 0.5u2+0.08sin(4u). On the unmatched saturating flux, median amplitude-OOD error is 0.005199, versus 0.04113/0.1839/0.1165 for SINDy/direct-KAN/MLP. Conservative models preserve mass to approximately 10−16; the direct networks drift by 10−3–10−2.
The stronger mixed-constitutive gate uses ut=νuxx−∂xF(u)+R(u) and assigns a separate copy of the atlas to the flux and reaction edges. The compiled columns are −ϕj′(u)ux for the former and ϕj(u) for the latter. Across three seeds the mixed atlas reaches median/worst amplitude-OOD nRMSE 2.205×10−5/2.979×10−5, compared with median 0.1531, 0.02890, 0.2544, and 0.2719 for mixed cardinal, polynomial/SINDy, direct KAN, and MLP models. Median resolution-OOD error is 5.762×10−7. The flux 0.5u2+0.06sin(4u) is selected consistently to at least six coefficient digits; the reaction function is recovered to median 7.91×10−4 nRMSE.
Functional identifiability does not imply symbolic identifiability. The reaction atom lists vary across seeds because u and sinu are nearly collinear over the narrow training amplitude interval. They synthesize nearly the same reaction on the tested range, but their individual coefficients cannot be interpreted uniquely. Broader excitation, group sparsity by pole family, or annihilator-based incoherence constraints are required before the mixed model can support a symbolic reaction-law claim.
Weak compilation under observation noise.
For a separable test η(t)ψ(x), periodic integration by parts gives −∬ηtψu=ν∬ηψxxu+∬ηψxF(u)+∬ηψR(u). No derivative acts on the noisy observation. In the implementation, the temporal test uses the exact discrete adjoint of the centered time-difference operator, while sine/cosine spatial tests supply analytic first and second derivatives. Across three seeds at 1% additive observation noise, median amplitude-OOD nRMSE is 0.001401 for weak pole identification, 0.03844 for pointwise pole identification, and 0.003194 for weak polynomial/SINDy. At 2%, weak pole remains far stronger than pointwise (0.006581 versus 0.1464) but slightly trails weak polynomial (0.005207). The flux function remains below 10−3 median error through 2% noise; the reaction function loses identifiability first. Thus weak compilation delays, but does not remove, the variance cost of the larger pole dictionary.
Annihilator-designed excitation.
If u(x,t)=a(t) is spatially constant, then ∂xF(u)=0 for every flux law. Such trajectories are therefore operator-null-space probes of the reaction edge. Merely replacing half the generic trajectories by four constant probes is a negative: it never recovers the exact support in three seeds. Randomly broadening generic amplitudes is also unreliable (one of three exact recoveries). The successful design uses four generic flux trajectories plus a balanced 24-level constant sweep over [−1.5,1.5]. It recovers the exact support {F:u2,sin(4u);R:u,u3} in all three seeds. Relative to eight narrow generic trajectories, median reaction-function nRMSE falls from 0.002197 to 6.569×10−7 and amplitude-OOD rollout nRMSE from 2.689×10−5 to 2.049×10−7. This converts the earlier functional identifiability into repeatable symbolic identifiability by experimental design, without supervising either edge separately.
Nonseparable closure and quotient dictionaries.
For the additional closure C(u,ux)=γsin(2u)sin(ux), scalar flux and reaction edges are structurally misspecified. Adding sparse tensor products of pole atoms lowers three-seed median amplitude-OOD nRMSE at γ=0.05 from 0.005192 for the scalar compiler, 0.1012 for the direct KAN, and 0.07287 for the MLP to 0.0002971. The unrestricted dictionary is not identifiable, however, because a(u)ux=∂xA(u),A′(u)=a(u). It can move a conservative flux term into the generic interaction edge without changing the PDE residual. The naive fit does exactly this and has median interaction-function nRMSE 23.71. Defining the interaction dictionary on the quotient that removes all atoms linear in ux restores flux attribution and reduces interaction error to 0.2246, with unchanged rollout.
Guarded scalar-to-tensor growth.
A validation router selects the quotient tensor atlas only when at least one interaction survives and held-out residual error falls. Across three seeds and five mismatch strengths, it retains the scalar model in all three γ=0 cases and grows the tensor model in all 12 nonzero cases. At zero mismatch it prevents a worst-case tensor amplitude-OOD error of 0.02601; at γ=0.05 it lowers median error from 0.006218 to 6.197×10−6. This supplies a concrete KAN-like growth rule: add product structure only after operator residuals demonstrate scalar-edge mismatch. Occasional worst-seed errors up to 0.004567 show that sparse selection inside the enlarged tensor space still needs a stability penalty.
Validation-selected tensor order.
The surplus-interaction failure is not intrinsic to the product atlas. At γ=0.05, a six-term quotient fit selects exactly one interaction atom in all three seeds. Relative to the loose eight-term fit, median/worst amplitude-OOD nRMSE improves from 2.971×10−4/3.204×10−4 to 2.625×10−5/7.994×10−5, while median interaction-function nRMSE falls from 0.2246 to 0.006863. A five-term control is worse (4.221×10−4/1.740×10−3 median/worst rollout error), because one small atom is needed to absorb the centered-difference bias in addition to the generating support. A non-oracle selector fits budgets on six trajectories and chooses the sparsest model within 10% of the minimum residual on two held-out trajectories. It selects six in all three strong-interaction seeds and reproduces the oracle result after a full refit. For γ=0.005–0.02, accurate prediction does not guarantee interaction recovery: in some seeds the closure signal is comparable to discretization bias, so functional attribution remains unidentifiable.
Weak-residual routing is not rollout routing.
A final control uses held-out weak moments to choose between pole and degree-five polynomial weak models across six noise levels and three seeds. It selects the amplitude-OOD rollout winner in only 9 of 18 cases, fails all three seeds at both 1% and 4% noise, and has worst rollout regret 4.95×. The two libraries can have nearly identical integrated residuals but different errors after recursive deployment. Thus the weak form is an effective estimator, but its regression residual is not by itself a safe architecture-selection objective; the validation functional must include rollout stability or a provable surrogate for it.
Short low-pass rollout validation raises the development decision accuracy to 14 of 18, including all three 1% cases, but does not solve the problem: near-tied scores cause a worst regret of 5.32×. A robust alternative is to average models when their identity is below the noise-resolution limit. We estimate relative white-noise amplitude from the upper spatial half-band, keep the pole law below 1.5%, and above that threshold deploy the convex law
Because averaging occurs before the known outer operators, conservation and the compiled calculus are unchanged. After freezing the rule on ten development seeds, ten new seeds confirm the result. At 2% noise, mean/worst amplitude-OOD nRMSE is 0.006288/0.01107 for the blend, versus 0.007666/0.01424 for pole and 0.007699/0.02382 for polynomial. At 4%, blend median/worst is 0.005424/0.009171, versus 0.008142/0.01449 and 0.007678/0.02368. The practical rule is therefore continuous as well as operator-aware: use hard architectural growth when the validation gap is resolved, but use an operator-compatible model average when the data cannot support that decision.
Two-dimensional jet-space experimental design.
Consider the anisotropic extension ut=νΔu−∂xFx(u)−∂yFy(u)+R(u). Constants annihilate both conservative edges, x-only fields annihilate the y-flux, and y-only fields annihilate the x-flux. These probes make the inverse problem block triangular: estimate R from constants, subtract it from each directional balance, and then estimate the corresponding flux. There is an additional jet-space requirement. A zero-mean wave does not adequately span (u,ux) because its largest state values occur where ux=0. We therefore use offset directional waves u(x)=c+v(x) and sweep c, independently varying the constitutive argument and its multiplying derivative.
This design converts the 2D experiment from approximate prediction to exact law recovery. The hierarchical atlas uses 10,200 compressed scalar rows and recovers Fx(u)=0.5u2+0.05sin(3u),Fy(u)=−0.3u2+0.04sin(5u),R(u)=0.2u−0.15u3 to seven–eight coefficient digits in every one of three seeds. On unseen 2D fields of amplitude 1.35, beyond the designed value range, median/worst amplitude-OOD nRMSE is 5.695×10−9/5.739×10−9, while median resolution- and joint-OOD errors are 7.404×10−9 and 5.702×10−9. The generic 2D control consumes 204,800 rows—a 20.1× larger scalar data matrix—yet has median/worst amplitude-OOD error 0.03235/0.06712; randomly broadening its amplitudes gives 0.04093/0.09698. Earlier zero-offset directional probes remain at 10−3–10−2 rollout error. The resulting design rule is more precise than ``excite broadly'': span the jet variables that multiply each compiled operator and choose null-space probes that triangularize the unknown edges.
Noisy 2D derivative compilation.
The exact clean recovery does not survive raw temporal differencing under measurement noise. We therefore test two theoretically matched preprocessing steps: projection onto the transverse-invariant subspace of each directional probe, and a Savitzky–Golay local-polynomial filter along time before applying the fourth-order difference. Window selection itself requires a held-out audit. A seven-sample window looks best at 1% on three development seeds, but on five untouched seeds a 21-sample window is more robust and also remains stable at 2%.
With width 21, temporal filtering alone lowers held-out median/worst amplitude-OOD nRMSE from 0.06918/0.07640 to 0.003980/0.006052 at 1% noise, and from 0.1183/0.1355 to 0.005420/0.02143 at 2%. The symmetry hypothesis is only partly useful. Projection alone has median errors 0.07136 and 0.09785 and therefore does not repair differentiated noise. Combining projection and temporal filtering gives median 0.002678 at 1% but 0.007249 at 2%, so it is not uniformly better than temporal filtering alone. The supported rule is to put the reproducing regularizer on the coordinate that will be differentiated; null-space symmetry averaging may reduce variance but is not the causal robustness mechanism here.
The same construction reduces the inverse problem itself. A generic N×N trajectory contributes O(N2) scalar rows to a joint 45-column design. A transverse-invariant directional trajectory contributes only O(N) distinct rows, and triangularization means that only one 15-column edge block is materialized at a time. At N=64, the benchmark therefore uses 786,432 versus 19,008 rows (41.4× fewer) and 270 MiB versus 1.05 MiB of logical peak design storage (256× lower). Identification is 2.7–2.9× faster across three seeds even before fused or GPU kernels. At N=16–24, separate block-fit overhead dominates, with the measured runtime crossover near N=32; the asymptotic storage gain is already present.
Coupled vector systems.
The triangular principle also applies across field channels. We test ut=νuuxx−∂xF(u)+C(v),vt=νvvxx−∂xG(v)+D(u), where all four constitutive edges are unknown. Constant paired states identify the two cross-reactions. Their contributions are then subtracted from offset directional-wave balances before fitting the two conservative fluxes. This remains valid while the fields evolve and drive each other; the separation is by typed operator action, not by freezing the other channel.
Across three seeds, the resulting atlas recovers all seven generating atoms and their edge functions to 10−8–1.5×10−7. Median/worst amplitude-OOD rollout nRMSE is 2.419×10−9/2.964×10−9; median resolution- and joint-OOD errors are 1.650×10−9 and 2.202×10−9. The generic joint fit has median/worst amplitude-OOD error 0.002016/0.004428, selects between 7 and 12 atoms, and has median functional error 0.4453 on the cubic cross-reaction. Thus the relevant object is a typed graph of operator-compiled edges: null-space probes can order that graph into identifiable blocks even when its state dynamics remain coupled.
The exact result has two jointly necessary causes. Keeping the jet probes and triangular solve fixed while deleting only the true frequency-five atom raises hierarchical median/worst amplitude-OOD nRMSE from 5.695×10−9/5.739×10−9 to 0.01853/0.02243 and median y-flux function error to 0.1303. Removing every sinusoidal pole raises median rollout and y-flux errors to 0.05825 and 0.3471. The probes make the restricted design full rank, but cannot synthesize an absent exponential mode; conversely, a correct pole library is not identifiable without jet coverage. This is the experimental form of the reproduction/identifiability factorization.
Finite-dimensional identifiability criterion.
Let Se denote the active reproduction atoms for edge e, and let Xqe(P) be the compiled response of those atoms in equation q under a probe family P. If the probes can be ordered so that XS(P)=X11X21⋮Xm10X22⋱⋯⋯⋱⋱Xm,m−10⋮0Xmm, then the active coefficients are uniquely identifiable exactly when every diagonal block Xee has full column rank, modulo the explicit gauge quotient. This follows directly by block forward substitution; if a diagonal block is rank deficient, a nonzero coefficient perturbation in its null space produces the same measurements. Pole inclusion guarantees that the truth lies in the column span, annihilators create the zero blocks, and offset jet probes supply rank to the diagonal blocks. The three experimental ablations separately remove each condition: missing poles create approximation bias, generic probes destroy triangular attribution, and zero-offset waves lose jet rank near state extrema.
Continuous pole profiling.
The same design makes learnable exponential-spline poles tractable. For F(u)=au2+bsin(ωu), fixing ω leaves a two-column compiled design, so (a,b) are eliminated by a ridge solve. We evaluate the resulting scalar profile objective on a coarse frequency grid and refine its best interval by bounded scalar minimization. This is variable projection: the difficult pole is not optimized jointly with the linear edge coefficients.
Across five noninteger frequencies from 1.7 to 5.6 and three seeds, six offset jet probes recover the pole with median/worst absolute error 2.158×10−9/1.355×10−8 and median/worst amplitude-OOD rollout nRMSE 2.581×10−10/9.446×10−10. The same profiled model on generic trajectories has worst pole error 0.02548 and worst rollout 0.001606. Fixed integer poles give median/worst rollout 0.001574/0.01527, and degree-five polynomial gives 0.03169/0.1191.
Pole estimation is still statistically fragile. With a 21-sample temporal polynomial prefilter and 24 replicated offset probes, median rollout errors are 0.000709, 0.002612, and 0.008512 at 0.1%, 0.5%, and 1% observation noise. These improve strongly over six noisy probes and over the fixed/polynomial offset controls, but generic continuous-pole trajectories are competitive or better; median pole error reaches 0.1052 at 1%. Variable projection removes coefficient–pole optimization coupling, while weak-form or probabilistic inference is still required to control pole variance.
For multiple poles, sequential pursuit is not reliable: after profiling the linear coefficients, a greedy second-pole insertion can enter a wrong basin even as the true separation increases. The operator calculus supplies a finite-dimensional repair. Evaluate every candidate pole column once, form its Gram matrix and target correlations, and score every coarse pole pair by a constant-size profiled solve. Joint continuous refinement is then initialized from several distinct low-residual pairs. Offset jet experiments recover grid-aligned separations from 0.6 down to 0.025 to machine precision across three seeds. The normalized active-design condition number increases from 7.36 to 178.2 over that range. Off-grid tests separate prediction from symbol recovery: at gap 0.058 rollout remains near 10−7 with pole error near 10−2, whereas at gap 0.025 individual poles become unstable even though their combined flux is accurate. The relevant resolution certificate is therefore the conditioned active Gram, not convergence of a nonconvex optimizer alone.
Noise requires compiling the profile in weak form, not differentiating a prefiltered trajectory. For a separable test η(t)ψ(x), conservation gives the pole column directly as ∫∫η(t)ψx(x)sin(ωu(x,t))dxdt, while temporal and diffusive derivatives act only on η and ψ. With offset jets, a frozen 40-sample window and four spatial modes, median amplitude-OOD nRMSE at 0.1%, 0.5%, and 1% noise is respectively 1.266×10−4, 5.232×10−4, and 7.639×10−4, versus 3.548×10−3, 3.502×10−2, and 4.426×10−2 for matched pointwise profiling. Weak fixed-integer and degree-seven polynomial controls are also worse. However, the median maximum pole error rises to 0.0658, 0.2687, and 2.7. Consequently the weak profile is a robust predictive estimator beyond the pole-identification regime; noisy symbolic claims require uncertainty sets, replicated excitation, or an explicit minimum-separation prior.
The same Gram geometry suggests D-optimal probes: differentiate compiled columns with respect to their coefficients and poles, then maximize a nominal sensitivity log determinant. A strict control shows why this apparently natural rule must not be accepted without matching excitation support. Its initial advantage over offsets in [−0.75,0.75] disappears when both arms use the same [−1.2,1.2] range and identical wave phases. We also compile a second D-optimal rule from the exact weak sensitivity Gram under a nominal simulator. Across ten deterministic-phase seeds at 0.5% noise, broad uniform offsets attain median rollout/flux/pole errors 2.875×10−4/6.984×10−4/0.05319, versus 3.418×10−4/9.788×10−4/0.1052 for narrow uniform. Geometric improvement factors are 1.43, 1.59, and 2.27, with bootstrap 95% intervals above one. Neither D-optimal rule significantly improves on broad uniform. Thus broad jet coverage is causal; the tested local Fisher surrogates are not. A one-versus-two profiled BIC is likewise only a conservative symbolic flag, not a rollout selector. Measurement design, attribution, and deployment require separate validation criteria.
A weak-form cardinal edge provides the direct KAN-inspired control. Compact B-spline columns with negligible weak sensitivity must first be pruned; otherwise their standardized coefficients diverge. Even after pruning and ridge tuning, a 33-knot edge has median amplitude-OOD error about 0.105 at 0.5% noise. Supplying the correct quadratic conservative base and using the spline only as an innovation lowers this to 0.0371, still approximately 71× above the continuous-pole result. Local support reduces parameter interference, but cannot substitute for the correct global reproduction space when extrapolation and weak observability are decisive.
The profiled estimator provides a local resolution certificate without ground truth. Append to the weak linear design the pole-sensitivity columns bj∂ωj∂∫∫ηψxsin(ωju)=bj∫∫ηψxucos(ωju), and estimate covariance from the weak residual variance and the pseudoinverse of this full Jacobian Gram. Across 20 independent broad-jet trials per noise level, nominal 95% intervals jointly cover both true poles in 20/20, 20/20, and 18/20 cases at 0.1%, 0.5%, and 1% noise; marginal coverage at 1% is 95% for each pole. Median half-widths grow from [0.0555,0.0697] to [0.2706,0.3970] and [0.6464,0.6921]. The certificate requiring both half-widths below 0.1 accepts every low-noise case and rejects every medium/high-noise case. It therefore distinguishes resolved symbolic poles from an accurate but non-identifiable combined flux.
The same compiler supports genuine complex exponential-spline roots. For F(u)=0.5u2+0.04exp(σu)sin(ωu), fixing (σ,ω) again leaves only a small linear least-squares problem. A two-dimensional Gram profile followed by local refinement recovers both positive and negative real parts. Across nine clean offset-jet cases, median/worst amplitude-OOD nRMSE is 1.486×10−10/2.827×10−10, and the maximum error in either pole component is 2.23×10−11. Imaginary-only and degree-seven polynomial controls have median rollout 0.01020 and 0.002269.
In weak form the new feature is compiled without differentiating the noisy state, ∫∫ηψxexp(σu)sin(ωu)dxdt. Three-seed median amplitude-OOD errors are 5.458×10−5, 2.059×10−4, and 5.167×10−4 at 0.1%, 0.5%, and 1% noise. The imaginary-only weak model remains near 1.15×10−2 and weak polynomial closure near 4.2×10−3. At 1% noise, median absolute errors in (σ,ω) are (0.00717,0.00122). This establishes that the learned carrier need not be a Fourier atom: its real and imaginary pole parts can both be recovered from noisy conservation-law data.
For a complex pole the profile Jacobian appends both carrier sensitivities, b∫∫ηψxuexp(σu)sin(ωu),b∫∫ηψxuexp(σu)cos(ωu). Across 20 new trials at every noise level, the resulting joint 95% intervals cover (σ,ω) in 20/20 cases at 0.1%, 0.5%, and 1%. Median half-widths grow from [0.00325,0.00199] to [0.01664,0.01011] and [0.03251,0.01992]; median component errors remain smaller at [0.000666,0.000373], [0.00458,0.00156], and [0.00581,0.00338]. Median/worst rollout at 1% is 4.264×10−4/1.120×10−3. Thus the local covariance is conservative on the nominal-amplitude task; a maximum-half-width threshold of 0.05 accepts all 60 trials, and must next be challenged by weakening the carrier rather than by retroactively changing the threshold.
The amplitude sweep verifies that behavior. At 1% noise, amplitude 0.03 has median half-widths [0.0441,0.0271] and an 8/10 acceptance rate. At amplitudes 0.02, 0.01, and 0.005, median real-part half-widths become 0.0661, 0.1331, and 0.2638, and the fixed certificate accepts 0/5 in each group. Joint coverage is 100% throughout and median rollout remains below 5.5×10−4, explicitly separating prediction from symbol resolution. Five-seed tests at σ=−0.75 and +0.75 retain 100% coverage and about 5×10−4 median rollout, excluding proximity to the profile bounds as the cause.
Local cardinal support becomes useful when the law contains a localized constitutive defect, but only with an attribution constraint. In a controlled test the true flux is the complex carrier plus one compact cardinal cubic. Simultaneously fitting the pole and a dense local dictionary reduces prediction error yet lets the nuisance dictionary absorb pole perturbations. We instead compute each probe's carrier weak residual, refit the carrier on the least-mismatched half of the offset jets, freeze it, and estimate the local innovation from all probes. At 0.1% noise over five seeds this residual- routed hybrid has median amplitude-OOD error 1.927×10−4, compared with 6.902×10−3 for the global carrier, 6.292×10−4 for a cardinal-only edge, and 1.025×10−2 for degree seven. Its median (σ,ω) errors are (2.86×10−4,1.14×10−3). At 0.5% noise the routed median is 7.737×10−4, compared with 6.986×10−3 and 8.747×10−4 for carrier-only and cardinal-only, while pole errors remain (1.13×10−3,2.78×10−3). This triangular scheme gives a precise role to the KAN idea: a local edge is a routed nuisance correction around an identified operator carrier, not a free competitor for the same signal.
For deployment we sparsify the nuisance step. Each fixed-grid cardinal atom is scored after residual routing, and an extended BIC charges for both its linear amplitude and the search over centers. In five defect-present and five defect-absent trials at each of 0.1% and 0.5% noise, this gate accepts all 10 true defects, rejects all 10 nulls, and selects the exact center u=0.4 in every accepted run. Present-case median rollout changes from 5.652×10−3 to 9.386×10−5 at 0.1% and from 5.634×10−3 to 3.293×10−4 at 0.5%. In null cases the gate returns the original all-probe carrier, so median rollout is exactly unchanged at 3.965×10−5 and 2.039×10−4. This is the operative synthesis: exponential-spline reproduction supplies global extrapolation, residual routing protects pole attribution, and a sparse cardinal edge supplies conditional local repair with no null-case accuracy tax.
Fixed cardinal centers introduce a separate quantization boundary. Defects at grid centers −0.6, 0, and 0.8 are accepted and localized exactly in 9/9 trials at 0.5% noise. With true center 0.35, however, fixed pursuit chooses 0.3 or 0.4 and obtains median/worst rollout 2.034×10−3/7.140×10−3. Retaining cardinal screening but refining only the winning center by bounded scalar variable projection recovers 0.34774, 0.35014, and 0.34989 and reduces median/worst error to 2.715×10−4/4.632×10−4. The extra center degree of freedom is charged in the extended BIC; 5/5 new null cases are rejected with exact carrier fallback.
The offset-jet design remains essential. Under generic zero-centered probes, the same off-grid task has median carrier rollout 0.02291 and adaptive-hybrid rollout 0.02420, with large pole errors and scattered selected centers. The evidence criterion can identify a residual feature in the observed state band, but cannot manufacture the missing state-space separation. Local refinement and annihilator/jet excitation solve different parts of the inverse problem.
The same independence issue controls local order. Applying ordinary EBIC to every overlapping weak row can add a spurious boundary atom in a one-defect case. A probe-block criterion counts non-overlapping temporal windows times orthogonal spatial tests and then adds the combinatorial support penalty. On five null, five one-defect, and five two-defect trials at 0.5% noise, it selects order 0, 1, and 2 correctly in all 15 cases, always recovering the exact nonempty support {0.4} or {−0.6,0.4}. Median rollout is 1.754×10−4, 5.048×10−4, and 4.330×10−4, compared with carrier-only medians 7.106×10−3 and 8.368×10−3 in the nonnull groups. Thus hierarchical local growth is feasible, but its likelihood must be defined on independent experiment blocks.
The remaining bridge is discretization bias. In a matched three-seed audit at 0.5% noise, spectral-data training yields sparse-hybrid median/worst rollout 4.238×10−4/1.005×10−3. A separately implemented centered conservative finite-volume generator transfers partially at 1.405×10−3/1.745×10−3, still below its carrier-only median 6.206×10−3. A Rusanov generator fails: median/worst error becomes 7.414×10−3/3.533×10−2 and pole estimates are biased because the compiler interprets numerical viscosity as physical constitutive signal. Mesh doubling improves matched Rusanov and centered cases only to 2.733×10−3 and 1.464×10−3. Hence weak differentiation removes observation-noise amplification, but it does not remove misspecification of the discrete diffusion operator.
A scalar calibration profile partly repairs this mismatch without using rollout labels. Scanning the viscosity inside the compiler and selecting by weak residual chooses νeff=0.08 in 3/3 Rusanov trials. Median rollout falls from 7.414×10−3 to 2.100×10−3, a 3.5× reduction, and the selected value matches the rollout-oracle grid choice in two trials. Median real-pole error remains about 0.08, however. Scalar operator calibration captures an average modified-equation viscosity; state-dependent numerical diffusion remains a nuisance operator and prevents symbolic attribution.
For centered spatial discretization the continuous–discrete bridge can be exact on every retained Fourier test. Replace the continuum derivative symbols by k1=Δxsin(kΔx),k2=Δx24sin2(kΔx/2).
Discrete-adjoint exactness.
Let Dh be any periodic translation-invariant spatial stencil with discrete Fourier symbol dh(k), and let ⟨⋅,⋅⟩h denote the grid inner product. For every retained Fourier test ψk and sampled field v, one has exactly ⟨ψk,Dhv⟩h=⟨Dh∗ψk,v⟩h=dh(k)⟨ψk,v⟩h. Consequently a weak constitutive design compiled with dh(k) matches the semidiscrete generator on those modes without differentiating the observations. The statement follows because circulant stencils are diagonal in the discrete Fourier basis and their adjoints conjugate the eigenvalues. Continuum weak forms are recovered as dh(k)→(ik)r; using continuum symbols at finite h instead introduces a deterministic bridge bias that no increase in the constitutive dictionary can remove.
Equivalently, rescale ψx by sin(kΔx)/(kΔx) and ψxx by [sin(kΔx/2)/(kΔx/2)]2. This acts only on the analytic tests. Across three centered-volume trials it reduces median/worst rollout from 1.405×10−3/1.745×10−3 to 4.197×10−4/1.003×10−3, matching the spectral-data median 4.238×10−4.
For constant-speed Rusanov, the modified equation adds exactly λΔx/2 to viscosity. The resulting νeff=0.07927 is selected by weak residual in 3/3 paired trials; combined with the discrete symbols it gives median/worst rollout 6.332×10−4/1.607×10−3 and median pole errors (0.00127,0.00253). Without symbol correction the median is 1.939×10−3. For state-dependent Rusanov, exact centered symbols improve the calibrated median only from 2.100×10−3 to 1.849×10−3 and real-pole error remains about 0.081. This isolates the unresolved term: not observation derivatives or stencil mismatch, but the state-dependent numerical-viscosity operator itself.
That nuisance can be removed by one operator fixed point. Initialize from the scalar-calibrated flux, evaluate its state derivative to obtain Rusanov's local face speed, construct the corresponding numerical-viscosity right-hand side, and subtract its projection using the already stored weak test weights. Refit the physical-viscosity carrier and local edge afterward; observations are never differentiated. Across three paired state-dependent Rusanov trials, median/ worst rollout becomes 3.774×10−4/8.934×10−4, versus 1.849×10−3/1.852×10−3 after scalar and exact-symbol calibration and 7.414×10−3/3.533×10−2 without correction. All three recover the true local center, with median (σ,ω) errors (0.00317,0.00890). The discrete nuisance operator can therefore be inferred, compiled out, and separated from the physical constitutive law. The expanded audit selects local order one in all 10 defect cases, with median/worst rollout 4.025×10−4/1.259×10−3 and median pole errors (0.00558,0.00667). All five fresh no-defect controls select order zero and have median/worst rollout 2.210×10−4/2.581×10−4. At 1% observation noise, 5/5 further trials recover the exact one-atom support with median/worst rollout 6.248×10−4/1.173×10−3 and median pole errors (0.00235,0.00459). With two local defects at 0.5% noise, 5/5 trials select exact order two and support {−0.6,0.4}, with median/worst rollout 6.143×10−4/1.100×10−3 and pole errors (0.00700,0.00363). Hence the discrete nuisance projection is stable to both statistical and sparse structural scaling in this controlled solver-transfer test.
The data stencil need not be supplied as an oracle, but it should not be selected by the constitutive residual that it also changes. Full-fit residual selection chooses the correct continuous or exact-centered compiler in only 18/20 cases. A two-probe holdout remains unstable: a frozen 2% preference margin gets all five untouched spectral trials but only one of five centered trials. Raw holdout is 10/10 on those new trials but already missed one pilot. This is a useful negative—operator and constitutive selection are coupled at this sample size.
An independent dispersion fingerprint resolves the coupling. Apply small- amplitude single-mode probes, estimate each mode's exponential amplitude decay and unwrapped phase rate, and profile the unknown diffusion and transport coefficients under either (k,k2) or the centered symbols above. With 64 cells and modes 1–12, the lower normalized modal residual identifies the generator in 80/80 trials at 0.5–1% noise. The identifiability boundary is set by Fourier-symbol separation: at 1% noise, maximum modes 2, 3, and 4 give only 20/40, 23/40, and 28/40 pooled correct decisions, while modes 6 and 8 give 40/40. Modes through 12 give 40/40 at 128 cells and 38/40 at 256 cells; extending the 256-cell probe to mode 20 restores 40/40. Across these cases the reliable design has approximately kmaxΔx≥0.5; below it, the continuous and discrete symbols converge faster than noise permits their separation.
The resulting two-stage protocol is non-oracle: fingerprint the numerical measurement operator, compile its exact adjoint, then identify the nonlinear constitutive edge. On ten seeds per generator, this selection yields median amplitude-OOD rollout 3.053×10−4 for spectral data and 3.074×10−4 for centered data, versus 1.137×10−3 and 1.331×10−3 with the mismatched bridges. The short calibration is therefore not a preprocessing convenience; it is an identifiability experiment that prevents discretization artifacts from being reported as physical poles.
State offsets provide the second excitation axis needed to recognize nonlinear artificial diffusion. Under the Rusanov hypothesis we constrain modal decay to a shared physical viscosity plus ∣F′(u0)∣Δx/2, taking F′(u0) from the phase speed fitted independently at each small-amplitude offset jet. Across spectral, centered, and state-dependent Rusanov generators, five offsets and modes 1–12 give 120/120 correct three-way classifications at 0.5–1% noise. At 1% noise the Rusanov fit recovers physical viscosity with median absolute error 2.72×10−4. A single offset is genuinely nonidentifiable: centered versus Rusanov selection is correct in only 19/40 pooled trials because either model can absorb one effective decay rate into viscosity. Two separated offsets restore 40/40 decisions and median Rusanov viscosity error 2.12×10−4; three offsets also give 40/40. Hence Fourier-mode diversity identifies the derivative stencil, while state-offset diversity identifies its nonlinear viscosity. Nuisance separation comes from orthogonal axes of excitation, not from enlarging the constitutive dictionary.
Offset identifiability.
Linearize a conservative flux about constant states uj and write cj=F′(uj). For a centered semidiscretization, the mode-k eigenvalue is λjkC=−νk2(k)−icjk1(k), whereas local Lax–Friedrichs/Rusanov flux adds, to first order in probe amplitude, λjkR=−(ν+∣cj∣Δx/2)k2(k)−icjk1(k). At one offset the added term is exactly confounded with an unknown ν. At two offsets with ∣c1∣=∣c2∣, the centered hypothesis requires equal decay divided by k2, while the Rusanov hypothesis requires their difference to equal (∣c1∣−∣c2∣)Δx/2. Phase identifies the cj independently through k1. Thus, given one nonzero retained mode and noiseless linearized rates, the two hypotheses and physical viscosity are identifiable from two such offsets. Multiple modes provide noise averaging and distinguish the continuous from discrete symbols; the preceding sweeps quantify the finite-noise bandwidth required for that second distinction.
Finite amplitude exposes the expected bias–variance compromise. Three-way selection remains 30/30 through 5% relative noise and for amplitudes from 0.025 to 0.4, but median Rusanov physical-viscosity error increases from 2.72×10−4 at amplitude 0.025 to 4.88×10−3 at 0.4 because the local linearization is no longer exact. With fixed absolute noise 5×10−4, amplitudes 0.005, 0.01, and 0.025 give 23/30, 29/30, and 30/30 correct decisions; their Rusanov viscosity errors are respectively 1.04×10−4, 1.09×10−4, and 2.73×10−4. Hence an adaptive calibration should increase amplitude only until symbol separation is statistically decisive, then stop before nonlinear bias dominates.
The calibration can replace the final oracle in the nonlinear pipeline. We freeze the independent Rusanov estimate ν=0.040272 and use it in the exact-symbol fixed-point nuisance compiler, rather than resetting to the simulator's true ν=0.04. Across ten new 0.5%-noise constitutive trials, the resulting pipeline selects the one-atom support in 10/10 and has median/worst rollout 5.222×10−4/9.080×10−4. The paired true-viscosity oracle gives 5.252×10−4/9.296×10−4; the median paired error ratio is 1.018. Median pole errors are (0.00462,0.00383) with estimated viscosity and (0.00464,0.00384) with the oracle. Thus the measured modal/offset response supplies all numerical- operator quantities required by the fixed-point constitutive discovery stage.
The candidate stencil can itself be removed. Represent an unknown periodic translation-invariant first derivative by the odd symbol da(k)=Δx1r=1∑Rarsin(krΔx) and its even diffusive part by ℓb(k)=Δx22r=1∑Rbr{1−cos(krΔx)}. The offset-by-mode phase matrix is rank one in the linearized regime, so its right singular vector supplies the shape of da. Phase speed and symbol scale have a gauge; first-derivative consistency removes it through ∑rrar=1. Modal decay identifies the physical coefficients b directly. Stencil radius is then an ordinary held-out model- selection problem.
The companion experiment fits ten calibration records and validates on ten untouched records. Radius one is selected with score 2.641×10−5, versus 2.650×10−5, 2.663×10−5, and 2.668×10−5 for radii two through four; the learned odd and even physical coefficients are 1.0 and 0.0400019. On ten new centered-data constitutive trials, compiling these learned symbols gives median/worst amplitude-OOD rollout 5.6696×10−4/1.1341×10−3, numerically identical to the analytic centered-stencil oracle 5.6689×10−4/1.1342×10−3. The continuum compiler gives 1.446×10−3/2.107×10−3. If the rank-one symbol is instead normalized arbitrarily by d(1)=1, median/worst error is only 9.553×10−4/1.406×10−3. This negative control establishes that continuum consistency, not curve fitting alone, closes the unknown- stencil bridge.
To exclude a three-point coincidence, we repeat the procedure with an independently implemented fourth-order five-point generator. The one-standard- error rule selects radius two at both 0.5% and 1% calibration noise. At 32 cells the recovered odd coefficients are (1.333332,−0.166666), compared with (4/3,−1/6), and the even physical coefficients are (0.0533288,−0.00333008), compared with (0.0533333,−0.00333333). On ten new coarse-grid, eight-mode constitutive trials at 0.1% noise, the learned-symbol compiler yields median/worst rollout 6.480×10−5/2.974×10−4, the analytic-symbol oracle yields 6.349×10−5/2.945×10−4, and the continuum compiler yields 4.323×10−4/7.364×10−4. At 0.5% noise, learned and oracle remain matched at about 5.168×10−4/1.525×10−3, whereas the continuum approximation has a slightly smaller median 4.503×10−4 but a larger maximum 1.927×10−3 and worse pole errors. This is the honest statistical boundary: exact adjoints are required for attribution and control deterministic bias, but a biased low-bandwidth model can occasionally regularize a noisy finite sample.
Calibration cost can be compressed spectrally. With four candidates (continuous, three-point centered, fourth-order centered, and Rusanov), two offsets, 12 separate modes, and 80 steps give 80/80 correct decisions at 1% noise, requiring 24 trajectories. A single multisine can carry all 12 modes: one amplitude-0.005 trajectory at each of two offsets, observed for 160 steps, also gives 80/80. At 0.5% noise and 80 steps it gives 79/80. The shorter 1%-noise, 80-step audit falls to 69/80, demonstrating a temporal-aperture boundary. Replacing log-amplitude/phase regression by a forward–backward complex AR(1) estimate is a negative at 43/80 despite one successful pilot; noise in both consecutive Fourier coefficients invalidates that shortcut. Thus broadband excitation reduces trajectory count by a factor of 12, while time duration remains governed by modal-rate signal to noise.
The unknown-stencil fit can use the same compression. One two-offset multisine record estimates coefficients and two further records validate radius, for six trajectories total. The one-standard-error rule selects the true radius two in 6/6 disjoint three-record groups. Freezing one such calibration and applying it to ten new 32-cell/eight-mode constitutive trials at 0.1% noise gives median/worst rollout 1.215×10−4/3.308×10−4, compared with 4.323×10−4/7.364×10−4 for continuum compilation, 6.480×10−5/2.974×10−4 for the larger learned-symbol calibration, and 6.349×10−5/2.945×10−4 for analytic symbols. Thus six broadband trials already obtain a 3.6× median gain without a named stencil; additional calibration reduces coefficient variance toward the oracle limit.
Sparse instrumentation introduces a third, purely sampling-theoretic boundary. For K=12 multisine modes, we recover modal coefficients by least squares from fixed spatial sensors. Twenty-four uniformly distributed sensors are rank deficient and produce only 29/80 correct four-way decisions. At the exact 2K+1=25 real-sample threshold, the audit gives 80/80 at 0.5% noise and 77/80 at 1%; extending the latter from 160 to 240 time steps restores 80/80. Random placement does not inherit the same conditioning: at 0.5% noise, 25, 32, 40, 48, and 56 random sensors yield respectively 20, 53, 75, 77, and 79 decisions out of 80. This is the cardinal sampling requirement in an experimental-design role: spatial geometry must make the trigonometric frame stable before temporal modal rates can identify the discrete operator.
Uniform time sampling is not required. With full spatial readout, 40 random time stamps over a 240-step aperture retain 80/80 four-way decisions at 1% noise; 40 samples over only 160 steps give 73/80. Under joint sparsity, 25 uniform sensors and 40 random times give 74/80, while 80 random times restore 80/80. Raising the sensor count to 32 but keeping only 40 times gives 75/80, confirming that the temporal rate estimate is then limiting. The successful joint protocol consumes 2×25×80=4000 scalar observations, versus 2×64×241=30848 under full space–time sampling. Above the spatial sampling threshold, identifiability depends on temporal aperture and count, not on a uniform clock.
The final composition learns the stencil from the jointly sparse records and then transfers it across resolution. With 25 uniform sensors and 80 irregular times, a ten-record fit and ten-record validation split selects radius two and recovers odd coefficients (1.33540,−0.16770) and even physical coefficients (0.053235,−0.003306). A stencil coefficient vector is grid independent, whereas its sampled Fourier symbol is not. We therefore evaluate the learned trigonometric polynomial anew on the 32-cell constitutive grid. Across ten new 0.1%-noise nonlinear trials, this compiler selects the exact one-atom support in 10/10 and attains median/worst amplitude-OOD rollout 1.510×10−4/2.707×10−4, versus 1.738×10−4/2.505×10−4 for the analytic-stencil oracle and 4.331×10−4/8.078×10−4 for the continuum compiler. Its paired median error ratio to the oracle is 0.996. Directly reusing the 64-cell sampled factors at 32 cells instead gives median 4.081×10−4, a 2.7× regression. The learned operator must therefore cross resolutions as coefficients and be recompiled at the target mesh; this is the discrete counterpart of transporting a cardinal spline by its generator rather than by samples tied to one grid.
This sparse design also exposes a replication and experimental-design boundary. The initial symmetric-offset panel yielded 80/80 classifications, but a disjoint matched panel yields only 35/40 (38/40 with twice as many retained times). At offsets (−0.4,+0.4) the two values of ∣F′(u0)∣ are too similar, so state-dependent numerical viscosity is nearly coherent with the shared physical-viscosity column. Changing only the offsets to (−0.4,+0.8) raises a larger validation panel from 72/80 to 77/80 at the same measurement budget, while four wide/asymmetric pairs each give 40/40 in the pilot. If perturbed sensor coordinates are supplied to the Fourier frame, even one-cell jitter remains 74/80; if a 0.025-cell perturbation is unmodelled, accuracy falls to 58/80. We therefore define the normalized selection margin (S(2)−S(1))/S(2) and freeze a threshold 0.15 from the pilot. It accepts 210 of 320 cases in the subsequent multi-regime validation and all 210 accepted decisions are correct. The resulting rule is operator-theoretic: choose offsets that reduce column coherence, include sensor geometry in the analysis operator, and abstain whenever the candidate quotient is not separated.
The same excitation can identify the analysis geometry. Let qj(x) denote the known initial multisine of probe j and let xˉm be the nominal sensor position. We estimate its displacement independently by
xm=arg∣x−xˉm∣≤ρΔxminj∑{yjm(0)−qj(x)}2,
then build the trigonometric sampling matrix at xm. Independent probe phases make the local code increasingly injective as j grows. At 1% noise and unknown 0.1-cell jitter, two, three, and five probes give 33/40, 37/40, and 38/40 classifications, compared with 19/40 under nominal coordinates. In an 80-case five-probe validation, nominal, self-calibrated, and known coordinates yield 31/80, 75/80, and 79/80. Median coordinate error is 0.00370 cells and the median trial-wise maximum is 0.01493; median Rusanov viscosity error improves from 3.74×10−4 nominal to 1.14×10−4, with known-position error 6.50×10−5. The frozen margin gate accepts 55 self-calibrated cases and all 55 are correct. Thus the excitation first self-surveys the sampling operator and only then identifies the discrete differential operator.
For harmonic generators the shift law eliminates even the nonlinear local search. Two static fields give (yms,ymc)=A(sinKxm,cosKxm)+ϵm, so
xm=K1atan2(yms,ymc)(mod2π/K),
with the branch nearest xˉm selected. Its small-noise position variance scales as K−2, motivating a locator mode above the dynamics band. Raising K from 12 to 28 reduces pilot median coordinate error from 0.00406 to 0.00170 cells. On 80 new trials with 0.1-cell unknown jitter, nominal, quadrature-calibrated, and known coordinates give 24/80, 78/80, and 79/80 classifications. Quadrature calibration has median coordinate error 0.00178 cells and median trial-wise maximum 0.00602; its median Rusanov viscosity error is 8.82×10−5, between the known-position 6.56×10−5 and nominal-position 1.11×10−3. The frozen margin gate accepts 68 cases and all 68 are correct. Only two 25-value static snapshots are added to two 25×80 dynamics records, for 4050 scalar measurements total. This is a direct algorithmic consequence of exponential reproduction: translation becomes phase, so the sampling operator can be self-calibrated before the differential operator is learned.
The fine phase has branch radius nx/(2K) in grid-cell units. For K=28 this is 1.14 cells: the single-frequency locator changes from 36/40 correct at one-cell jitter to only 8/40 at 1.25 cells, with errors of one phase period (2.29 cells). A coarse-to-fine construction first decodes mode 8 and chooses the mode-28 branch nearest that estimate. It maintains median and median- maximum coordinate errors of approximately 0.0017 and 0.0057 cells through two-cell jitter, and every confidence-gated downstream decision is correct. The remaining loss follows conditioning of the irregular spatial frame, even when positions are known. At two-cell jitter, 64 sensors and 80 times give 40/40 known and 39/40 self-calibrated classifications, while 40 sensors and 160 times give 39/40 and 38/40. The former lowers the median sampling-frame condition number to about 3.06. Thus multi-frequency exponential reproduction sets the coordinate range, and spatial oversampling controls the subsequent modal inversion variance.
The composed experiment removes both geometry and named-stencil oracles. Twenty quadrature-self-calibrated five-point records, each with 25 sensors and 80 irregular times under hidden 0.1-cell jitter, select radius two and recover odd coefficients (1.33978,−0.16989) and even physical coefficients (0.053134,−0.003260). We transfer these coefficients and re-evaluate their symbols on a 32-cell grid. Across ten new nonlinear discovery trials, exact local support is selected in 10/10. Median/worst amplitude-OOD rollout is 2.091×10−4/4.544×10−4 for the learned compiler, 1.498×10−4/4.012×10−4 for the analytic oracle, and 4.603×10−4/9.046×10−4 for continuum compilation. Learned beats continuum in every paired trial and by a factor 2.2 in median; its median penalty relative to the oracle is 1.37. Hence the sampling geometry, discrete adjoint, target-grid symbol, and sparse constitutive innovation can all be identified sequentially from designed measurements, with a quantified final calibration-variance cost.
The clean nonseparable result has a stricter noise boundary. Pointwise tensor selection, temporal plus spatial smoothing, ridge variation over seven orders, and increasing the generic trajectory count from 8 to 64 all leave the interaction-function error near one. Prescribed initial jets remove regressor noise but require a boundary time derivative; even 256 averaged bursts do not make that route symbolic. We therefore compile a space–time weak tensor design in which the temporal, diffusion, and conservative derivatives act on analytic tests, while only the irreducible nonlinear sin(ux) factor is evaluated from the state. Although one pilot recovers 0.04725sin(2u)sin(ux) for a true coefficient 0.05, a 45-case audit does not reproduce atom identity reliably. The robust result is predictive: a development-frozen mixture with 25% weak tensor and 75% weak scalar beats the scalar in 10/10 new γ=0.05, 0.5%-noise trials, reducing median/worst amplitude-OOD rollout from 3.776×10−3/8.016×10−3 to 3.203×10−3/5.558×10−3. Thus the operator-compatible tensor edge contributes below the symbolic resolution threshold, but must be averaged rather than interpreted.
Repeated observation reveals a coherence rather than a variance floor. Over five fixed clean ensembles, increasing independent noisy replicates from one to 32 decreases median interaction nRMSE from 0.1299 to 0.0204 and median rollout error from 7.20×10−4 to 1.79×10−4, but exact atom recovery saturates at four of five; 64–512 replicates do not remove the failure. Increasing random trajectories from 8 to 64 is non-monotone (2/5, 3/5, 4/5, and 3/5 exact). More importantly, a coefficient-frequency certificate developed over repeated noise is externally falsified: one of five new clean ensembles stably certifies two false surrogate interactions. Stability to measurement noise is not identifiability of the physical law.
Weak-feature excitation design nevertheless provides a strong intermediate result. Selecting eight of sixteen candidate trajectories by greedy D-optimality improves exact recovery from 3/10 to 7/10, reduces median/worst interaction nRMSE from 0.997/19.16 to 0.0385/1.009, and reduces median/worst rollout from 2.730×10−3/1.342×10−1 to 1.125×10−3/1.532×10−2, with eight paired rollout wins. This does not close the symbolic gate: selecting 8 or 12 from 32 candidates retains a shared catastrophic seed, and quotienting scalar features before D-optimal selection worsens exact recovery from 4/5 to 3/5. The conclusion is that designed excitation is the right control variable, but determinant volume and replicate confidence alone do not resolve nuisance coherence.
The failure atoms identify the missing design coordinate: coverage of the nonlinear argument. On the original amplitude-0.65 trajectories, sin(2u) remains coherent after weak projection with higher trigonometric surrogates. Holding the weak compiler and sparse selector fixed, amplitude 0.95 yields 4/5 exact recoveries, whereas amplitude 1.20 with D-optimal selection yields 5/5. A frozen ten-seed validation then gives 9/10 exact for random wide-amplitude trajectories and 10/10 for D-optimal wide-amplitude trajectories. The D-optimal coefficient lies in [0.04945,0.05045] for truth 0.05, with median/worst interaction nRMSE 0.00400/0.01096. Its median rollout is neutral relative to random (8.62×10−4 versus 8.13×10−4, five paired wins), but its worst rollout improves from 4.03×10−3 to 1.42×10−3. A no-averaging audit is stronger: D-optimal wide-amplitude design recovers the exact atom in 5/5 pilot and 10/10 untouched single-observation seeds, compared with 5/5 and 9/10 for random wide-amplitude design. Its validation median/worst interaction nRMSE is 0.00674/0.04524 versus 0.01620/0.66393, and median/worst rollout is 6.77×10−4/8.18×10−4 versus 7.54×10−4/2.64×10−3. This closes the noisy symbolic gate: operator compilation removes derivative noise, argument-range coverage separates nonlinear atoms, and D-optimal selection controls the remaining support tail. Replication refines coefficients but is not the source of identifiability.
Without replicate averaging, a five-seed-per-level noise sweep gives exact support in 5/5 cases at 1%, 2%, and 3% noise. The corresponding median/worst interaction nRMSE values are 0.0308/0.0568, 0.0730/0.1143, and 0.1967/0.2720, so quantitative accuracy degrades before atom identity. At 5%, 7.5%, and 10%, support drops to 4/5, 3/5, and 1/5, median interaction nRMSE rises to 0.661, 1.663, and 4.060, and median rollout rises to 7.55×10−3, 3.82×10−2, and 6.88×10−2. Selected clean trajectories cover approximately u∈[−1.24,1.22]. Hence the useful quantitative operating regime is about 2% noise or less; support alone at 3% must not be read as an accurate recovered law.
At fixed 0.5% raw noise, lowering the true interaction coefficient to 0.02, 0.01, and 0.005 gives exact support in 5/5, 4/5, and 2/5 cases, with median/worst interaction nRMSE 0.0387/0.1093, 0.1246/1.0, and 1.0/1.392. Rollout medians nevertheless stay near 8×10−4 because the omitted term becomes dynamically small. This separates a discovery threshold near coefficient 0.02 from a much weaker prediction threshold and shows why rollout agreement alone cannot certify a learned physical edge.
Configurable target tests establish both transfer and a parity boundary. With the same wide-amplitude D-optimal single-observation protocol, sin(3u)sin(ux) and cos(2u)sin(ux) are each recovered exactly in 5/5 new seeds, with median/worst interaction nRMSE 0.00329/0.00875 and 0.00991/0.01346. In contrast, sin(2u)cos(ux) is exact in only 2/5 with median nRMSE 0.436: near zero gradient, its even factor has a reaction-like constant component. IC frequency scaling by two or three yields only 0/5 and 2/5; at scale two, shortening weak windows to 8, 16, or 24 steps yields 0/3 at every setting. Hence the method transfers over state-side poles and odd-gradient factors, but even-gradient terms need an explicit gauge quotient or controlled gradient offset rather than indiscriminate high-frequency excitation.
The appropriate repair is an operator quotient. Replacing each even-gradient interaction by a(u)[cos(qux)−1] assigns its null-gradient component to the reaction edge and leaves only irreducible gradient dependence in the tensor edge. This centered basis is exact in 5/5 pilot seeds. On ten untouched paired seeds, the ordinary basis is exact in only 4/10 with median/worst interaction nRMSE 0.441/0.453; the quotient basis is exact in 10/10 with 0.00805/0.0408. Median/worst rollout improves from 5.64×10−4/7.95×10−4 to 2.48×10−4/3.88×10−4. Thus hierarchical edge ownership can be compiled algebraically: annihilate each higher-order atom at the reference jet of every lower-order edge before selection. This is an identifiability operation rather than a numerical preconditioner. With this quotient fixed, random excitation is likewise exact in 10/10 (median/worst interaction nRMSE 0.0133/0.0303), so the algebra itself closes the support gate. D-optimal selection chiefly controls the dynamic tail, reducing median/worst rollout from 3.19×10−4/1.12×10−3 to 2.48×10−4/3.88×10−4. The quotient transfers to cos(2u)[cos(ux)−1] and sin(3u)[cos(2ux)−1], each exact in 5/5 seeds, with median/worst interaction nRMSE 0.00994/0.0301 and 0.0467/0.0685. The higher-gradient case has median/worst rollout 8.09×10−4/1.26×10−3. The state-cosine case reveals a residual hierarchy boundary: two fits misidentify the lower-order reaction component, producing worst rollout 0.1156 despite exact tensor attribution. Thus the quotient is reusable, but its receiving lower-order edge also requires robust staged identification. A one-pass lower-first scheme is a negative control: selecting five base atoms before one interaction fails in 0/5, with median/worst interaction nRMSE 0.898/0.909 and rollout 8.25×10−3/8.60×10−2. Freezing an early surrogate makes the hierarchy irreversible. The next solver must use alternation or hierarchical group constraints within a joint objective.
The negative results determine the architecture. The plain cardinal model is excellent in interpolation and resolution transfer but has median amplitude- OOD errors of order 0.16–0.19. A cubic carrier only partly repairs the non-polynomial cases. Once the correct pole carrier is selected, fitting a local spline innovation is neutral or harmful. Hence the current hypothesis is narrower and stronger than generic KAN substitution: discover a sparse operator reproduction space for the unknown constitutive edge, compile exact outer calculus, and introduce local spline innovations only when held-out evidence demonstrates residual mismatch.
Limitations
The current implementation establishes the algebraic path but is not yet a complete scientific library. Important limitations remain:
E-spline basis functions are represented in simplified benchmark-specific forms.
FRI moment inputs are exact in the current verification scripts; noisy measurement pipelines remain future work.
Boundary conditions must be formalized per operator: periodic, finite interval, causal, or corrected by null-space constraints.
The first multidimensional tensor-product Hermite solver is implemented, but it remains a synthetic periodic validation rather than a calibrated external CFD benchmark.
Fourth-order structural mechanics is validated on a manufactured clamped plate, but it still requires dedicated biharmonic preconditioning and external benchmark comparison.
Benchmark comparisons against external PINN/SIREN baselines must be standardized and repeated under controlled conditions.
The operator-compiled constitutive result now includes controlled 2D anisotropic and coupled two-field laws, noisy weak multipole prediction, and identified noisy bivariate interactions, but noisy pole-level identifiability, broader continuous interaction dictionaries, high-dimensional interaction selection, and external trajectories remain open validation gates.
Conclusion
OSNR is best understood as a continuous-domain signal-processing architecture that adopts the interface of neural representations while rejecting their blind optimization core. For operator-bound fields, the correct spline dictionary collapses training into stable coefficient recovery. For sparse non-Gaussian fields, FRI localization and matched sparse dictionaries avoid the grid leakage and coherence traps of uniform frames. The external sparse-assimilation results add a practical systems role: OSNR can act as a deterministic test-time correction layer on top of physical or neural priors, with the Darcy U-Net adapter, the PDEBench Test 28 temporal FNO rescue, and the station-gated Test 28 neural-prior DST high-pass ladder showing the same mechanism in static elliptic and time-dependent vorticity settings. The engineering rule is strict: continuous-domain exactness only survives when the discrete bridge is correct. That bridge consists of calibrated knot maps, stable bases, proper inner products, cross-Gram coupling, boundary-aware solvers, and regularized inverses where identifiability fails.
appendix
Hermite Block-Circulant Inner Products
This appendix records the coefficient-domain inner-product calculus for the second-order Hermite tier. It is included explicitly because the OSNR use of the full autocorrelation tensor is a library-level construction rather than a theorem that can be cited as a pre-existing implementation recipe.
Multichannel synthesis
Let the Hermite generator be Φ(t)=[ϕ0(t)ϕ1(t)ϕ2(t)]⊤, with suppϕp⊂[−1,1]. A coefficient sequence is c[k]=[c0[k]c1[k]c2[k]]⊤. The continuous field is f(t)=k∈Z∑c[k]⊤Φ(t−k)=p=0∑2k∈Z∑cp[k]ϕp(t−k).
Autocorrelation tensor
The continuous L2 energy expands into nine generator-pair channels: ∥f∥L22=∫Rf(t)2dt=p=0∑2q=0∑2k∈Z∑m∈Z∑cp[k]cq[m]∫Rϕp(t−k)ϕq(t−m)dt. Define Γpq[n]=∫Rϕp(t)ϕq(t+n)dt. Changing variables gives ∫Rϕp(t−k)ϕq(t−m)dt=Γpq[k−m]. Therefore ∥f∥L22=p=0∑2q=0∑2k,m∑cp[k]Γpq[k−m]cq[m]. The tensor Γpq is the Hermite analogue of the scalar spline autocorrelation filter used in periodic spline inner-product calculus [badoual2016inner,badoual2018periodic].
Because every ϕp is supported on [−1,1], Γpq[n]=0 for ∣n∣≥2 in the non-periodic infinite-line case. Only the shifts n∈{−1,0,1} can contribute. This local support is the algebraic reason the spatial-domain block Gram is sparse before periodization.
Block Toeplitz and block-circulant matrices
For a finite coefficient vector with M knots, define the M×M block [Γpq]k,m=Γpq[k−m]. On the infinite line or with finite non-periodic truncation, these blocks are Toeplitz away from boundary corrections. Under periodic boundary conditions, indices are taken modulo M, and each block becomes circulant: [Γpqper]k,m=Γpqper[(k−m)modM], where Γpqper[n]=ℓ∈Z∑Γpq[n+ℓM]. The full Hermite Gram matrix is the 3M×3M block matrix ΓH=Γ00Γ10Γ20Γ01Γ11Γ21Γ02Γ12Γ22. The transpose symmetry follows directly from the definition: Γpq[n]=Γqp[−n],Γpq=Γqp⊤. Thus the diagonal blocks are symmetric, while the off-diagonal blocks need not be symmetric individually. In particular, the slope channel is odd/asymmetric for the standard Hermite construction, so value-slope and curvature-slope blocks encode directional cross-talk.
Fourier block diagonalization
Let Γpq[ℓ] be the M-point DFT of the first column of the circulant block Γpqper. The DFT simultaneously diagonalizes all nine circulant blocks: Γpqper=F−1diag(Γpq[0],…,Γpq[M−1])F. After applying the DFT to the knot dimension of each Hermite channel, the large 3M×3M system decouples into M independent 3×3 Hermite channel systems: Γ[ℓ]=Γ00[ℓ]Γ10[ℓ]Γ20[ℓ]Γ01[ℓ]Γ11[ℓ]Γ21[ℓ]Γ02[ℓ]Γ12[ℓ]Γ22[ℓ]. For a right-hand side with Hermite-channel DFT coefficients b[ℓ]∈C3, the exact periodic normal-equation solve is c[ℓ]=(Γ[ℓ]+γI3)−1b[ℓ],ℓ=0,…,M−1. The ridge γ is optional for strictly Riesz-stable settings but mandatory in finite precision whenever boundary constraints, redundant channels, or composite dictionaries create near-null directions. The complexity is O(3MlogM) for the channel FFTs plus O(27M) for the M dense 3×3 solves, instead of O((3M)3) for a dense inversion.
Physical meaning for OSNR Tier 3
The block matrix is not a bookkeeping artifact. It is the exact continuous L2 metric for value, slope, and curvature streams. For a Hermite neural operator, a branch encoder may output coefficient tensors, but the comparison of predicted and target fields should be performed through ⟨f,g⟩L2=p,q=0∑2cf,p⊤Γpqcg,q, not through a sampled coordinate loss unless sampling is required by the measurement model. Boundary clamping is similarly direct: at a boundary knot kb, Dirichlet, Neumann, and curvature data are imposed by assigning c0[kb], c1[kb], and c2[kb]. The Hermite tier therefore converts soft boundary penalties into coefficient constraints and converts continuous PDE energies into block-circulant linear algebra.
2D tensor-product block-circulant calculus
For the 2D tensor-product Hermite generator, index the nine channels by a=(px,py),b=(qx,qy),px,py,qx,qy∈{0,1,2}. The generator pair is Ha(x,y)=ϕpxx(x)ϕpyy(y). The 2D cross-correlation filter between channels a and b is Γab2D[nx,ny]=∬R2Ha(x,y)Hb(x+nx,y+ny)dxdy. By separability, Γab2D[nx,ny]=Γpxqxx[nx]Γpyqyy[ny]. Since each one-dimensional Hermite generator is supported on [−1,1], the nonzero spatial shifts satisfy nx,ny∈{−1,0,1}. For a periodic Mx×My grid, every channel pair defines a block-circulant-with-circulant-blocks matrix. The full 2D Hermite Gram is ΓH,2D=Γ00⋮Γ80⋯⋱⋯Γ08⋮Γ88, where each Γab is the circulant 2D convolution operator associated with Γab2D.
Let Γab2D[νy,νx] be the 2D DFT of the first column/filter of the (a,b) block. Applying a 2D DFT to the spatial dimensions of all coefficient channels yields, for every frequency coordinate (νy,νx), the local 9×9 system Γ2D[νy,νx]c[νy,νx]=b[νy,νx], with Γ2D[νy,νx]=Γ002D[νy,νx]⋮Γ802D[νy,νx]⋯⋱⋯Γ082D[νy,νx]⋮Γ882D[νy,νx]. Thus the global dense inverse is replaced by c[νy,νx]=(Γ2D[νy,νx]+γI9)−1b[νy,νx]. In code, this is the torch.fft.fft2 path implemented by the 2D block-Fourier solver. The spatial complexity is governed by the FFTs, O(9MxMylog(MxMy)), while each frequency bin performs a constant-size 9×9 complex solve. This is the algebraic mechanism behind the measured 1.28 ms per-frame 2D tensor-Hermite fluid validation and the fourth-order biharmonic structural-shell validation.
\documentclass[11pt]{article}
\usepackage[margin=1in]{geometry}
\usepackage{amsmath,amssymb,amsthm,mathtools}
\usepackage{bm}
\usepackage{booktabs}
\usepackage{graphicx}
\usepackage{longtable}
\usepackage{hyperref}
\usepackage{enumitem}
\usepackage{xcolor}
\usepackage{float}
\hypersetup{
colorlinks=true,
linkcolor=blue!50!black,
citecolor=blue!50!black,
urlcolor=blue!50!black
}
\newtheorem{definition}{Definition}
\newtheorem{proposition}{Proposition}
\newtheorem{remark}{Remark}
\newcommand{\R}{\mathbb{R}}
\newcommand{\Z}{\mathbb{Z}}
\newcommand{\C}{\mathbb{C}}
\newcommand{\dd}{\mathrm{d}}
\newcommand{\calL}{\mathcal{L}}
\newcommand{\calN}{\mathcal{N}}
\newcommand{\bs}{\boldsymbol}
\newcommand{\A}{\mathbf{A}}
\newcommand{\G}{\mathbf{G}}
\newcommand{\I}{\mathbf{I}}
\newcommand{\cvec}{\mathbf{c}}
\newcommand{\xvec}{\mathbf{x}}
\newcommand{\yvec}{\mathbf{y}}
\newcommand{\wvec}{\mathbf{w}}
\newcommand{\zvec}{\mathbf{z}}
\newcommand{\uvec}{\mathbf{u}}
\DeclareMathOperator*{\argmin}{arg\,min}
\title{Operator-Spline Neural Representations: \\
Continuous-Domain Operator Bases for Efficient Physical Learning}
\author{OSNR Project Notes}
\date{\today}
\begin{document}
\maketitle
\begin{abstract}
Operator-Spline Neural Representations study how known continuous structure
can be represented and manipulated through discrete coefficients. For
admissible constant-coefficient operators, polynomial and exponential
B-splines provide compact local generators, operator-null-space reproduction
and analog-to-digital filtering identities. Hermite generators provide
explicit value and derivative coordinates. This document develops those
continuous--discrete connections and records the implementation and empirical
conditions under which they are useful.
The central computational objects are continuous function, derivative and
cross-basis Grams. Fixed generators and grids permit their reuse across
coefficient updates; projection, differential energies and adjoints then
become structured linear algebra. Periodic equal-spacing settings admit
circulant or block-circulant realizations, whereas unequal-grid cross-Grams,
finite-boundary operators and data-dependent statistics need their actual
structure respected. The implementations address canonical endpoint handling,
stable reciprocal-root factorizations, positive spectra, banded/bordered
solves, exact polynomial-piece integration and local matrix/Horner execution.
Grid refinement preserves a function only with the required subspace
embedding; changing a feature family does not in general preserve sufficient
statistics for unseen directions.
Recent extensions apply this calculus to compact causal physical models.
For fixed spline directions and an unclipped regression model, a stable
first-order actuator permits a joint constrained quadratic fit after absorbing
its pole-dependent scale into static spline coefficients.
For monotone fitted laws, explicit saturation also admits finitely many
ordered data regions with constrained quadratic subproblems; batched
quadratic lower bounds prune most solves while retaining numerical
optimality checks. This is not a claim that the clipped objective is
globally convex or that every fit meets a fixed real-time bound.
Combining bounded-function products with stable
exponential response sums yields an exact infinite-response inner product
and a compact source-bank descriptor. This is exponential response calculus
around a cardinal cubic model, not an exponential B-spline neural activation
experiment or a new universal-policy theorem.
A thirty-pool temporal-task transfer experiment provides a further boundary:
cardinal support planning solves 161/480 composed missions versus MLP 242/480,
with sequential reuse matching both. Exact Bernstein specialization preserves
audited development outputs while reducing planning time by
$2.10$--$2.46\times$, but informative state acquisition and useful composition
do not follow from correct local calculus. This experiment uses cubic fields,
not the full exponential/Hermite physical-operator toolbox.
A measured CT extension combines exact rectangular field queries, finite-domain
hierarchical product matrices, incremental likelihood information and full
Gaussian block-design solves. Its strongest classical control answers 47 of
72 late questions without confident mistakes, but fails the fixed coverage
target and gains no measurement saving from targeted acquisition. Exact-output
memory compilation does not by itself establish useful uncertainty or a
spline-specific inspection advantage.
The statistical and behavioral claims remain conditional. Generalized
increments filter continuous innovations and need not be independent.
Operator matching alone does not guarantee lower prediction risk or sparse
support adaptation by ridge. Pooled fixed-feature statistics preserve a
fitting objective, not every old prediction or private datum; protected
supports and immutable archived programs provide different preservation
contracts. Physical identification also does not imply useful closed-loop
recovery: the accompanying controlled experiments include both gains and
failed restoration or comparator criteria. The resulting foundation is an
auditable set of representation, calculus and compilation mechanisms, with
explicit boundaries on observability, model mismatch, numerical conditioning,
memory and downstream performance.
\end{abstract}
\tableofcontents
\section{Introduction}
Coordinate networks have become a standard tool for representing continuous signals and fields. A typical implicit neural representation (INR) maps coordinates to field values through a multilayer perceptron,
\[
f_\theta : \xvec \mapsto y,
\]
where the weights $\theta$ are optimized by stochastic gradient descent. SIRENs improve high-frequency representation by using sinusoidal activations, while PINNs add differential-equation residuals to the training objective.
The OSNR thesis is that this approach is structurally misaligned for physical systems governed by known operators. It is inspired by the operator-based spline signal-processing program of Unser, Blu, Vetterli, and collaborators \cite{unser1993bspline1,unser1993bspline2,unser2005cardinal1,unser2005cardinal2,vetterli2002fri}. If a field is constrained by
\[
L\{s\} = r,
\]
then the representation should be built from the operator $L$ itself. The role of learning or numerical inversion should be reduced to coefficient recovery in an already appropriate continuous function space.
OSNR therefore replaces black-box nonlinear layers by operator-matched spline dictionaries. The spline coefficients play the role of the latent representation. The architecture is not a generic neural network with a different activation; it is a continuous-domain inverse-problem engine presented in neural-representation form, aligned with continuous-domain inverse-problem representer theorems \cite{gupta2018continuous,debarre2019hybrid,debarre2021composite}. The same principle also suggests a deterministic alternative to the trunk side of branch-trunk neural operator models such as DeepONet and MIONet \cite{lu2021deeponet,jin2022mionet}: the coordinate-to-field map can be a spline synthesis operator with exact calculus rather than a learned MLP.
The empirical record tests this thesis rather than establishing universal
superiority. \S\ref{sec:grown-topology} covers closed-form identification,
grown-topology control and continual-learning comparisons;
\S\ref{sec:matched-vs-sindy} compares operator-matched identification with
SINDy and neural ODEs, including a negative blind-discovery control.
\S\ref{sec:pde-vs-fno} studies PDE comparisons with neural operators, regime
changes and partially known physics. Their measured gains depend on the
information supplied, comparator, accuracy and computational accounting:
knowing an operator family is not by itself a sufficient condition for an
advantage. The later exact battery-response experiment, for example, admits
a nearly equally fast training-free classical control; the real battery-aging
fit loses to its matched polynomial comparator.
\S\ref{sec:ssp-view} develops the sparse-stochastic-process connection under
its stated assumptions, not a universal statistical dominance theorem.
\S\ref{sec:rsi} records recursive-improvement experiments, while subsequent
claim audits distinguish pooled fitting statistics, protected function
supports, immutable archives and retained task performance. None alone
establishes unrestricted self-improvement without forgetting or drift.
The constructive result is a tested representation and calculus toolbox;
independent-task benefit against strong controls remains the capability test.
\section{Mathematical Background}
\subsection{Cardinal spline representation}
The classical cardinal spline model represents a continuous signal as \cite{unser1993bspline1,unser1993bspline2}
\[
s(t) = \sum_{k \in \Z} c[k] \beta(t-k),
\]
where $\beta$ is a spline generator and $c[k]$ are discrete coefficients. The essential point is that a continuous function space is controlled by a discrete sequence. This makes exact digital processing of continuous objects possible, provided the generator is stable and the correct coefficient-domain filters are used.
\subsection{Linear differential operators and null spaces}
Let $L$ be a constant-coefficient differential operator with characteristic roots or poles
\[
\bs{\alpha} = (\alpha_1,\ldots,\alpha_N).
\]
The null space is
\[
\calN_L = \{s : L\{s\}=0\}
= \mathrm{span}\{e^{\alpha_m t}\}_{m=1}^N
\]
with polynomial factors included for repeated roots. For the Helmholtz operator
\[
L = \frac{\dd^2}{\dd x^2} + k^2,
\]
the poles are $\alpha=\pm j k$, and the null space consists of sinusoidal modes.
\subsection{Cardinal exponential splines}
Cardinal exponential splines are compactly supported spline generators matched to $\bs{\alpha}$ \cite{unser2005cardinal1}. Their Fourier-domain form is
\[
\widehat{\beta}_{\bs{\alpha}}(\omega)
=
\prod_{m=1}^N
\frac{1 - e^{\alpha_m - j\omega}}{j\omega-\alpha_m}.
\]
Integer shifts of $\beta_{\bs{\alpha}}$ generate a stable spline space under appropriate Riesz conditions. The basis reproduces exponential polynomials and supports exact operator calculus in the coefficient domain.
\begin{definition}[Operator-matched OSNR dictionary]
Given an operator $L$ with pole vector $\bs{\alpha}$ and physical knot spacing $T$, an OSNR dictionary is a matrix of samples
\[
\A_{n,k} = \beta_{\bs{\alpha}}\!\left(\frac{x_n}{T}-k\right),
\]
augmented when necessary by explicit null-space columns. The continuous field is represented by
\[
s(x_n) \approx (\A \cvec)_n.
\]
\end{definition}
\subsection{Exponential Hermite splines}
Cardinal E-splines attach the operator to a scalar coefficient stream. For higher-order boundary value problems, OSNR requires a vector-valued cardinal system whose coefficients store not only function values but also derivatives. The second-order exponential-polynomial Hermite generator of Schmitter, Badoual, Uhlmann, Fageot, and Unser \cite{schmitter2016hermite} provides exactly this structure.
Let
\[
\Phi(t)=
\begin{bmatrix}
\phi_0(t) & \phi_1(t) & \phi_2(t)
\end{bmatrix}^{\top}.
\]
The Hermite spline expansion is
\[
f(t)
=
\sum_{k\in\Z}
\left(
c_0[k]\phi_0(t-k)
+c_1[k]\phi_1(t-k)
+c_2[k]\phi_2(t-k)
\right),
\]
with coefficient vectors
\[
\mathbf{c}[k]=
\begin{bmatrix}
c_0[k] \\ c_1[k] \\ c_2[k]
\end{bmatrix}
=
\begin{bmatrix}
f(k) \\ f'(k) \\ f''(k)
\end{bmatrix}
\]
for exactly interpolated data. The interpolation conditions are
\[
\phi_p^{(r)}(k)=\delta_{pr}\delta_k,
\qquad p,r\in\{0,1,2\},\quad k\in\Z.
\]
On the interval $[0,1]$, each channel has the exponential-polynomial form
\[
\phi_p(t)
=
A_p+B_p t+C_p t^2+D_p t^3
+E_p e^{j\omega_0 t}+F_p e^{-j\omega_0 t},
\qquad
\omega_0=\frac{2\pi}{M}.
\]
The constants $A_p,\ldots,F_p$ are determined by the six Hermite endpoint constraints
\[
\phi_p^{(r)}(0)=\delta_{pr},
\qquad
\phi_p^{(r)}(1)=0,
\qquad r=0,1,2.
\]
The negative branch is fixed by the Hermite parity
\[
\phi_0(-t)=\phi_0(t),\qquad
\phi_1(-t)=-\phi_1(t),\qquad
\phi_2(-t)=\phi_2(t),
\]
so the generator is compactly supported on $[-1,1]$. This construction is $C^2$, reproduces polynomials up to cubic degree, and reproduces the trigonometric modes $\sin(\omega_0 t)$ and $\cos(\omega_0 t)$ through the coefficient samples of the function and its first two derivatives.
For OSNR Tier 3, the crucial change is semantic: Dirichlet, Neumann, and curvature boundary data become direct coefficient assignments. A clamped boundary at knot $k_b$, for example, is imposed by setting
\[
\mathbf{c}[k_b]
=
\begin{bmatrix}
0 \\ 0 \\ 0
\end{bmatrix},
\]
rather than by adding a soft loss term. The same vector-valued structure also produces the block-circulant Hermite Gram system derived in Appendix~\ref{app:hermite-block-gram}.
\section{OSNR Architecture}
\subsection{Master core apparatus}
The current OSNR research code separates into three production tiers. Each tier has a different admissible function space, solver topology, and verification target.
\begin{center}
\scriptsize
\begin{tabular}{@{}p{0.3\linewidth}p{0.3\linewidth}p{0.3\linewidth}@{}}
\toprule
\multicolumn{3}{c}{\textbf{OSNR Master Packaging Core}} \\
\midrule
\textbf{Tier 1: Steady-State} &
\textbf{Tier 2: Adaptive} &
\textbf{Tier 3: Neural Operator} \\
\midrule
Uniform E-splines &
Non-uniform shift-splines &
2D tensor-product Hermite \\
$O(M\log M)$ circulant FFT &
TLS matrix-pencil FRI &
9-stream local calculus \\
Fixed pole-locking inversion &
Oblique cross-Gram shield &
Parallel $9\times9$ DFT solver \\
\midrule
Profile: $314.86$ dB, $4.21$ ms &
Profile: $118.78$ dB, $97.7\%$ sparsity &
Profile: $183.77$ dB, $1.28$ ms 2D CFD \\
\bottomrule
\end{tabular}
\end{center}
\subsection{Tier 1: deterministic operator-bound systems}
Tier 1 targets smooth physical fields that lie in, or close to, the null space of a known linear operator. Examples include Helmholtz waves, damped oscillators, and linear constant-coefficient PDE components. The solver pipeline is:
\begin{enumerate}[leftmargin=2em]
\item identify $L$ and its poles $\bs{\alpha}$;
\item construct calibrated E-spline or null-space dictionaries;
\item recover coefficients through a stable linear solve or circulant FFT inversion;
\item compute differential quantities through operator identities rather than autograd.
\end{enumerate}
\subsection{Tier 2: adaptive sparse continua}
Tier 2 targets composite fields with sparse discontinuities or shocks:
\[
s = s_{\mathrm{smooth}} + s_{\mathrm{sparse}}.
\]
Uniform grids are not sufficient for non-bandlimited discontinuities at sub-grid coordinates. The sparse tier must first identify finite-rate innovations, then adapt the dictionary to those coordinates:
\begin{enumerate}[leftmargin=2em]
\item estimate innovation locations using FRI or matrix-pencil methods;
\item snap sparse knots to the recovered coordinates;
\item solve a cross-Gram-coupled sparse-plus-smooth inverse problem;
\item debias with scale-invariant ridge stabilization.
\end{enumerate}
\subsection{Tier 3: higher-order neural operators}
Tier 3 targets families of PDE solutions rather than a single fitted field. Neural operators such as DeepONet and MIONet learn maps between function spaces by pairing branch networks, which encode input functions or boundary data, with trunk networks, which encode query coordinates \cite{lu2021deeponet,jin2022mionet}. Spline-PINN shows a related but distinct path: a CNN predicts Hermite spline coefficients, and a continuous Hermite spline layer evaluates PDE residuals without finite-difference losses \cite{wandel2022splinepinn}. OSNR adopts the continuous Hermite idea but removes the black-box coordinate trunk where the governing operator and boundary calculus are known.
Let
\[
\Phi(t) =
\begin{bmatrix}
\phi_0(t) & \phi_1(t) & \phi_2(t)
\end{bmatrix}^{\top}
\]
be a second-order Hermite generator compactly supported on $[-1,1]$. The generalized higher-order Hermite construction of Schmitter, Badoual, Uhlmann, Fageot, and Unser stores value, slope, and curvature data at each knot and interpolates all three channels exactly \cite{schmitter2016hermite}:
\[
\phi_0^{(r)}(k)=\delta_{r0}\delta_k,\qquad
\phi_1^{(r)}(k)=\delta_{r1}\delta_k,\qquad
\phi_2^{(r)}(k)=\delta_{r2}\delta_k,
\qquad r=0,1,2.
\]
The synthesized field is
\[
f(t)
=
\sum_{k\in\Z}
c_0[k]\phi_0(t-k)
+
c_1[k]\phi_1(t-k)
+
c_2[k]\phi_2(t-k),
\]
where
\[
\mathbf{c}[k]
=
\begin{bmatrix}
c_0[k] \\ c_1[k] \\ c_2[k]
\end{bmatrix}
=
\begin{bmatrix}
f(k) \\ f'(k) \\ f''(k)
\end{bmatrix}
\]
for exactly interpolated data. The Hermite generator in \cite{schmitter2016hermite} is piecewise polynomial-exponential, $C^2$, compactly supported, and reproduces cubic polynomials as well as trigonometric functions. This is the missing boundary mechanism for a neural-operator OSNR tier: Dirichlet, Neumann, and curvature constraints are coefficient assignments, not soft loss penalties.
For a PDE solution operator
\[
\mathcal{S}: u \mapsto v,
\]
the branch-side network or analytic encoder should output Hermite coefficient tensors
\[
\{\mathbf{c}[k;u]\}_{k\in\Z},
\]
while the trunk side is replaced by deterministic Hermite synthesis. In multidimensional domains, tensor products of one-dimensional Hermite generators give mixed value/derivative channels, matching the construction used by Spline-PINN for continuous PDE residuals \cite{wandel2022splinepinn}. Unlike Spline-PINN, OSNR evaluates continuous energies and cross-correlations in coefficient space through the block-Gram calculus derived in Appendix~\ref{app:hermite-block-gram}, avoiding Monte Carlo residual quadrature whenever the operator and boundary model admit exact inner products.
\section{Continuous-Discrete Calibration}
The most important implementation condition is the cardinal coordinate map, which preserves the shift-invariant structure required by spline filtering calculus \cite{unser2005cardinal1,unser2005cardinal2}.
\[
v_k(x) = \frac{x}{T} - k.
\]
Here $T$ is the physical knot spacing. For a domain $[0,D]$ and a compact generator of support order $q$, a stable finite dictionary uses
\[
T = \frac{D}{M-q}.
\]
Then
\[
v_k(x+T) = v_k(x)+1,
\]
so one physical knot step maps to one cardinal interval.
\begin{proposition}[Failure of uncalibrated grid refinement]
If the map is implemented as $v_k(x)=x+b_k$ with biases distributed over a fixed interval while $M$ changes, increasing $M$ decreases the relative shift between adjacent columns without changing physical support. The resulting dictionary columns become nearly collinear, and the Gram matrix $\A^\top \A$ develops near-zero eigenvalues.
\end{proposition}
\begin{proof}[Sketch]
Let adjacent atoms be sampled as $\phi(x+b_k)$ and $\phi(x+b_{k+1})$. If $b_{k+1}-b_k=O(1/M)$ while the support width of $\phi$ is fixed, then a first-order expansion gives
\[
\phi(x+b_{k+1}) = \phi(x+b_k) + O(1/M).
\]
Thus adjacent columns converge to each other as $M$ increases. The Gram matrix approaches rank deficiency and the pseudoinverse amplifies roundoff along small singular directions.
\end{proof}
This explains why a derivative residual can be zero while reconstruction fails. If the derivative dictionary is defined algebraically by $\A_{d2}=-k^2 \A$, then the PDE residual
\[
\A_{d2}\cvec + k^2 \A \cvec
\]
is identically zero regardless of whether $\A$ is a well-conditioned reconstruction basis.
\subsection{2D tensor-product Hermite expansion}
The 2D Tier 3 engine is obtained by tensorizing the one-dimensional second-order Hermite streams. Let
\[
h_i^x(x),\qquad h_j^y(y),\qquad i,j\in\{0,1,2\},
\]
denote the value, slope, and curvature Hermite generators along the two axes. The tensor-product basis functions are
\[
h_{i,j}(x,y)
=
h_i^x(x)h_j^y(y),
\qquad i,j\in\{0,1,2\}.
\]
For a grid node $(k,\ell)$, OSNR stores a nine-stream local state
\[
\mathbf{c}_{k,\ell}
=
\begin{bmatrix}
f & \partial_x f & \partial_{xx}f &
\partial_y f & \partial_{xy}f & \partial_{xxy}f &
\partial_{yy}f & \partial_{xyy}f & \partial_{xxyy}f
\end{bmatrix}_{(k,\ell)}^{\top}.
\]
The synthesized field is
\[
f(x,y)
=
\sum_{k,\ell}
\sum_{i=0}^{2}\sum_{j=0}^{2}
c_{i,j}[k,\ell]\,
h_i^x\!\left(\frac{x}{T_x}-k\right)
h_j^y\!\left(\frac{y}{T_y}-\ell\right).
\]
This formula is the tensor-product analogue of the Schmitter et al. Hermite generator and the multidimensional counterpart of the periodic inner-product calculus of Badoual, Schmitter, and Unser.
The practical consequence is that the spatial calculus of 2D physical fields is a forward coefficient operation. In the stream-function formulation for incompressible flow,
\[
v_x=\partial_y a_z,\qquad v_y=-\partial_x a_z,
\]
and therefore
\[
\nabla\cdot \mathbf{v}
=
\partial_x\partial_y a_z-\partial_y\partial_x a_z
=
0
\]
up to the commutation error of the discrete finite-difference stencil. The nonlinear transport term is evaluated as
\[
(\mathbf{v}\cdot\nabla)\mathbf{v}
=
\begin{bmatrix}
v_x\partial_x v_x+v_y\partial_y v_x\\
v_x\partial_x v_y+v_y\partial_y v_y
\end{bmatrix},
\]
and viscous diffusion as
\[
\Delta\mathbf{v}
=
\begin{bmatrix}
\partial_{xx}v_x+\partial_{yy}v_x\\
\partial_{xx}v_y+\partial_{yy}v_y
\end{bmatrix}.
\]
All operators appearing in these expressions are evaluated by shift-invariant finite-difference ladders on the Hermite coefficient streams during the forward pass. No backward-mode automatic differentiation tape is constructed; the measured PyTorch autograd graph allocation in all Tier 3 validations is therefore $0.00$ bytes.
\section{Autograd-Free Differential Calculus}
\label{sec:autograd-free}
For an exponential spline with pole vector $\bs{\alpha}$, applying a first-order operator $(D-\alpha_m)$ reduces the order of the spline \cite{unser2005cardinal1,delgadogonzalo2012exponential}:
\[
(D-\alpha_m)\beta_{\bs{\alpha}}(t)
=
\beta_{\bs{\alpha}\setminus \alpha_m}(t)
-
e^{\alpha_m}
\beta_{\bs{\alpha}\setminus \alpha_m}(t-1).
\]
This is the finite-difference ladder. Differential fields can be evaluated by filtering coefficients or by applying deterministic lower-order dictionary maps. For the Helmholtz null space,
\[
\frac{\dd^2}{\dd x^2} s(x) = -k^2 s(x)
\]
inside the smooth spans, with boundary innovations handled separately.
This removes the need for backward-mode automatic differentiation in Tier 1 PDE residual evaluation. The differential operator is encoded in the basis and coefficient algebra.
\section{Circulant and FFT Solvers}
When the dictionary is shift-invariant and periodized, the Gram matrix is circulant. This is the coefficient-domain form of the exact periodic inner-product calculus for spline curves and functions \cite{badoual2016inner,badoual2018periodic}:
\[
\G =
\begin{bmatrix}
g_0 & g_{M-1} & \cdots & g_1 \\
g_1 & g_0 & \cdots & g_2 \\
\vdots & \vdots & \ddots & \vdots \\
g_{M-1} & g_{M-2} & \cdots & g_0
\end{bmatrix}.
\]
Such matrices are diagonalized by the discrete Fourier transform:
\[
\G = \mathbf{F}^{-1} \Lambda \mathbf{F}.
\]
Solving $\G \cvec = \mathbf{b}$ reduces to
\[
\widehat{\cvec}[\ell] = \frac{\widehat{\mathbf{b}}[\ell]}{\lambda_\ell}.
\]
This converts $O(M^3)$ dense solves into $O(M\log M)$ FFT operations.
\section{Tomographic Radiance Fields as Spline Inverse Problems}
The failed direct ray-kernel DL3DV experiment clarifies an important modeling boundary. A nontrivial view-synthesis scene is not naturally a smooth map from ray origin and direction to RGB. The physically shared quantity is a latent field in space, observed through line or ray measurements. The spline tomography literature gives the appropriate replacement model: represent the unknown continuous field by shifted basis functions, push those basis functions through the forward projector, and solve the coefficient inverse problem with fast adjoint and normal operators \cite{nilchian2013fast,mccann2016fast,donati2018multiscale,haouchat2025generalized}. Jin et al. make the complementary point that inverse problems whose normal operators are convolutional admit physics-aware direct inversions before any learned artifact-removal stage \cite{jin2017deep}; in OSNR, the direct inverse is the primary object, and any neural residual must remain secondary.
For a first linearized density or opacity stage, write
\[
\sigma(\mathbf{x})=\sum_{\mathbf{k}} a[\mathbf{k}]\,\varphi_\sigma(\mathbf{x}-\Lambda\mathbf{k}),
\qquad
\mathbf{g}=H\mathbf{a}+\boldsymbol{\eta},
\]
where $H$ samples line or ray integrals of the spline density field. The regularized inverse update is
\[
\mathbf{a}^{\star}
=
\arg\min_{\mathbf{a}}
\frac12\|W^{1/2}(H\mathbf{a}-\mathbf{g})\|_2^2
+\lambda R(\mathbf{a}),
\]
with $W$ encoding ray confidence or frequency reliability, as in weighted phase-retrieval formulations \cite{bostan2016variational}. For quadratic $R(\mathbf{a})=\|L\mathbf{a}\|_2^2$, the normal equation is
\[
(H^\top W H+\lambda L^\top L)\mathbf{a}
=
H^\top W\mathbf{g}.
\]
The computational opportunity is the McCann--Donati normal-operator identity. For shift-invariant basis functions and locally stationary projection blocks, $H^\top H$ acts as a discrete convolution:
\[
(H^\top H\mathbf{a})[\mathbf{k}]
=
(\mathbf{a}*\mathbf{r})[\mathbf{k}],
\]
where $\mathbf{r}$ is the sampled autocorrelation of the projected basis function. Thus the expensive repeated normal-operator application inside conjugate gradients or ADMM becomes a Fourier-domain multiplication. Multiscale basis functions then provide a controlled coarse-to-fine path that is robust to pose and angular uncertainty \cite{donati2018multiscale}.
Fast forward projection and sparse acquisition variants can be imported from the same lineage: Arcadu et al. use Fourier regridding with minimal oversampling for efficient forward projectors, while Donati et al. show how randomized STEM sampling can be coupled to regularized tomographic recovery \cite{arcadu2016forward,donati2017compressed}. This does not make full NeRF rendering linear. The volume-rendering equation contains transmittance
\[
C(r)=\int T(t)\sigma(r(t))c(r(t),\mathbf{d})\,dt,
\qquad
T(t)=\exp\!\left(-\int_0^t\sigma(r(s))\,ds\right),
\]
so exact RGB fitting remains nonlinear in $\sigma$. The OSNR route is therefore staged: first recover coarse density/support with a tomographic spline inverse solve; then refine the density multiscale; then solve color/radiance coefficients on the recovered support; and finally apply visibility-weighted nonlinear corrections. Positivity, support constraints, total-variation, Hessian, and sparse-innovation priors enter naturally through the constrained ADMM machinery developed for spline tomography \cite{nilchian2013constrained,nilchian2015spline}.
The controlled runner \texttt{apps\_industrial\_breakthrough/spline\_tomographic\_radiance\_solver.py} validates only the linearized operator claim. It builds a periodic synthetic density field, samples $48$ discrete projection directions, constructs $H^\top H$ explicitly once from a delta impulse, and then replaces all subsequent normal-operator applications by FFT convolution. On a $96\times96$ field, the convolutional normal operator matches explicit $H^\top H$ with relative error $2.4462\times10^{-7}$. One explicit normal-operator application costs $2.4956$ ms, while the FFT version costs $0.0969$ ms. Solving the same ridge-regularized inverse problem by conjugate gradients takes $188.1746$ ms with explicit normals and $4.9666$ ms with FFT normals, yielding a reconstruction PSNR of $22.3578$ dB under noisy sparse projections. This is not yet a NeRF result; it is the isolated mathematical validation that the spline-tomographic normal operator can be diagonalized as the literature predicts.
The next controlled runner, \texttt{apps\_industrial\_breakthrough/spline\_ray\_operator\_validation.py}, validates the more fundamental Haouchat-style requirement: the forward ray operator and its adjoint must be matched before any real-scene radiance experiment is meaningful. On a $36\times36$ coefficient grid with $2688$ parallel rays, the script constructs two explicit small operators for auditability: a pixel basis and a quadratic tensor-product spline basis. The spline data are generated by the spline operator itself, and both models solve the same noisy inverse problem with conjugate gradients. The adjoint identity $\langle Hc,p\rangle=\langle c,H^\top p\rangle$ holds to $1.5672\times10^{-15}$ relative error for the spline operator and $2.7427\times10^{-16}$ for the pixel operator. At the same coefficient count, the spline inverse reconstructs the continuous rendered target at $54.0869$ dB, while the pixel model reaches only $27.0125$ dB. This is an intentionally controlled operator test: it proves that the coefficient-domain ray basis and adjoint are now correctly formulated, not that the full DL3DV visibility problem is solved.
The basis choice itself is not incidental. Following the exponential-spline construction of Delgado-Gonzalo, Thevenaz, and Unser \cite{delgadogonzalo2012exponential}, a one-dimensional cardinal exponential B-spline associated with poles $\boldsymbol{\alpha}=(\alpha_1,\ldots,\alpha_N)$ has Fourier-domain form
\[
\widehat{\beta}_{\boldsymbol{\alpha}}(\omega)
=
\prod_{m=1}^{N}
\frac{1-\exp(\alpha_m-j\omega)}{j\omega-\alpha_m}.
\]
The two-dimensional smooth tier then uses the tensor-product generator
\[
\varphi_{\boldsymbol{\alpha}_x,\boldsymbol{\alpha}_y}(x,y)
=
\beta_{\boldsymbol{\alpha}_x}(x)\,
\beta_{\boldsymbol{\alpha}_y}(y),
\]
so the basis can reproduce the local modes implied by the operator rather than merely interpolate samples. The controlled sweep \texttt{apps\_industrial\_breakthrough/exponential\_spline\_basis\_sweep.py} tests this precision-first hypothesis on the same $1920$ rays and $784$ coefficients while changing only the tensor-product basis. The data are generated by a damped-harmonic exponential spline with poles $(-\lambda,-\lambda+j\omega,-\lambda-j\omega)$, and all candidate operators reuse the identical ray geometry. The matched damped-harmonic basis reaches $56.3837$ dB, compared with $52.1420$ dB for a monotone exponential-decay basis, $50.0620$ dB for an undamped harmonic basis, $49.9084$ dB for a quadratic polynomial spline, and $28.7579$ dB for a pixel box basis. This isolates the important design rule for the next radiance-field stage: the highest precision should come from tensor-product operator splines whose poles are matched to the expected local dynamics, with sparsity and compression applied after the operator basis is correct.
The follow-up runner \texttt{apps\_industrial\_breakthrough/exponential\_spline\_pole\_sweep.py} turns this from a hand-picked basis comparison into a deterministic pole-selection problem. It fixes the target field, ray geometry, coefficient count, noise level, ridge parameter, and conjugate-gradient budget, then sweeps damped-harmonic exponential splines of orders $2$, $3$, and $4$ over $\lambda\in\{0.18,0.28,0.35,0.42,0.56,0.72\}$ and periods $\{8,10,12,16\}$. The target is generated by the order-$3$ pole set $(-0.42,-0.42+j2\pi/10,-0.42-j2\pi/10)$. The pole sweep correctly ranks that matched operator first at $55.2077$ dB on $1280$ rays and $576$ coefficients. The best order-$4$ candidate reaches $52.4859$ dB, and the best order-$2$ candidate reaches only $42.1900$ dB. This is the first automated evidence that pole selection is a meaningful OSNR model-selection axis: increasing support/order blindly does not dominate, while matching the operator poles controls reconstruction precision.
The support/regularity runner \texttt{apps\_industrial\_breakthrough/exponential\_spline\_support\_regularization\_sweep.py} then isolates the opposite regime: a field generated by the shortest first-order Green-matched spline with pole $\alpha=-0.42$. The one-pole basis has support length $1$ and no continuity guarantee, but it exactly matches the local Green mode. It reaches $65.7918$ dB with operator density $0.0362$ and an average of $20.86$ active coefficients per ray. The smooth order-$3$ mixed basis $[0,0,\alpha]$ reaches only $27.1299$ dB and has density $0.1084$ with $62.44$ active coefficients per ray. This confirms a second design rule: if the modeled object is a Green response or sparse innovation, the shortest matched spline can be both more accurate and more localized than a smoother high-order basis. Regularity should be introduced because the signal class requires it, not by default.
Finally, \texttt{apps\_industrial\_breakthrough/exponential\_spline\_basis\_selection\_map.py} evaluates the full pole-multiset selection problem. It tests seven target regimes against fourteen candidate bases: first-order Green response, repeated real poles, two distinct real poles, zero-augmented smooth operators, damped oscillators, pure polynomial splines, and support-$4$ mixed bases. Across all regimes, the exact pole multiset ranks first. This is the strongest evidence so far that OSNR basis design should be formulated as operator pole selection rather than degree selection. Order controls support and regularity; pole multiplicity and location control the reproduced null-space modes. Higher order improves asymptotic approximation power for smooth functions, but it is not a substitute for matching the operator that generated the signal.
The final synthetic step, \texttt{apps\_industrial\_breakthrough/exponential\_spline\_operator\_inference.py}, removes access to rendered target PSNR during model selection. Each candidate basis is fitted on $75\%$ of the rays and scored on held-out rays using a normalized validation residual plus small support and density penalties. This exposes a practical distinction between generative pole matching and predictive operator selection. Under sparse-ray training, the oracle PSNR basis is sometimes a smoother support-$4$ model rather than the exact generating basis, because the added regularity improves interpolation across unobserved rays. The validation score still selects the oracle basis in four of seven regimes and stays within $0.2039$ dB of oracle in all cases, with mean PSNR loss $0.0465$ dB. Thus the operational rule becomes: use the differential operator poles as the first prior, then choose among nearby pole augmentations by held-out measurement prediction rather than training residual.
We then stress-test the same idea in a mixed local-operator field with \texttt{apps\_industrial\_breakthrough/exponential\_spline\_local\_operator\_adaptation.py}. The synthetic field assigns different pole multisets to different spatial regions and compares a single global basis, a held-out-ray local block selector, and an oracle local label map. The oracle local solve reaches $33.5291$ dB, while the best global held-out model reaches $28.6559$ dB. This proves that local operator adaptation has substantial headroom. However, a blind ray-only greedy block selector reaches only $28.8074$ dB with $25\%$ block-label accuracy. Global line measurements make small independent block labels weakly identifiable unless the selection objective includes stronger spatial priors, localized measurements, or joint segmentation/coefficient optimization.
The sparse-innovation remedy is tested in \texttt{apps\_industrial\_breakthrough/exponential\_spline\_fri\_region\_adaptation.py}. Instead of selecting arbitrary blocks, the method first reconstructs a global proxy, computes derivative-energy profiles, extracts sparse FRI-style transition proposals, expands them into a small boundary lattice, and then jointly scores region geometry and region pole choices on held-out rays. The raw derivative peaks locate approximate boundaries at $x=(-1.4149,1.4149)$ and $y=3.0319$; held-out refinement moves them to $x=(-3.4149,3.4149)$ and $y=4.0319$, yielding $100\%$ region-label agreement at the coefficient-grid resolution. The resulting FRI-region adaptive model reaches $34.2000$ dB, outperforming both the global model and the nominal oracle-region pole assignment. This is the first successful local adaptation mechanism: sparse innovation proposals supply the missing spatial prior that held-out global rays alone could not provide.
The robustness sweep \texttt{apps\_industrial\_breakthrough/exponential\_spline\_fri\_region\_stress\_sweep.py} repeats the experiment over three projection-angle budgets and three additive noise levels. To keep the sweep diagnostic rather than combinatorial, it uses a compact region-basis candidate set containing the physical oracle family, the single-case selected family, a smoother support-$4$ family, and the best global family; the exhaustive $5^4$ region-basis search remains available as an optional mode. Across all nine stress cases, the FRI-region model improves over the best global held-out basis. The gain increases with measurement density, from a mean $+1.1667$ dB at $12$ angles to $+5.0277$ dB at $24$ angles, while the recovered region accuracy rises from $89.50\%$ to $100.00\%$. This confirms that the sparse-innovation step is not a one-off artifact: as the inverse problem receives enough projections to identify the transition set, local operator adaptation becomes reliably beneficial.
We then deliberately break the axis-aligned assumption with \texttt{apps\_industrial\_breakthrough/exponential\_spline\_fri\_curve\_adaptation.py}. The target field is generated by three smooth transition curves: two vertical pole-boundary curves and one top-interface curve. A row/column FRI-style derivative tracker fits low-order curve proposals from the global proxy and refines them by held-out rays. This harder test exposes the current bottleneck. The global held-out basis reaches $32.2114$ dB, while the true curved local operator assignment reaches $41.5032$ dB, proving that curved local operators have large headroom. However, the detected curved selector reaches only $31.1969$ dB despite $91.00\%$ region-label agreement and boundary RMSE $1.4606$. The failure is not the absence of local operator advantage; it is the scoring layer. Under curved imperfect labels, the held-out ray residual prefers smoother surrogate pole assignments rather than the physical local pole map. The next algorithmic step is therefore joint geometry--basis--coefficient refinement, or a region-contrastive validation score that prevents the local operator assignment from collapsing to a globally smooth surrogate.
The joint refinement runner \texttt{apps\_industrial\_breakthrough/exponential\_spline\_fri\_curve\_joint\_refinement.py} implements this next correction. It expands the curve-offset lattice near the best FRI proposal, fits coefficients for each candidate local operator assignment, and augments the held-out residual with an edge-consistency contrast term that rewards reconstructions whose gradient energy concentrates on the proposed sparse transition curves. This converts the curved selector from a negative result into a partial recovery: the selected joint model reaches $34.7216$ dB, a $+2.5102$ dB gain over the global held-out basis. The best candidate present in the searched family reaches $35.8310$ dB, while the true-curve oracle remains at $41.5032$ dB. Thus the scoring fix is directionally correct but incomplete. The remaining gap now separates two effects: curve localization error and the limited pole-assignment candidate family. This is the cleanest current target for further algorithmic work.
Finally, \texttt{apps\_industrial\_breakthrough/exponential\_spline\_geometry\_reproduction.py} isolates the more geometric point raised by the exponential-spline curve and surface literature \cite{delgadogonzalo2012exponential}. The target is a closed harmonic curve with modes up to order three, represented from only twelve parameter samples and eight control points. The matched compact harmonic E-spline uses the pole set $\{0,\pm j2\pi/M,\pm j4\pi/M,\pm j6\pi/M\}$ and reaches dense curve RMSE $5.4457\times10^{-6}$. With the same number of control points and the same samples, a generic cubic polynomial spline reaches only $1.0256\times10^{-2}$ RMSE, and a piecewise-linear polygon reaches $3.8693\times10^{-2}$ RMSE. This is a small but important result: if the expected geometry is known to be harmonic, elliptic, spherical, cylindrical, or otherwise parametrizable by a known exponential-polynomial family, then OSNR should place that family directly in the geometric span rather than recover it indirectly through a generic volumetric grid. For NeRF-like scenes this suggests a patch-based route: segment or initialize object surfaces with geometry-reproducing parametric E-splines, then fit texture/radiance on those surfaces. For PDE domains it suggests an even cleaner route: represent both boundary geometry and the solution field in operator-matched spline spaces.
The follow-up optimizer \texttt{apps\_industrial\_breakthrough/eggroll\_spline\_shape\_optimizer.py} tests whether this geometric advantage can be used when the target shape is not known in closed form. Motivated by Schmitter and Unser's continuous-domain shape projectors and functional PCA construction \cite{schmitter2018landmark}, the experiment represents a closed spline curve by $16$ control points but restricts the learned geometric search to an $8$-dimensional continuous shape subspace. This is the analogue of replacing an arbitrary coordinate-field parameter vector by a learned spline-shape chart. The stochastic search component is motivated by the EGGROLL low-rank evolution-strategy result \cite{sarkar2026eggroll}: rank-one perturbations can be evaluated as hardware-friendly low-rank updates, but the experiment separates this hardware trick from the geometric prior itself.
The result is deliberately diagnostic. Full Gaussian ES over all $32$ control coordinates reaches dense RMSE $2.9097\times10^{-2}$ after $67{,}200$ forward evaluations, while rank-one EGGROLL-style perturbations applied directly to the raw control matrix reach $3.3517\times10^{-2}$. Thus low-rank noise alone does not solve geometry discovery. When the same evaluation budget is spent inside the Schmitter-style spline subspace, the error drops to $6.9268\times10^{-3}$; rank-one EGGROLL perturbations inside that subspace reach a comparable $7.4862\times10^{-3}$. The algebraic continuous-subspace projection oracle reaches $3.5689\times10^{-4}$ in $0.293$ ms, exposing the remaining optimization gap. The practical implication is precise: the promising route is not blind evolution over arbitrary OSNR coefficients, but variable projection. Use stochastic low-rank search only for nonlinear geometry, visibility, and knot variables; solve the linear radiance or texture coefficients algebraically once a candidate geometry is proposed.
The next controlled runner, \texttt{apps\_industrial\_breakthrough/eggroll\_adjoint\_variable\_projection.py}, implements that variable-projection step explicitly. A one-dimensional spline boundary partitions a $40\times40$ radiance field into two continuous regions. For each candidate boundary, the code constructs the ray operator $H(\varphi)$ by projecting masked smooth atoms, eliminates the linear radiance coefficients by the adjoint normal equation
\[
c^\star(\varphi)=\left(H(\varphi)^\top H(\varphi)+\lambda I\right)^{-1}H(\varphi)^\top y,
\]
and scores the resulting field on acquisition directions not used in the coefficient solve. This is the minimal inverse-problem analogue of a NeRF geometry/radiance separation: nonlinear geometry is searched, while linear radiance is solved in closed form.
The experiment clarifies both the opportunity and the bottleneck. The mean-geometry variable-projection baseline reaches $16.8968$ dB on held-out projections. Full Gaussian ES over raw boundary controls improves to $18.9082$ dB, rank-one EGGROLL over raw controls reaches $19.9034$ dB, and rank-one EGGROLL in the six-dimensional spline subspace reaches $20.1043$ dB. However, the true-geometry variable-projection oracle reaches $33.1874$ dB with a field RMSE of $3.3521\times10^{-2}$. Thus the adjoint variable-projection mechanism is working, but stochastic boundary discovery remains underidentified from the current projection residual alone. This is an important negative constraint for the NeRF/SIREN campaign: the next improvement must add stronger geometry evidence, such as FRI edge measurements, silhouette consistency, epipolar visibility constraints, or a learned continuous shape prior. More low-rank perturbation budget alone is unlikely to close the oracle gap.
The Haouchat-matched follow-up \texttt{apps\_industrial\_breakthrough/haouchat\_matched\_variable\_projection.py} performs that correction. Instead of using a primitive projection mask, it builds the inner ray operator from the same quadrature-evaluated tensor-product spline basis used in the $52$--$56$ dB matched-ray experiments. The target data are generated by a damped-harmonic exponential spline ray operator. Candidate geometries still define a boundary-dependent masked coefficient dictionary, but the forward map is now
\[
H_\varphi = H_{\beta_L}\,\Phi(\varphi),
\]
where $H_{\beta_L}$ is the Haouchat-style ray projector for the selected tensor-product basis and $\Phi(\varphi)$ applies the boundary-dependent coefficient atoms. The stochastic outer loop is also given a weak FRI-like edge observation of the boundary, so the score combines held-in projection residual and edge consistency. This is the first experiment in this line that combines all three ingredients: matched spline rays, adjoint variable projection, and sparse boundary evidence.
The result changes the interpretation sharply. With the matched damped-harmonic operator, the true-geometry projection oracle reaches $96.9888$ dB on held-out rays, confirming that the inner ray/inverse model itself is not the limiting factor. The edge-aware rank-one spline-subspace EGGROLL search reaches $42.2339$ dB, up from the mean-geometry baseline of $33.3709$ dB, with boundary RMSE reduced from $9.5169\times10^{-2}$ to $2.7066\times10^{-2}$. The exponential-decay candidate also benefits from edge evidence, improving from $40.6125$ dB without the edge term to $41.0644$ dB with it. This supports the current thesis: OSNR does not need a dense NeRF-style MLP to represent the radiance once the operator is matched; the hard remaining problem is physically constrained geometry and visibility discovery.
The follow-up \texttt{apps\_industrial\_breakthrough/haouchat\_fri\_edge\_variable\_projection.py} removes the remaining artificial part of the edge-aware score. Instead of injecting a boundary hint directly from the true geometry, it forms a measurement-derived FRI proxy: training rays are backprojected through the matched adjoint, row-wise derivatives of the normalized adjoint image are localized, and the resulting peak track is smoothed into a candidate boundary. This is not a complete multidimensional FRI surface solver, but it is a measurement-only sparse-transition proposal. In the default run, the synthetic edge hint has boundary RMSE $1.3676\times10^{-2}$, while the adjoint-derived edge track has RMSE $1.9117\times10^{-2}$.
Using this measurement-derived edge evidence, the matched damped-harmonic model reaches $41.5632$ dB on held-out rays, compared with $33.3709$ dB for mean geometry and $96.9888$ dB for the true-geometry oracle. The cleaner synthetic edge hint reaches $45.9174$ dB in the same runner. The gap between $41.56$ dB and $45.92$ dB is useful: it quantifies the price of deriving geometry evidence from measurements rather than providing it externally. The experiment therefore validates the direction without hiding the remaining work. The next real-scene version should replace row-wise adjoint peaks by multi-view epipolar FRI proposals and visibility-aware surface clustering.
The next controlled refinement, \texttt{apps\_industrial\_breakthrough/haouchat\_multiray\_fri\_edge\_projection.py}, replaces the row-wise adjoint peak estimate by a multi-ray residual selection loop. The adjoint boundary is used only as initialization. Each boundary control point is perturbed over a small local offset lattice, and candidates are scored by the held-in variable-projection residual plus curvature and anchor penalties:
\[
\mathcal{S}(\varphi)=
\|H_{\beta_L}\Phi(\varphi)c^\star(\varphi)-y_{\mathrm{score}}\|_2^2
+\eta\|\Delta^2\varphi\|_2^2
+\rho\|\varphi-\varphi_{\mathrm{adj}}\|_2^2.
\]
This converts the crude measurement edge into a ray-consistent FRI boundary proposal without accessing the hidden target geometry.
The improvement is large. The row/adjoint edge has boundary RMSE $1.9117\times10^{-2}$; the multi-ray refined edge has RMSE $1.4848\times10^{-3}$, better than the noisy synthetic edge control ($1.3676\times10^{-2}$). With the matched damped-harmonic operator, direct variable projection on the multi-ray FRI boundary reaches $66.1488$ dB on held-out rays and field RMSE $2.5205\times10^{-3}$. The true-geometry oracle remains $96.9888$ dB, so the experiment is still controlled rather than a final SOTA benchmark, but it shows that the geometry bottleneck can be attacked algebraically by ray-consistent sparse innovation refinement. Interestingly, applying the stochastic EGGROLL search after this refined geometry is worse ($40.3169$ dB), so the current best path is not more random search but better deterministic boundary proposal.
The robustness sweep \texttt{apps\_industrial\_breakthrough/haouchat\_multiray\_fri\_robustness\_sweep.py} repeats the deterministic part of the experiment across two boundary variants, projection-angle counts $\{12,18,24\}$, and additive noise levels $\{0,3\times10^{-4},10^{-3}\}$. The mean-geometry baseline averages $32.5922$ dB over the $18$ cases; direct row/adjoint edge projection averages $38.0604$ dB; multi-ray FRI variable projection averages $64.0770$ dB; and the true-geometry oracle averages $97.9883$ dB. The multi-ray method remains above $61.76$ dB in every tested case. This confirms that the $66.15$ dB result is not a one-off numerical accident, but a stable consequence of selecting boundary offsets by held-in ray residuals.
\begin{table}[h]
\centering
\small
\begin{tabular}{lcc}
\toprule
Quantity & Explicit normal & FFT convolution normal \\
\midrule
Single $H^\top H$ application & $2.4956$ ms & $0.0969$ ms \\
CG inverse solve & $188.1746$ ms & $4.9666$ ms \\
Normal-operator relative error & \multicolumn{2}{c}{$2.4462\times10^{-7}$} \\
Reconstruction PSNR & \multicolumn{2}{c}{$22.3578$ dB} \\
\bottomrule
\end{tabular}
\caption{Controlled synthetic validation of the spline-tomographic normal-operator identity. The experiment isolates the linear inverse-problem component needed before returning to nonlinear radiance-field rendering.}
\label{tab:spline-tomographic-radiance}
\end{table}
\begin{table}[h]
\centering
\small
\begin{tabular}{lcc}
\toprule
Quantity & Pixel basis & Quadratic spline basis \\
\midrule
Matched-adjoint relative error & $2.7427\times10^{-16}$ & $1.5672\times10^{-15}$ \\
CG inverse solve & $19.5137$ ms & $21.9203$ ms \\
Coefficient RMSE & $4.1936\times10^{-2}$ & $7.5915\times10^{-3}$ \\
Rendered reconstruction PSNR & $27.0125$ dB & $54.0869$ dB \\
\bottomrule
\end{tabular}
\caption{Controlled validation of the matched spline ray operator $H_\varphi$ and adjoint $H_\varphi^\top$ on $2688$ rays and $1296$ coefficients. The result establishes the correct operator foundation before porting the method to DL3DV camera rays with visibility weights.}
\label{tab:spline-ray-operator-validation}
\end{table}
\begin{table}[h]
\centering
\small
\begin{tabular}{lcccc}
\toprule
Basis profile & PSNR & Adjoint error & CG solve & Coefficient RMSE \\
\midrule
Pixel box & $28.7579$ dB & $1.8997\times10^{-16}$ & $4.8182$ ms & $1.8724\times10^{-1}$ \\
Quadratic spline & $49.9084$ dB & $3.0959\times10^{-15}$ & $4.1478$ ms & $1.8928\times10^{-1}$ \\
Exponential decay & $52.1420$ dB & $2.0145\times10^{-16}$ & $4.7577$ ms & $1.1161\times10^{-2}$ \\
Harmonic exponential & $50.0620$ dB & $2.2573\times10^{-16}$ & $4.1359$ ms & $1.5208\times10^{-2}$ \\
Damped-harmonic exponential & $56.3837$ dB & $0.0000$ & $3.9541$ ms & $6.8398\times10^{-3}$ \\
\bottomrule
\end{tabular}
\caption{Tensor-product basis sweep for a matched ray inverse problem on $1920$ rays and $784$ coefficients. The target field is generated by the damped-harmonic exponential spline; all candidate bases reuse the same rays, regularization, and conjugate-gradient inverse solve.}
\label{tab:exponential-spline-basis-sweep}
\end{table}
\begin{table}[h]
\centering
\small
\begin{tabular}{lcccc}
\toprule
Best candidate class & Pole profile & PSNR & CG solve & Coefficient RMSE \\
\midrule
Order $2$ & $\lambda=0.35$, period $16$ & $42.1900$ dB & $2.1103$ ms & $1.6017\times10^{-1}$ \\
Order $3$ & $\lambda=0.42$, period $10$ & $55.2077$ dB & $1.9365$ ms & $7.5642\times10^{-3}$ \\
Order $4$ & $\lambda=0.35$, period $16$ & $52.4859$ dB & $2.2178$ ms & $4.9702\times10^{-2}$ \\
\bottomrule
\end{tabular}
\caption{Deterministic pole-selection sweep for damped-harmonic tensor-product exponential splines. The target is generated by the order-$3$, $\lambda=0.42$, period-$10$ operator; the matched pole set ranks first across the tested grid.}
\label{tab:exponential-spline-pole-sweep}
\end{table}
\begin{table}[h]
\centering
\small
\begin{tabular}{lccccc}
\toprule
Pole multiset & Regularity & PSNR & CG solve & Density & Active/ray \\
\midrule
$[\alpha]$ & $C^{-1}$ & $65.7918$ dB & $2.1925$ ms & $0.0362$ & $20.86$ \\
$[0]$ & $C^{-1}$ & $29.1917$ dB & $2.1154$ ms & $0.0362$ & $20.86$ \\
$[0,\alpha]$ & $C^0$ & $26.8630$ dB & $2.1370$ ms & $0.0724$ & $41.71$ \\
$[\alpha,\alpha]$ & $C^0$ & $26.5164$ dB & $2.1558$ ms & $0.0724$ & $41.71$ \\
$[0,0,\alpha]$ & $C^1$ & $27.1299$ dB & $2.1973$ ms & $0.1084$ & $62.44$ \\
$[\alpha,\alpha,\alpha]$ & $C^1$ & $26.8975$ dB & $2.0574$ ms & $0.1084$ & $62.44$ \\
\bottomrule
\end{tabular}
\caption{Support-versus-regularity sweep for a first-order Green target with $\alpha=-0.42$ on $1280$ rays and $576$ coefficients. The shortest matched basis wins because the target is Green-like rather than smooth.}
\label{tab:exponential-spline-support-regularity}
\end{table}
\begin{table}[h]
\centering
\small
\begin{tabular}{lccccc}
\toprule
Target regime & Best candidate & Best PSNR & Matched PSNR & Support & Regularity \\
\midrule
Green $[\alpha]$ & $[\alpha]$ & $52.9697$ dB & $52.9697$ dB & $1$ & $C^{-1}$ \\
Repeated $[\alpha,\alpha]$ & $[\alpha,\alpha]$ & $45.6373$ dB & $45.6373$ dB & $2$ & $C^0$ \\
Two-real $[\alpha,\beta]$ & $[\alpha,\beta]$ & $44.6760$ dB & $44.6760$ dB & $2$ & $C^0$ \\
Smooth $[0,0,\alpha]$ & $[0,0,\alpha]$ & $49.4038$ dB & $49.4038$ dB & $3$ & $C^1$ \\
Damped oscillator & $[-\lambda,-\lambda\pm j\omega]$ & $48.4246$ dB & $48.4246$ dB & $3$ & $C^1$ \\
Polynomial $[0,0,0]$ & $[0,0,0]$ & $49.9511$ dB & $49.9511$ dB & $3$ & $C^1$ \\
Mixed $[0,0,\alpha,\alpha]$ & $[0,0,\alpha,\alpha]$ & $52.2576$ dB & $52.2576$ dB & $4$ & $C^2$ \\
\bottomrule
\end{tabular}
\caption{Pole-multiset basis-selection map on $768$ rays and $400$ coefficients. The exact pole multiset ranks first in every tested target regime, confirming that operator matching and pole multiplicity are distinct from simply increasing spline order.}
\label{tab:exponential-spline-basis-selection-map}
\end{table}
\begin{table}[h]
\centering
\small
\begin{tabular}{lcccc}
\toprule
Target regime & Oracle basis & Inferred basis & Oracle PSNR & Inferred PSNR \\
\midrule
Green $[\alpha]$ & $[\alpha]$ & $[\alpha]$ & $35.3689$ dB & $35.3689$ dB \\
Repeated $[\alpha,\alpha]$ & $[0,0,\alpha,\alpha]$ & $[0,0,0,0]$ & $33.1575$ dB & $33.1202$ dB \\
Two-real $[\alpha,\beta]$ & $[0,0,\alpha,\alpha]$ & $[0,0,0,0]$ & $32.2104$ dB & $32.0064$ dB \\
Smooth $[0,0,\alpha]$ & $[0,0,\alpha,\alpha]$ & $[0,0,\alpha,\alpha]$ & $34.9124$ dB & $34.9124$ dB \\
Damped oscillator & $[0,0,\alpha,\alpha]$ & $[0,0,\alpha,\alpha]$ & $34.6249$ dB & $34.6249$ dB \\
Polynomial $[0,0,0]$ & $[0,0,0,0]$ & $[0,0,\alpha,\alpha]$ & $34.9254$ dB & $34.8408$ dB \\
Mixed $[0,0,\alpha,\alpha]$ & $[0,0,\alpha,\alpha]$ & $[0,0,\alpha,\alpha]$ & $36.1621$ dB & $36.1621$ dB \\
\bottomrule
\end{tabular}
\caption{Measurement-driven operator inference with a held-out ray split. The selected basis is chosen without rendered target access; it remains within $0.2039$ dB of the oracle PSNR basis across all tested regimes.}
\label{tab:exponential-spline-operator-inference}
\end{table}
\begin{table}[h]
\centering
\small
\begin{tabular}{@{}p{0.22\linewidth}p{0.34\linewidth}p{0.27\linewidth}p{0.10\linewidth}@{}}
\toprule
Model & PSNR & Selection signal & Block-label accuracy \\
\midrule
Global held-out basis & $28.6559$ dB & Held-out rays & n/a \\
Local greedy basis & $28.8074$ dB & Held-out rays & $25.00\%$ \\
Oracle local labels & $33.5291$ dB & Ground-truth region labels & $100.00\%$ \\
\bottomrule
\end{tabular}
\caption{Local operator adaptation stress test on a mixed pole-multiset field. The oracle gap confirms that local bases can matter, while the weak greedy-label recovery identifies the next algorithmic bottleneck.}
\label{tab:exponential-spline-local-operator-adaptation}
\end{table}
\begin{table}[h]
\centering
\small
\begin{tabular}{lccc}
\toprule
Model & Boundary source & Region accuracy & PSNR \\
\midrule
Global held-out basis & none & n/a & $28.6559$ dB \\
Blind greedy blocks & fixed grid & $25.00\%$ & $28.8074$ dB \\
Oracle region labels & ground truth & $100.00\%$ & $33.5291$ dB \\
FRI-region adaptation & derivative peaks + held-out refinement & $100.00\%$ & $34.2000$ dB \\
\bottomrule
\end{tabular}
\caption{FRI-guided local operator adaptation. Sparse-innovation boundary proposals convert the weak blockwise selection problem into a region-level operator-selection problem and recover the local-basis advantage.}
\label{tab:exponential-spline-fri-region-adaptation}
\end{table}
\begin{table}[h]
\centering
\small
\begin{tabular}{lcccc}
\toprule
Projection angles & Mean global PSNR & Mean FRI PSNR & Mean gain & Region accuracy \\
\midrule
$12$ & $28.2410$ dB & $29.4078$ dB & $+1.1667$ dB & $89.50\%$ \\
$18$ & $28.7692$ dB & $31.9501$ dB & $+3.1809$ dB & $86.00\%$ \\
$24$ & $28.9903$ dB & $34.0180$ dB & $+5.0277$ dB & $100.00\%$ \\
\bottomrule
\end{tabular}
\caption{FRI-region stress sweep averaged over additive noise levels $\{0,5\times10^{-4},2\times10^{-3}\}$. The sparse-transition prior becomes more valuable as the ray geometry provides enough measurements to localize region boundaries.}
\label{tab:exponential-spline-fri-region-stress-sweep}
\end{table}
\begin{table}[h]
\centering
\small
\begin{tabular}{lccc}
\toprule
Model & Boundary model & Region accuracy & PSNR \\
\midrule
Global held-out basis & none & n/a & $32.2114$ dB \\
FRI curved selector & fitted curves & $91.00\%$ & $31.1969$ dB \\
Oracle curved labels & ground-truth curves & $100.00\%$ & $41.5032$ dB \\
\bottomrule
\end{tabular}
\caption{Curved-interface local operator adaptation. The oracle gap confirms large local-operator headroom, while the selected model identifies the next bottleneck: held-out ray residuals alone are not sufficient to choose physical pole assignments under imperfect curved segmentation.}
\label{tab:exponential-spline-fri-curve-adaptation}
\end{table}
\begin{table}[h]
\centering
\small
\begin{tabular}{lccc}
\toprule
Model & Selection signal & Region accuracy & PSNR \\
\midrule
Global held-out basis & held-out rays & n/a & $32.2114$ dB \\
FRI curved selector & held-out rays & $91.00\%$ & $31.1969$ dB \\
Joint curve refinement & residual + edge contrast & $92.25\%$ & $34.7216$ dB \\
Best searched candidate & hidden PSNR oracle & n/a & $35.8310$ dB \\
Oracle curved labels & ground-truth curves & $100.00\%$ & $41.5032$ dB \\
\bottomrule
\end{tabular}
\caption{Joint curved-interface refinement. Adding edge-consistency contrast to the measurement score recovers a useful local-operator gain, but the remaining oracle gap shows that curved sparse-innovation geometry and pole assignment still need joint refinement.}
\label{tab:exponential-spline-fri-curve-joint-refinement}
\end{table}
\begin{table}[h]
\centering
\small
\begin{tabular}{lccc}
\toprule
Geometry basis & Parameters & Dense RMSE & Max error \\
\midrule
Matched harmonic E-spline, $L=3$ & $16$ & $5.4457\times10^{-6}$ & $9.2024\times10^{-6}$ \\
Generic cubic polynomial spline & $16$ & $1.0256\times10^{-2}$ & $1.9079\times10^{-2}$ \\
Piecewise-linear polygon & $16$ & $3.8693\times10^{-2}$ & $1.6177\times10^{-1}$ \\
\bottomrule
\end{tabular}
\caption{Geometry-reproduction benchmark for a closed harmonic curve from twelve parameter samples. The matched exponential-spline pole set places the target geometry in the span; generic polynomial and polygonal bases require more parameters to reach the same precision.}
\label{tab:exponential-spline-geometry-reproduction}
\end{table}
\begin{table}[h]
\centering
\small
\begin{tabular}{lcccc}
\toprule
Optimizer & Search dimension & Dense RMSE & Runtime & Structural compression \\
\midrule
Full Gaussian ES on controls & $32$ & $2.9097\times10^{-2}$ & $46.708$ ms & $0.00\%$ \\
Rank-one EGGROLL on controls & $32$ & $3.3517\times10^{-2}$ & $39.477$ ms & $0.00\%$ \\
Schmitter spline-subspace ES & $8$ & $6.9268\times10^{-3}$ & $35.385$ ms & $75.00\%$ \\
Rank-one EGGROLL in spline subspace & $8$ & $7.4862\times10^{-3}$ & $35.693$ ms & $75.00\%$ \\
Continuous subspace projection oracle & $8$ & $3.5689\times10^{-4}$ & $0.293$ ms & $75.00\%$ \\
\bottomrule
\end{tabular}
\caption{Low-rank stochastic geometry search on a continuous spline shape family. The result separates the EGGROLL hardware mechanism from the Schmitter-style geometric prior: raw rank-one perturbations are not enough, while low-dimensional continuous spline shape coordinates produce the large error reduction.}
\label{tab:eggroll-spline-shape-optimizer}
\end{table}
\begin{figure}[h]
\centering
\includegraphics[width=\linewidth]{../apps_industrial_breakthrough/eggroll_spline_shape_optimizer_outputs/eggroll_spline_shape_optimizer.png}
\caption{Controlled spline-shape discovery benchmark. Red points are sparse noisy observations, gray curves mark the target where shown, and black curves show the recovered continuous spline shape for each optimizer.}
\label{fig:eggroll-spline-shape-optimizer}
\end{figure}
\begin{table}[h]
\centering
\small
\begin{tabular}{lccccc}
\toprule
Method & Search dim. & Held-out PSNR & Field RMSE & Boundary RMSE & Search time \\
\midrule
Mean geometry + variable projection & $0$ & $16.8968$ dB & $2.2344\times10^{-1}$ & $1.4287\times10^{-1}$ & $0.00$ ms \\
Full Gaussian ES controls & $18$ & $18.9082$ dB & $1.7413\times10^{-1}$ & $3.9663\times10^{-1}$ & $7132.59$ ms \\
Rank-one EGGROLL controls & $18$ & $19.9034$ dB & $1.5539\times10^{-1}$ & $3.4238\times10^{-1}$ & $7333.54$ ms \\
Rank-one EGGROLL spline subspace & $6$ & $20.1043$ dB & $1.7349\times10^{-1}$ & $3.2396\times10^{-1}$ & $7177.82$ ms \\
True geometry projection oracle & $6$ & $33.1874$ dB & $3.3521\times10^{-2}$ & $0.0000$ & $0.00$ ms \\
\bottomrule
\end{tabular}
\caption{Adjoint variable-projection geometry benchmark. Each candidate boundary defines $H(\varphi)$, radiance coefficients are eliminated by a ridge normal solve, and quality is measured on held-out projection directions. The oracle gap shows that coefficient elimination is not enough; physical boundary evidence must be strengthened.}
\label{tab:eggroll-adjoint-variable-projection}
\end{table}
\begin{figure}[h]
\centering
\includegraphics[width=\linewidth]{../apps_industrial_breakthrough/eggroll_adjoint_variable_projection_outputs/eggroll_adjoint_variable_projection.png}
\caption{Variable-projection radiance benchmark. Left panel is the target field; subsequent panels show recovered fields for mean geometry, full ES, rank-one control EGGROLL, rank-one spline-subspace EGGROLL, and true-geometry oracle. Red curves mark the recovered boundary.}
\label{fig:eggroll-adjoint-variable-projection}
\end{figure}
\begin{table}[h]
\centering
\small
\begin{tabular}{llccc}
\toprule
Basis & Method & Held-out PSNR & Field RMSE & Boundary RMSE \\
\midrule
Polynomial quadratic & Mean geometry & $33.3249$ dB & $3.8306\times10^{-1}$ & $9.5169\times10^{-2}$ \\
Polynomial quadratic & Edge-aware EGGROLL & $41.7015$ dB & $3.8036\times10^{-1}$ & $2.3912\times10^{-2}$ \\
Polynomial quadratic & True-geometry oracle & $48.9766$ dB & $3.7926\times10^{-1}$ & $0.0000$ \\
Exponential decay & Mean geometry & $33.3627$ dB & $6.6277\times10^{-2}$ & $9.5169\times10^{-2}$ \\
Exponential decay & Edge-aware EGGROLL & $41.0644$ dB & $3.3052\times10^{-2}$ & $2.4452\times10^{-2}$ \\
Exponential decay & True-geometry oracle & $52.4398$ dB & $5.1297\times10^{-3}$ & $0.0000$ \\
Damped harmonic & Mean geometry & $33.3709$ dB & $6.6046\times10^{-2}$ & $9.5169\times10^{-2}$ \\
Damped harmonic & EGGROLL, no edge term & $39.9315$ dB & $4.0312\times10^{-2}$ & $5.0651\times10^{-2}$ \\
Damped harmonic & Edge-aware EGGROLL & $42.2339$ dB & $2.9005\times10^{-2}$ & $2.7066\times10^{-2}$ \\
Damped harmonic & True-geometry oracle & $96.9888$ dB & $3.2065\times10^{-5}$ & $0.0000$ \\
\bottomrule
\end{tabular}
\caption{Haouchat-matched variable projection. The inner operator is a quadrature-evaluated tensor-product spline ray projector, while the outer loop searches boundary geometry. Matching the damped-harmonic exponential basis restores the very high oracle ceiling, and adding sparse edge evidence moves the stochastic search into the $42$ dB held-out regime.}
\label{tab:haouchat-matched-variable-projection}
\end{table}
\begin{figure}[h]
\centering
\includegraphics[width=\linewidth]{../apps_industrial_breakthrough/haouchat_matched_variable_projection_outputs/haouchat_matched_variable_projection.png}
\caption{Haouchat-matched variable projection for the damped-harmonic basis. Panels show the target, mean geometry, no-edge EGGROLL, edge-aware EGGROLL, and true-geometry oracle.}
\label{fig:haouchat-matched-variable-projection}
\end{figure}
\begin{table}[h]
\centering
\small
\begin{tabular}{llccc}
\toprule
Basis & Method & Held-out PSNR & Field RMSE & Boundary RMSE \\
\midrule
Polynomial quadratic & Measurement FRI edge & $41.2085$ dB & $3.8043\times10^{-1}$ & $2.5332\times10^{-2}$ \\
Polynomial quadratic & True-geometry oracle & $48.9766$ dB & $3.7926\times10^{-1}$ & $0.0000$ \\
Exponential decay & Measurement FRI edge & $40.5459$ dB & $3.4998\times10^{-2}$ & $4.8217\times10^{-2}$ \\
Exponential decay & True-geometry oracle & $52.4398$ dB & $5.1297\times10^{-3}$ & $0.0000$ \\
Damped harmonic & Mean geometry & $33.3709$ dB & $6.6046\times10^{-2}$ & $9.5169\times10^{-2}$ \\
Damped harmonic & Synthetic edge hint & $45.9174$ dB & $2.1096\times10^{-2}$ & $1.9775\times10^{-2}$ \\
Damped harmonic & Measurement FRI edge & $41.5632$ dB & $3.3397\times10^{-2}$ & $4.2715\times10^{-2}$ \\
Damped harmonic & True-geometry oracle & $96.9888$ dB & $3.2065\times10^{-5}$ & $0.0000$ \\
\bottomrule
\end{tabular}
\caption{Measurement-derived FRI edge variable projection. The sparse boundary proposal is estimated from adjoint backprojection and derivative peak localization, not from the hidden target geometry. It recovers most of the useful edge-aware gain but remains below the cleaner synthetic edge hint, identifying multi-view FRI geometry extraction as the next bottleneck.}
\label{tab:haouchat-fri-edge-variable-projection}
\end{table}
\begin{figure}[h]
\centering
\includegraphics[width=\linewidth]{../apps_industrial_breakthrough/haouchat_fri_edge_variable_projection_outputs/haouchat_fri_edge_variable_projection.png}
\caption{Measurement-derived FRI edge variable projection for the damped-harmonic basis. Panels show target, mean geometry, synthetic-edge EGGROLL, measurement-edge EGGROLL, and true-geometry oracle.}
\label{fig:haouchat-fri-edge-variable-projection}
\end{figure}
\begin{table}[h]
\centering
\small
\begin{tabular}{llccc}
\toprule
Basis & Method & Held-out PSNR & Field RMSE & Boundary RMSE \\
\midrule
Polynomial quadratic & Multi-ray FRI variable projection & $48.6908$ dB & $3.7930\times10^{-1}$ & $1.4848\times10^{-3}$ \\
Polynomial quadratic & True-geometry oracle & $48.9766$ dB & $3.7926\times10^{-1}$ & $0.0000$ \\
Exponential decay & Multi-ray FRI variable projection & $51.9659$ dB & $5.6990\times10^{-3}$ & $1.4848\times10^{-3}$ \\
Exponential decay & True-geometry oracle & $52.4398$ dB & $5.1297\times10^{-3}$ & $0.0000$ \\
Damped harmonic & Mean geometry & $33.3709$ dB & $6.6046\times10^{-2}$ & $9.5169\times10^{-2}$ \\
Damped harmonic & Row/adjoint FRI edge & $37.7852$ dB & $4.7983\times10^{-2}$ & $6.8656\times10^{-2}$ \\
Damped harmonic & Multi-ray FRI variable projection & $66.1488$ dB & $2.5205\times10^{-3}$ & $1.4848\times10^{-3}$ \\
Damped harmonic & EGGROLL after multi-ray edge & $40.3169$ dB & $3.6051\times10^{-2}$ & $3.9707\times10^{-2}$ \\
Damped harmonic & True-geometry oracle & $96.9888$ dB & $3.2065\times10^{-5}$ & $0.0000$ \\
\bottomrule
\end{tabular}
\caption{Multi-ray FRI edge refinement. Local boundary offsets are selected by held-in ray residuals after variable projection. The direct refined geometry nearly closes the oracle gap for the polynomial and exponential-decay bases and raises the matched damped-harmonic profile to $66.15$ dB without synthetic edge injection.}
\label{tab:haouchat-multiray-fri-edge-projection}
\end{table}
\begin{figure}[h]
\centering
\includegraphics[width=\linewidth]{../apps_industrial_breakthrough/haouchat_multiray_fri_edge_projection_outputs/haouchat_multiray_fri_edge_projection.png}
\caption{Multi-ray FRI edge refinement for the damped-harmonic basis. The direct multi-ray FRI geometry, not the subsequent EGGROLL search, gives the dominant quality gain.}
\label{fig:haouchat-multiray-fri-edge-projection}
\end{figure}
\begin{table}[h]
\centering
\small
\begin{tabular}{lccc}
\toprule
Method & Mean PSNR & Minimum PSNR & Mean boundary RMSE \\
\midrule
Mean geometry & $32.5922$ dB & $31.5426$ dB & $1.0261\times10^{-1}$ \\
Row/adjoint edge & $38.0604$ dB & $32.5913$ dB & $3.8393\times10^{-2}$ \\
Multi-ray FRI variable projection & $64.0770$ dB & $61.7645$ dB & $1.9800\times10^{-3}$ \\
True-geometry oracle & $97.9883$ dB & $83.1976$ dB & $0.0000$ \\
\bottomrule
\end{tabular}
\caption{Robustness sweep for multi-ray FRI edge refinement over $18$ cases: two boundary variants, three projection-angle budgets, and three noise levels. The multi-ray FRI geometry remains consistently high quality and closes most of the gap between crude adjoint peaks and the oracle.}
\label{tab:haouchat-multiray-fri-robustness}
\end{table}
\begin{figure}[h]
\centering
\includegraphics[width=\linewidth]{../apps_industrial_breakthrough/haouchat_multiray_fri_robustness_outputs/haouchat_multiray_fri_robustness_sweep.png}
\caption{Representative robustness-sweep preview for the base boundary at the largest projection budget.}
\label{fig:haouchat-multiray-fri-robustness}
\end{figure}
\section{Sparse Stochastic Innovation Models}
The sparse stochastic framework defines a process by \cite{unser2014sparse1,unser2014sparse2}
\[
L\{s\}=w,
\]
where $L$ is a whitening operator and $w$ is white innovation noise. Gaussian $w$ produces dense least-sparse processes; non-Gaussian Levy noise produces sparse or impulsive innovations. The operator controls correlation and physics, while the Levy measure controls sparsity.
The discrete-domain theory shows that matched B-spline filters convert continuous innovations into discrete generalized increments \cite{unser2014sparse2}. MAP and MMSE estimators for these priors are developed in \cite{bostan2013sparse,amini2013bayesian,kamilov2013mmse}. For OSNR, this means sparse parameters should not be arbitrary dense neural weights. They should be coefficient-domain innovations induced by the correct operator.
\subsection{Controlled SPDE validation: advection--diffusion with Levy innovations}
To convert the sparse stochastic theory into a PINN/weather-facing experiment, we implemented \texttt{apps\_industrial\_breakthrough/spde\_operator\_spline\_benchmark.py}. The controlled PDE is a periodic one-dimensional advection--diffusion--reaction model over a space--time block,
\begin{equation}
\mathcal{L}u
=
\left(\partial_t + a\partial_x-\nu\partial_{xx}+\lambda\right)u
= w(x,t),
\label{eq:spde-advection-diffusion}
\end{equation}
where $w$ is not restricted to be Gaussian. Following the Unser--Tafti sparse process model, the operator $\mathcal{L}$ fixes the correlation and propagation physics, while the innovation law determines the forcing morphology. We test four innovation profiles: a smooth periodic source, a Gaussian stochastic source, a compound-Poisson sparse impulse source, and a mixed weather-like source containing smooth waves, Gaussian background, and sparse jump events.
On the periodic grid, \eqref{eq:spde-advection-diffusion} has the Fourier-domain symbol
\begin{equation}
\widehat{\mathcal{L}}(\omega_t,\omega_x)
=
j\omega_t + ja\omega_x+\nu\omega_x^2+\lambda,
\end{equation}
so the OSNR state-free solve is the diagonal complex division
\begin{equation}
\widehat{u}(\omega_t,\omega_x)
=
\frac{\overline{\widehat{\mathcal{L}}(\omega_t,\omega_x)}}{|\widehat{\mathcal{L}}(\omega_t,\omega_x)|^2+\epsilon}
\widehat{w}(\omega_t,\omega_x).
\label{eq:spde-fft-solve}
\end{equation}
This is the SPDE analogue of an operator-matched exponential spline solve: the Green structure is built into the inverse operator, and no neural coordinate residual or automatic-differentiation tape is required. A low-pass Fourier reconstruction is included as a spectral-bias baseline; it mimics what happens when a smooth model family cannot carry non-Gaussian sparse innovations.
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{lcccccc}
\toprule
Profile & OSNR PSNR & Low-pass PSNR & Innovation RMSE & Event error & Hard zeros & Solve time \\
\midrule
Smooth periodic, $128^2$ & $138.0840$ dB & $85.4051$ dB & $1.2801\times10^{-4}$ & n/a & $0.00\%$ & $0.1383$ ms \\
Gaussian SPDE, $128^2$ & $90.9460$ dB & $38.7232$ dB & $2.9458\times10^{-4}$ & n/a & $0.00\%$ & $0.1370$ ms \\
Poisson sparse, $128^2$ & $69.9296$ dB & $34.4731$ dB & $3.6954\times10^{-3}$ & $0.0000$ px & $99.78\%$ & $0.1319$ ms \\
Mixed Levy weather, $128^2$ & $99.5862$ dB & $60.4309$ dB & $3.7530\times10^{-3}$ & $7.0438$ px & $0.00\%$ & $0.1414$ ms \\
Poisson sparse, low diffusion & $62.1205$ dB & $30.7575$ dB & $9.2609\times10^{-3}$ & $0.0000$ px & $99.41\%$ & $0.1353$ ms \\
Mixed Levy, low diffusion & $80.6083$ dB & $47.5993$ dB & $9.2813\times10^{-3}$ & $3.2817$ px & $0.00\%$ & $0.1472$ ms \\
Poisson sparse, $192^2$ & $71.4692$ dB & $34.3896$ dB & $2.4829\times10^{-3}$ & $0.0000$ px & $99.83\%$ & $0.5018$ ms \\
Mixed Levy weather, $192^2$ & $104.2661$ dB & $63.6906$ dB & $2.5052\times10^{-3}$ & $5.1575$ px & $0.00\%$ & $0.5070$ ms \\
\bottomrule
\end{tabular}
\caption{Controlled SPDE operator-spline benchmark for advection--diffusion with Gaussian and sparse Levy innovations. The OSNR solve is the direct FFT inversion of \eqref{eq:spde-fft-solve}; the low-pass row is a smooth spectral-bias baseline.}
\label{tab:spde-operator-spline}
\end{table}
\begin{figure}[h]
\centering
\includegraphics[width=\linewidth]{../apps_industrial_breakthrough/spde_operator_spline_outputs_main/spde_operator_spline_profiles.png}
\caption{SPDE profile comparison on the $128^2$ benchmark. Each row shows the target field, the OSNR state-free FFT reconstruction, and the low-pass smooth baseline. The sparse and mixed rows expose why a Gaussian/smooth-only surrogate is not enough for weather-like fronts and impulses.}
\label{fig:spde-operator-spline}
\end{figure}
Table~\ref{tab:spde-operator-spline} gives the current interpretation. The state-free operator solve is essentially exact for all four innovation laws and remains below one millisecond even at $192^2$. The pure compound-Poisson case recovers event coordinates exactly at the tested grid resolutions, validating the sparse innovation view. The mixed case is more realistic and more difficult: the field reconstruction remains excellent, but raw top-$K$ event localization degrades because the smooth and Gaussian components overlap the sparse impulses in the recovered innovation. This is not a failure of the operator inverse; it identifies the next algorithmic requirement. A weather-grade OSNR solver should add the same sparse-plus-smooth oblique innovation sieve used elsewhere in this paper, but now applied to $\mathcal{L}u$ rather than to the field $u$ itself.
We therefore added an explicit innovation-domain sieve to the same benchmark. Given the recovered innovation $\tilde{w}=\mathcal{L}\tilde{u}$, the sieve first estimates a smooth background $w_{\mathrm{sm}}=G_\sigma\ast \tilde{w}$ and then extracts sparse events from the residual
\begin{equation}
w_{\mathrm{sp}}=\mathcal{T}(\tilde{w}-w_{\mathrm{sm}}),
\end{equation}
where $\mathcal{T}$ is either a known-cardinality top-$K$ selector or an adaptive median-absolute-deviation threshold. This is not a field smoother; it acts after applying the physical operator and is therefore an innovation prior in the sense of Unser and Tafti. On the $128^2$ mixed Levy/weather case, raw top-$K$ localization has mean event error $7.0438$ px. The known-cardinality sieve reduces this to $0.0000$ px. The same result holds for the low-diffusion stress case, where raw localization is $3.2817$ px, and for the $192^2$ case, where raw localization is $5.1575$ px. The adaptive MAD sieve with threshold $8$ also recovers the mixed-weather events exactly without being told the number of events; it selects $36$ sparse sites in the mixed case and yields $0.0000$ px event error. In the pure Poisson case it selects a larger sparse support ($602$ sites at $128^2$) because the Gaussian smoothing residual leaves a local halo around each impulse, but nearest-event localization is still exact. Thus the next refinement is amplitude/support debiasing, not event detection. The practical weather implication is encouraging: OSNR can solve the stochastic PDE block globally and then separate sparse front/impulse innovations from smooth meteorological background in the physically meaningful residual domain.
We also tested the immediate nonlinear extension in \texttt{apps\_industrial\_breakthrough/forced\_burgers\_spde\_benchmark.py}. The model is a periodically forced viscous Burgers equation,
\begin{equation}
u_t + uu_x-\nu u_{xx}=f_{\mathrm{smooth}}(x,t)+f_{\mathrm{sp}}(x,t),
\end{equation}
where $f_{\mathrm{sp}}$ is a sparse set of localized Gaussian events. A high-resolution spectral RK4 rollout is treated as the reference trajectory; compressed OSNR rollouts retain only a fixed number of Fourier/operator modes. The key inverse-problem distinction is that the forcing innovation must be estimated by applying the nonlinear physical operator to the observed trajectory,
\begin{equation}
\tilde f(x,t)=u_t+u u_x-\nu u_{xx},
\end{equation}
not by thresholding the difference between a coarse rollout and the reference. The latter is mostly a truncation and phase-defect diagnostic. The former is the nonlinear analogue of the operator-domain innovation extraction used in the linear SPDE experiment.
We therefore report both sparse support recovery and compact event-atom recovery. The point sieve thresholds $\tilde f-G_\sigma\ast \tilde f$ and measures whether each true event overlaps the recovered sparse support. The weak-form atom score integrates $\tilde f$ against anisotropic Gaussian test functions matched to the injected event scale and then applies non-maximum suppression. We also apply a local centroid debiasing step around each detected atom. This approximates
\begin{equation}
\eta_m=\langle \tilde f,\varphi_m\rangle,
\end{equation}
where $\varphi_m$ is a compact adjoint/test atom. On clean synthetic forcing, direct operator-domain support recovery is sharper than the weak score; the weak form is expected to become more useful once observations are noisy or irregular.
\begin{table}[h]
\centering
\small
\begin{tabular}{lccccc}
\toprule
Coarse modes & Rollout PSNR & Support error & Atom error & Centroid error & Defect ratio \\
\midrule
$8$ & $39.5805$ dB & $0.0000$ px & $1.0357$ px & $0.7917$ px & $0.2290$ \\
$18$ & $68.7114$ dB & $0.0000$ px & $1.0357$ px & $0.7917$ px & $0.0098$ \\
$32$ & $108.0643$ dB & $0.0000$ px & $1.0357$ px & $0.7917$ px & $0.0001$ \\
\bottomrule
\end{tabular}
\caption{Forced Burgers SPDE diagnostic after correcting the inverse-problem residual. Applying the nonlinear operator to the observed trajectory recovers every sparse forcing support location at the tested grid resolution. The atom-center error is about one pixel because the injected events are finite-width Gaussian blobs and overlapping events shift local maxima; local centroid debiasing reduces this to $0.7917$ px. The defect ratio reports the norm of the coarse-rollout phase/truncation defect relative to the physical innovation norm.}
\label{tab:forced-burgers-spde}
\end{table}
\begin{figure}[h]
\centering
\includegraphics[width=0.92\linewidth]{../apps_industrial_breakthrough/forced_burgers_spde_weak_outputs_main/forced_burgers_spde.png}
\caption{Forced Burgers SPDE diagnostic at $18$ retained modes. The corrected operator innovation $\tilde f=u_t+u u_x-\nu u_{xx}$ exposes the sparse forcing structure directly. The support map recovers the event locations, while the weak-form score produces compact event atoms within about one pixel.}
\label{fig:forced-burgers-spde}
\end{figure}
The conclusion is important for the weather/PINN program. Linear operator-matched SPDEs are already a home-turf win for OSNR: exact global solves, sparse Levy innovations, and sub-millisecond runtime. The corrected nonlinear Burgers diagnostic shows that sparse forcing can also be recovered when the physical operator is applied in the right domain. A denser stress case with $56$ injected events still yields $0.0000$ px support error and $0.8912$ px centroid error. The remaining bottleneck is not event detection but support and amplitude debiasing for finite-width/overlapping events, especially under noisy or partially observed fields. The next layer should estimate sparse innovations through an adjoint weak form,
\begin{equation}
\langle f,\varphi_m\rangle
=
\langle u_t+uu_x-\nu u_{xx},\varphi_m\rangle,
\end{equation}
with test functions $\varphi_m$ matched to the operator and the expected front scale, plus a local centroid/amplitude debiasing step. An operator-splitting scheme that alternates deterministic nonlinear advection with a sparse forcing inverse problem is the natural production path before claiming weather-grade nonlinear SPDE recovery.
\paragraph{Direct PINN home-turf challenger.}
We added a more direct PINN-facing control in \texttt{apps\_industrial\_breakthrough/pinn\_operator\_home\_turf\_challenger.py}. The benchmark is a periodic two-dimensional Helmholtz/Poisson problem,
\begin{equation}
(-\Delta+\lambda)u(x,y)=f(x,y),
\end{equation}
where $u$ is a mixed-frequency smooth field and $f$ is obtained by applying the known operator. The OSNR path solves the field by a single FFT-domain division. The baseline is a SIREN-style coordinate PINN trained with Adam on data samples and automatic-differentiation residual collocation. On the quality-first $128^2$ run with $\lambda=6$, the full OSNR solve reaches $141.7829$ dB PSNR and RMSE $1.9704\times10^{-7}$ in $0.1695$ ms on CPU. A compressed low-mode OSNR profile retaining only $8.3557\%$ of Fourier bins still reaches $136.3778$ dB in $0.2935$ ms. The SIREN PINN baseline, after $1800$ epochs, reaches only $21.1458$ dB and RMSE $2.1203\times10^{-1}$ after $109.112$ s. This is not a noisy external-data claim; it is a clean operator-known PINN control. It demonstrates the central home-turf point: when the differential operator and boundary topology are known, structural inversion gives both higher accuracy and roughly $6.44\times10^5$ lower training latency than residual-learning the same field.
\begin{table}[h]
\centering
\small
\begin{tabular}{lcccc}
\toprule
Profile & PSNR & RMSE & Time & Active coefficients \\
\midrule
OSNR full spectral solve & $141.7829$ dB & $1.9704\times10^{-7}$ & $0.1695$ ms & $100.00\%$ \\
OSNR low-mode solve & $136.3778$ dB & $3.6712\times10^{-7}$ & $0.2935$ ms & $8.3557\%$ \\
SIREN PINN, $1800$ epochs & $21.1458$ dB & $2.1203\times10^{-1}$ & $109.112$ s & dense MLP \\
\bottomrule
\end{tabular}
\caption{Direct PINN home-turf challenger on a periodic $128^2$ Helmholtz/Poisson field. The OSNR rows are measured FFT/operator inversions; the SIREN PINN row is measured Adam training with autograd residual collocation.}
\label{tab:pinn-home-turf}
\end{table}
\begin{figure}[h]
\centering
\includegraphics[width=\textwidth]{../apps_industrial_breakthrough/pinn_operator_home_turf_outputs_n128_e1800/pinn_operator_home_turf_panel.png}
\caption{PINN home-turf visual panel. The full and low-mode OSNR inversions are visually indistinguishable from the target at the displayed scale, while the trained SIREN PINN remains visibly over-smoothed after the measured optimization budget.}
\label{fig:pinn-home-turf}
\end{figure}
The scale follow-up at $256^2$ confirms that the coefficient fraction improves with resolution when the operator spectrum is compact. With the same low-mode budget, OSNR reaches $143.7870$ dB in $0.6588$ ms for the full solve, and $136.8069$ dB in $1.1293$ ms while retaining only $2.0889\%$ of Fourier bins. A $900$-epoch PINN baseline on the same field reaches $19.4427$ dB after $59.916$ s. At $512^2$, the same low-mode budget retains only $0.5222\%$ of Fourier bins and still reaches $136.4969$ dB in $3.5055$ ms; the full solve reaches $143.3450$ dB in $2.3255$ ms, while a $300$-epoch PINN baseline reaches $19.0573$ dB after $19.664$ s. The purpose of these rows is not to claim a universal neural-operator benchmark victory; they isolate the regime where PINN residual learning is structurally the wrong computational tool.
We ran an additional observation-noise stress test to separate robust atom detection from brittle support thresholding. Gaussian observation noise is added to the trajectory before evaluating the nonlinear operator. At $0.1\%$ relative observation noise, raw pointwise support thresholding misses many events ($7.8618$ px support error), but ranked atom selection from the same operator residual remains accurate ($0.7801$ px after centroid refinement). Mild pre-operator smoothing restores support overlap ($0.0357$ px) but blurs atom centers ($1.4350$ px). At $0.5\%$ noise, the best tested atom setting uses $\sigma=0.75$ pre-smoothing and reaches $0.7975$ px centroid error, while binary support thresholding is unreliable. This confirms the correct noisy-weather design: detect a ranked set of operator-domain event atoms first, then run local amplitude/support debiasing rather than relying on a global hard threshold.
\subsection{Operator-symbol identification by variable projection}
The preceding SPDE experiments assume that the differential operator is known. The next weather/PINN question is whether OSNR can also learn a compact operator from data without falling back to a dense coordinate network. We therefore added \texttt{apps\_industrial\_breakthrough/operator\_pole\_identification\_benchmark.py}. The controlled model is the same advection--diffusion--reaction family
\begin{equation}
\left(\partial_t+a\partial_x-\nu\partial_{xx}+\lambda\right)u=w,
\end{equation}
but now the coefficients $(a,\nu,\lambda)$ are treated as unknown operator parameters. In Fourier space,
\begin{equation}
\widehat{w}
-
j\omega_t\widehat{u}
=
\left(ja\omega_x+\nu\omega_x^2+\lambda\right)\widehat{u},
\end{equation}
so the unknown operator coefficients enter linearly once the observed field and innovation are transformed. OSNR therefore identifies the operator by one complex ridge least-squares solve over selected frequency bins,
\begin{equation}
\widehat{\theta}
=
\arg\min_{\theta=(a,\nu,\lambda)}
\left\|
\mathbf{D}(\widehat{u})\theta
-
\left(\widehat{w}-j\omega_t\widehat{u}\right)
\right\|_2^2
+\epsilon\|\theta\|_2^2.
\end{equation}
This is a variable-projection step: linear field coefficients remain solved by the operator inverse, while the low-dimensional operator symbol is recovered directly from the data. A backpropagation baseline optimizes the same three parameters by Adam through the spectral residual for $800$ steps.
\begin{table}[h]
\centering
\small
\begin{tabular}{lcccccc}
\toprule
Profile & Method & $\hat a$ & $\hat\nu$ & $\hat\lambda$ & PSNR & Time \\
\midrule
Clean, all bins & OSNR LS & $0.730001$ & $0.021008$ & $0.168636$ & $76.9176$ dB & $1.4462$ ms \\
Clean, band $24$ & OSNR LS & $0.729999$ & $0.021000$ & $0.170013$ & $116.0950$ dB & $0.2236$ ms \\
Clean & Adam residual & $0.729998$ & $0.037428$ & $0.038625$ & $17.9194$ dB & $206.4690$ ms \\
$0.5\%$ noise, band $24$ & OSNR LS & $0.729844$ & $0.019821$ & $0.366630$ & $40.2696$ dB & $0.1575$ ms \\
$0.5\%$ noise, band $12$ & OSNR LS & $0.729962$ & $0.020995$ & $0.171116$ & $56.3422$ dB & $0.1897$ ms \\
$0.5\%$ noise & Adam residual & $0.723947$ & $0.030782$ & $0.044012$ & $20.8344$ dB & $205.5888$ ms \\
\bottomrule
\end{tabular}
\caption{Operator-symbol identification for the advection--diffusion--reaction family with true parameters $(a,\nu,\lambda)=(0.73,0.021,0.17)$. The OSNR row uses a single complex least-squares solve in the Fourier/operator domain; the baseline uses iterative backpropagation through the same residual. Conservative spectral fitting bands suppress derivative-amplified observation noise.}
\label{tab:operator-pole-identification}
\end{table}
\begin{figure}[h]
\centering
\includegraphics[width=0.92\linewidth]{../apps_industrial_breakthrough/operator_pole_identification_outputs_noise005_band12/operator_pole_identification.png}
\caption{Noisy operator-identification result at $0.5\%$ observation noise with a conservative fitting band. The recovered operator reconstructs the state at $56.3422$ dB after one structured coefficient solve.}
\label{fig:operator-pole-identification}
\end{figure}
Table~\ref{tab:operator-pole-identification} is the first explicit operator-learning result. In the clean case, a frequency band of $24$ modes recovers all three coefficients to near machine precision and improves the reconstruction from $76.9176$ dB to $116.0950$ dB by avoiding ill-conditioned bins. With $0.5\%$ observation noise, fitting too many frequencies corrupts the reaction estimate because derivative operators amplify high-frequency noise. Tightening the band to $12$ modes restores the coefficients to sub-percent relative error and yields $56.3422$ dB, while the Adam residual baseline remains near $20.8$ dB after $800$ gradient steps. The lesson is directly relevant to weather data: unknown physics should be learned as a compact, stability-constrained operator symbol with explicit spectral/noise control, not as an unconstrained dense coordinate network.
We then tested the harder field-only variant in \texttt{apps\_industrial\_breakthrough/blind\_operator\_sparsity\_identification.py}. Here $w$ is hidden: the search chooses the operator whose residual $\mathcal{L}_\theta u$ is most compressible as a low-pass smooth field plus a fixed number of sparse atoms. This is closer to unsupervised weather-model discovery, but it exposes an identifiability boundary. On four independent trajectories sharing the same true operator, the blind compressibility score selects $(\hat a,\hat\nu,\hat\lambda)=(0.91,0.021,0.26)$ instead of $(0.73,0.021,0.17)$, even though the sparse event locations are recovered exactly. A support-projected oracle that masks the true sparse event neighborhoods but does not know their amplitudes also fails to recover the reaction coefficient. The reason is structural: from $u$ alone, a wrong operator can be absorbed into a different smooth forcing background, so sparse-plus-smooth compressibility is not a unique operator identifier.
\begin{table}[h]
\centering
\small
\begin{tabular}{lccccc}
\toprule
Setting & $\hat a$ & $\hat\nu$ & $\hat\lambda$ & Event error & Time \\
\midrule
Blind sieve, $4$ clean trajectories & $0.910000$ & $0.021000$ & $0.260000$ & $0.0000$ px & $2596.29$ ms \\
Blind sieve, $4$ trajectories, $0.2\%$ noise & $0.910000$ & $0.009000$ & $0.260000$ & $0.0000$ px & $2594.86$ ms \\
Support-projected oracle, clean & $0.570370$ & $0.014276$ & $-14.202470$ & oracle support & $5.86$ ms \\
Support-projected oracle, $0.2\%$ noise & $0.314185$ & $0.000886$ & $1.900169$ & oracle support & $5.96$ ms \\
\bottomrule
\end{tabular}
\caption{Blind operator-discovery diagnostic. Sparse event geometry can be recovered from $\mathcal{L}_\theta u$, but field-only sparse-plus-smooth compressibility does not uniquely identify the true operator because operator mismatch can be reinterpreted as smooth forcing.}
\label{tab:blind-operator-sparsity}
\end{table}
\begin{figure}[h]
\centering
\includegraphics[width=0.92\linewidth]{../apps_industrial_breakthrough/blind_operator_sparsity_outputs_support_projected/blind_operator_sparsity_identification.png}
\caption{Blind operator-sparsity diagnostic. The selected residual preserves sparse event locations but corresponds to the wrong operator, demonstrating that fully blind field-only operator discovery needs additional physical anchors.}
\label{fig:blind-operator-sparsity}
\end{figure}
This negative result is useful. It says the breakthrough lane is not arbitrary unsupervised PDE discovery from a single scalar field. The credible path is semi-blind operator learning: use measured innovations, multiple observed state channels, conservation laws, boundary/flux constraints, or assimilation windows to anchor the smooth forcing ambiguity, then recover the compact operator symbol by structured least squares or variable projection.
The first semi-blind anchor test follows this prescription. In \texttt{apps\_industrial\_breakthrough/anchored\_operator\_identification.py}, only a random subset of the forcing samples is revealed. The field $u$ is observed everywhere, but the operator coefficients are fitted from the pointwise equations
\begin{equation}
u_t(t_i,x_i)+a u_x(t_i,x_i)-\nu u_{xx}(t_i,x_i)+\lambda u(t_i,x_i)=w(t_i,x_i)
\end{equation}
at the anchor sites. All derivatives are evaluated analytically by spectral/operator columns, and the three unknown coefficients are recovered by a real ridge least-squares solve.
\begin{table}[h]
\centering
\small
\begin{tabular}{lccccc}
\toprule
Setting & Anchors & $\hat a$ & $\hat\nu$ & $\hat\lambda$ & PSNR \\
\midrule
Clean, $0.25\%$ anchors & $41$ & $0.729829$ & $0.021105$ & $0.143701$ & $49.6630$ dB \\
Clean, $0.50\%$ anchors & $82$ & $0.729892$ & $0.021072$ & $0.159291$ & $58.4940$ dB \\
Clean, $5.00\%$ anchors & $819$ & $0.729956$ & $0.021002$ & $0.172572$ & $69.6683$ dB \\
$0.2\%$ noise, no denoise, $5.00\%$ anchors & $819$ & $0.732336$ & $0.008352$ & $2.355422$ & $29.6130$ dB \\
$0.2\%$ noise, band $12$, $5.00\%$ anchors & $819$ & $0.729782$ & $0.020778$ & $0.183559$ & $53.7874$ dB \\
\bottomrule
\end{tabular}
\caption{Semi-blind operator identification from sparse forcing anchors. Clean operator recovery is accurate with very few forcing samples. Under observation noise, derivative columns require spectral denoising; with a band-$12$ field prefilter, $5\%$ anchors recover the operator to below $8\%$ worst relative error and reconstruct the state at $53.7874$ dB.}
\label{tab:anchored-operator-identification}
\end{table}
\begin{figure}[h]
\centering
\includegraphics[width=0.92\linewidth]{../apps_industrial_breakthrough/anchored_operator_identification_outputs_noise002_denoise12_moreanchors/anchored_operator_identification.png}
\caption{Semi-blind noisy operator identification with sparse forcing anchors. A small set of pointwise forcing measurements breaks the field-only ambiguity exposed in Table~\ref{tab:blind-operator-sparsity}.}
\label{fig:anchored-operator-identification}
\end{figure}
Table~\ref{tab:anchored-operator-identification} is a more realistic weather/PINN direction than fully blind scalar discovery. It shows that a small number of physical anchors can make compact operator identification well-posed again. The nonmonotone noisy rows also identify the next engineering layer: anchors should be selected by leverage or derivative-energy criteria rather than uniformly at random.
We tested the simplest version of that idea by selecting anchors with the largest normalized derivative-column energy. This naive leverage rule is not sufficient. In the clean case it reaches only $52.1897$ dB at $5\%$ anchors, below the random-anchor $69.6683$ dB result. With $0.2\%$ observation noise and band-$12$ denoising, it degrades to $37.4777$ dB at $5\%$ anchors because high-leverage points are also the points where derivative noise is most amplified. The active-anchor rule must therefore combine derivative leverage with noise sensitivity and spatial diversity; selecting the largest rows of the design matrix is too brittle.
A follow-up robustification adds a trimmed ridge solve: after the first anchor fit, the largest pointwise residuals are discarded and the operator is refit on the lowest-residual fraction. Random anchors with trimming do not improve the best $5\%$ noisy row, but diverse leverage plus trimming uncovers a useful low-anchor operating point. With only $0.5\%$ anchors under $0.2\%$ observation noise, diverse-trimmed anchors estimate $(a,\nu,\lambda)=(0.731760,0.020692,0.175461)$, corresponding to only $3.2124\%$ worst relative parameter error and $49.5094$ dB reconstruction. The best state PSNR still comes from denser random anchors, but the active-trimmed result shows that carefully chosen anchors can reduce physical measurements by an order of magnitude while preserving an accurate compact operator.
\subsection{Nonlinear Burgers operator identification and held-out forecasting}
The next SOTA-facing PINN target is nonlinear forecasting rather than static reconstruction. We implemented \texttt{apps\_industrial\_breakthrough/burgers\_operator\_identification\_forecast.py}, which treats viscous Burgers dynamics
\begin{equation}
u_t + c\,u u_x = \nu u_{xx}
\end{equation}
as a compact operator-identification problem. From the observed training window, OSNR forms analytic derivative columns $(u_t,uu_x,u_{xx})$ and solves the two unknown coefficients $(c,\nu)$ by a tiny ridge system over sparse anchor samples. The learned operator is then rolled forward over the held-out future window. This is the nonlinear analogue of the semi-blind anchor experiments above, but the validation target is future prediction, which is the quantity that PINNs and neural operators usually report.
\begin{table}[h]
\centering
\small
\begin{tabular}{lccccc}
\toprule
Setting & Anchors & $\hat c$ & $\hat\nu$ & Future PSNR & ID time \\
\midrule
Clean, $0.25\%$ anchors & $35$ & $1.003308$ & $0.0045239$ & $70.7236$ dB & $0.4064$ ms \\
Clean, $0.50\%$ anchors & $69$ & $0.998324$ & $0.0044914$ & $77.6890$ dB & $0.2062$ ms \\
Clean, $1.00\%$ anchors & $138$ & $0.999828$ & $0.0044959$ & $88.6117$ dB & $0.1814$ ms \\
Adam residual, clean & all & $0.997542$ & $0.0044901$ & $74.8301$ dB & $82.15$ ms \\
$0.1\%$ noise, $0.50\%$ anchors & $69$ & $0.995605$ & $0.0044539$ & $66.3736$ dB & $0.2458$ ms \\
$0.2\%$ noise, $1.00\%$ anchors & $138$ & $0.995355$ & $0.0044588$ & $66.8131$ dB & $0.3160$ ms \\
Adam residual, $0.2\%$ noise & all & $0.979668$ & $0.0044440$ & $56.8246$ dB & $81.62$ ms \\
Wrong prior & n/a & $0.75$ & $0.00225$ & $30.7955$ dB & n/a \\
\bottomrule
\end{tabular}
\caption{Nonlinear Burgers operator identification and held-out forecasting. True parameters are $(c,\nu)=(1.0,0.0045)$, the training window is $45\%$ of the timeline, and the future window contains the remaining $90$ frames. OSNR identifies the operator in sub-millisecond time from sparse anchors and forecasts the future without a neural training loop.}
\label{tab:burgers-operator-id-forecast}
\end{table}
\begin{figure}[h]
\centering
\includegraphics[width=0.92\linewidth]{../apps_industrial_breakthrough/burgers_operator_identification_outputs_noise002_modes28_sub16/burgers_operator_identification_forecast.png}
\caption{Noisy Burgers operator-ID forecast at $0.2\%$ observation noise. OSNR recovers the nonlinear operator from $1\%$ training-window anchors and forecasts the held-out future at $66.8131$ dB.}
\label{fig:burgers-operator-id-forecast}
\end{figure}
This is the strongest nonlinear PINN-facing result so far. The clean $1\%$ anchor row reaches $88.6117$ dB future PSNR with a $0.1814$ ms identification solve, while the Adam residual fit is about $450\times$ slower and reaches only $74.8301$ dB. Under $0.2\%$ observation noise, OSNR still reaches $66.8131$ dB from $1\%$ anchors, outperforming the Adam residual fit by about $10$ dB. The wrong-prior row shows that the forecast is not trivially easy: incorrect physics collapses to $30.7955$ dB. This is the first result that directly combines nonlinear coefficient discovery, held-out forecasting, sparse physical measurements, and a clear optimization-speed gap.
\subsection{Coupled nonlinear shallow-water operator forecasting}
The scalar Burgers forecast is a necessary control, but the weather/PINN claim requires a coupled multi-field nonlinear system. We therefore implemented \texttt{apps\_industrial\_breakthrough/nonlinear\_shallow\_water\_operator\_forecast.py}. The state is $\mathbf q=(\eta,u,v)$ and the operator family is
\begin{align}
\eta_t &=
-H(u_x+v_y)-\beta\,\nabla\cdot(\eta(u,v))+\mu_h\Delta\eta,\\
u_t &=
-g\eta_x+f v-r u-\beta(u u_x+v u_y)+\nu\Delta u,\\
v_t &=
-g\eta_y-f u-r v-\beta(u v_x+v v_y)+\nu\Delta v.
\end{align}
The unknown physical vector is
\[
\theta=(H,\beta,g,f,r,\nu,\mu_h),
\]
covering mean depth, nonlinear transport strength, gravity, Coriolis coupling, damping, momentum viscosity, and height diffusion. OSNR forms the full analytic derivative library from the observed training window and solves a scaled sparse-anchor linear system for all seven coefficients at once. The fitted operator is then advanced over the held-out future window using the same spectral RK4 physics core. The comparison baseline fits the same residual equations by Adam over all training rows, so the quality comparison is not against a weak interpolant but against the standard differentiable residual-minimization path used by PINN-style methods.
\begin{table}[h]
\centering
\small
\begin{tabular}{lcccccc}
\toprule
Setting & Anchors & $\hat\beta$ & $\hat g$ & Max param. err. & Future PSNR & ID time \\
\midrule
$48^2$, clean, $0.25\%$ & $795$ & $0.7178$ & $0.8590$ & $1.8214\%$ & $68.7348$ dB & $14.24$ ms \\
$48^2$, clean, $1.00\%$ & $3180$ & $0.7181$ & $0.8590$ & $1.6975\%$ & $68.7314$ dB & $13.24$ ms \\
Adam residual, $48^2$ clean & all & $0.7186$ & $0.8590$ & $1.8572\%$ & $68.6608$ dB & $830.30$ ms \\
$48^2$, $0.1\%$ noise & $795$ & $0.7091$ & $0.8594$ & $21.347\%$ & $66.3867$ dB & $16.93$ ms \\
Adam residual, $0.1\%$ noise & all & $0.7171$ & $0.8587$ & $17.495\%$ & $66.0641$ dB & $1093.02$ ms \\
$48^2$, $0.2\%$ noise & $795$ & $0.7218$ & $0.8599$ & $34.417\%$ & $65.6597$ dB & $17.96$ ms \\
Adam residual, $0.2\%$ noise & all & $0.7156$ & $0.8585$ & $32.218\%$ & $63.4251$ dB & $1095.57$ ms \\
$64^2$, clean, $0.10\%$ & $713$ & $0.7191$ & $0.8594$ & $1.8368\%$ & $72.6945$ dB & $24.00$ ms \\
Adam residual, $64^2$ clean & all & $0.7191$ & $0.8593$ & $1.1991\%$ & $72.5755$ dB & $1103.88$ ms \\
Wrong prior, $64^2$ & n/a & $0.3240$ & $1.0750$ & n/a & $27.7993$ dB & n/a \\
\bottomrule
\end{tabular}
\caption{Coupled nonlinear shallow-water operator identification and held-out forecasting. True parameters are $(H,\beta,g,f,r,\nu,\mu_h)=(1.0,0.72,0.86,0.58,0.065,0.006,0.004)$. OSNR uses sparse derivative anchors; Adam optimizes the same residual over all training rows.}
\label{tab:nonlinear-shallow-water-forecast}
\end{table}
\begin{figure}[h]
\centering
\includegraphics[width=0.92\linewidth]{../apps_industrial_breakthrough/nonlinear_shallow_water_operator_forecast_outputs_121x64_clean/nonlinear_shallow_water_operator_forecast.png}
\caption{Coupled nonlinear shallow-water forecast at $121\times64^2$. With only $0.1\%$ sparse derivative anchors, OSNR identifies the seven-parameter nonlinear operator and forecasts the held-out future at $72.6945$ dB. The wrong-prior forecast falls to $27.7993$ dB, confirming that the high score is not a trivial smoothness artifact.}
\label{fig:nonlinear-shallow-water-forecast}
\end{figure}
This is the first multi-field nonlinear weather-core forecast result in the project. It preserves the key advantage seen in Burgers: the residual landscape can be collapsed into a small structured operator solve instead of optimized by thousands of neural/PINN gradient steps. On the $64^2$ run, OSNR uses only $713$ anchor equations out of the full derivative library and identifies the operator in $24.00$ ms, while the Adam residual fit takes $1103.88$ ms. Both methods converge to similar coefficients in the clean case because the library is correct, but OSNR reaches the solution in one scaled linear solve with about a $46\times$ identification-speed advantage and no neural training loop. Under observation noise, derivative bias still affects the weak damping/diffusion terms, but the future forecast remains above $65$ dB and stays ahead of Adam in the tested $0.2\%$ setting. The next moonshot is therefore not another scalar PDE; it is sparse/partial observation data assimilation for this same coupled nonlinear operator family.
\subsection{Adaptive sparse sensors for nonlinear shallow-water assimilation}
We then cross-pollinated the weather station-placement result with the coupled nonlinear shallow-water core. The runner \texttt{apps\_industrial\_breakthrough/nonlinear\_shallow\_water\_adaptive\_sensor\_assimilation.py} keeps the same no-backprop pipeline: sparse sensors reconstruct the observed training window by a closed-form Fourier-dictionary ridge solve, the seven-parameter nonlinear operator is identified by sparse least squares, and the state is rolled into the held-out future. The only changed variable is where the sparse sensors are placed. Sensor policies are computed from the training window only, never from held-out future frames. We compare random points, raw training-window variance/gradient/leverage scores, lattice-plus-score hybrids, residual-innovation hybrids, and centered/phase-shifted coverage policies.
\begin{table}[H]
\centering
\small
\begin{tabular}{lcccc}
\toprule
Observation setting & Modes & Sensor policy & Sensors & Future PSNR \\
\midrule
Clean & $4$ & lattice $3\%$ & $123$ & $19.842$ dB \\
Clean & $4$ & lattice+hybrid $10\%$ budget, $3\%$ total & $123$ & $19.822$ dB \\
Clean & $3$ & centered lattice $1.5\%$ & $61$ & $19.852$ dB \\
Clean & $3$ & centered lattice $2\%$ & $82$ & $19.942$ dB \\
Clean & $3$ & offset-best lattice $2\%$ & $82$ & $19.958$ dB \\
Clean & $3$ & random $20\%$ & $819$ & $19.920$ dB \\
$1\%$ sensor noise & $3$ & centered lattice $2\%$ & $82$ & $19.948$ dB \\
$1\%$ sensor noise & $3$ & random $20\%$ & $819$ & $19.906$ dB \\
\bottomrule
\end{tabular}
\caption{Adaptive sparse-sensor placement for nonlinear shallow-water assimilation and forecasting. All rows use the same $121\times64^2$ trajectory, closed-form sparse-window assimilation, sparse least-squares operator identification, and no backpropagation. Mode $3$ centered/offset coverage reaches dense-random forecast quality with $10\times$ fewer observations. The offset-best policy chooses the best phase among $16$ centered lattice shifts by training-window reconstruction PSNR only.}
\label{tab:nonlinear-shallow-water-adaptive-sensors}
\end{table}
\begin{figure}[H]
\centering
\includegraphics[width=0.92\linewidth]{../apps_industrial_breakthrough/nonlinear_shallow_water_adaptive_sensor_assimilation_outputs_full_modes3_offset_lattice/adaptive_sensor_panel.png}
\caption{Adaptive shallow-water sparse-sensor forecast panel for the mode-$3$ coverage-geometry run. Centered/offset lattice rows preserve the large-scale future height field with $2$--$3\%$ sensors, while dense random placement needs about $20\%$ sensors to reach the same forecast band.}
\label{fig:nonlinear-shallow-water-adaptive-sensors}
\end{figure}
This result is useful because it is positive and diagnostic. The naive high-information policies are not winners: variance, gradient, and hybrid-diverse placement overconcentrate sensors in active regions and can make the Fourier reconstruction ill-conditioned. The follow-up lattice-plus-information experiment confirmed the same boundary: at $3\%$ total sensors with mode $4$, lattice+gradient, lattice+hybrid, and lattice+residual $10\%$ allocation reach $19.786$, $19.822$, and $19.807$ dB, all below the pure lattice row at $19.828$ dB. The real improvement is coverage geometry plus basis order. Mode $3$ is the sparse bias-variance sweet spot; modes $1$--$2$ underfit and modes $5$--$6$ are underconstrained at low sensor counts. A centered lattice at $2\%$ sensors reaches $19.942$ dB clean future PSNR, above the same-run $20\%$ random reference at $19.920$ dB; with $1\%$ sensor noise, the same $2\%$ centered lattice row reaches $19.948$ dB versus noisy random $20\%$ at $19.906$ dB. The low-count sweep shows the transition: $0.5\%$ centered sensors fail ($15.628$ dB), $1\%$ is not yet dense-random quality ($19.158$ dB), $1.5\%$ approaches it ($19.852$ dB), and $2\%$ crosses it.
We then tested whether this was a single-trajectory phase artifact. The script now exposes initial roll and amplitude controls, and the offset-best policy chooses the best of $16$ lattice phases by training-window reconstruction only. Across four robustness variants, offset-best $2\%$ sensors remains at or above random $20\%$: roll $(7,11)$ gives $19.946$ versus $19.942$ dB, roll $(13,5)$ gives $19.955$ versus $19.919$ dB, amplitude scale $1.25$ gives $19.948$ versus $19.894$ dB, and a changed dynamics profile $(\beta,g,f)=(0.9,0.78,0.45)$ gives $19.981$ versus $19.932$ dB. Thus adaptive station placement is not only a terminal weather-assimilation trick. In a coupled nonlinear weather-core forecast, a coverage-aware sensor topology can reduce observations by $10\times$ while preserving future forecast quality, using only algebraic assimilation and operator identification.
Finally, we removed the known-family assumption after sparse sensing. The runner \texttt{apps\_industrial\_breakthrough/nonlinear\_shallow\_water\_sparse\_sensor\_library\_discovery.py} first reconstructs the training window from sparse sensors, then fits the $24$-column shallow-water library by sequential thresholded least squares, and forecasts from the assimilated last state. This is a harder test because the solver must reject decoy columns and no longer receives the seven-parameter operator family.
\begin{table}[H]
\centering
\scriptsize
\setlength{\tabcolsep}{3pt}
\begin{tabular}{@{}p{0.17\linewidth}p{0.13\linewidth}c c c c p{0.15\linewidth}@{}}
\toprule
Observation setting & Policy & Sensors & Threshold & Support $(TP,FP,FN)$ & Future PSNR & Reference \\
\midrule
Clean & offset-best $2\%$ & $82$ & $0.003$ & $(8,0,5)$ & $19.960$ dB & known-family $19.962$ dB \\
Clean & random $20\%$ & $819$ & $0.005$ & $(7,0,6)$ & $19.888$ dB & known-family $19.902$ dB \\
$1\%$ sensor noise & offset-best $2\%$ & $82$ & $0.003$ & $(8,0,5)$ & $19.965$ dB & known-family $19.964$ dB \\
$1\%$ sensor noise & random $20\%$ & $819$ & $0.003$ & $(7,0,6)$ & $19.885$ dB & known-family $19.912$ dB \\
\bottomrule
\end{tabular}
\caption{Sparse-sensor governing-equation discovery after Fourier assimilation. The discovered support is counted against the $13$ true library columns and $11$ decoys. Offset-best $2\%$ sensors recover an $8$-term true subset with zero decoys and match or exceed dense-random $20\%$ forecast quality.}
\label{tab:nonlinear-shallow-water-sparse-sensor-library}
\end{table}
The sparse-library result changes the interpretation. The $2\%$ offset-best row does not fully recover all weak nonlinear/damping terms, but it recovers the dominant conservative, pressure, Coriolis, and diffusion operators with no decoys and forecasts within about $0.002$ dB of the known-family coefficient fit. A lower threshold $0.001$ recovers $9$ true terms with one decoy and reaches $19.964$ dB clean, but the zero-decoy $8$-term threshold is the cleaner scientific claim. Thus the sensor topology is not merely helping a fixed PDE prior; it preserves enough operator information for sparse governing-equation discovery from partial observations.
The next step was to remove a flaw in the pointwise library: it thresholds every equation-specific term and every decoy on the same normalized scale, even though the physically meaningful shallow-water operators are shared typed groups. We therefore added a weak-form typed library in \texttt{nonlinear\_shallow\_water\_sparse\_sensor\_weakform\_discovery.py}. Each window enforces
\[
q(t_b)-q(t_a)\approx \int_{t_a}^{t_b}\mathcal{L}_j(q(t))\,dt
\]
and projects the balance onto low Fourier test modes. The $13$ true pointwise terms are then tied into seven shared physical groups $(H,\beta,\mu_h,g,f,r,\nu)$, while the $11$ nuisance columns receive a larger typed-selection threshold. This is not a neural loss or a backward pass: it is a weak-form operator balance followed by weighted sequential thresholded least squares.
\begin{table}[H]
\centering
\scriptsize
\setlength{\tabcolsep}{3pt}
\begin{tabular}{@{}p{0.17\linewidth}p{0.14\linewidth}c c c c p{0.17\linewidth}@{}}
\toprule
Observation setting & Policy & Sensors & $(\tau,\lambda_{\rm decoy})$ & Support $(TP,FP,FN)$ & Future PSNR & Reference \\
\midrule
Clean & random $2\%$ & $82$ & $(2{\times}10^{-6},50)$ & $(7,1,0)$ & unstable & known-family $16.055$ dB \\
Clean & offset-best $2\%$ & $82$ & $(5{\times}10^{-7},20)$ & $(7,0,0)$ & $19.961$ dB & weak dense $19.961$ dB \\
Clean & random $20\%$ & $819$ & $(10^{-6},50)$ & $(7,0,0)$ & $19.909$ dB & weak dense $19.909$ dB \\
$1\%$ sensor noise & offset-best $2\%$ & $82$ & $(5{\times}10^{-7},20)$ & $(7,0,0)$ & $19.956$ dB & weak dense $19.956$ dB \\
$1\%$ sensor noise & random $20\%$ & $819$ & $(10^{-6},50)$ & $(7,0,0)$ & $19.910$ dB & weak dense $19.910$ dB \\
\bottomrule
\end{tabular}
\caption{Typed weak-form sparse-sensor governing-equation discovery. Support is counted over seven shared physical operator groups and eleven decoys. The decoy multiplier $\lambda_{\rm decoy}$ applies only to nuisance columns. Offset-best $2\%$ sensors recover the complete seven-group shallow-water operator with zero decoys and match the dense weak-form solve, while using $10\times$ fewer observations than random $20\%$.}
\label{tab:nonlinear-shallow-water-typed-weak-sensor-discovery}
\end{table}
This closes the support-recovery gap left by Table~\ref{tab:nonlinear-shallow-water-sparse-sensor-library}. With typed weak-form rows, offset-best $2\%$ sensors recover all seven physical groups with zero false positives in both the clean and $1\%$ sensor-noise settings. The same row is also forecast-competitive: $19.961$ dB clean and $19.956$ dB noisy, above the corresponding typed random-$20\%$ rows ($19.909$ and $19.910$ dB). Random $2\%$ still fails despite selecting most physical groups, which confirms that the result is not merely a threshold artifact. The station geometry must preserve a well-conditioned weak operator balance; once it does, typed OSNR selection can recover the full coupled shallow-water operator from sparse partial observations without backpropagation.
The typed weak-form result also passed the first robustness sweep. At $1\%$ offset sensors, the selector already recovers $(7,0,0)$ support but only reaches $19.505$ dB, so full support and forecast-quality crossing are separate requirements. At $2\%$ offset sensors, all four trajectory variants recover $(7,0,0)$: roll $(7,11)$ gives $19.926$ dB, roll $(13,5)$ gives $19.948$ dB, initial scale $1.25$ gives $19.946$ dB, and changed dynamics $(\beta,g,f)=(0.9,0.78,0.45)$ gives $19.978$ dB. The random-$20\%$ typed weak-form rows for the same variants are $19.942$, $19.916$, $19.944$, and $19.977$ dB, respectively. Thus complete support recovery is robust in the tested variants; the $10\times$ forecast-quality advantage holds in three of four variants and narrowly fails on roll $(7,11)$.
We then tested whether the same typed weak-form mechanism learns a reusable operator rather than a trajectory-specific correction. The multi-trajectory runner trains one shared operator from sparse-observed variants \{base, roll $(7,11)$, scale $1.25$\} and forecasts unseen roll variants $(13,5)$ and $(5,17)$ from their held-out states. Offset-best $2\%$ sensors recover full support and reach a mean unseen-trajectory PSNR of $51.100$ dB; random $2\%$ is unstable even with nearly full support. Dense random $20\%$ also recovers full support and reaches $52.232$ dB, while offset-best $20\%$ reaches $62.902$ dB. This is the first sparse-observation cross-trajectory operator-learning result in this section. It is not a $10\times$ dense-quality win at $2\%$ observations, but it shows that the typed OSNR weak-form solver can learn a shared coupled operator from partial observations and transfer it to unseen initial conditions without backpropagation.
A follow-up sensor-fraction and placement sweep showed that the cross-trajectory coefficient bottleneck is primarily geometric. With offset-best placement and the same typed solver, $3\%$ sensors already recover $(7,0,0)$ and reach $53.282$ dB, exceeding the random-$20\%$ result; $5\%$ reaches $57.847$ dB; $10\%$ drops to $52.865$ dB; and $20\%$ reaches $62.902$ dB. Thus adding sensors is not monotone unless the station geometry remains well conditioned for the weak-form operator rows. A placement sweep at $3$--$10\%$ found that pointwise saliency policies (variance, gradient, and hybrid-diverse additions) consistently introduce decoy groups and degrade transfer. The best sparse clean-support result is a centered lattice with $5\%$ sensors and a stronger decoy multiplier: it recovers $(7,0,0)$ and reaches $58.829$ dB on the two unseen trajectories. If the two derivative-decoy groups are allowed, the same $5\%$ centered lattice reaches $60.775$ dB, but we treat this as a numerical correction rather than a clean governing-equation discovery. The practical conclusion is that operator-identifiability-balanced station geometry is more important than generic high-activity station placement.
We then made the geometry test explicit by adding row-conditioning and fixed lattice-phase sweeps. Unit-design, unit-joint, and clipped row normalizations all made the solver worse: they activated most decoys and collapsed transfer to roughly $22$--$30$ dB. This negative result is important because the weak-form row magnitudes carry physical operator information; flattening them destroys the balance rather than improving conditioning. In contrast, fixed quarter-phase lattice placement is a productive control variable. At $5\%$ sensors on the three-training-trajectory protocol, the best clean fixed phase $(0.75,0.75)$ reaches $60.365$ dB with exact $(7,0,0)$ support, and the result validates on fresh roll and scale variants with mean $60.365$ dB. Expanding the training set to eight sparse-observed trajectories raises the same clean $5\%$ phase result to $60.945$ dB. Most importantly, a focused sensor curve with this phase shows that $7.5\%$ sensors per training trajectory recover exact support and reach $65.072$ dB on four fresh test variants, exceeding the same-protocol $20\%$ phase reference of $64.149$ dB. The curve remains nonmonotone: $10\%$ drops to $55.014$ dB and $15\%$ to $59.252$ dB. Thus the current lesson is sharper than ``more sensors'': sparse OSNR operator learning can beat denser observation budgets when station geometry is phase-balanced for the weak operator, but station-count increases can still harm coefficient estimation if they alias the weak-form rows.
We also audited whether the phase can be selected without looking at the final test variants. Simple training-only proxies failed: physical-column condition number, physical--decoy coherence, dense weak residual, leave-one-training-trajectory weak residual, leave-one theta stability, and an inner training-window rollout score did not rank the best phases. A standard validation split, however, does. Selecting the phase on validation variants \{roll $(3,9)$, scale $0.75$\} chooses $(0.25,0.75)$, which then transfers to disjoint test variants \{roll $(11,4)$, scale $1.40$, roll $(19,2)$, scale $1.30$\}. The validation-selected $7.5\%$ phase recovers exact support and reaches $65.161$ dB on that disjoint test set, while the same phase with $20\%$ sensors reaches $64.283$ dB. This is the cleanest current sparse cross-trajectory result: the station geometry is selected on validation data, the test variants are unseen, and the learned seven-group operator still beats the denser observation budget without backpropagation.
Finally, we tested whether the nonmonotone phase-lattice curve could be repaired by replacing the lattice with low-discrepancy or jittered station families. It could not. On the same validation-selected test protocol, Sobol phase stations activated $7$--$10$ decoys and produced unstable forecasts at $7.5\%$, $10\%$, and $15\%$ sensors. Jittered phase lattices were stable but much weaker: $46.774$ dB at $7.5\%$, $44.902$ dB at $10\%$, and $58.205$ dB at $15\%$. The unjittered phase lattice remains the best clean geometry, with $65.161$ dB at $7.5\%$. Thus the current station rule is not generic space filling; it is a Fourier-compatible phase-balanced sampling rule.
We then promoted the validation split from phase selection to joint phase/count selection. The training set stayed fixed at eight sparse-observed variants \{base, roll $(7,11)$, scale $1.25$, roll $(13,5)$, roll $(5,17)$, roll $(2,19)$, scale $0.90$, scale $1.10$\}. The validation variants were again roll $(3,9)$ and scale $0.75$. The grid searched lattice phases $(0.25,0.75)$, $(0.75,0.25)$, $(0,0.5)$, $(0.75,0.75)$, $(0,0.75)$, and $(0.25,0.5)$ at sensor fractions $5\%$, $6.25\%$, $7.5\%$, $8.75\%$, $10\%$, $12.5\%$, and $15\%$, with threshold $5\times10^{-7}$, decoy multiplier $20$, no row normalization, window $7$, stride $2$, and a $75\%$ inner training-window diagnostic split. Validation selected the $8.75\%$ lattice with phase $(0.75,0.75)$: it recovered exact $(7,0,0)$ support and reached $67.996$ dB on the two validation variants. Without changing any hyperparameter, the selected row transferred to disjoint test variants \{roll $(11,4)$, scale $1.40$, roll $(19,2)$, scale $1.30$\}, reaching $67.841$ dB with exact support. Same-run references were $65.432$ dB at $6.25\%$, $65.161$ dB for the previous $7.5\%$ phase $(0.25,0.75)$ row, and $64.283$ dB for the same $20\%$ phase $(0.25,0.75)$ row. Thus held-out validation can now select both observation count and phase, and the selected sparse geometry uses only $358$ stations per training trajectory while outperforming $819$-station dense-phase references.
A final fine phase/count refinement around this winner exposed a sharper resonance. We searched $8.125\%$, $8.4375\%$, $8.75\%$, $9.0625\%$, and $9.375\%$ sensors with phases $(0.62,0.62)$, $(0.62,0.75)$, $(0.75,0.62)$, $(0.75,0.75)$, $(0.75,0.87)$, $(0.87,0.75)$, $(0.87,0.87)$, $(0.62,0.87)$, and $(0.87,0.62)$. The adjacent count bands $8.125\%$ and $8.4375\%$ were poor despite exact support, reaching only about $53$--$54$ dB; $9.0625\%$ recovered to $68.520$ dB at phase $(0.75,0.75)$, but the validation winner was again $8.75\%$, now with phase $(0.75,0.87)$. This row reached $70.774$ dB on validation and transferred to the disjoint test variants at $69.803$ dB, with exact $(7,0,0)$ support and a $1.31$ ms sparse solve. The same fine-test run reproduced the old $8.75\%$ phase $(0.75,0.75)$ result at $67.841$ dB and showed that moving the winning phase to $9.0625\%$ drops to $65.941$ dB. The result is therefore not a generic phase preference. It is a count-specific Fourier sampling geometry that materially improves coefficient accuracy while keeping the observation budget at $358$ stations per training trajectory.
To check whether the fine geometry was overfitting the two validation variants, we ran a broader fresh-variant audit with roll shifts $(1,23)$, $(23,1)$, $(31,17)$, $(17,31)$ and amplitude scales $0.60$, $1.60$, $0.50$, and $1.75$. The selected $8.75\%$ phase $(0.75,0.87)$ row reached $70.166$ dB across these eight variants with exact support. Same-count controls were $67.882$ dB for phase $(0.75,0.75)$ and $63.047$ dB for phase $(0.25,0.75)$, while $20\%$ references reached only $64.044$, $64.104$, and $65.107$ dB for the three tested phases. This robustness audit strengthens the interpretation: the selected sparse station geometry generalizes across unseen roll and amplitude perturbations and beats substantially denser station budgets because it better identifies the weak operator coefficients, not because it sees more observations.
Because the resonance was phase-sharp, we then ran a local phase-only refinement at the fixed $8.75\%$ count. The validation grid swept $x$ phases $0.70,0.72,0.75,0.78,0.80$ and $y$ phases $0.84,0.87,0.90,0.93$ around the previous winner. Most rows were much weaker even with exact support; for example $x=0.80$ remained below $60$ dB and $(0.75,0.93)$ dropped to $66.390$ dB. The validation winner was $(0.75,0.90)$ at $73.568$ dB. Tested on the union of the four disjoint variants and the eight broad-audit variants, this row reached $73.379$ dB with exact support and a $1.17$ ms solve. On the same $12$-variant audit, $(0.75,0.87)$ reached $70.045$ dB and $(0.75,0.75)$ reached $67.868$ dB. This is the current best clean shallow-water result: validation-selected sparse station geometry with $358$ observations per training trajectory beats both same-count neighboring phases and all tested $819$-station references by a large margin.
One more one-percent refinement around $(0.75,0.90)$ saturated rather than improved the result. Sweeping $x\in\{0.73,0.74,0.75,0.76,0.77\}$ and $y\in\{0.88,0.89,0.90,0.91,0.92\}$ at the same $8.75\%$ count again selected $(0.75,0.90)$; $(0.75,0.91)$ tied it because the rounded station set is effectively equivalent. Nearby rows drop quickly: $(0.75,0.89)$ gives $72.614$ dB, $(0.75,0.92)$ gives $68.431$ dB, $x=0.76$--$0.77$ with $y=0.90$--$0.91$ gives $70.367$ dB, and $x=0.73$--$0.74$ remains near $62$--$63$ dB. Thus the station-design frontier appears locally saturated at this lattice resolution; the next improvement must come from a different station family, a richer validation criterion, or a stronger operator/library model rather than sub-percent phase nudging.
We next tested whether more sparse-observed training trajectories improve the shared operator. They do not automatically help. On a fresh test set \{roll $(9,27)$, roll $(27,9)$, roll $(15,29)$, roll $(29,15)$, scale $0.70$, scale $1.50$, scale $0.40$, scale $1.90$\}, the current eight-training-variant row with $8.75\%$ phase $(0.75,0.90)$ reaches $73.381$ dB. Adding four more sparse-observed training variants \{roll $(1,23)$, roll $(23,1)$, scale $0.60$, scale $1.60$\} while keeping the same per-trajectory station budget and solver drops the same fresh-test mean to $70.316$ dB. The support remains exact, but the coefficient vector shifts, especially in the nonlinear and damping terms. Thus the next operator-learning lever is not simply more trajectories; training variants must be selected or weighted so that assimilation bias from scale-extreme trajectories does not distort the shared weak-form coefficients.
The isolating controls confirm that the degradation is not caused by one family alone. Adding only the two extra roll variants to the eight-variant training set gives $71.321$ dB on the same fresh test set; adding only the two scale-extreme variants gives $70.131$ dB. Both retain exact support, but both move the coefficients away from the high-PSNR eight-variant estimate. This suggests that the original eight sparse-observed trajectories already form a good coefficient-calibration design for this station phase. Additional trajectories should enter only through validation-selected weights or subset selection, not by unweighted concatenation.
We implemented that weighting hook in the runner as \texttt{--train\_variant\_weights}, multiplying each variant's weak-form rows and targets by the square root of its weight before the closed-form solve. Downweighting the four rejected variants improves over unweighted concatenation but still does not beat the eight-variant subset: weights $0.25$, $0.10$, $0.03$, and $0.01$ on the four added variants yield $72.490$, $73.035$, $73.280$, and $73.346$ dB, respectively, on the same fresh test set, versus $73.381$ dB for weight zero. Thus the validation-selected action for these candidates is rejection. The useful research conclusion is that the no-backprop operator learner can support neuromodulatory-style reliability weights, but the first weighted audit says the next gain requires discovering better candidate trajectories or operator features, not softly retaining known harmful variants.
We then audited three alternative explanations before changing the operator model. First, a threshold/decoy-pressure sweep around the $8.75\%$ phase $(0.75,0.90)$ frontier used thresholds $10^{-7}$, $2\times10^{-7}$, $5\times10^{-7}$, $10^{-6}$, and $2\times10^{-6}$ with decoy multipliers $5$, $10$, $20$, $50$, and $100$. Low decoy penalties admitted false positives and dropped validation to about $66$ dB, but every exact-support row gave the same $73.568$ dB validation score; applying the validation-selected $10^{-7}$, multiplier-$50$ row to the $12$-variant fresh audit reproduced $73.379$ dB. Thus the frontier is not limited by the sparse threshold once decoys are suppressed.
Second, we implemented phase-preserving score-mixed station policies such as \texttt{lattice\_phase75\_90\_gradient01}. These keep the tuned phase lattice as the backbone and replace only $1$--$5\%$ of the station budget with diverse high-gradient, high-variance, or hybrid-score sites. This fairer saliency audit was decisively negative. At the fixed $352$--$358$ station scale, generic score-mixed policies tied to the untuned lattice collapsed to $49.713$ dB or worse, and even the phase-preserving variants degraded monotonically: gradient replacement at $1\%$, $2\%$, $3\%$, and $5\%$ gave $68.468$, $65.134$, $62.046$, and $52.616$ dB; variance replacement gave $55.350$, $51.061$, $46.917$, and $42.942$ dB; hybrid replacement gave $58.324$, $54.872$, $50.958$, and $45.586$ dB. The conclusion is that pointwise saliency is not an adequate station objective for this weak operator learner. The lattice points themselves carry Fourier conditioning, and replacing even a few of them damages the coefficient estimate despite exact support in several rows.
The positive improvement came from exact decimation of the phase lattice. Scanning the integer station counts $350$ through $361$ at phase $(0.75,0.90)$ found a new validation winner at $352$ stations, i.e. sensor fraction $352/4096=0.0859375$. This row recovers exact $(7,0,0)$ support and reaches $74.442$ dB on the validation variants, compared with $73.568$ dB for the previous $358$-station row and $73.267$ dB for the complete $19\times19$ grid with $361$ stations. Nearby counts are sharply worse: $350$--$351$ give about $66$ dB, $353$--$354$ give $68.795$--$69.939$ dB, $356$ gives $66.661$ dB, and $359$--$360$ give $70.570$--$72.844$ dB. A local phase refinement at the $352$-station count confirmed $(0.75,0.90)$, with $(0.75,0.91)$ tied by an effectively equivalent rounded station set. On the $12$-variant fresh audit, the validation-selected $352$-station row reaches $73.925$ dB with exact support and a $1.21$ ms sparse solve, improving the previous $358$-station fresh frontier of $73.379$ dB while using fewer observations. This is now the cleanest sparse shallow-water operator-learning result in the manuscript: progress came not from more data, saliency replacement, or threshold tuning, but from validation-selected Fourier-compatible station decimation.
We added an exact \texttt{--sensor\_counts} option and widened the decimation sweep to counts $320$--$380$ at the same phase. This exposed an even sharper sparse resonance at $334$ stations, i.e. $334/4096=0.08154296875$ observations per training trajectory. The $334$-station row reaches $75.851$ dB on the validation variants with exact $(7,0,0)$ support, while nearby counts again fluctuate strongly: $328$--$330$ sit near $71$--$72$ dB, $332$ activates false support, $335$ gives $71.424$ dB, $336$ activates two false positives, and the entire $362$--$380$ side-$20$ band stays below $68$ dB except for false-support rows. A fresh $12$-variant audit of the locked $334$-station row reaches $75.841$ dB with exact support and a $1.18$ ms sparse solve. Local phase refinement at count $334$ again selects $(0.75,0.90)$, with $(0.75,0.91)$ tied by the rounded station set. This supersedes the $352$-station checkpoint: the validation-selected operator now improves the broad fresh audit by $2.462$ dB over the previous $358$-station frontier while using $6.7\%$ fewer observations.
We also checked whether the same phase contains an even lower-count resonance. A validation sweep over exact station counts $220$--$319$ at phase $(0.75,0.90)$ was negative. The best row in that band is $318$ stations at only $68.349$ dB, and most rows sit near $55$--$66$ dB, with occasional false-support failures such as counts $232$, $235$, $240$, $297$, and $306$. Thus the current sparse optimum is not simply ``as few stations as possible.'' For this Fourier dictionary and weak-form window, the useful resonance appears to start near the high end of the side-$19$ decimation family, with $334$ stations as the current validated minimum-quality sweet spot.
The next audit asked whether the $334$-station geometry was limited by the Fourier assimilation basis itself. Holding the training variants, validation variants, station count, phase $(0.75,0.90)$, window, threshold, and decoy pressure fixed, we swept the reconstruction basis from modes $2$ through $6$. Mode $2$ still selected exact support but underfit the observed window and biased the nonlinear coefficient, reaching only $52.464$ dB validation PSNR. The previous mode-$3$ row reached $75.851$ dB validation and $75.841$ dB on the locked $12$-variant fresh audit. Mode $4$ gives a small but clean improvement: it reaches $76.474$ dB on validation, exact $(7,0,0)$ support, and a $1.23$ ms sparse solve; the disjoint $12$-variant audit reaches $76.444$ dB with exact support and a $1.30$ ms solve. Modes $5$ and $6$ regress to $75.486$ and $74.903$ dB, respectively, despite exact support. Repeating the exact-count sweep $320$--$380$ under mode $4$ again selects $334$ stations; count $352$ rises to $74.894$ dB but stays below the $334$-station row, and the side-$20$ band remains weaker or false-support. The current interpretation is therefore a two-axis resonance: the best sparse operator learner is not maximal observation count or maximal basis bandwidth, but the mode-$4$, $334$-station, phase-balanced Fourier geometry.
We then revisited the weak temporal projection itself. The committed rows used a window of $7$ frames, stride $2$, and test-mode radius $3$. At the locked mode-$4$, $334$-station geometry, shortening the window is a major coefficient-calibration lever. With stride $2$ and test-mode radius $3$, validation PSNR rises from $76.474$ dB at window $7$ to $76.131$ dB at window $6$, $78.346$ dB at window $5$, $78.244$ dB at window $4$, $79.204$ dB at window $3$, and $79.582$ dB at window $2$, all with exact $(7,0,0)$ support. The projection radius is sharp: at window $3$, radius $2$ collapses to $59.003$ dB and radius $4$ drops to $69.496$ dB; at window $5$, radius $2$ and $4$ give $59.126$ and $70.595$ dB. The lower boundary and stride controls also reject a trivial ``shorter is always better'' rule: window $1$ gives $78.707$ dB, while window $2$ with stride $1$ and $3$ gives $79.189$ and $78.717$ dB. The selected weak setting is therefore window $2$, stride $2$, radius $3$. On the locked $12$-variant fresh audit this reaches $79.509$ dB, exact support, and a $1.26$ ms sparse solve, improving the previous mode-$4$ fresh frontier by $3.065$ dB and the old $358$-station frontier by $6.130$ dB. The inferred vector $(H,\beta,g,f,r,\nu,\mu_h)=(0.99980,0.72348,0.86033,0.58257,0.06370,0.005998,0.004022)$ is now close to the true operator across all seven groups. The lesson is precise: long weak windows were smearing the sparse-assimilated trajectory balance; a short, non-overdense weak window better matches the local truncation and assimilation error scale.
Re-sweeping station counts after the weak-window correction shows that this is not a new count-search problem. With mode $4$, window $2$, stride $2$, radius $3$, phase $(0.75,0.90)$, and the same validation variants, counts $300$--$360$ again select $334$ stations at $79.582$ dB. The low-count side remains far below the frontier: the best sub-$320$ count is $318$ at $68.956$ dB. The old $352$-station checkpoint improves from $74.894$ dB to $76.376$ dB under the shorter weak window, and the previous $358$-station phase row rises to $74.607$ dB, but both remain clearly below $334$. Secondary bumps such as count $328$ at $72.560$ dB and count $343$ at $73.083$ dB do not change the ordering. Thus the weak-window correction improves coefficient calibration at fixed geometry, while the station-count resonance itself remains locked.
We also ported exact count and phase-station support into the sparse weak-form library-discovery runner, then audited the harder raw and grouped libraries at the locked geometry. This uses the same mode-$4$, $334$-station, phase $(0.75,0.90)$, window-$2$ weak system, but asks the selector to reject nuisance terms rather than assuming the seven physical groups. The raw typed weak solve at threshold $5\times10^{-7}$ and decoy penalty $5$ recovers all $13$ physical columns with zero decoys, active support $(13,0,0)$, in $0.54$ ms. Its grouped typed counterpart recovers all seven shared physical groups with zero decoys, active support $(7,0,0)$, in $0.56$ ms. Raising the threshold to $5\times10^{-6}$ prunes weak true terms, and thresholds $5\times10^{-5}$ or larger over-prune the operator. The forecast values in this single-trajectory discovery runner remain near $20$ dB because the rollout starts from the sparse-assimilated state whose reconstruction PSNR is only about $21$ dB; this row should therefore be read as a support-identifiability result, not as the high-quality multitrajectory forecast frontier above.
We then inserted the same raw-vs-grouped choice into the high-quality multitrajectory forecast protocol. This separates support recovery from long-horizon transfer. At the locked mode-$4$, $334$-station, phase $(0.75,0.90)$, window-$2$, stride-$2$, radius-$3$ setting, the raw typed library again recovers exact support $(13,0,0)$ at threshold $5\times10^{-7}$, but its two-variant validation forecast is only $73.676$ dB. The grouped typed operator recovers $(7,0,0)$ and reproduces the $79.582$ dB frontier. Thus the seven-group collapse is not merely a reporting convention. Enforcing the shared physical coefficients $(\beta,g,f,r,\nu)$ across their equation-specific columns is a strong structural regularizer for rollout quality, even when the ungrouped raw support is exactly correct.
We also repeated the station and weak-projection controls under the improved window-$2$ setting. A $20$-policy local phase grid with $x\in\{0.65,0.70,0.75,0.80,0.85\}$ and $y\in\{0.80,0.85,0.90,0.95\}$ again selects phase $(0.75,0.90)$ at $79.582$ dB; the nearest strong neighbor is $(0.75,0.80)$ at $79.179$ dB, while many exact-support phases fall into the $60$--$69$ dB range. Phase-preserving score replacement remains decisively negative at the locked count: replacing only $2$--$15\%$ of the $(0.75,0.90)$ lattice by gradient, variance, hybrid, or leverage stations never improves the frontier. The best mixed row is gradient-$2\%$ at $66.530$ dB, and larger gradient replacements introduce decoys or drop below $54$ dB; variance, hybrid, and leverage replacements are similarly weaker. A targeted low-amplitude train-weighting control also loses: weights $(1,1,0.5,1,1,1,2,1)$ on \{base, roll $(7,11)$, scale $1.25$, roll $(13,5)$, roll $(5,17)$, roll $(2,19)$, scale $0.90$, scale $1.10$\} reach $79.153$ dB, below the equal-weight row. Finally, the radius sweep closes the weak-test-mode axis for window $2$: radius $1$ diverges, radius $2$ gives $58.987$ dB, radius $3$ gives $79.582$ dB, radius $4$ gives $69.786$ dB, and radius $5$ gives $65.683$ dB. The selected radius is therefore not arbitrary; it is the unique tested projection scale that balances sparse-assimilation bias and weak-form identifiability.
As a no-backprop control-theory follow-up, we added optional sparse-sensor rollout calibration to the multitrajectory runner via \texttt{--theta\_calibration\_steps}. Starting from the weak-form coefficient vector, the routine performs coordinate search over the seven physical parameters and accepts changes that reduce a held-out training-tail loss measured only at the observed sparse station locations. This is biologically and control-theoretically plausible in the sense that it uses forward rollouts and local observation residuals, not reverse-mode differentiation or full-field labels. On the current frontier, however, it is a hard negative result. With steps $1\%$, $0.3\%$, and $0.1\%$, the sparse sensor-tail loss decreases only from $0.2092457$ to $0.2092395$, while validation PSNR collapses from $79.582$ dB to $57.751$ dB. The calibrated vector moves to $(0.99780,0.73364,0.85514,0.58257,0.06281,0.005914,0.003970)$. The conclusion is useful: direct sparse-sensor replay is an overfitting objective for this problem. The weak-form grouped solve generalizes because it optimizes an operator balance, not because it best replays a short sparse observation tail.
We then attacked the same bottleneck from the reconstruction and weak-row side. First, temporal smoothing of the assimilated training sequence before weak integration is negative: centered binomial-$3$, binomial-$5$, and box-$3$ smoothing reduce validation to $54.490$, $50.174$, and $51.966$ dB. The smoothed sequences have nearly the same reconstruction PSNR as the raw assimilated fields, but their weak coefficients are biased. Second, concatenating additional weak views is also negative. At the locked mode-$4$, $334$-station, phase $(0.75,0.90)$ geometry, a $50/50$ window-$2$/window-$3$ weak-row mixture reaches only $79.341$ dB, an $80/20$ mixture reaches $79.470$ dB, and a $90.9/9.1$ window-$2$/window-$4$ mixture reaches $79.344$ dB. A lightly weighted radius-$4$ test-mode view also degrades the diagnostic validation run. Thus the selected window-$2$, radius-$3$ weak projection is not merely underdetermined; adding nearby valid projections injects biased rows rather than averaging out the reconstruction error.
The next reconstruction audit clarifies the direction. Raising the Fourier assimilation dictionary from mode $4$ to modes $5$, $6$, and $7$ lowers the training-window reconstruction PSNR and drops validation to $76.850$, $75.503$, and $71.546$ dB, even though exact $(7,0,0)$ support is still recovered. In contrast, applying a rectangular low-pass filter to the mode-$4$ assimilated fields before forming the weak rows gives a tiny but clean improvement when the keep radius is $3$. Filter radius $2$ underfits and reaches only $76.787$ dB; radius $4$ is effectively the unfiltered baseline at $79.583$ dB. Radius $3$ reaches $79.587$ dB on validation and $79.519$ dB on the $12$-variant fresh audit with exact support and a $1.19$ ms sparse solve. The coefficient vector is $(0.99980,0.72336,0.86033,0.58258,0.06371,0.005998,0.004022)$. This supersedes the unfiltered same-ridge fresh row at $79.511$ dB and the previous unfiltered frontier at $79.509$ dB, but only by about $0.01$ dB. The scientific value is therefore diagnostic rather than headline: the weak learner is now limited by aliasing and sparse-reconstruction bias at the operator-balance level.
We then made this anti-aliasing more operator-specific. The runner now supports \texttt{--weak\_filter\_application} and \texttt{--assim\_spatial\_filter\_shell\_weight}. The selected row keeps the weak target and all linear operator columns on the unfiltered mode-$4$ assimilated sequence, replaces only the true quadratic flux/advection columns by their mode-$3$ low-pass values, and retains the first excluded Fourier shell with weight $0.05$. This is a term-local weak-form filter: it does not smooth the rollout state, does not alter the sparse station observations, and does not use held-out future fields. On the validation variants this \texttt{nonlinear\_terms} row reaches $79.588$ dB with exact $(7,0,0)$ support; on the locked $12$-variant fresh audit it reaches $79.521$ dB, exact support, and a $1.24$ ms sparse solve, with coefficients $(0.99980,0.72336,0.86033,0.58258,0.06371,0.005998,0.004022)$. Filtering nonlinear terms and decoys gives the same validation score; filtering all columns with the same shell reaches only $79.519$ dB fresh. Finally, the same nonlinear-only filter does not make higher assimilation bandwidth safe: mode $5$ and mode $6$ validation runs fall to $76.693$ and $75.301$ dB despite exact support. The interpretation is sharper: the main aliasing source is the quadratic product library, but high-bandwidth assimilated states also bias the linear weak balance and coefficient calibration. The next material step should therefore be an operator-aware station or anti-aliasing objective that keeps the stable mode-$3$ nonlinear products while preserving the useful mode-$4$ linear state information.
A decomposition audit then checked whether the filter should act after products are formed or only on one nonlinear physical channel. Product-column filtering is too late: filtering the already-formed nonlinear product columns reaches only about $79.583$ dB validation, essentially the unfiltered row. Momentum-advection-only state filtering is also negative at $79.581$ dB. Mass-flux-only state filtering is the strongest validation row, reaching $79.591$ dB when only the continuity-equation $\beta$ column is formed from the mode-$3$ filtered state. However, this validation gain does not transfer: the hard mass-flux row reaches $79.520$ dB on the $12$-variant fresh audit, and shell weights $0.05$ and $0.10$ also reach only $79.520$ dB. Thus the fresh frontier remains the all-nonlinear state-prefiltered row above. The useful conclusion is methodological: two validation variants can over-rank continuity-specific anti-aliasing, so the next selector must use a broader validation design or a physically derived anti-aliasing criterion rather than a two-trajectory validation score alone.
We therefore widened the selector itself before running further filter searches. The broad validation set contains eight additional roll and amplitude variants, \texttt{roll9\_27}, \texttt{roll27\_9}, \texttt{roll15\_29}, \texttt{roll29\_15}, \texttt{scale070}, \texttt{scale150}, \texttt{scale040}, and \texttt{scale190}, while the eight sparse-observed training variants remain fixed. Under this V8 selector, the unfiltered row scores $79.513$ dB, all-column keep-$3$ filtering scores $79.521$ dB, all nonlinear-state filtering scores $79.523$ dB, the all-nonlinear shell-$0.05$ row scores $79.523$ dB, mass-flux-only hard filtering scores $79.522$ dB, and mass-flux-only shell-$0.05$ filtering scores $79.522$ dB, all with exact $(7,0,0)$ support. The broader selector therefore chooses the same all-nonlinear shell-$0.05$ rule that transferred best to the $12$-variant fresh audit, rather than the continuity-only row that won the narrow two-variant validation. This is the current robust selection rule: keep the mode-$4$ assimilated state for targets and linear operators, form all true nonlinear state products from the mode-$3$ state with a $0.05$ first-shell taper, and validate across both phase rolls and amplitude extremes.
Two follow-up geometry audits closed the obvious remaining local axes. Re-sweeping exact station counts $318,328,334,343,352,358$ under the V8 selector and the nonlinear shell rule again selects $334$ stations. The tested rows all recover exact support, but coefficient calibration is sharply count-dependent: the V8 means are $68.898$, $71.748$, $79.523$, $73.055$, $75.713$, and $74.287$ dB. A nearby phase grid at count $334$ is even sharper. The nine policies \texttt{lattice\_phase70\_85} through \texttt{lattice\_phase80\_95} all recover exact support, but only \texttt{lattice\_phase75\_90} reaches the frontier. The other phase rows range from $58.980$ to $68.973$ dB. Thus the station objective is not ``recover the seven groups''; it is to preserve the Fourier-compatible weak-row geometry that calibrates the nonlinear and damping coefficients.
We also tested whether the weaker high-amplitude validation rows could be fixed by adding amplitude-extreme trajectories to the training weak solve. On a disjoint V8b holdout, the fixed eight-trajectory training set scores $79.536$ dB. Adding \texttt{scale070} and \texttt{scale150} with equal weights drops to $76.522$ dB; giving those added variants only $0.1$ weight still drops to $79.317$ dB. A new whole-trajectory \texttt{--train\_block\_normalization} control was added to test block-level scaling without row-wise physics destruction. Target-RMS block normalization on the same augmented set activates all $11$ decoys and drops to $65.019$ dB, while the dense true-support coefficient row is still only $77.358$ dB. The result is negative but useful: broad amplitude coverage is not automatically a regularizer for this sparse weak system. The next structural lever should be an operator-domain anti-aliasing or station-design criterion, not more amplitude variants or block/row rescaling.
The first useful station-design criterion came from the operator domain rather than from reconstruction or row conditioning. We added \texttt{--operator\_reference\_mode dense\_train}, which fits a dense weak-form reference operator on the training window only and scores each sparse station geometry by the relative distance between its sparse-derived $\theta$ and this dense-training $\theta$. The dense-training reference for the current eight-trajectory set is $(1.00037,0.71970,0.86033,0.58005,0.06454,0.006010,0.004011)$, within $4.59\times10^{-4}$ relative error of the known simulator coefficients. This training-only objective selects \texttt{lattice\_phase75\_90} in the nine-phase grid and selects the same count-$334$, phase-$75\_90$ row in a local $3\times3$ count/phase grid. The selected row has operator-reference error $0.00283$ and V8 PSNR $79.523$ dB; the next closest local candidate is count $352$, phase $75\_90$ with error $0.01088$ and $75.713$ dB. This is not a new quality frontier, but it is a more principled path to station placement: design sparse observation geometry to reproduce the dense training operator, then evaluate transfer on disjoint futures.
We then stress-tested that rule as an actual selector over a broader $30$-candidate grid: station counts $318,328,334,343,352,358$ crossed with phases \texttt{70\_90}, \texttt{75\_85}, \texttt{75\_90}, \texttt{75\_95}, and \texttt{80\_90}. On the V8 holdout, both held-out PSNR and the training-only dense-operator objective select count $334$, phase \texttt{75\_90}, with $79.523$ dB and operator-reference error $0.00283$. The top held-out alternatives are count $352$, phase \texttt{75\_90} at $75.713$ dB and count $358$, phase \texttt{75\_90} at $74.287$ dB. Repeating the same grid on a disjoint V8b holdout gives the same top row: $79.536$ dB for count $334$, phase \texttt{75\_90}; the next held-out rows are count $352$, phase \texttt{75\_90} at $75.834$ dB and count $358$, phase \texttt{75\_90} at $74.347$ dB. The caveat is that operator-reference error is a successful top-$1$ selector here, not yet a calibrated total ordering: for example, count $358$, phase \texttt{75\_85} has the second-lowest reference error but only about $68$ dB. The next station-design objective should therefore keep dense-operator matching as the primary constraint, but add a training-only stability term that penalizes geometries whose coefficient match is fragile under small phase, count, or trajectory perturbations.
A follow-up diagnostic showed that the failure mode is not local instability but coefficient scaling. The original dense-operator score is an absolute relative L2 distance, so it is dominated by the large coefficients $(H,\beta,g,f)$ and can underweight small but rollout-sensitive coefficients such as damping and viscosity. We therefore added \texttt{--operator\_reference\_metric} with \texttt{absolute\_l2}, \texttt{relative\_l2}, and \texttt{relative\_linf} options. The \texttt{relative\_linf} metric scores the maximum coefficient-wise relative error to the dense training operator. On the same V8 $30$-candidate grid, \texttt{relative\_linf} again selects count $334$, phase \texttt{75\_90}, with score $0.01294$ and $79.523$ dB. The row count $358$, phase \texttt{75\_85}, which was second-best under absolute L2, is demoted to score $0.04446$ because its small-coefficient errors are large. This makes the selector more physically balanced: it still does not perfectly rank all held-out PSNR values, but it removes the most obvious scale artifact in dense-operator matching.
Finally, a broader relative-\(\ell_\infty\) selector search crossed eight counts, $300,318,328,334,343,352,358,372$, with the full $3\times3$ phase neighborhood from \texttt{70\_85} through \texttt{80\_95}. Across all $72$ rows, both held-out PSNR and the balanced dense-operator score again select count $334$, phase \texttt{75\_90}, with $79.523$ dB and score $0.01294$. The best non-frontier held-out row is count $352$, phase \texttt{75\_90} at $75.713$ dB; count $372$ never exceeds $67.353$ dB. This closes the obvious count/phase lattice search. The next lever should be the weak-form estimator itself: multi-window test functions, coefficient-balanced row construction, or model-bias correction for the nonlinear product columns.
We then held the selected station geometry fixed and tested whether concatenated weak views could reduce estimator bias. They did not. The baseline single view, window $2$ and test radius $3$, remains $79.523$ dB with balanced operator score $0.01294$. Equal-weight windows $(1,2,3)$ drop to $79.334$ dB; equal-weight windows $(2,3,4)$ drop to $78.918$ dB; a conservative $90/10$ mix of windows $(2,3)$ still drops to $79.476$ dB. Spatial test-mode concatenation is worse: modes $(2,3)$ drop to $71.673$ dB and modes $(3,4)$ drop to $74.947$ dB. This also validates the balanced score: the $(1,2,3)$ row has a smaller absolute operator distance than the baseline, but worse balanced relative error and worse rollout. Thus the useful weak-form view is now sharply identified as window $2$, test radius $3$; the remaining bias is not fixed by naive multi-scale row concatenation.
Finally, we tested whether the remaining gap is fundamental or merely a calibration-objective problem. A diagnostic forward-only coordinate search was run from the current sparse coefficient vector, using full-field training-tail rollout residuals over the eight training variants and no reverse-mode differentiation. This is not a sparse-only claim, because the calibration objective uses full-field training tails; it is a boundary diagnostic. The result is decisive: the sparse vector $(0.99980,0.72336,0.86033,0.58258,0.06371,0.005998,0.004022)$ gives $79.523$ dB on the disjoint V8 holdout, while the tracked full-field calibration runner \texttt{nonlinear\_shallow\_water\_theta\_calibration\_diagnostic.py} reaches $(1.00000,0.72076,0.85999,0.58002,0.06489,0.006001,0.004000)$, $104.088$ dB on V8, and $104.319$ dB on the disjoint V8b holdout. The accepted moves mainly correct $\beta$, $f$, $r$, $H$, and $\mu_h$ toward the simulator values. In contrast, rolling out the dense-training weak-form reference vector, although closer to the simulator coefficients, gives only $78.583$ dB. Therefore the $79.5$ dB frontier is not a support, station-count, or weak-window ceiling; it is a coefficient-calibration objective ceiling.
The follow-up calibration audit isolates what information is missing. If the calibration starts from the exact full training-tail state but observes only the $334$ sparse station values over the tail, the same forward-only coordinate search reaches the same vector and the same V8/V8b values, $104.088/104.319$ dB. Sparse station values therefore contain enough coefficient information once the hidden state at the calibration start is correct. The failure is the state anchor: using the sparse-assimilated tail-start state makes even a full-field tail loss collapse to $55.875$ dB, and optimizing either the sparse station tail, the assimilated pseudo-field tail, or the weak residual of the assimilated tail drives the coefficients in the wrong direction. Dense-training weak residual and dense-operator relative-$\ell_\infty$ objectives are also not sufficient by themselves: they reduce their training losses but reach only $78.997$ and $79.294$ dB on V8.
We then tested sparse or partial state-anchor repairs. A low-mode forward-sensitivity correction fitted from station history is too weak: mode-$2$, mode-$3$, and mode-$4$ anchors improve full-tail PSNR by only $0.00029$, $0.00043$, and $0.00084$ dB, respectively, and the calibrated mode-$2$ row still collapses to $63.281$ dB. A stronger but more privileged dynamics anchor, obtained by propagating the known full initial training state to the calibration split with the sparse OSNR operator, gives a much better training-tail state ($75.778$ dB tail PSNR and station loss $5.98\times10^{-7}$). Unconstrained station-tail calibration from this anchor still overfits, dropping to $73.021$ dB after $39$ accepted moves, but a one-accepted-move trust-region update raises only the damping coefficient, $r:0.0637088\mapsto0.0643459$, and improves V8/V8b to $80.123/80.139$ dB. A two-move variant immediately accepts an overlarge Coriolis correction and falls to $76.831$ dB. Thus the first legitimate direction beyond the $79.5$ dB row is not blind coordinate search; it is trust-region, state-aware station calibration.
The next selector audit made that trust-region rule explicit in the tracked diagnostic runner. The new one-step selector enumerates all single-coordinate candidates at the configured step sizes, ranks them using training-side objectives only, and evaluates V8/V8b only after the selected move is fixed. Station-tail loss alone is a negative control: it selects a depth decrease, $H:0.9998017\mapsto0.9988019$, because that gives the largest station-tail improvement, but the held-out scores collapse to $74.987/74.999$ dB. Adding a dense training-tail weak-residual consistency gate changes the selected move. Among candidates that improve both the model-anchor station tail and the dense training weak residual, the selector chooses the small Coriolis correction $f:0.5825801\mapsto0.5808324$; only after that training-only selection do the disjoint diagnostics evaluate to \textbf{$82.194/82.221$ dB} on V8/V8b. This is a larger lift than the previous damping-only trust move, but it is still diagnostic rather than sparse-only, because the auxiliary gate uses dense training-tail weak rows. Multi-accept variants do not improve the result: accepting a subsequent viscosity move gives $82.050/82.079$ dB, and an additional dense-operator relative-$\ell_2$ gate still keeps the best state at the first accepted $f$ move. The useful conclusion is sharper: sparse station replay supplies candidate moves but cannot select them safely by itself; a training-side operator-consistency gate can reject destructive station overfits and select a real coefficient correction. The next publishable step is to replace the dense weak-residual gate with a station-observable or assimilated operator-consistency surrogate while preserving the one-move trust-region discipline.
That replacement attempt is now also informative. We added station-observable selector objectives based on sparse-assimilated weak residuals, station finite-difference RHS residuals, one-step station-increment replay, held-out station splits, and separate higher-mode gate reconstructions. None is a safe substitute for the dense weak gate. The sparse weak gate selects the same destructive $H$ move and gives $74.987/74.999$ dB; station-RHS and station-one-step primaries select $\beta:0.7233627\mapsto0.7161291$ and slightly reduce V8/V8b to $79.489/79.509$ dB; held-out station one-step replay selects $g:0.8603310\mapsto0.8517277$ and collapses to $55.525/55.525$ dB; higher-mode sparse weak gates at modes $5$ and $6$, with and without temporal smoothing, still admit the destructive $H$ move. A clean sparse weak-solve hyperparameter audit over shell weights $0,0.025,0.075,0.10$ and sensor ridges $10^{-5},10^{-4}$ also fails to move the frontier, staying at $79.520$--$79.523$ dB. The only positive replacement so far is structural rather than learned: restrict the selector to momentum coefficients $(f,r,\nu)$ and to small trust steps $(0.003,0.001)$. With the same model-anchor station primary, this selects the same $f:0.5825801\mapsto0.5808324$ move and reaches $82.194/82.221$ dB without dense weak rows; on two fresh eight-variant audit lists it moves $79.521/79.524$ dB to $82.191/82.194$ dB. This is a cleaner diagnostic than the dense-gated selector, but it is still not a sparse-only claim because the selector primary uses the model-propagated full initial training state. When the primary is changed to the fully sparse assimilated station tail, the same small-trust momentum selector chooses $\nu:0.0059976\mapsto0.0059796$ and drops to $78.351/78.360$ dB. The next real problem is therefore state anchoring: station observations contain the coefficient signal, but the current sparse-assimilated calibration-start state distorts the selector enough that station-local objectives prefer wrong coefficient directions.
The state-anchor follow-up gives the first clean sparse calibration lift. Instead of using the privileged full initial training state, we propagate only the sparse-assimilated initial frame to the calibration split with the learned sparse OSNR operator, then run the same one-move small-trust momentum selector on station-tail loss. This \emph{sparse-model anchor} is still a low-PSNR field in full space (mean start/tail PSNR $21.069/20.386$ dB over the training variants), but it is dynamically consistent with the learned operator and improves the station-tail selector geometry. With primary objective \texttt{sparse\_model\_station\_tail}, coordinate subset $(f,r,\nu)$, and steps $(0.003,0.001)$, the selector chooses $f:0.5825801\mapsto0.5808324$ and moves V8/V8b from $79.523/79.536$ dB to \textbf{$82.194/82.221$ dB}, without dense weak rows, without full-field calibration tails, and without the full-initial-state model anchor. The same selected move transfers on two fresh eight-variant audit lists, $79.521/79.524\mapsto82.191/82.194$ dB. The constraints are sharp: allowing the large $0.01$ step without validation oversteps to $f=0.5767543$ and falls to $76.540/76.547$ dB; allowing a second small momentum move falls to $80.905/80.925$ dB; allowing all coordinates even at small trust selects $\beta:0.7233627\mapsto0.7255328$ and falls to $79.213/79.248$ dB.
We then added a training-variant split to the sparse-model station objective. The selector can now use \texttt{sparse\_model\_station\_tail\_fit} as the primary loss and require improvement on \texttt{sparse\_model\_station\_tail\_val}. This removes the hand-coded step scale within the momentum subspace: with candidate steps $(0.01,0.003,0.001)$, the validation split rejects the destructive $0.01$ overstep and selects the same small $f:0.5825801\mapsto0.5808324$ move, preserving $82.194/82.221$ dB. A second split-validated refinement from that point, using only $f$ and smaller steps $(0.001,0.0003,0.0001,0.00003)$, accepts $f:0.5808324\mapsto0.5806582$ and improves V8/V8b to \textbf{$82.278/82.305$ dB}; a fresh disjoint two-list audit gives $82.274/82.278$ dB. A subsequent $(f,r,\nu)$ pass finds no eligible move. However, the split does \emph{not} learn the coordinate mask: all-coordinate split validation at small trust selects $H$ and collapses to $66.444/66.443$ dB, all-coordinate large-trust split validation selects $\beta$ and gives $78.174/78.270$ dB, and adding a sparse weak-residual gate selects $g$ and collapses to $66.058/66.061$ dB. The publishable statement is therefore narrow but stronger than before: a sparse station-derived dynamic anchor, a physics-motivated momentum coordinate mask, and a training-split trust rule give a reproducible $+2.75$ dB held-out lift over the $334$-station exact-support frontier. The next missing piece is a station-observable coordinate-confidence rule, likely based on operator-block sensitivity or adjoint/Fisher geometry, that rejects the compensatory $H,\beta,g$ moves without using dense labels or held-out futures.
\subsection{Sparse governing-equation discovery for nonlinear shallow water}
The preceding experiment assumes that the correct seven-column operator family is known. The more ambitious physics-learning problem is to discover the governing equation itself from a larger nonlinear library. We implemented \texttt{apps\_industrial\_breakthrough/nonlinear\_shallow\_water\_library\_discovery.py}, which expands the residual library to $24$ candidate columns: the $13$ true equation-specific terms for height and both velocity channels, plus $11$ decoys including raw fields, quadratic field products, and misplaced height-gradient terms. OSNR applies a sequential thresholded least-squares solve on only $0.25\%$ of the training residual rows. The discovered support is then collapsed back into the shared physical vector $(H,\beta,g,f,r,\nu,\mu_h)$ and used for held-out future forecasting.
\begin{table}[h]
\centering
\small
\begin{tabular}{lcccccc}
\toprule
Setting & Support $(TP,FP,FN)$ & Future PSNR & Discovery time & Adam support & Adam PSNR & Adam time \\
\midrule
Clean, STLS threshold $5\cdot10^{-4}$ & $(13,0,0)$ & $72.7835$ dB & $1.0432$ ms & $(10,11,3)$ & $43.8385$ dB & $1587.90$ ms \\
$0.1\%$ noise, band-$18$ & $(13,0,0)$ & $69.2189$ dB & $1.4878$ ms & $(10,10,3)$ & $33.8352$ dB & $1632.81$ ms \\
$0.2\%$ noise, band-$16$ & $(12,0,1)$ & $64.5354$ dB & $1.6853$ ms & $(12,0,1)$ & $66.7134$ dB & $1634.47$ ms \\
Oracle true-library LS, clean & n/a & $72.6779$ dB & $0.2817$ ms & n/a & n/a & n/a \\
\bottomrule
\end{tabular}
\caption{Sparse nonlinear shallow-water governing-equation discovery from a $24$-term overcomplete library. OSNR recovers the exact clean PDE support and remains robust at $0.1\%$ observation noise. The Adam library baseline uses the same anchor rows and $20{,}000$ optimization steps with an $\ell_1$ penalty.}
\label{tab:nonlinear-shallow-water-library-discovery}
\end{table}
\begin{figure}[h]
\centering
\includegraphics[width=0.92\linewidth]{../apps_industrial_breakthrough/nonlinear_shallow_water_library_discovery_outputs_thr0005_adam20k/nonlinear_shallow_water_library_discovery.png}
\caption{Sparse governing-equation discovery for the coupled nonlinear shallow-water core. The discovered PDE recovers the clean future at $72.7835$ dB after selecting all true terms and no decoys from the overcomplete library.}
\label{fig:nonlinear-shallow-water-library-discovery}
\end{figure}
This is the first result in the project that begins to look like a genuine PINN/SINDy-class breakthrough rather than only a fast solver. In the clean case, OSNR selects every true nonlinear PDE term and rejects every decoy, then slightly outperforms the oracle true-library forecast. The comparable Adam library run is not just slower; after $20{,}000$ gradient steps it keeps $11$ false-positive decoys, misses $3$ true terms, and loses almost $29$ dB of forecast quality. The wall-clock ratio for discovery is about $1522\times$ in favor of OSNR. At $0.1\%$ observation noise, the same sparse support is still recovered exactly and the speed ratio remains above $1000\times$. At $0.2\%$ noise, spectral denoising preserves zero false positives but one weak damping term drops below threshold; the optimized Adam library catches up in forecast quality only after paying the full $1.6$ s optimization cost. This identifies the next hard technical layer: noise-aware thresholding or group sparsity for weak physical terms, not larger neural networks.
\subsection{Canonical PDE discovery: Burgers and Kuramoto--Sivashinsky}
To reduce the risk that the shallow-water result is viewed as a repository-specific construction, we added \texttt{apps\_industrial\_breakthrough/canonical\_pde\_discovery\_benchmark.py}. It evaluates two standard equation-discovery controls: viscous Burgers,
\[
u_t=-u u_x+\nu u_{xx},
\]
and the chaotic Kuramoto--Sivashinsky equation,
\[
u_t=-u u_x-u_{xx}-u_{xxxx}.
\]
Both are discovered from the same $10$-term library
\[
\{u,u^2,u_x,u u_x,u^2u_x,u_{xx},u u_{xx},u^3,u_{xxx},u_{xxxx}\},
\]
using only $128$ sampled residual rows. The Kuramoto--Sivashinsky trajectory is generated with the standard ETDRK4 spectral integrator; the discovery stage is independent of that generator and sees only the sampled field values.
\begin{table}[h]
\centering
\small
\begin{tabular}{lccccc}
\toprule
Equation & OSNR support & OSNR PSNR & OSNR time & Adam support & Adam time \\
\midrule
Burgers, clean & $(2,0,0)$ & $43.1187$ dB & $0.3690$ ms & $(2,0,0)$ & $1265.03$ ms \\
Kuramoto--Sivashinsky, clean & $(3,0,0)$ & $12.2751$ dB & $0.0972$ ms & $(3,0,0)$ & $1244.63$ ms \\
\bottomrule
\end{tabular}
\caption{Canonical PDE discovery controls. Support is reported as $(TP,FP,FN)$ against the known governing equation. Adam uses the same sampled residual rows and $20{,}000$ $\ell_1$-regularized optimization steps.}
\label{tab:canonical-pde-discovery}
\end{table}
\begin{figure}[h]
\centering
\includegraphics[width=0.92\linewidth]{../apps_industrial_breakthrough/canonical_pde_discovery_outputs_clean_thr01/canonical_pde_discovery.png}
\caption{Canonical Burgers and Kuramoto--Sivashinsky discovery from a shared overcomplete library. Both OSNR and Adam recover the clean support, but OSNR does it via a millisecond-scale sparse solve rather than a long gradient-optimization loop.}
\label{fig:canonical-pde-discovery}
\end{figure}
The canonical control confirms that the sparse operator-discovery mechanism is not confined to the shallow-water generator. On Burgers, OSNR recovers exactly $\{u u_x,u_{xx}\}$ and is about $3428\times$ faster than the Adam library optimizer. On Kuramoto--Sivashinsky, OSNR recovers exactly $\{u u_x,u_{xx},u_{xxxx}\}$ and is about $12809\times$ faster. The chaotic KS forecast PSNR is naturally low over the held-out horizon because small coefficient and phase errors amplify quickly; for this control, support recovery and coefficient recovery are the meaningful scientific-discovery metrics. At $0.1\%$ direct observation noise, both OSNR and Adam pick decoys under simple pointwise derivative regression, which confirms that the next publishable robustness layer must be weak-form or group-sparse denoised discovery rather than more gradient steps.
\subsection{Weak-form canonical PDE discovery under observation noise}
The pointwise canonical experiment exposes the correct failure mode: differentiating noisy data directly creates spurious high-frequency library columns. We therefore implemented the weak-form variant in \texttt{apps\_industrial\_breakthrough/canonical\_pde\_weakform\_discovery.py}. Instead of regressing $u_t$ at individual grid points, OSNR integrates the PDE over temporal windows and projects the resulting balance onto low-frequency spatial Fourier test functions:
\[
u(t_b)-u(t_a)=\int_{t_a}^{t_b}\Theta(u(t))\,\xi\,dt.
\]
This is the operator-spline analogue of weak-form PDE discovery: the test functions absorb observation noise before sparse regression sees the library.
\begin{table}[h]
\centering
\small
\begin{tabular}{lcccccc}
\toprule
Equation/noise & OSNR support & OSNR PSNR & OSNR time & Adam support & Adam PSNR & Adam time \\
\midrule
Burgers, $0.1\%$ & $(2,0,0)$ & $49.8283$ dB & $0.4965$ ms & $(2,0,0)$ & $47.9135$ dB & $1560.30$ ms \\
Burgers, $0.5\%$ & $(2,0,0)$ & $49.3129$ dB & $0.2580$ ms & $(2,0,0)$ & $40.7179$ dB & $1537.55$ ms \\
Burgers, $1.0\%$ & $(2,0,0)$ & $48.8199$ dB & $0.2103$ ms & $(2,0,0)$ & $45.2210$ dB & $1539.72$ ms \\
KS, $0.1\%$ & $(3,0,0)$ & $12.6155$ dB & $0.2870$ ms & $(3,0,0)$ & $12.7014$ dB & $1549.60$ ms \\
KS, $0.5\%$ & $(3,0,0)$ & $12.5214$ dB & $0.3026$ ms & $(3,0,0)$ & $11.9874$ dB & $1560.78$ ms \\
KS, $1.0\%$ & $(3,0,0)$ & $12.6433$ dB & $0.3148$ ms & $(3,0,0)$ & $11.7123$ dB & $1529.31$ ms \\
\bottomrule
\end{tabular}
\caption{Weak-form canonical PDE discovery under observation noise. The support tuple is $(TP,FP,FN)$. Both methods use the same weak rows and library, while Adam uses $20{,}000$ $\ell_1$-regularized optimization steps.}
\label{tab:weak-canonical-pde-discovery}
\end{table}
\begin{figure}[h]
\centering
\includegraphics[width=0.92\linewidth]{../apps_industrial_breakthrough/canonical_pde_weakform_outputs_thr002/canonical_pde_weakform_discovery.png}
\caption{Weak-form Burgers and Kuramoto--Sivashinsky discovery under noisy observations. The weak operator rows recover the correct governing support through $1\%$ noise while avoiding the pointwise derivative decoys observed in the direct regression control.}
\label{fig:weak-canonical-pde-discovery}
\end{figure}
This is the strongest canonical scientific-ML result so far. The weak-form OSNR solver recovers the exact Burgers and KS support at every tested noise level up to $1\%$. On Burgers, it is also materially more accurate than Adam in forecast quality, improving the $0.5\%$ noise row by $8.5950$ dB. On KS, both methods recover the same support, but OSNR reaches the solution roughly $4{,}858\times$ to $5{,}400\times$ faster for the threshold-$0.002$ profile. This converts the earlier ``mostly speed'' canonical result into a robustness result: integral operator rows eliminate noisy derivative decoys while preserving millisecond-scale discovery.
\subsection{Weak-form two-dimensional Navier--Stokes vorticity discovery}
The next CFD-facing step is a genuinely two-dimensional incompressible flow operator rather than a scalar one-dimensional PDE. We implemented \texttt{apps\_industrial\_breakthrough/navier\_stokes\_weakform\_discovery.py}, which generates periodic vorticity trajectories and discovers the vorticity equation
\[
\omega_t = -u\omega_x-v\omega_y+\nu\Delta\omega,
\qquad
u=\psi_y,\quad v=-\psi_x,\quad -\Delta\psi=\omega.
\]
The discovery stage is not told the two-term equation. It sees a $12$-term library containing the advective term, the Laplacian, and ten decoys built from raw vorticity, velocity, first derivatives, and nonlinear products. As in the canonical weak-form experiment, OSNR integrates over time windows and projects the balance onto low-frequency two-dimensional Fourier test functions,
\[
\omega(t_b)-\omega(t_a)=\int_{t_a}^{t_b}\Theta(\omega(t),u(t),v(t))\,dt,
\]
then applies a scaled sequential thresholded solve. The Adam control optimizes the same weak rows for $20{,}000$ $\ell_1$-regularized steps. A first high-resolution attempt at $160^2$ with the coarse timestep became numerically unstable, so the retained scaled run tightens the timestep and reference substepping rather than hiding the CFL boundary.
\begin{table}[h]
\centering
\small
\begin{tabular}{lcccccc}
\toprule
Grid/noise & OSNR support & OSNR coefficients $(c_{\rm adv},\nu)$ & OSNR PSNR & OSNR time & Adam support & Adam time \\
\midrule
$96^2$, clean & $(2,0,0)$ & $(0.9999865,\;0.0015000)$ & $135.3978$ dB & $0.6477$ ms & $(2,11,0)$ & $1958.15$ ms \\
$96^2$, $0.1\%$ & $(2,0,0)$ & $(0.9999592,\;0.0014999)$ & $121.0461$ dB & $0.5487$ ms & $(2,11,0)$ & $2013.22$ ms \\
$96^2$, $0.5\%$ & $(2,0,0)$ & $(1.0007806,\;0.0015010)$ & $99.6485$ dB & $0.3822$ ms & $(2,8,0)$ & $1966.74$ ms \\
$128^2$, clean & $(2,0,0)$ & $(0.9999944,\;0.0015000)$ & $146.9002$ dB & $0.7231$ ms & $(2,11,0)$ & $2322.75$ ms \\
$128^2$, $0.1\%$ & $(2,0,0)$ & $(1.0000714,\;0.0015000)$ & $126.0550$ dB & $0.4781$ ms & $(2,11,0)$ & $2324.12$ ms \\
\bottomrule
\end{tabular}
\caption{Weak-form two-dimensional Navier--Stokes vorticity discovery. The support tuple is $(TP,FP,FN)$ relative to the two true terms $\{-u\omega_x-v\omega_y,\Delta\omega\}$. Adam is the same weak-library regression optimized by backpropagation, not a full neural Navier--Stokes model.}
\label{tab:navier-stokes-weakform-discovery}
\end{table}
\begin{figure}[h]
\centering
\includegraphics[width=0.92\linewidth]{../apps_industrial_breakthrough/navier_stokes_weakform_outputs_128_stable/navier_stokes_weakform_discovery.png}
\caption{Scaled $128^2$ Navier--Stokes weak-form discovery. OSNR identifies the exact advection--diffusion vorticity operator and forecasts the held-out future from the recovered coefficients, while the gradient-optimized sparse regression admits many decoys.}
\label{fig:navier-stokes-weakform-discovery}
\end{figure}
This result is the first high-impact two-dimensional CFD discovery benchmark in the repository. It is still a controlled periodic vorticity system, not a direct DeepMind weather-model comparison. The important claim is narrower and stronger: for a known candidate library and noisy observations, weak-form OSNR recovers the exact incompressible Navier--Stokes vorticity support and coefficients at $96^2$ and $128^2$ resolution, remains stable through $0.5\%$ noise in the $96^2$ run, and solves the sparse operator identification in less than a millisecond. The Adam control uses the same rows and library but remains thousands of times slower and selects many decoy terms. The next step is therefore to move from periodic vorticity discovery to partial-observation assimilation and forced/stochastic Navier--Stokes, where the sparse innovation machinery can be tested on genuinely unknown forcing rather than only coefficient recovery.
\subsection{Forced Navier--Stokes sparse innovation assimilation}
We then tested the more realistic assimilation problem in \texttt{apps\_industrial\_breakthrough/navier\_stokes\_sparse\_forcing\_assimilation.py}. The governing operator is assumed known, but the trajectory is driven by hidden sparse spatiotemporal forcing:
\[
\omega_t = -u\omega_x-v\omega_y+\nu\Delta\omega+f(t,x,y),
\qquad
f(t,x,y)=\sum_{r=1}^R a_r\,\varphi_t(t-\tau_r)\varphi_x(x-x_r,y-y_r).
\]
This is closer to weather and flow data assimilation than coefficient discovery: the unknowns are localized forcing events, not just scalar PDE coefficients. OSNR first denoises the observed trajectory spectrally, applies the known Navier--Stokes operator to form the innovation residual
\[
\widehat f(t+\tfrac12)=\frac{\omega(t+\Delta t)-\omega(t)}{\Delta t}
-\left[-u\omega_x-v\omega_y+\nu\Delta\omega\right]_{t+1/2},
\]
then runs a weak three-dimensional matched atom detector: the residual is convolved with the separable spatial--temporal Gaussian test function associated with the forcing atom before non-maximum suppression. A final small least-squares amplitude debiasing step fits the active atoms to the residual. The comparison baselines are an unforced rollout and a smooth low-pass residual forcing field.
\begin{table}[h]
\centering
\small
\begin{tabular}{lcccccc}
\toprule
Run & Events & Noise & Event error & Recall@3 & Traj. PSNR & Assimilation time \\
\midrule
$96^2\times121$ & $48$ & $0.2\%$ & $0.7416$ & $97.92\%$ & $81.9812$ dB & $225.76$ ms \\
$128^2\times161$ & $96$ & $0.2\%$ & $1.2556$ & $94.79\%$ & $83.0769$ dB & $851.00$ ms \\
$128^2\times161$ & $96$ & $0.5\%$ & $2.6500$ & $87.50\%$ & $79.5815$ dB & $883.36$ ms \\
$128^2\times161$ & $96$ & $1.0\%$ & $5.2118$ & $73.96\%$ & $74.5790$ dB & $905.40$ ms \\
$192^2\times201$ & $160$ & $1.0\%$ & $5.0134$ & $78.75\%$ & $79.1944$ dB & $4086.68$ ms \\
$192^2\times201$ & $240$ & $1.0\%$ & $3.8576$ & $82.50\%$ & $77.9266$ dB & $6200.45$ ms \\
$192^2\times201$ & $240$ & $2.0\%$ & $4.6710$ & $77.50\%$ & $70.9027$ dB & $6158.36$ ms \\
$192^2\times201$ & $240$ & $3.0\%$ & $4.8846$ & $77.08\%$ & $67.3531$ dB & $6367.19$ ms \\
$192^2\times201$ & $240$ & $5.0\%$ & $8.1801$ & $60.83\%$ & $61.7619$ dB & $6334.94$ ms \\
\bottomrule
\end{tabular}
\caption{Forced two-dimensional Navier--Stokes sparse innovation assimilation. Event error is measured in joint $(t,x,y)$ grid units against the injected forcing centers. The weak three-dimensional atom detector keeps the sparse residual useful through $5.0\%$ observation noise and improves trajectory reconstruction over both unforced and low-pass residual rollouts.}
\label{tab:navier-stokes-sparse-forcing}
\end{table}
\begin{figure}[h]
\centering
\includegraphics[width=0.92\linewidth]{../apps_industrial_breakthrough/navier_stokes_sparse_forcing_outputs_192_events240_noise030_smooth/navier_stokes_sparse_forcing.png}
\caption{Forced Navier--Stokes sparse innovation assimilation on the $192^2\times201$ stress run with $240$ hidden events at $3.0\%$ observation noise. The recovered weak-form sparse atoms preserve the assimilated trajectory substantially better than an unforced model and better than a smooth low-pass residual forcing field.}
\label{fig:navier-stokes-sparse-forcing}
\end{figure}
The $128^2$ dense-event case recovers $94.79\%$ of the hidden forcing events within three grid units at $0.2\%$ noise and reaches $83.0769$ dB trajectory PSNR, compared with $77.8847$ dB for low-pass forcing and $71.1900$ dB for the unforced operator. After replacing the point detector with the weak three-dimensional matched atom score, the same dense case remains useful at $0.5\%$ and $1.0\%$ noise: at $1.0\%$ noise it recovers $73.96\%$ of events within three grid units and reaches $74.5790$ dB, compared with $69.8860$ dB for the low-pass residual and $67.8846$ dB for the unforced operator. The larger $192^2\times201$ stress run with $240$ hidden events at $1.0\%$ noise recovers $82.50\%$ of events within three grid units and reaches $77.9266$ dB, compared with $72.5998$ dB for the low-pass residual and $70.5923$ dB for the unforced model. With scale-adjusted weak atom smoothing, the same $240$-event stress run remains ahead of the low-pass residual at $2.0\%$, $3.0\%$, and $5.0\%$ observation noise. At $3.0\%$ noise, it recovers $77.08\%$ of forcing events and improves trajectory quality by $3.8774$ dB over low-pass; at $5.0\%$ noise, it still recovers $60.83\%$ of events and keeps a $2.0572$ dB trajectory advantage. This is a meaningful step beyond coefficient discovery: sparse OSNR innovations can assimilate unknown localized forcing in a nonlinear two-dimensional flow under noisy observations. The remaining bottleneck is now external benchmark standardization and heavy-overlap amplitude calibration, not basic sparse forcing recovery.
To compare against a trained coordinate-field alternative, we added \texttt{apps\_industrial\_breakthrough/navier\_stokes\_neural\_forcing\_baseline.py}. The neural baseline receives the same innovation residual as OSNR and fits a Fourier-feature MLP $g_\theta(t,x,y)$ with AdamW, using a sample distribution biased toward high residual magnitude so that sparse events are not hidden by uniform sampling. The learned forcing is then rolled through the same Navier--Stokes solver. On the $128^2\times161$ case with $96$ hidden events and $3.0\%$ observation noise, a $5$-layer, $128$-hidden-unit Fourier MLP trained for $5{,}000$ steps reaches only $42.9701$ dB trajectory PSNR. OSNR reaches $64.6633$ dB from the same residual, while the low-pass residual baseline reaches $60.8749$ dB. The recovery step takes $0.9165$ s for OSNR versus $52.1354$ s for the neural training loop, a measured $56.9\times$ speed advantage before rollout. The conclusion is not that this small MLP is a definitive neural SOTA baseline; rather, it isolates the key mechanism: dense coordinate-field training smooths or misallocates sparse innovations, while the operator-sparse residual directly preserves the hidden forcing events.
\begin{figure}[h]
\centering
\includegraphics[width=0.92\linewidth]{../apps_industrial_breakthrough/navier_stokes_neural_forcing_outputs_128_noise030_steps5000/navier_stokes_neural_forcing.png}
\caption{Forced Navier--Stokes sparse forcing recovery against a trained Fourier-feature neural residual field. The neural field is trained directly on the same residual observations but remains much less accurate in the downstream flow rollout.}
\label{fig:navier-stokes-neural-forcing-baseline}
\end{figure}
Finally, we converted the high-noise forced-flow result into a replicated stress suite in \texttt{apps\_industrial\_breakthrough/navier\_stokes\_high\_noise\_suite.py}. The protocol repeats the $192^2\times201$, $240$-event experiment across three independent random seeds and reports aggregate gains over low-pass residual assimilation. At $3.0\%$ observation noise, OSNR reaches a mean trajectory PSNR of $67.0929$ dB versus $63.4397$ dB for low-pass, a mean gain of $3.6532$ dB with a worst-seed gain of $3.3038$ dB. At $5.0\%$ observation noise with the high-noise weak-atom smoothing profile, OSNR reaches $61.6653$ dB versus $59.6858$ dB for low-pass, a mean gain of $1.9795$ dB with a worst-seed gain of $1.7582$ dB. This establishes that the forced-flow advantage is not a single-seed artifact.
\begin{table}[h]
\centering
\small
\begin{tabular}{lccccc}
\toprule
Noise & Seeds & Mean recall & Mean OSNR & Mean low-pass & Worst gain \\
\midrule
$3.0\%$ & $3$ & $73.33\%$ & $67.0929$ dB & $63.4397$ dB & $3.3038$ dB \\
$5.0\%$ & $3$ & $59.44\%$ & $61.6653$ dB & $59.6858$ dB & $1.7582$ dB \\
\bottomrule
\end{tabular}
\caption{Replicated high-noise forced Navier--Stokes sparse innovation assimilation at $192^2\times201$ with $240$ hidden forcing events. The reported gain is OSNR trajectory PSNR minus low-pass residual trajectory PSNR.}
\label{tab:navier-stokes-high-noise-suite}
\end{table}
\begin{figure}[h]
\centering
\includegraphics[width=0.78\linewidth]{../apps_industrial_breakthrough/navier_stokes_high_noise_suite_outputs_main/navier_stokes_high_noise_suite.png}
\caption{Replicated high-noise forced-flow suite. OSNR remains ahead of smooth low-pass residual assimilation across all tested seeds at $3.0\%$ noise and after retuning the weak atom scale at $5.0\%$ noise.}
\label{fig:navier-stokes-high-noise-suite}
\end{figure}
To probe whether the effect survives across a broader operating envelope, we added the ``destroyer'' matrix \texttt{apps\_industrial\_breakthrough/navier\_stokes\_destroyer\_protocol.py}. It evaluates $24$ forced-flow cases across four grid families ($96^2$, $128^2$, $160^2$, $192^2$), event counts from $48$ to $240$, noise levels from $1.0\%$ to $5.0\%$, and two random seeds per configuration. OSNR wins $22/24$ cases against the low-pass residual baseline, with mean trajectory gain $+4.2384$ dB and mean event recall $74.24\%$. At $3.0\%$ noise, OSNR wins all $12/12$ cases with mean gain $+2.9985$ dB. At larger grids ($128^2$, $160^2$, and $192^2$), OSNR wins every tested case; the only two losses occur in the smallest $96^2$ grid with the densest $96$-event, $5.0\%$ noise setting, where event overlap exceeds the available spatial resolution.
\begin{table}[h]
\centering
\small
\begin{tabular}{lcccc}
\toprule
Slice & Cases & Wins & Mean gain & Minimum gain \\
\midrule
All destroyer cases & $24$ & $22$ & $+4.2384$ dB & $-0.7136$ dB \\
$1.0\%$ noise & $4$ & $4$ & $+13.6026$ dB & $+12.7441$ dB \\
$3.0\%$ noise & $12$ & $12$ & $+2.9985$ dB & $+0.7201$ dB \\
$5.0\%$ noise & $8$ & $6$ & $+1.4161$ dB & $-0.7136$ dB \\
$128^2$--$192^2$ grids & $16$ & $16$ & $+4.3012$ dB & $+1.7768$ dB \\
\bottomrule
\end{tabular}
\caption{Destroyer forced Navier--Stokes sparse assimilation matrix. Gains are OSNR trajectory PSNR minus low-pass residual trajectory PSNR.}
\label{tab:navier-stokes-destroyer-protocol}
\end{table}
\begin{figure}[h]
\centering
\includegraphics[width=0.92\linewidth]{../apps_industrial_breakthrough/navier_stokes_destroyer_protocol_outputs_main/navier_stokes_destroyer_protocol.png}
\caption{Destroyer matrix summary. Bars show mean OSNR gain over low-pass residual forcing for each grid/event/noise configuration; labels show mean event recall.}
\label{fig:navier-stokes-destroyer-protocol}
\end{figure}
\subsection{External PDEBench/FNO weather and fluid assimilation audits}
\paragraph{Test~28 stabilizer boundary.}
To move beyond internally generated forced-flow fields, we audited the hosted prediction tensors from the external \texttt{pdebench-fno-audit/fno-predictions} artifact. The target case is Test~28, a $512^2$ incompressible Navier--Stokes vorticity--Poisson benchmark. Each chunk stores FNO vorticity predictions $\hat\omega$, target vorticity $\omega$, and the published velocity-space nRMSE obtained by solving the Dirichlet Poisson problem
\[
-\Delta\psi=\omega,\qquad \mathbf{v}=(\partial_y\psi,-\partial_x\psi),
\]
then comparing velocity fields after the first ten input frames. We implemented the same DST-I Poisson recovery in \texttt{apps\_industrial\_breakthrough/pdebench\_fno\_test28\_stabilizer.py} and verified that the recomputed FNO velocity nRMSE on chunk~00 matches the stored metric to within expected numerical drift.
The OSNR diagnostic applies a deterministic spectral-viscosity operator to the FNO vorticity field,
\[
\hat\omega_{\mathrm{osnr}} = \mathcal{F}^{-1}\!\left[M_K(k_x,k_y)\mathcal{F}\hat\omega\right],
\]
where $M_K$ is a compact rectangular frequency support. This is a blind post-processing stabilizer: it does not use held-out targets. We also tested a diagonal spectral transfer calibrated from five samples and applied to the remaining five samples; this is reported only as an assimilation diagnostic because it uses calibration targets.
\begin{table}[h]
\centering
\small
\begin{tabular}{lcc}
\toprule
Profile & Vorticity nRMSE & Velocity nRMSE \\
\midrule
FNO artifact, chunk~00 & $1.390145$ & $0.243828$ \\
OSNR spectral viscosity, $K=64$ & $0.772164$ & $0.243647$ \\
OSNR best vorticity filter, $K=16$ & $0.542349$ & $0.244357$ \\
Diagonal spectral calibration, held-out & $0.497082$ & $0.266637$ \\
\midrule
Oracle replace low modes, $K=8$ & -- & $0.057275$ \\
Oracle replace low modes, $K=32$ & -- & $0.023451$ \\
Oracle replace high modes, $K=32$ & -- & $0.243831$ \\
\bottomrule
\end{tabular}
\caption{External PDEBench/FNO Test-28 stabilizer audit on chunk~00. The OSNR spectral operator strongly suppresses vorticity outliers, but the official velocity-space metric changes only marginally and aggressive vorticity filtering can hurt velocity. The calibrated diagonal transfer is not a blind forecast result and is included to expose the metric boundary.}
\label{tab:pdebench-fno-test28-stabilizer}
\end{table}
\begin{figure}[h]
\centering
\includegraphics[width=0.92\linewidth]{../apps_industrial_breakthrough/pdebench_fno_test28_outputs/test28_vorticity_panel.png}
\caption{External PDEBench Test-28 vorticity panel for a held-out chunk sample. Spectral OSNR filtering removes large high-frequency FNO vorticity spikes, but this does not automatically translate into a large improvement in the benchmark velocity-space nRMSE.}
\label{fig:pdebench-fno-test28-stabilizer}
\end{figure}
This external audit is a useful boundary result. It confirms that operator-spline spectral structure can repair raw vorticity instability in a real hosted FNO artifact, but it also prevents an overclaim: the official velocity metric is dominated by low-frequency phase and Poisson-integrated velocity structure. The oracle rows make this precise. Replacing only the lowest spectral modes of the FNO prediction with the target reduces velocity nRMSE from $0.243828$ to $0.057275$ at $K=8$ and $0.023451$ at $K=32$, whereas replacing high modes while leaving the low modes unchanged barely moves the metric. A SOTA-facing improvement on this benchmark therefore requires a velocity-aware low-mode dynamics corrector or Poisson-adjoint training objective, with sparse OSNR machinery reserved for high-frequency vorticity stabilization.
We tested three follow-up low-mode correction families on held-out chunk~02 after calibrating on chunks~00--01. A diagonal vorticity-space transfer improved held-out vorticity nRMSE from $1.3716$ to $1.0479$ but worsened velocity nRMSE from $0.2303$ to $0.2491$. Direct velocity-space diagonal and mean-residual transfers also worsened the held-out metric, reaching best velocity nRMSE $0.2436$. Finally, a compact MPS-trained low-resolution velocity CNN fit the training loss but evaluated at $0.2488$ nRMSE on chunk~02. These negative results are informative: the low-mode error is not a stationary spectral bias and not solved by a small framewise image corrector. It is a sample-specific dynamical phase error. The next external benchmark attempt must either learn a genuine temporal low-mode evolution operator from the input history or select a benchmark where sparse/operator innovations, not phase drift, dominate the published metric.
\paragraph{Sparse-station assimilation.}
We therefore reframed the external FNO artifacts as sparse-station data assimilation problems, which is closer to operational weather and fluid monitoring. In \texttt{apps\_industrial\_breakthrough/pdebench\_weather\_sparse\_station\_assimilation.py}, the FNO forecast is treated as a neural dynamical prior. At each forecast step, a small number of station observations are used to solve either a closed-form per-channel affine correction or a low-rank DCT residual correction. On the three external Test~29 four-channel forecast configurations, this post-processing consistently improves the FNO forecast. The strongest case, \texttt{M01\_Eta01}, drops from mean nRMSE $0.00461$ to $0.000947$ with a rank-$8$ DCT residual fit from $512$ stations, a $79.5\%$ relative reduction. The \texttt{M10\_Eta01} case drops from $0.00551$ to $0.00159$ ($71.1\%$ reduction), while the harder \texttt{M10\_Eta001} case drops from $0.01204$ to $0.00908$ ($24.6\%$ reduction). We also tested the harder $512^2\times101$ vorticity artifact in \texttt{apps\_industrial\_breakthrough/pdebench\_vorticity\_sparse\_station\_assimilation.py}. Per-step affine station calibration reduces mean full-window vorticity nRMSE from $1.1802$ to $0.6872$ with $1024$ sparse stations, a $41.8\%$ relative reduction, but the velocity-space audit above shows that affine-only vorticity calibration is not the right final correction family.
\begin{table}[h]
\centering
\small
\begin{tabular}{lccc}
\toprule
External forecast artifact & FNO mean nRMSE & Sparse-station OSNR nRMSE & Relative gain \\
\midrule
Test~29 \texttt{M01\_Eta01} & $0.004611$ & $0.000947$ & $79.5\%$ \\
Test~29 \texttt{M10\_Eta01} & $0.005507$ & $0.001594$ & $71.1\%$ \\
Test~29 \texttt{M10\_Eta001} & $0.012038$ & $0.009076$ & $24.6\%$ \\
Test~28 vorticity chunk~00 & $1.180217$ & $0.687232$ & $41.8\%$ \\
\bottomrule
\end{tabular}
\caption{External PDEBench/FNO sparse-station assimilation. The neural forecast is kept as the dynamical prior, while operator/dictionary corrections are solved from sparse observations without backpropagation.}
\label{tab:pdebench-weather-station-assimilation}
\end{table}
The stronger Test~28 result comes from making the station correction operator-aware in the published metric. The runner \texttt{apps\_industrial\_breakthrough/pdebench\_vorticity\_dct\_station\_assimilation.py} fits, at each forecast step, an affine vorticity calibration followed by a rank-$32$ DCT residual from $1024$ contemporaneous vorticity stations. Unlike the affine-only station correction, this low-mode residual directly repairs the Poisson-integrated velocity structure. On all three available Test~28 chunks, the fixed profile improves both vorticity and the recomputed velocity metric:
\begin{table}[h]
\centering
\small
\begin{tabular}{lcccc}
\toprule
Chunk & FNO velocity & DCT-station velocity & FNO vorticity & DCT-station vorticity \\
\midrule
00 & $0.243826$ & $0.191702$ & $1.389743$ & $0.429227$ \\
01 & $0.247869$ & $0.188312$ & $1.421356$ & $0.408249$ \\
02 & $0.230292$ & $0.189971$ & $1.371190$ & $0.435880$ \\
\midrule
Mean & $0.240662$ & $0.189995$ & $1.394096$ & $0.424452$ \\
\bottomrule
\end{tabular}
\caption{External PDEBench/FNO Test~28 DCT station assimilation. Metrics are computed after the first ten input frames. The correction uses $1024$ lattice stations, affine calibration, and a rank-$32$ DCT residual solve at each forecast step. Mean velocity nRMSE drops by $21.0\%$ and mean vorticity nRMSE drops by $69.5\%$ across the three local chunks.}
\label{tab:pdebench-vorticity-dct-stations}
\end{table}
This per-frame result was the first Test~28 improvement in the project that moved the official velocity-space metric rather than only suppressing raw vorticity outliers. The next run added temporal structure to the station adapter. In \texttt{apps\_industrial\_breakthrough/pdebench\_vorticity\_temporal\_osnr\_station\_rescue.py}, chunk~00 selects the profile and chunks~01--02 are held out. For each trajectory, the frozen FNO rollout remains the neural dynamical prior, sparse contemporary vorticity stations are observed at each future frame, and one separable spatiotemporal OSNR residual is fitted over the whole forecast window,
\[
\hat\omega(y,x,t)=a_t\hat\omega_{\mathrm{FNO}}(y,x,t)+b_t
+\sum_{p,q,r} c_{pqr}\,\phi_p(y)\phi_q(x)\tau_r(t),
\]
where $\phi$ are spatial DCT atoms and $\tau$ is a temporal DCT basis over the post-input frames. The metric is again the Poisson-recovered velocity nRMSE plus raw vorticity nRMSE.
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{lccccc}
\toprule
Method on held-out chunks~01--02 & Stations/frame & Velocity nRMSE & Gain & Vorticity nRMSE & Gain \\
\midrule
Frozen FNO prior & $0$ & $0.238631$ & -- & $1.396631$ & -- \\
Previous per-frame DCT, rank $32$ & $1024$ & $0.189142$ & $20.7\%$ & $0.422065$ & $69.8\%$ \\
Temporal OSNR, rank $32\times8$ & $2048$ & $0.112993$ & $52.6\%$ & $0.349838$ & $75.0\%$ \\
Temporal OSNR, rank $32\times12$ & $4096$ & $0.106656$ & $55.3\%$ & $0.328363$ & $76.5\%$ \\
Temporal OSNR, rank $48\times12$ & $4096$ & $0.103622$ & $56.6\%$ & $0.327553$ & $76.5\%$ \\
\bottomrule
\end{tabular}
\caption{External PDEBench/FNO Test~28 temporal OSNR station rescue. The final row uses $4096/262144=1.5625\%$ of grid sites per frame and fits one spatiotemporal residual over each forecast trajectory. Hyperparameters are selected on chunk~00 and reported on held-out chunks~01--02.}
\label{tab:pdebench-vorticity-temporal-stations}
\end{table}
This became the strongest closed-form Test~28 external FNO result in the workspace. Across all three local chunks, the rank-$48\times12$ temporal adapter reduces mean velocity nRMSE from $0.240406$ to $0.103368$ and mean vorticity nRMSE from $1.394469$ to $0.330206$. It is still an assimilation result, not a blind forecast: the method uses contemporary sparse measurements. That distinction is important, but it is also exactly the operational setting where station, buoy, radar, and satellite observations are available and a neural forecast acts as the dynamical prior. The broader mechanism is now clearer than in the first Test~28 audit: OSNR can act as a closed-form, low-rank test-time correction layer on top of a frozen neural PDE forecaster, and the same plug-in idea already transferred from external Darcy sparse assimilation to time-dependent vorticity forecasts.
We then ran the harder SOTA-facing comparison: a trained sparse neural assimilator under the same station protocol. The runner \texttt{apps\_industrial\_breakthrough/pdebench\_vorticity\_neural\_sparse\_assimilation\_baseline.py} trains on chunks~00--01 and reports held-out chunk~02. Its inputs are the frozen FNO vorticity frame, the same-station affine calibration, sparse target and residual maps, a station mask, coordinates, and forecast time. A compact $770{,}241$-parameter U-Net predicts a $128^2$ residual that is upsampled to the full $512^2$ grid before both vorticity and Poisson velocity metrics are computed. This is not a no-backprop OSNR result; it is the competent trained sparse neural baseline that the closed-form adapter must be compared against.
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{lcccc}
\toprule
Method on held-out chunk~02 & Stations/frame & Velocity nRMSE & Gain vs FNO & Vorticity nRMSE \\
\midrule
Frozen FNO prior & $0$ & $0.230295$ & -- & $1.371558$ \\
FNO + temporal OSNR & $256$ & $0.198508$ & $13.8\%$ & $0.551600$ \\
Sparse neural assimilator & $256$ & $0.073335$ & $68.2\%$ & $0.191289$ \\
FNO + temporal OSNR & $512$ & $0.147680$ & $35.9\%$ & $0.461464$ \\
Sparse neural assimilator & $512$ & $0.031716$ & $86.2\%$ & $0.154792$ \\
FNO + temporal OSNR & $1024$ & $0.134717$ & $41.5\%$ & $0.416209$ \\
Sparse neural assimilator & $1024$ & $0.019973$ & $91.3\%$ & $0.139996$ \\
Sparse neural assimilator, repeat seed & $1024$ & $0.021120$ & $90.8\%$ & $0.140282$ \\
FNO + temporal OSNR & $4096$ & $0.110239$ & $52.1\%$ & $0.338437$ \\
Sparse neural assimilator & $4096$ & $\mathbf{0.015279}$ & $\mathbf{93.4\%}$ & $0.155159$ \\
\bottomrule
\end{tabular}
\caption{PDEBench Test~28 trained sparse neural assimilation on held-out chunk~02. The $1024$-station row uses only $1024/262144=0.390625\%$ of grid sites per frame and is stable under one seed repeat. The neural model is trained with backpropagation and is included as the relevant SOTA-facing sparse-assimilation comparator.}
\label{tab:pdebench-vorticity-neural-stations}
\end{table}
This shifts the Test~28 frontier. The $1024$-station neural row reduces velocity error by about $91\%$ against the frozen FNO and by about $85\%$ against the matched closed-form FNO+OSNR row. The $4096$-station row reaches the best velocity value, $0.015279$, but the $1024$ row is the cleaner observation-efficiency result. The fixed FNO-tuned temporal OSNR adapter does not transfer unchanged onto the trained neural prior: at $1024$ stations it worsens the neural velocity row from $0.019973$ to $0.075423$, and at $4096$ from $0.015279$ to $0.048526$. This negative adapter result is useful. Once the neural model has learned the low-frequency station-conditioned correction, the next OSNR layer must be selected specifically for a neural prior, with identity/gating/high-ridge/low-rank candidates or an orthogonalized residual space. Reusing the FNO prior's adapter is not valid.
The active-station follow-up then exposed that the original closed-form gap was
mostly geometric. We extended the same runner with centered, interior,
space-filling, and rounded phase-lattice station policies while keeping the
train/test split, neural architecture, epochs, and Poisson velocity metric
fixed. The old lattice includes boundary-heavy samples; a centered or
interiorized lattice spends the same budget on Fourier-compatible interior
coverage. Table~\ref{tab:pdebench-vorticity-active-stations} shows the result
on held-out chunk~02.
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{lccccc}
\toprule
Station policy & Stations & FNO+OSNR velocity & FNO+OSNR vorticity & Neural velocity & Neural vorticity \\
\midrule
Old edge lattice & $512$ & $0.147680$ & $0.461464$ & $0.031716$ & $0.154792$ \\
Centered lattice & $512$ & $0.036161$ & $1.289455$ & $0.042066$ & $1.085325$ \\
Space filling & $512$ & $0.142717$ & $0.461791$ & $0.031003$ & $0.176815$ \\
\midrule
Old edge lattice & $1024$ & $0.134717$ & $0.416209$ & $0.019973$ & $0.139996$ \\
Centered lattice & $1024$ & $0.020786$ & $1.281390$ & $0.034668$ & $1.077794$ \\
Best rounded phase $(0.50,0.25)$ & $1024$ & $\mathbf{0.020563}$ & $1.285525$ & $0.038600$ & $1.078815$ \\
Space filling & $1024$ & $0.118799$ & $0.350398$ & $0.017219$ & $0.176978$ \\
\midrule
Old edge lattice & $2048$ & $0.119788$ & $0.364925$ & $0.020052$ & $0.147939$ \\
Space filling & $2048$ & $0.108844$ & $0.336009$ & $0.016828$ & $0.212521$ \\
Old edge lattice & $4096$ & $0.110239$ & $0.338437$ & $\mathbf{0.015279}$ & $0.155159$ \\
Space filling & $4096$ & $0.106388$ & $0.340152$ & $0.015478$ & $0.221071$ \\
\bottomrule
\end{tabular}
\caption{PDEBench Test~28 active station-geometry audit on held-out chunk~02. Regular interior station geometry nearly closes the velocity gap between closed-form temporal OSNR and the trained sparse neural assimilator at $1024$ stations, but does not repair raw vorticity. Space-filling stations improve the trained neural velocity curve at $1024$--$2048$ stations but worsen vorticity and do not beat the old $4096$-station neural velocity frontier.}
\label{tab:pdebench-vorticity-active-stations}
\end{table}
The best closed-form phase row reduces FNO velocity nRMSE from $0.230295$ to
$0.020563$ using only $1024/262144=0.390625\%$ contemporary station sites per
frame. This is an $84.7\%$ reduction relative to the old same-budget
edge-lattice OSNR row and is only about $3\%$ worse than the trained
$1024$-station neural row. It also beats one repeat seed of that neural row
($0.021120$). The caveat is just as important: the same regular/phase lattice
rows leave vorticity near $1.28$, so the win is a Poisson-velocity low-mode
correction, not full vorticity reconstruction. Space-filling gives the
complementary behavior: the neural model reaches new $1024$- and
$2048$-station velocity-efficiency rows, $0.017219$ and $0.016828$, but with
worse vorticity and no improvement over the old $4096$-station velocity
frontier. A local phase refinement around $(0.50,0.25)$ was locally saturated
and quantized by integer-grid rounding. A naive one-mask hybrid that concatenates
$50\%$ or $75\%$ centered/interior lattice stations with space-filling fill
points was decisively negative: the best $1024$-station hybrid neural velocity
was only $0.131524$, and closed-form hybrid velocity was worse than the frozen
FNO. Thus the next Test~28 step should be a vorticity-aware two-geometry or
two-head adapter: keep Fourier-compatible interior stations for the low-mode
velocity correction, add a separate high-frequency/vorticity residual mechanism,
and gate any OSNR residual against the trained neural prior rather than reusing
the FNO-prior adapter blindly.
The two-head follow-up made this decomposition explicit. Because the benchmark
velocity is recovered by a DST-I Poisson solve, the useful fusion basis is not a
generic DCT split but the same sine basis that diagonalizes the reported metric.
The runner \texttt{pdebench\_vorticity\_two\_head\_frequency\_adapter.py} keeps
the phase-lattice OSNR head for low modes, adds a separate space-filling OSNR
vorticity head for high modes, selects the DST cutoff and scalar weight on
chunk~01, and reports chunk~02. Table~\ref{tab:pdebench-vorticity-dst-two-head}
summarizes the resulting Pareto rows.
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{lccc}
\toprule
Method & Station observations/frame & Velocity nRMSE & Vorticity nRMSE \\
\midrule
FNO & $0$ & $0.230295$ & $1.371558$ \\
Phase low head & $1024$ & $0.020563$ & $1.285525$ \\
Space-filling high head & $1024$ & $0.118799$ & $0.350398$ \\
DST two-head & $1024+1024$ & $0.023612$ & $0.192818$ \\
DST two-head, larger high head & $1024+2048$ & $0.023712$ & $0.195319$ \\
Cached neural comparator & $1024$ & $0.017219$ & $0.176978$ \\
Corrected neural repeat & $1024$ & $0.017219$ & $0.176978$ \\
Neural + DST high-pass & $1024$ & $0.017060$ & $0.113794$ \\
Corrected neural repeat & $2048$ & $0.016828$ & $0.212521$ \\
Neural + DST high-pass & $2048$ & $0.016403$ & $0.097739$ \\
Corrected neural repeat & $4096$ & $0.015478$ & $0.221071$ \\
Neural + DST high-pass & $4096$ & $\mathbf{0.014927}$ & $\mathbf{0.097201}$ \\
Corrected neural repeat, seed $20260611$ & $1024$ & $0.020466$ & $0.177627$ \\
Neural + DST high-pass, seed $20260611$ & $1024$ & $0.020210$ & $0.095205$ \\
\bottomrule
\end{tabular}
\caption{PDEBench Test~28 DST two-head and neural-prior high-pass adapters on held-out chunk~02. Separate station counts indicate distinct low-mode and high-mode observation sets for the closed-form rows. Each neural high-pass row uses the same deterministic space-filling station set as its neural comparator, selects the DST split on chunk~01, and improves both reported metrics on chunk~02.}
\label{tab:pdebench-vorticity-dst-two-head}
\end{table}
The closed-form DST row is the first Test~28 adapter in the workspace that
substantially improves both sides of the earlier closed-form tradeoff: vorticity
drops from the standalone space-filling value $0.350398$ to $0.192818$, while
velocity remains close to the phase-lattice value ($0.023612$ versus
$0.020563$). Increasing the high-frequency head to $2048$ stations does not
improve the fused row, so the bottleneck is not simply high-head station count.
The neural-prior row is the clean frontier. The first repeat used the correct
station geometry but the wrong statistics-sampling seed; after matching the
baseline convention, the runner exactly reproduces the cached $1024$-station
neural row. A validation-selected DST high-pass OSNR layer with cutoff $80$ and
weight $0.2$ then improves held-out velocity from $0.017219$ to $0.017060$ and
raw vorticity from $0.176978$ to $0.113794$, using the same deterministic
space-filling station set. The follow-up ladder strengthens the claim. At
$2048$ stations, the vorticity-aware selector chooses cutoff $128$ and weight
$0.1$, improving the neural row from $0.016828$/$0.212521$ to
$0.016403$/$0.097739$. At $4096$ stations, cutoff $128$ and weight $0.05$
improve $0.015478$/$0.221071$ to $0.014927$/$0.097201$. A second $1024$-station
seed repeats the pattern, improving $0.020466$/$0.177627$ to
$0.020210$/$0.095205$. This reverses the earlier negative neural+OSNR result,
where an FNO-tuned residual damaged the trained neural prior: the useful adapter
is neural-prior-specific and orthogonalized into the DST high-frequency space.
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{lcccc}
\toprule
Gate row & Fit/gate stations & Choices & Neural vel/vort & Gated vel/vort \\
\midrule
$1024$, seed $20260610$ & $768/256$ & $112{:}10,\ 128{:}0$ & $0.017219/0.176978$ & $0.016922/0.096622$ \\
$2048$, seed $20260610$ & $1536/512$ & $112{:}10,\ 128{:}0$ & $0.016828/0.212521$ & $0.016403/0.099564$ \\
$4096$, seed $20260610$ & $3072/1024$ & $112{:}6,\ 128{:}4$ & $0.015478/0.221071$ & $0.014927/0.098693$ \\
$1024$, seed $20260611$ & $768/256$ & $112{:}10,\ 128{:}0$ & $0.020466/0.177627$ & $0.020210/0.097617$ \\
\bottomrule
\end{tabular}
\caption{No-leakage station-heldout gate for the Test~28 neural-prior DST adapter. The gate chooses identity, cutoff $112$, or cutoff $128$ from held-out station residuals only, then refits the chosen correction on all available stations for the reported field. Identity is never selected in these runs.}
\label{tab:pdebench-vorticity-station-gate}
\end{table}
The gate table removes the remaining hand-picked-selector weakness. The
decision uses only contemporary station values: $25\%$ of stations are withheld
from the gate fit, candidate residuals are scored on those stations, and the
selected candidate is then refit on the full station set. This standard
cross-validation pattern is operationally different from using full-field
validation metrics. The gate is conservative relative to the full-field oracle,
which would choose cutoff $128$ in all four runs; it often chooses cutoff $112$
instead. The price is a small vorticity gap versus the oracle, but the
no-leakage rows still improve both velocity and vorticity over the trained
neural prior at every tested budget and on the second seed. A fit-only ablation
that permanently discards the gate stations damages the velocity metric, so the
deployable protocol is station-heldout selection followed by all-station refit.
The more aggressive follow-up removes the assimilation advantage entirely. The
blind diffusion-refiner runner sees only the first ten true vorticity frames of
each held-out Test~28 trajectory at inference time. It never reads the FNO
prediction, never observes future stations, and uses the cached FNO tensor only
as an evaluation comparator. The model treats forecasting as an iterative
refinement-time PDE,
\[
u_{k+1}=P\!\left(u_k+\eta F_\theta(u_k,\mathrm{history},\tau,k)\right),
\]
where $P$ is either the identity or a DST spectral-viscosity projection. After
the first run showed a clean failure mode--excellent vorticity but weaker
Poisson velocity--we added a light Poisson-weighted DST coefficient loss and
then a differentiable low-mode Poisson-velocity loss.
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{lcc}
\toprule
Blind-from-history row on chunk~02 & Velocity nRMSE & Vorticity nRMSE \\
\midrule
FNO comparator & $0.230295$ & $1.371558$ \\
Persistence from frame 9 & $0.693986$ & $1.175046$ \\
Linear extrapolation from frames 8/9 & $1.497893$ & $4.620379$ \\
Pure learned refiner, $r128/e10$ & $0.303547$ & $0.414996$ \\
DST projected refiner, $r128/e10$ & $0.297367$ & $0.415566$ \\
DST + Poisson loss, seed $20260612$ & $0.261985$ & $0.405375$ \\
DST + Poisson loss, seed $20260613$ & $0.261936$ & $0.400889$ \\
DST + Poisson loss, seed $20260614$ & $0.252706$ & $0.396079$ \\
DST + Poisson loss, seed $20260615$ & $0.281126$ & $0.409406$ \\
DST + Poisson + velocity loss, seed $20260614$ & $0.246451$ & $0.396002$ \\
DST + Poisson + velocity loss, seed $20260613$ & $0.265005$ & $0.405434$ \\
DST + Poisson + velocity loss, seed $20260612$ & $0.268111$ & $0.413803$ \\
Uniform ensemble, seeds $20260614/13/12$ & $0.233286$ & $0.374172$ \\
Uniform ensemble, seeds $20260614/13/12/11/15$ & $\mathbf{0.228610}$ & $\mathbf{0.366712}$ \\
Validation-locked subset, seeds $20260614/11/15$ & $0.233002$ & $0.370837$ \\
Validation-locked weighted, seeds $20260614/12/11/15$ & $0.231505$ & $0.370884$ \\
Validation-locked weighted, seeds $20260614/13/12/11/15$ & $0.231374$ & $0.371718$ \\
\bottomrule
\end{tabular}
\caption{Blind Test~28 from-initial-history refiner. The OSNR/DST rows do not
use FNO predictions or future observations at inference time. They are trained
from chunks~00/01 and evaluated on held-out chunk~02; the FNO row is a frozen
external comparator.}
\label{tab:pdebench-vorticity-blind-refiner}
\end{table}
The direct velocity objective tightened the single-seed frontier from
$0.252706$ to $0.246451$ while preserving the vorticity win. More importantly,
seed diversity exposed an ensemble effect rather than a single lucky run. A
three-seed uniform average nearly closes the FNO velocity gap, and the five-seed
uniform ensemble becomes the first blind from-initial-history Test~28 row in
this project to beat the hosted FNO comparator on both reported metrics:
velocity improves from $0.230295$ to $0.228610$ ($0.73\%$), while vorticity
drops from $1.371558$ to $0.366712$ ($73.3\%$). We then froze an explicit
validation-locked model-selection protocol in
\texttt{pdebench\_vorticity\_blind\_locked\_ensemble.py}: train candidate
seeds on chunk~00, select a uniform seed subset on validation chunk~01 with the
predeclared score velocity plus $0.02$ times vorticity, refit only the selected
seeds on chunks~00/01, and evaluate chunk~02 once. That protocol selects
seeds $20260614/11/15$ and reaches $0.233002$ velocity and $0.370837$
vorticity on held-out chunk~02. The locked row therefore confirms the blind
raw-vorticity result--$73.0\%$ lower vorticity than FNO--but it does not yet
confirm the exploratory velocity edge, missing FNO velocity by $1.18\%$. The
scientific signal is sharp but narrower than the frontier row: without FNO input
or future observations, OSNR/DST refinement has a defensible no-leakage
mechanism for collapsing raw vorticity, while the Poisson-velocity win still
requires a better locked selector or dynamics model.
We then tested whether the missing velocity margin was simply a selector issue.
The weighted locked runner,
\texttt{pdebench\_vorticity\_blind\_locked\_weighted\_ensemble.py}, builds
validation-only low-resolution prediction quadratics and searches a convex
$0.05$ simplex grid with the predeclared score mean velocity plus $0.02$ times
mean vorticity plus $0.25$ times velocity p90 plus $0.05$ times maximum
velocity. Allowing four or five active seeds selects weights
$(0.20,0.15,0.20,0.45)$ on seeds $20260614/12/11/15$ and reaches $0.231505$
velocity, $0.370884$ vorticity on chunk~02. Forcing all five seeds active
selects weights $(0.20,0.05,0.10,0.15,0.50)$ and reaches the best strict
locked velocity row so far: $0.231374$ velocity and $0.371718$ vorticity. This
narrows the locked velocity gap from $1.18\%$ to $0.47\%$ relative to FNO, but
still does not beat the FNO velocity comparator. The vorticity result remains
stable at roughly $72.9\%$ lower error. Thus the next velocity gain is unlikely
to come from selector-only sweeps; it likely requires a stronger cross-mode or
horizon-conditioned dynamics model.
The negative controls are also informative. A no-backprop per-mode DST ridge
dynamics runner reaches only $0.352961$ velocity and $0.555332$ vorticity at
$r128/K64$; the validation-selected cross-mode kernel dynamics follow-up is
worse still at $0.868662$ velocity and $0.887061$ vorticity; naive $256$-grid
scaling reaches only $0.320404$ velocity and $0.462010$ vorticity after
rescaling the Poisson loss; and two $20$-epoch schedules improve training loss
without improving held-out velocity ($0.270341$ and $0.267655$). Thus the
current blind lesson is specific: iterative learned refinement plus light
Poisson-weighted OSNR/DST structure and seed diversity are the useful branch,
while per-mode closed-form spectral extrapolation, naive cross-mode kernels,
naive resolution scaling, and longer single-seed training do not close the
single-seed velocity gap.
This is not a claim of beating end-to-end global weather systems such as GraphCast or GenCast. It is a more precise and immediately defensible claim: operator-spline station assimilation can dramatically improve external neural PDE forecasts when a small number of contemporary observations are available. That setting is scientifically meaningful because real forecasting systems already assimilate sparse stations, buoys, sondes, radar, and satellite products; the OSNR contribution is a closed-form, low-rank correction layer that can sit on top of a neural forecaster without retraining it.
We also tested a local compact-kernel variant in \texttt{apps\_industrial\_breakthrough/pdebench\_weather\_rbf\_station\_assimilation.py}. Here the station residual after affine calibration is interpolated through a Gaussian RBF kernel, which better matches spatially local forecast-error structure than a global DCT basis at low station counts. On \texttt{M01\_Eta01}, affine+RBF improves the best result further from $0.000947$ to $0.000819$, an $82.2\%$ reduction from the FNO baseline. On \texttt{M10\_Eta01}, RBF reaches $0.001640$ and already obtains a $52.9\%$ reduction with only $16$ stations, while the global DCT correction remains slightly better at the highest station budget. On \texttt{M10\_Eta001}, DCT remains the better choice. The conclusion is algorithmic rather than cosmetic: the assimilation layer should select its correction dictionary from the forecast-error geometry. Smooth global biases favor low-rank DCT; localized residual structure favors compact RBF/operator-spline kernels.
The follow-up hybrid runner \texttt{apps\_industrial\_breakthrough/pdebench\_weather\_hybrid\_station\_assimilation.py} combines the two correction families in one sequential closed-form layer: affine calibration, rank-$8$ DCT residual fitting, and station-centered Gaussian RBF residual interpolation. A low-budget targeted sweep shows that blindly mixing dictionaries can underperform the best single family because station equations are split across redundant atoms. At the higher $512$-station random budget, however, the hybrid layer improves all three external forecasts: \texttt{M01\_Eta01} drops to $0.000707$ ($84.7\%$ reduction), \texttt{M10\_Eta01} drops to $0.001438$ ($73.9\%$ reduction), and \texttt{M10\_Eta001} drops to $0.008584$ ($28.7\%$ reduction). This became the random-station reference for the adaptive placement tests. The design rule is now clearer: use a single compact dictionary at very sparse station counts, but switch to a hybrid global--local operator dictionary once the observation budget is high enough to identify both smooth bias and localized forecast residuals.
\paragraph{Adaptive Test~29 station placement.}
The next experiment, \texttt{apps\_industrial\_breakthrough/pdebench\_weather\_adaptive\_station\_assimilation.py}, tests whether station placement can reduce the observation budget. The correction dictionary is kept fixed and only the station policy changes. Forecast-gradient, DCT-leverage, and two-stage pilot-residual policies are not reliable: they oversample high-variation or high-error regions and leave the RBF/DCT normal equations poorly covered. A farthest-point space-filling policy is better because it improves interpolation coverage and dictionary conditioning. The first adaptive sweep nearly matched the $512$-random hybrid with $256$ stations. A targeted follow-up then retuned only the RBF length scale and showed that $256$ space-filling stations are enough to beat the $512$-random reference on all three external Test~29 files: \texttt{M01\_Eta01} reaches $0.000705$, \texttt{M10\_Eta01} reaches $0.001377$, and \texttt{M10\_Eta001} reaches $0.007649$. The same policy at $288$ stations improves all three again, and $384$ stations gives the strongest current Test~29 assimilation results: $0.000646$, $0.001267$, and $0.007274$, respectively. Thus the useful high-information criterion for this external artifact is not local forecast-error magnitude; it is conditioned spatial coverage for the global--local operator dictionary. The result is a concrete observation-efficiency gain: half as many stations now beat the former $512$-station hybrid reference, without retraining the FNO and without backpropagating through the assimilation layer.
For reproducibility, this weather result should be read as a frozen-neural-prior plus deterministic assimilation architecture, not as a newly trained neural network. The neural component is the external FNO prediction tensor already present in the hosted audit artifact. The OSNR code never changes the FNO weights and does not run reverse-mode differentiation. Each Test~29 file contains arrays \texttt{preds} and \texttt{targets} with shape $10\times128\times128\times21\times4$: ten held-out forecast samples, a $128^2$ spatial grid, $21$ forecast frames, and four weather/PDE channels. For each sample $n$, time index $\tau$, and channel $c$, the layer treats the FNO forecast $p_c(x)$ as a dynamical prior and receives contemporary station observations $y_c(x_i)$ at a selected set $S$ of grid sites. The correction is the three-block operator
\[
\mathcal{A}_{S,r,\ell}(p,y)
=
\mathcal{R}_{S,\ell}\!\left(
\mathcal{D}_{S,r}\!\left(
\mathcal{C}_{S}(p,y),y\right),y\right),
\]
where $\mathcal{C}_S$ is per-channel affine calibration, $\mathcal{D}_{S,r}$ is a low-rank DCT residual solve, and $\mathcal{R}_{S,\ell}$ is a compact Gaussian-RBF station residual solve. The first block solves
\[
(a_c,b_c)=\arg\min_{a,b}\sum_{i\in S}(a\,p_c(x_i)+b-y_c(x_i))^2,
\qquad
q_c^{(0)}(x)=a_c p_c(x)+b_c.
\]
The DCT block builds a tensor-product dictionary $\Phi_r\in\mathbb{R}^{HW\times r^2}$ with normalized atoms
\[
\phi_{k_y,k_x}(i,j)
=
Z^{-1}_{k_y,k_x}
\cos\!\left(\frac{\pi(i+1/2)k_y}{H}\right)
\cos\!\left(\frac{\pi(j+1/2)k_x}{W}\right),
\qquad 0\leq k_y,k_x<r,
\]
and solves one ridge system shared across channels,
\[
B_c=(\Phi_S^\top\Phi_S+\lambda I)^{-1}\Phi_S^\top
\bigl(y_c(S)-q_c^{(0)}(S)\bigr),
\qquad
q_c^{(1)}(x)=q_c^{(0)}(x)+\Phi_r(x)B_c .
\]
The local residual block then places Gaussian atoms at the same station sites,
\[
K_{ij}=\exp\!\left(-\frac{\|x_i-x_j\|_2^2}{2\ell^2}\right),
\qquad
\alpha_c=(K_{SS}+\lambda I)^{-1}
\bigl(y_c(S)-q_c^{(1)}(S)\bigr),
\]
and evaluates $q_c^{(2)}(x)=q_c^{(1)}(x)+K_{xS}\alpha_c$ on the full grid. All reported rows use $\lambda=10^{-3}$ and double-precision NumPy linear algebra. The metric is per-sample relative $L^2$ error over all forecast pixels, times, and channels,
\[
\mathrm{nRMSE}_n=
\frac{\|\hat{Y}_n-Y_n\|_2}{\|Y_n\|_2+10^{-20}},
\]
with the table reporting the mean over the ten samples.
The station-placement rule in the winning rows is also deterministic given the script seed. The grid is the integer lattice $\{0,\ldots,127\}^2$. For \texttt{space\_filling}, the runner draws one initial grid point from \texttt{np.random.default\_rng(seed)} and then repeatedly adds the point whose squared distance to the already selected set is largest. For each file and hyperparameter row the seed is
\[
20260531+1009\,i_{\mathrm{file}}+131\,m+17\,r+\lfloor\ell\rfloor,
\]
because the winning sweeps use only the \texttt{space\_filling} strategy. Here $m$ is the station count, $r$ is the DCT rank, and $\ell$ is the RBF length scale. The exact runs that produced Table~\ref{tab:pdebench-weather-adaptive-stations} are:
\begin{verbatim}
venv/bin/python apps_industrial_breakthrough/\
pdebench_weather_adaptive_station_assimilation.py \
--sensor_counts 256 --dct_ranks 4,8 --length_scales 12,16,18 \
--strategies space_filling --candidate_pool 4096 \
--output_dir apps_industrial_breakthrough/\
pdebench_weather_adaptive_station_outputs_256_ell16
venv/bin/python apps_industrial_breakthrough/\
pdebench_weather_adaptive_station_assimilation.py \
--sensor_counts 288,320,384 --dct_ranks 4,8 \
--length_scales 16,20,24 --strategies space_filling \
--candidate_pool 4096 \
--output_dir apps_industrial_breakthrough/\
pdebench_weather_adaptive_station_outputs_sub512_targeted
\end{verbatim}
The runner uses only \texttt{numpy} and \texttt{PIL}; the adaptive station-placement results in this table were executed on CPU, not on CUDA or MPS. This is acceptable for the claim being made because the layer is a small closed-form assimilation operator rather than a trained network. GPU acceleration would mainly speed up the dense station-kernel solves and repeated grid evaluations; it is not responsible for the reported accuracy.
We next implemented a GPU-ready no-backprop OSNR network runner,
\texttt{apps\_industrial\_breakthrough/pdebench\_weather\_gpu\_osnr\_network.py}, to test whether the assimilation block should become a trained operator network rather than a fixed dictionary layer. The architecture keeps the same frozen FNO prior and station-conditioned affine/DCT/RBF heads, but adds a learned residual POD dictionary fitted from training samples by closed-form SVD. No PyTorch autograd graph is constructed; the entire run uses \texttt{torch.inference\_mode()}, and the script selects MPS/CUDA when the environment exposes those backends. The first bounded probe used training samples $0$--$3$, held-out samples $4$--$5$, forecast frames $10$--$13$, rank-$8$ DCT atoms, RBF length scale $16$, and $128$ or $256$ space-filling stations. The local environment reported \texttt{cuda\_available=false} and \texttt{mps\_available=false}, so this probe executed on CPU despite the GPU-capable code path. The result is informative but not a new headline: on \texttt{M01\_Eta01}, $256$ stations with eight learned POD atoms per channel slightly improves the probe nRMSE from $0.000665$ to $0.000663$; on \texttt{M10\_Eta01} and \texttt{M10\_Eta001}, the learned atoms slightly worsen the pure DCT/RBF head. Thus the next serious no-backprop weather-network target is not generic residual PCA; it is a trained station policy, operator gate, or dictionary-selection controller that preserves the conditioned spatial coverage responsible for Table~\ref{tab:pdebench-weather-adaptive-stations}.
A subsequent MPS audit exposed the true high-budget behavior of the same frozen-prior assimilation network. The Apple-MPS backend is visible in the project \texttt{venv} through the \texttt{runpy} invocation path, and the confirmed run reports \texttt{device=mps}, \texttt{torch\_version=2.12.0}, and \texttt{autograd=disabled}. The protocol is deliberately stricter than the earlier all-sample adaptive table: for each Test~29 file, samples $0$--$3$ fit any closed-form residual statistics, while samples $4$--$9$ are held out; all $21$ forecast frames and all four channels are evaluated. The best rows use no learned POD atoms, no joint residual atoms, sequential affine+DCT+RBF correction, rank-$8$ DCT atoms, RBF ridge $5\times10^{-4}$, and deterministic space-filling stations. The only swept variables in the headline rows are the number of stations and the RBF length scale. At $2048$ stations, which is $12.5\%$ of the $128^2$ grid per frame, the layer reached the first strong external-weather frontier: \texttt{M01\_Eta01} dropped to $3.19\times10^{-4}$, \texttt{M10\_Eta01} to $4.90\times10^{-4}$, and \texttt{M10\_Eta001} to $1.40\times10^{-3}$. We then pushed the dense local operator harder under the same MPS memory guard. At $3072$ stations the same architecture reached $2.57\times10^{-4}$, $3.84\times10^{-4}$, and $7.53\times10^{-4}$, respectively. At $4096$ stations, or $25\%$ of the $128^2$ grid per frame, it reached the current external Test~29 frontier: $2.21\times10^{-4}$, $3.46\times10^{-4}$, and $6.66\times10^{-4}$. Relative to the previous $2048$-station frontier, the $4096$-station rows improve the three held-out means by $30.8\%$, $29.4\%$, and $52.5\%$, respectively. The cost is the expected dense-kernel bottleneck: the $4096$ rows use a $320$ MB station kernel and take about $185$ s per file on MPS.
The exact metric artifacts are:
\begin{verbatim}
apps_industrial_breakthrough/
pdebench_weather_gpu_osnr_network_outputs_s2048_rbf0005_mps/metrics.json
pdebench_weather_gpu_osnr_network_outputs_s3072_eta01_mps/metrics.json
pdebench_weather_gpu_osnr_network_outputs_s3072_m10eta001_mps/metrics.json
pdebench_weather_gpu_osnr_network_outputs_s4096_eta01_mps/metrics.json
pdebench_weather_gpu_osnr_network_outputs_s4096_m10eta001_mps/metrics.json
\end{verbatim}
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{lcccccc}
\toprule
External held-out Test~29 file & FNO & $512$ & $1024$ & $2048$ & $3072$ & $4096$ stations \\
\midrule
\texttt{M01\_Eta01} & $0.004959$ & $0.000734$ & $0.000541$ & $0.000319$ & $0.000257$ & $\mathbf{0.000221}$ \\
\texttt{M10\_Eta01} & $0.006773$ & $0.001151$ & $0.000761$ & $0.000490$ & $0.000384$ & $\mathbf{0.000346}$ \\
\texttt{M10\_Eta001} & $0.013097$ & $0.005390$ & $0.002843$ & $0.001403$ & $0.000753$ & $\mathbf{0.000666}$ \\
\bottomrule
\end{tabular}
\caption{MPS high-budget sparse-station scaling on external PDEBench/FNO Test~29 weather artifacts. The protocol holds out samples $4$--$9$, evaluates all $21$ forecast frames and four channels, and keeps the FNO forecast frozen. All rows use \texttt{torch.inference\_mode()}, MPS tensors, rank-$8$ DCT correction, no learned residual POD atoms, no backpropagation, RBF ridge $5\times10^{-4}$, and a deterministic space-filling station policy. The $4096$-station frontier gives relative held-out mean reductions of $95.6\%$, $94.9\%$, and $94.9\%$, respectively, but the dense RBF block now costs about $185$ s per file on MPS.}
\label{tab:pdebench-weather-mps-high-budget}
\end{table}
This MPS sweep also falsified several tempting station-selection and operator-network variants. Residual-energy, forecast-gradient, observability, neuro-coverage, dynamic-lattice, dynamic-attention, reward-gated, time-local POD, global POD, and joint-atom variants did not beat the best conditioned space-filling baseline on the full held-out protocol. We also tested two stronger control-style placement ideas. A two-stage dynamic pilot-innovation policy first spends a station subset on a topographic pilot, diffuses the observed innovation magnitude through an RBF field, and then places the remaining stations from that inferred regional error signal. On a bounded \texttt{M10\_Eta001} smoke it underperformed plain space-filling. A training-loss station-seed search is valid and mildly useful: with $512$ stations on the hard held-out file it improves the best mean from $0.005390$ to $0.005359$, but the gain is too small to explain the frontier. The useful mechanism is therefore not merely ``look where the forecast is large,'' ``follow dopamine-like innovation,'' or ``add more learned atoms.'' It is the numerical conditioning of a global--local operator dictionary under sparse contemporary observations. The next weather-scale research target is a subquadratic or partitioned station solver, because the current dense RBF block scales like $O(m^3)$ in the station count and becomes the runtime bottleneck exactly where accuracy is best.
A subsequent sub-$512$ MPS audit confirms this conclusion under the same strict held-out protocol. With rank-$8$ DCT, no learned POD or joint atoms, RBF ridge $5\times10^{-4}$, and all $21$ frames evaluated, a $256/384$ station sweep over the three Test~29 files finds best means $0.000791$, $0.001301$, and $0.006318$. The first is not a new \texttt{M01\_Eta01} frontier, but the latter two beat the old $512$-random hybrid references for \texttt{M10\_Eta01} and \texttt{M10\_Eta001}. A focused hard-file sweep then shows that $320$ stations already reduce \texttt{M10\_Eta001} to $0.007178$, below the old $512$-random $0.008584$ reference, while $448$ space-filling stations with length scale $9$ reach $0.005784$. Training-loss station-seed search helps some rows (\texttt{M01\_Eta01} and \texttt{M10\_Eta01}) but hurts or misses the best \texttt{M10\_Eta001} rows. Thus the high-information station-placement rule is still conditioned coverage plus length-scale matching, not residual hot-spot chasing.
We then tested the subquadratic solver implied by this bottleneck. A purely spatial partition-of-unity RBF head fails: on \texttt{M10\_Eta001}, side-$2$ local blocks over $448/640/768$ stations give best mean $0.009167$, and adding a $128$-station coarse global scaffold before the local residual solve gives best mean $0.009540$. Both are much worse than the dense $448$-station $0.005784$ result, so the weather residual is not a set of independent local patches. It requires the global covariance geometry of the RBF kernel. Replacing the dense station kernel by a global inducing-center RBF dictionary is the useful compression. With all station observations retained as regression rows but only $384$ globally distributed RBF centers as columns, rank-$8$ DCT, no POD or joint atoms, length scale $7$, and \texttt{torch.inference\_mode()} on MPS, the three-file run reaches $0.000661$, $0.001029$, and $0.004823$ at $1024$ stations for \texttt{M01\_Eta01}, \texttt{M10\_Eta01}, and \texttt{M10\_Eta001}; increasing to $512$ inducing centers, $1536$ stations, and length scale $5$ improves these to $0.000596$, $0.000883$, and $0.003766$; increasing again to $768$ centers and $2048$ stations gives $0.000494$, $0.000690$, and $0.002736$; using $1024$ centers, $3072$ stations, and length scale $4$ gives $0.000408$, $0.000561$, and $0.002092$; using $1536$ centers, $4096$ stations, and length scale $3$ gives $0.000329$, $0.000444$, and $0.001368$. A rank-$2048$ inducing profile with $6144$ stations and length scale $2.5$ reaches $0.000240$, $0.000383$, and $0.001034$ with relative reductions from the frozen FNO of $95.2\%$, $94.3\%$, and $92.1\%$. The recorded peak kernel/design estimate for this row is $192.0$ MB versus $528.0$ MB for a dense $6144$-station kernel, and the row evaluates in about $58.8$ s on MPS. We also fixed the runner so that inducing mode no longer materializes the unused dense station kernel before the low-rank solve; the subquadratic memory estimate now matches the executed branch. A first inducing-center audit shows that center geometry, not just center count, matters: on the hard file at rank $2048$, $6144$ stations, and length scale $2.5$, the historical farthest-point station prefix reaches $0.001040$, independent global space-filling centers reach $0.001059$, station-stride centers reach $0.001253$, and late station-tail centers collapse to $0.005219$. A second score-weighted audit keeps a coverage scaffold and spends the remaining centers on training residual geometry: on the hard file, $75\%$ coverage plus residual centers improves to $0.000966$, while $50\%$ coverage gives $0.000998$, $90\%$ gives $0.001023$, and a broader neuro-score field gives $0.000976$. The all-file rank-$2048$ residual-center validation reaches $0.000253$, $0.000366$, and $0.001012$ for \texttt{M01\_Eta01}, \texttt{M10\_Eta01}, and \texttt{M10\_Eta001}. Retuning its length scale over $2,2.25,2.5,2.75,3$ selects $\ell=2.75$ on the hard file with $0.000957$; all-file validation at this scale reaches $0.000255$, $0.000361$, and $0.000993$. A bottleneck isolation then rejects merely adding observations at fixed rank: rank $2048$ with $8192$ stations gives only $0.000972$ on \texttt{M10\_Eta001}, worse than the $6144$-station row. Increasing the inducing dictionary instead is decisive. With rank $3072$, $6144$ stations, $75\%$ coverage plus residual centers, and $\ell=2.5$, the MPS all-file validation reaches $0.000208$, $0.000306$, and $0.000634$ for \texttt{M01\_Eta01}, \texttt{M10\_Eta01}, and \texttt{M10\_Eta001}, with FNO reductions of $95.8\%$, $95.5\%$, and $95.2\%$. The kernel/design estimate is $300.0$ MB versus $528.0$ MB for the dense $6144$-station kernel. Rank $4096$ remains viable under a $512$ MB guardrail; a hard-file sweep over $\ell=2,2.25,2.5$ selects $\ell=2$ with mean $0.000540$, and all-file validation reaches $0.000183$, $0.000290$, and $0.000534$ with FNO reductions of $96.3\%$, $95.7\%$, and $95.9\%$. Its kernel/design estimate is $416.0$ MB versus $528.0$ MB dense and each row costs about $202$ s on MPS. A rank-$4096$ center-allocation audit keeps $\ell=2$ and varies only the center score: $50\%$ residual coverage gives $0.000547$ on \texttt{M10\_Eta001}, $87.5\%$ residual coverage gives $0.000552$, and a neuro-score center field at $75\%$ coverage gives $0.00054037$, fractionally worse than the $75\%$ residual-energy row at $0.00054017$. Ridge retuning then shows a shallow cross-file tradeoff: $\lambda_{\rm rbf}=2{\times}10^{-4}$ gives the best hard-file row at $0.000531$ but worsens \texttt{M10\_Eta01} to $0.000292$, $\lambda_{\rm rbf}=10^{-4}$ regresses the hard file to $0.000539$, and the best single-profile macro mean is $\lambda_{\rm rbf}=3{\times}10^{-4}$, reaching $0.000182$, $0.000290$, and $0.000532$ across \texttt{M01\_Eta01}, \texttt{M10\_Eta01}, and \texttt{M10\_Eta001}; $\lambda_{\rm rbf}=4{\times}10^{-4}$ lands slightly worse at $0.000182$, $0.000290$, and $0.000533$. Intermediate compression points preserve the same geometry with smaller dictionaries: rank $3584$ chooses $\ell=2.25$ and gives $0.000192$, $0.000295$, and $0.000564$ at $357.0$ MB and about $159$ s per row, while rank $3840$ also chooses $\ell=2.25$ and improves the tradeoff to $0.000189$, $0.000292$, and $0.000552$ at $386.25$ MB and about $180$ s per row. Finally, we added an \texttt{--eval\_split} hook to audit profile selection without held-out leakage. Training-split losses over $\lambda_{\rm rbf}\in\{2,3,5\}\times10^{-4}$ select $5{\times}10^{-4}$ for \texttt{M01\_Eta01}, $2{\times}10^{-4}$ for \texttt{M10\_Eta001}, and $5{\times}10^{-4}$ for \texttt{M10\_Eta01}; this misses the tiny held-out \texttt{M01\_Eta01} optimum but selects both M10 held-out winners, giving held-out values $0.000183$, $0.000290$, and $0.000531$, better in macro mean than any single ridge. Thus the scalable weather solver should be low-rank global RBF/Nystr\"om geometry with enough inducing capacity and a balanced global/residual center budget, not spatial partitioning or additional stations at a fixed insufficient rank; beyond rank $4096$, the next question is compression or per-regime regularization because memory approaches dense parity.
After this profile-selection audit, we retuned the rank-$4096$ frontier more finely. A control-style, station-held reward-gain branch first tested whether a terminal dopamine-like scalar could improve the correction magnitude. The runner was patched so that reward-gated assimilation uses the same inducing RBF operator as the frontier path; with $10\%$ held stations and candidate gains $\{0.75,0.9,1.0,1.1,1.25\}$, the hard \texttt{M10\_Eta001} file selected average gain $0.94$ but worsened to $0.000542$, so scalar terminal gain is a negative branch. The positive branch is earlier operator regularity. A fine length-scale audit at $6144$ stations, rank-$4096$, residual-center coverage $0.75$, and $\lambda_{\rm rbf}=2{\times}10^{-4}$ over $\ell\in\{1.5,1.75,2,2.125,2.25\}$ selects $\ell=2.125$; all-file validation reaches $0.000182$, $0.000526$, and $0.000289$ for \texttt{M01\_Eta01}, \texttt{M10\_Eta001}, and \texttt{M10\_Eta01}. Retuning $\lambda_{\rm rbf}$ at this length scale gives a hard-file knee at $5{\times}10^{-5}$: the all-file single-profile row reaches $0.0001816$, $0.0005214$, and $0.0002914$ with the same $416$ MB inducing design estimate. A no-leakage train-split selector between $2{\times}10^{-4}$ and $5{\times}10^{-5}$ at $\ell=2.125$ selects the hard-file and \texttt{M10\_Eta01} held-out winners, missing only the tiny \texttt{M01\_Eta01} preference; the train-selected held-out profile is approximately $0.000182$, $0.000521$, and $0.000289$. Applying the same lower ridge to the rank-$3840$ compression point shows that compressed profiles prefer a slightly broader kernel: $\ell=2.25$, $\lambda_{\rm rbf}=5{\times}10^{-5}$ reaches $0.0001857$, $0.0005369$, and $0.0002919$ at $386.25$ MB, improving the earlier rank-$3840$ tradeoff while remaining below rank-$4096$ quality. The practical conclusion is that the next weather lever is per-regime operator regularity and compression, not terminal scalar reward gain.
The next lower compression point confirms the rank curve. Rank $3584$ with the same lower ridge and $\ell=2.25$ reaches $0.0001896$, $0.0005505$, and $0.0002958$ at $357$ MB and about $159$ s per row. This improves the old rank-$3584$ row, but rank $3840$ is the cleaner quality--memory compromise before the rank-$4096$ frontier.
Pushing the opposite direction gives the current quality frontier while still respecting the memory guard. Rank $4608$ remains below the dense $6144$-station kernel, with a $477$ MB design estimate. A hard-file sweep at $\lambda_{\rm rbf}=5{\times}10^{-5}$ selects $\ell=2.25$ with \texttt{M10\_Eta001} mean $0.0005048$. The all-file validation reaches $0.0001805$, $0.0005048$, and $0.0002837$ for \texttt{M01\_Eta01}, \texttt{M10\_Eta001}, and \texttt{M10\_Eta01}, respectively, at about $251$ s per row. This improves the rank-$4096$ single-profile frontier on every Test~29 weather file while remaining below dense-kernel memory.
The edge-of-guardrail rank-$4800$ profile improves the frontier again while still staying below dense memory: the design estimate is $500.39$ MB versus $528$ MB dense. A hard-file sweep over $\ell\in\{2.125,2.25,2.375\}$ selects $\ell=2.25$ and reaches $0.0004994$ on \texttt{M10\_Eta001}, the first confirmed sub-$0.0005$ hard weather result in this campaign. The all-file validation reaches $0.0001797$, $0.0004994$, and $0.0002822$ for \texttt{M01\_Eta01}, \texttt{M10\_Eta001}, and \texttt{M10\_Eta01}, at about $271$ s per row. This is the current quality frontier under the sub-dense memory guard.
The edge-rank check uses rank $4864$, which still fits under the guard at $508.25$ MB. With $\ell=2.25$, $\lambda_{\rm rbf}=5{\times}10^{-5}$, and a $75\%$ inducing-center coverage scaffold, the all-file validation reaches $0.0001797$, $0.0004973$, and $0.0002820$ for \texttt{M01\_Eta01}, \texttt{M10\_Eta001}, and \texttt{M10\_Eta01}, at about $278$ s per row. The matching train-split audit gives $0.0001188$, $0.0006095$, and $0.0003040$, supporting selection of the edge-rank profile from training trajectories. A fine scaffold retune at the same rank and memory then improves the hard frontier. The first hard-file sweep over coverage fractions $0.50,0.625,0.875,1.00$ gives $0.0005017$, $0.0004967$, $0.0005087$, and $0.0005129$. A matched-station-seed sweep around the knee gives $0.0004966$, $0.0004950$, $0.00049323$, $0.00049266$, and $0.0004939$ for coverage $0.5625,0.600,0.625,0.650,0.6875$. Promoting the $0.650$ scaffold to all three files gives the current sub-dense quality frontier, $0.00017950$, $0.00049266$, and $0.00028253$, still at $508.25$ MB versus $528$ MB dense and about $277$ s per row. This improves all three files versus the $0.625$ row and improves the hard file by about $0.92\%$ relative to the old $75\%$ scaffold. The no-leakage train-split audit for the $0.650$ scaffold gives $0.00011875$, $0.00061174$, and $0.00030438$, so raw train loss would still prefer the $75\%$ scaffold for the hard and \texttt{M10\_Eta01} files. Thus the $0.650$ row is a real held-out frontier, but robust train-selectable per-regime control remains unsolved; the remaining rank headroom before dense parity is only a few MB, so further progress should come from algorithmic compression or better no-leakage controllers rather than raw rank escalation.
The first post-frontier compression audit therefore varied the inducing-center allocation instead of the raw rank. At rank $3840$, $6144$ stations, $\ell=2.25$, $\lambda_{\rm rbf}=5{\times}10^{-5}$, rank-$8$ DCT, and no POD/joint atoms, pure residual-energy centers with no global coverage scaffold validate at $0.0001845$, $0.0005313$, and $0.0002962$ for \texttt{M01\_Eta01}, \texttt{M10\_Eta001}, and \texttt{M10\_Eta01}. This slightly improves the rank-$3840$ macro mean over the $75\%$ coverage scaffold row $(0.0001857,0.0005369,0.0002919)$ at the same $386.25$ MB design estimate, but worsens \texttt{M10\_Eta01}. Pure global space-filling centers are worse on the hard file ($0.0005612$), broad neuro-score centers are also worse ($0.0005413$), and an intermediate $25\%$ coverage scaffold reaches only $(0.0001853,0.0005359,0.0002945)$. At rank $4864$, the pure-residual endpoint worsens the hard file to $0.0004998$ versus the $75\%$ scaffold frontier $0.0004973$. Thus center allocation is not a universal scalar setting: lower-rank compression benefits from more aggressive innovation-driven centers, while the edge-rank frontier still needs a coverage scaffold for conditioning. A no-leakage selector audit exposed the next bottleneck. Raw train-split losses at rank $3840$ give cov0 $(0.0001202,0.0006838,0.0003135)$ and cov75 $(0.0001208,0.0006750,0.0003132)$; this selects the held-out winners for \texttt{M01\_Eta01} and \texttt{M10\_Eta01} but misses the hard file. A one-sample inner validation split, fitting on samples $0$--$2$ and scoring sample $3$, gives cov0 $(0.0001339,0.0008521,0.0003611)$ and cov75 $(0.0001343,0.0008527,0.0003614)$, selecting cov0 for all three and therefore missing the held-out \texttt{M10\_Eta01} scaffold preference. A two-sample inner validation split, fitting on samples $0$--$1$ and scoring samples $2,3$, gives cov0 $(0.0001387,0.0010027,0.0004333)$ and cov75 $(0.0001393,0.0009815,0.0004329)$; it selects cov75 for the hard file even though held-out hard prefers cov0. We also implemented a station-held center-coverage gate that scores candidate center profiles on held observed stations inside each forecast slice and then refits the selected profile on all observed stations. At rank $3840$ with candidates $\{0,0.75\}$ and $10\%$ held stations, \texttt{M10\_Eta001} reaches $0.0005348$ with mean selected coverage $0.268$, better than static cov75 but worse than static cov0; \texttt{M10\_Eta01} reaches only $0.0003009$ with mean selected coverage $0.298$, worse than both static profiles. A cheap train-only proxy audit over residual-score entropy, grid coverage, station coverage, and center nearest-neighbor spacing is also inconclusive: cov0 and cov75 have nearly identical coverage means (about $0.81$ grid units) and nearest-neighbor statistics. We then patched the dynamic station-placement path so that dynamic gradient, neuro-attention, and pilot-innovation policies use the same low-rank inducing RBF solver as the static frontier. The full hard-file pilot-innovation row at $6144$ stations and rank $3840$ reaches only $0.0005426$ in $706$ s, worse than static cov0 ($0.0005313$). A short prefix scan over frames $0$--$4$ was invalid because those frames have zero baseline forecast error in the artifact. On the meaningful frame-$5$--$9$ prefix, dynamic neuro attention with a $75\%$ lattice scaffold improves the local comparator from $0.0003793$ to $0.0003679$, but the promoted full late-horizon hard-file run rejects the signal: static neuro coverage $0.90$ reaches $0.0006612$ in $137$ s, whereas dynamic neuro coverage $0.75$ reaches $0.0006768$ in $491$ s and coverage $0.90$ reaches $0.0007066$ in $459$ s. The compression result is real, but robust per-regime controller selection remains open; raw training loss, tiny validation windows, simple geometry proxies, repeated high-rank station-held gates, and the current greedy dynamic station selector are too noisy or too expensive for this knob.
We also ran the harder replacement test: remove the frozen FNO prior and forbid target-time stations. The new runner \texttt{apps\_industrial\_breakthrough/pdebench\_weather\_blind\_osnr\_dynamics\_mps.py} identifies a blind dynamics model from observed trajectories only and then rolls held-out samples forward from their history frames. Its cell has two coupled no-backprop components. The local PDE library uses current fields, velocity memory, first derivatives, Laplacians, biharmonic terms, Laplacian velocity, quadratic products, self-advection, and cross-channel products, with coefficients solved by ridge systems. A low-mode spectral liquid cell adds DCT coefficients, finite differences, stable pole traces $\{0.25,0.50,0.75,0.90,0.97\}$, and tanh coefficient features. All fits run under \texttt{torch.inference\_mode()} on MPS; no FNO prediction, no future observation, and no autograd graph are used. The result is a clear boundary rather than a win. A global fit with three history frames is unstable and worse than persistence. A sample-online fit with per-sample history clamps and five observed frames becomes stable but only improves persistence by $0.30\%$, $0.42\%$, and $0.41\%$ on \texttt{M01\_Eta01}, \texttt{M10\_Eta001}, and \texttt{M10\_Eta01}, while remaining $2.63\times$, $21.20\times$, and $12.06\times$ worse than FNO. With eight history frames the persistence gains rise only to $0.83\%$, $0.88\%$, and $0.88\%$, with FNO gaps of $1.45\times$, $20.53\times$, and $5.75\times$. Thus the current weather result is assimilation, not autonomous GraphCast/GenCast replacement. Replacing the entire neural forecaster requires substantially more trajectory diversity or a stronger endogenous PDE/ODE neural architecture; the present local library mostly learns a small homeostatic correction to persistence.
The follow-up autonomous-weather loop raised the history to $16$ frames and changed the local cell rather than only retuning ridge constants. The \texttt{temporal} feature set adds previous-frame state, velocity derivatives, previous Laplacian, saturating nonlinearities, and current--velocity products. The \texttt{memory2} feature set adds a second-order temporal state: older frame, prior velocity, acceleration, acceleration derivatives/Laplacian, older Laplacian, and acceleration cross-products. Finally, a partitioned operator layout fits separate PDE matrices on a $2\times2$ spatial grid but shrinks them toward the shared trajectory-level operator,
\[
W_{b}=(1-s)W_{\rm shared}+sW_{{\rm local},b},
\qquad s=0.25,
\]
which preserves global conditioning while allowing regional deviations. Table~\ref{tab:pdebench-blind-weather-autonomous} reports the current best rows. The important result is not yet a weather-model victory: the autonomous cell beats frozen FNO only on the two easy regimes where persistence is already very strong, and it remains about $9.5\times$ worse than FNO on the hard \texttt{M10\_Eta001} regime. The positive scientific signal is narrower but real: second-order memory plus light regional shrinkage improves the hard no-FNO row from $0.130594$ to $0.121345$, and a finer $8\times8$ partition with a global observed-history rollout selector improves the hard row further to $0.115310$. This selector is not trained on held-out future frames: it scores candidate ridge/update settings on training-sample observed-history tails, then applies the chosen setting to held-out histories. Its selected settings are $(\lambda,\gamma)=(300,0.70)$ for \texttt{M01\_Eta01}, $(500,0.75)$ for hard \texttt{M10\_Eta001}, and $(300,0.70)$ for \texttt{M10\_Eta01}. The partition-resolution audit also sets a boundary: $4\times4$ reaches hard $0.116181$, $8\times4$ reaches $0.115653$, $8\times8$ fixed reaches $0.115466$, but $16\times8$ and $8\times16$ regress to about $0.11587$ while doubling runtime; $8\times8$ shrinkage $s=0.20$ and $s=0.40$ regress to $0.115805$ and $0.116198$. Negative controls were decisive. A multiscale smoothed-context dictionary worsened hard error to $0.241980$, pure $2\times2$ blocks worsened to $0.124196$, temporal blocks worsened to $0.126452$, explicit coordinate augmentation of the memory2 cell reached only $0.116311$, compact per-channel global moment modulators reached only $0.116603$ and a damped version $0.117145$, nearby $8\times8$ shrinkage $s=0.25$ and $s=0.35$ reached only $0.115562$ and $0.115525$, a p90-robust forced row reached $0.115537$, and top-$3$/top-$5$ history-tail ensembles reached only $0.115345$/$0.115357$ on the hard file. Tail-length validation also closed around the current setting: observed-history rollout tails of $2$, $3$, $5$, and $6$ steps reached $0.118317$, $0.118317$, $0.115710$, and $0.115949$, so the four-step tail remains best. A wider channel-coupled cross-operator dictionary was especially diagnostic: its lower-ridge run improved the observed-history tail score from $0.283297$ to $0.280035$ but worsened held-out future error to $0.117344$; the stronger-ridge run reached $0.117920$. The observed-history selector can therefore be fooled by high-capacity local dictionaries. Low-mode spectral blending also remained dangerous: an ungated conservative spectral search would choose blend $0.20$ and degrade hard held-out error to $0.135724$. The same mean/p90 spectral gate now blocks that row because the observed-tail mean gain over zero blend is only $4.68\%<5\%$, restoring the zero-blend $0.115310$ frontier. Thus the next autonomous-weather step should not be more spectral blending, coordinate tags, compact global moments, scalar shrinkage tuning, tail-length retuning, top-k row averaging, or wider local dictionaries without a stronger validation guard; it should use a batched/subquadratic partitioned solver and a structurally different endogenous low-mode/state mechanism with no-leakage model selection.
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{lcccc}
\toprule
Autonomous no-FNO cell, history $16$ & \texttt{M01\_Eta01} & \texttt{M10\_Eta001} & \texttt{M10\_Eta01} & Macro mean \\
\midrule
FNO artifact & $0.006846$ & $0.012127$ & $0.009781$ & $0.009585$ \\
Persistence & $0.001090$ & $0.206921$ & $0.008352$ & $0.072121$ \\
Temporal shared OSNR & $0.000861$ & $0.130594$ & $0.006783$ & $0.046079$ \\
Memory2 shared OSNR & $0.000897$ & $0.123513$ & $0.007313$ & $0.043907$ \\
Memory2 $2\times2$ shrink OSNR & $0.000905$ & $0.121345$ & $0.007127$ & $0.043126$ \\
Selector: temporal easy, $2\times2$ hard & $0.000861$ & $0.121345$ & $0.006783$ & $0.043330$ \\
Memory2 $8\times8$ history-tail global OSNR & $0.000851$ & $0.115310$ & $0.007224$ & $0.041128$ \\
\bottomrule
\end{tabular}
\caption{Autonomous no-backprop weather replacement audit on external PDEBench/FNO Test~29 artifacts. The model observes only the first $16$ true frames of each held-out trajectory and predicts the remaining five frames. It does not read FNO predictions, target-time stations, or future labels at evaluation time. All rows use per-trajectory closed-form ridge identification under \texttt{torch.inference\_mode()} on MPS. The last row uses the formal \texttt{history\_tail\_global} selector: one ridge/update setting is chosen from training-sample observed-history rollout error for each file, then applied to held-out histories.}
\label{tab:pdebench-blind-weather-autonomous}
\end{table}
\begin{table}[h]
\centering
\small
\begin{tabular}{lccccc}
\toprule
External forecast artifact & FNO & $256$ sf tuned & $288$ sf & $384$ sf & $512$ random ref. \\
\midrule
Test~29 \texttt{M01\_Eta01} & $0.004611$ & $0.000705$ & $0.000686$ & $0.000646$ & $0.000707$ \\
Test~29 \texttt{M10\_Eta01} & $0.005507$ & $0.001377$ & $0.001334$ & $0.001267$ & $0.001438$ \\
Test~29 \texttt{M10\_Eta001} & $0.012038$ & $0.007649$ & $0.007639$ & $0.007274$ & $0.008584$ \\
\bottomrule
\end{tabular}
\caption{Adaptive station placement on external PDEBench/FNO Test~29 artifacts. All adaptive rows use farthest-point space-filling placement with the same closed-form affine+DCT+RBF assimilation layer; only the station count, DCT rank, and RBF length scale are selected by the bounded sweep. The $512$-station random hybrid reference is beaten by $256$ tuned space-filling stations on all three files, and the $384$-station row gives the current best Test~29 means.}
\label{tab:pdebench-weather-adaptive-stations}
\end{table}
\subsection{Direct no-backprop PDEBench sequence continuation}
The next non-Fashion test removes both the frozen FNO prior and target-time stations, but keeps a short true history of the same trajectory. The runner \texttt{apps\_industrial\_breakthrough/pdebench\_1d\_direct\_spectral\_forecaster\_mps.py} is a no-backprop direct sequence forecaster for the 1D PDEBench audit artifacts. For a target tensor $Y\in\mathbb{R}^{N\times X\times T\times C}$, history length $h$, and normalized spectral matrix $\Psi_K\in\mathbb{R}^{X\times K}$, it forms coefficient tokens
\[
a_{n,t}=\Psi_K^\top y_{n,t}\in\mathbb{R}^{K C}.
\]
The per-sample input feature is
\[
z_n =
\bigl[
1,\,
a_{n,0:h-1},\,
\Delta a_{n,0:h-2},\,
a_{n,h-1},\,
\bar a_n,\,
\operatorname{std}(a_n),\,
\tanh(0.5a_{n,h-1})
\bigr],
\]
flattened over time, modes, and channels. The future coefficients are obtained by one closed-form ridge solve,
\[
W_{\lambda,K}=(Z^\top Z+\lambda I)^{-1}Z^\top A_{h:T-1},
\qquad
\hat A_{h:T-1}=Z W_{\lambda,K}.
\]
The low-mode field is reconstructed by $\Psi_K\hat A$, and a high-frequency identity residual from the last observed frame is added with validation-selected decay $\rho$:
\[
\hat y_{n,t}=\Psi_K\hat a_{n,t}
+\rho\left(y_{n,h-1}-\Psi_K\Psi_K^\top y_{n,h-1}\right),
\qquad t\ge h.
\]
No FNO prediction, no future observation, and no reverse-mode graph are used by the model. The updated runner also validates a small no-backprop expert set---direct spectral, persistence, and linear extrapolation---so the reported ``selected'' row is the validation-only choice among physically simple history-conditioned mechanisms. Runs use \texttt{torch.inference\_mode()} on Apple MPS through the project \texttt{venv/bin/python -c "... runpy.run\_path(...)"} route, because direct script execution can hide the MPS backend in this environment.
The latest autonomous extension adds an explicit physical-cell feature grid while keeping the default base model unchanged. For each low-mode coefficient history, the liquid feature set forms causal pole traces
\[
p_t^{(\alpha)}=\alpha p_{t-1}^{(\alpha)}+(1-\alpha)a_t,
\qquad \alpha\in\{0.10,0.30,0.55,0.75,0.90,0.98\},
\]
then appends acceleration, last velocity, bounded \texttt{softsign}/\texttt{tanh} coordinates, pole innovations $a_{h-1}-p_{h-1}^{(\alpha)}$, low-mode quadratic products, and leading-mode spline hinges $\max(a-\kappa,0)$ with validation-selected spline width. A mixed \texttt{base,liquid} grid lets validation route each PDE file to the simpler DCT history model or to the richer liquid/spline cell. A separate \texttt{autoregressive} mode trains the same closed-form map on all sliding history windows and rolls forward without backprop; it is used only as a long-horizon diagnostic below.
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{lcccl}
\toprule
PDEBench audit file & FNO future nRMSE & no-backprop selected & Gain vs FNO & Selected expert \\
\midrule
\texttt{test\_01} & $0.004911$ & $0.004824$ & $1.8\%$ & base spectral \\
\texttt{test\_02} & $0.001513$ & $0.076136$ & $-4931.1\%$ & liquid spectral \\
\texttt{test\_03} & $0.001801$ & $0.001412$ & $21.6\%$ & persistence \\
\texttt{test\_07} & $0.003583$ & $0.002440$ & $31.9\%$ & base spectral \\
\texttt{test\_08} & $0.004866$ & $0.003091$ & $36.5\%$ & liquid spline \\
\texttt{test\_09} & $0.005087$ & $0.009067$ & $-78.2\%$ & liquid spectral \\
\texttt{test\_10} & $0.004532$ & $0.264953$ & $-5746.4\%$ & autoreg spectral \\
\texttt{test\_11} & $0.003982$ & $0.003134$ & $21.3\%$ & base spectral \\
\texttt{test\_13} & $0.000451$ & $0.001221$ & $-171.0\%$ & persistence \\
\texttt{test\_16} & $0.000976$ & $0.000938$ & $3.9\%$ & liquid spline \\
\bottomrule
\end{tabular}
\caption{MPS no-backprop direct sequence continuation on scalar 1D PDEBench/FNO audit artifacts after the liquid/pole/spline routing extension. The model observes the first eight true frames and predicts all remaining frames (thirteen frames for the 21-step scalar rows; thirty-three frames for the 41-step \texttt{test\_10} row). All rows use training samples $0$--$849$, validation samples $850$--$899$, and held-out test samples $900$--$999$. The result is a genuine non-Fashion no-backprop hit but not a universal replacement: six scalar rows beat the cached FNO future error, several failures improve materially, and \texttt{test\_10}/\texttt{test\_13} remain unresolved.}
\label{tab:pdebench-direct-sequence}
\end{table}
The boundary is equally important. The liquid/pole/spline extension improves, but does not solve, the coupled three-channel artifacts. The old direct runner gave \texttt{test\_05}=0.112862 and \texttt{test\_06}=0.052174; the mixed liquid/spline grid with validation top-3 averaging improves these to 0.092994 and 0.046179, respectively, versus FNO 0.003861 and 0.006197. The coupled gain is real (about 17.6\% and 11.5\% relative error reduction against the old no-backprop direct rows), but the remaining FNO gap is still too large for a SOTA claim. The next autonomous audit asked whether this was a validation-split, temporal-decoder, or locality problem. It was not. A shuffled validation split on the same scalar runner still failed to rescue the selector: on \texttt{test\_13}, FNO is 0.000451 while validation selects persistence at 0.001221 although the direct row is 0.001416; on \texttt{test\_03}, validation again selects persistence at 0.001412 while the direct row is 0.001413. Adding compressed future-time DCT decoding through \texttt{--future\_time\_ranks} gives \texttt{test\_10}=0.271725 at rank $8$, still worse than the previous autoregressive hard-row frontier near 0.264953, and a six-row scalar audit selects no temporal compression at all. A global RBF/kernel memory over observed histories is worse again on the hard row, reaching only 0.300106. Finally, the new local runner \texttt{apps\_industrial\_breakthrough/pdebench\_1d\_local\_direct\_forecaster\_mps.py} predicts each spatial point from a periodic local stencil, velocity, local derivatives, liquid pole states, spline hinges, global low-mode context, and optional online per-trajectory PDE coefficient tokens. Its best hard-row run reaches 0.272812 on \texttt{test\_10}; on the coupled rows it reaches 0.096646 and 0.085312, improving persistence but trailing the earlier global spectral/liquid rows. We then made the proposed regime idea explicit: the same local runner can fit separate ridge operators for quantile bins of local shock, curvature, velocity, or transport score, and can also identify per-trajectory local PDE-library coefficients from the observed frames and roll them forward autonomously. On the full \texttt{test\_10} hard-row confirmation, validation still selects the original local map at 0.282656 validation and 0.273743 held-out test; the best shock-regime row is slightly worse at 0.283498 validation, and the best online PDE rollout is much worse at 0.319173 validation. The conclusion is now stronger: the missing mechanism is not simply random validation leakage, a compressed future basis, kernel memory, local finite-difference tokens, quantile shock partitioning, or stepwise local PDE coefficient rollout. The current direct solver needs a genuinely conservative flux/discontinuity operator or a different endogenous architecture, not another ridge feature expansion.
We also added \texttt{apps\_industrial\_breakthrough/pdebench\_2d\_direct\_spectral\_forecaster\_mps.py}, which extends the architecture to 2D tensor-DCT tokens,
\[
a_{n,t,k,\ell,c}=\sum_{i,j}\Psi^{(H)}_{i,k}\Psi^{(W)}_{j,\ell}Y_{n,i,j,t,c},
\]
and uses the same closed-form future-coefficient solve plus identity residual expert gate. On \texttt{test\_27} with history $12$, rank $32$, train/validation/test split $70/10/20$, the 2D direct model improves as rank grows but still reaches only $0.009609$ against FNO $0.001856$; porting the 1D liquid pole/quadratic feature grid to the 2D tensor-DCT runner leaves validation on the original base row. On \texttt{test\_26}, the same 2D grid is worse: direct spectral reaches 0.554028 and validation falls back to linear extrapolation at 0.080726 versus FNO 0.002635. Additional negative controls on \texttt{test\_10} show that higher DCT rank, explicit real Fourier bases, more observed history, sliding-window autoregressive ridge training, global shift/transport extrapolation, and projected PDE-library features all plateau near 0.265--0.267, far from FNO. The same mixed base/liquid/PDE feature grid beats persistence but remains far from FNO on additional vector rows: \texttt{test\_17}=0.037810 versus FNO 0.001431, \texttt{test\_19}=0.111576 versus FNO 0.008693, and \texttt{test\_20}=0.041286 versus FNO 0.004741. Thus the current publishable claim is narrow and precise: no-backprop OSNR-style spectral sequence continuation can beat cached FNO on several scalar PDEBench dynamics and improve hard failures, but shock-like long-horizon rows, coupled systems, and 2D fields require a stronger endogenous architecture than one global ridge map from history tokens to future coefficients. The next operator-learning target should combine this direct spectral solver with local conservation-law blocks, channel-coupled low-mode poles, and validation-stable regional expert gates.
\subsection{External PDEBench Darcy representation compression audit}
The same hosted FNO artifact contains five static Darcy-flow rows, Tests~21--25. Since the artifact exposes only FNO predictions and targets, not the Darcy coefficient input fields, we do not claim blind PDE solving in this audit. Instead, \texttt{apps\_industrial\_breakthrough/pdebench\_darcy\_compression\_audit.py} measures a representation question: how compactly can an OSNR-style global operator dictionary encode the target solution fields compared with the error of the trained FNO prediction?
For each target field $u$, we retain a compact rectangular set of Fourier/operator coefficients,
\[
\hat u_K = \mathcal{F}^{-1}\!\left[M_K(k_x,k_y)\mathcal{F}u\right],
\]
and report relative $L^2$ nRMSE. Table~\ref{tab:pdebench-darcy-compression} shows that the Darcy targets are extremely compressible. On the hardest row, Test~21, the hosted FNO prediction has nRMSE $0.2670$, while OSNR target representation reaches $0.2039$ with only $0.096\%$ of spectral bins, $0.0730$ with $0.385\%$, and $0.0396$ with $0.865\%$. For the easier Darcy rows, FNO is already strong, but OSNR still passes the FNO error level with a small coefficient budget: $K=8$ for Test~23 and $K=16$ for Tests~24--25.
\begin{table}[h]
\centering
\small
\begin{tabular}{lcccc}
\toprule
PDEBench row & FNO nRMSE & First OSNR $K$ beating FNO & Coefficient fraction & OSNR nRMSE \\
\midrule
Darcy Test~21 & $0.267029$ & $2$ & $0.096\%$ & $0.203930$ \\
Darcy Test~22 & $0.117649$ & $4$ & $0.385\%$ & $0.075482$ \\
Darcy Test~23 & $0.027689$ & $8$ & $1.538\%$ & $0.027178$ \\
Darcy Test~24 & $0.011536$ & $16$ & $6.154\%$ & $0.009421$ \\
Darcy Test~25 & $0.009500$ & $16$ & $6.154\%$ & $0.009427$ \\
\bottomrule
\end{tabular}
\caption{External PDEBench Darcy target-representation compression audit. This table measures compact representation of the target solution fields, not blind prediction from Darcy coefficients.}
\label{tab:pdebench-darcy-compression}
\end{table}
\begin{figure}[h]
\centering
\includegraphics[width=0.92\linewidth]{../apps_industrial_breakthrough/pdebench_darcy_compression_outputs/darcy_test21_compression_panel.png}
\caption{PDEBench Darcy Test~21 compression panel. A tiny low-frequency OSNR coefficient set captures the smooth elliptic solution structure more accurately than the hosted FNO prediction error, but this is a target-representation result rather than a complete Darcy solver.}
\label{fig:pdebench-darcy-compression}
\end{figure}
This result identifies a strong but precise opportunity. For elliptic PDE outputs, the operator-spline basis has excellent compression power on real external benchmark tensors. To turn this into a SOTA solving claim, the next experiment must ingest Darcy coefficient fields and solve or learn the coefficient-to-solution operator directly; target compression alone is not enough.
We began this coefficient-to-solution step using the PhysArena/PDEBench Darcy Parquet mirror, which exposes both the diffusion coefficient and the flow target. A simple finite-volume CG solve of $-\nabla\cdot(a\nabla u)=0.01$ with homogeneous Dirichlet boundaries recovers the spatial shape of many samples very accurately after an oracle scalar alignment: on a $120$-sample pilot, oracle-scaled relative error averages $0.0473$ with median $0.0243$. However, the required scalar varies strongly with coefficient geometry. A CNN trained on coefficient images to predict this scalar improves the median held-out error but leaves large outliers, with robust training still giving mean error about $0.305$ versus oracle $0.060$. A larger hybrid attempt using $3000$ samples and an MPS residual CNN from $(a,u_{\mathrm{CG}},x,y)$ to $u$ also failed, increasing held-out error to $1.239$ compared with $0.317$ for the raw CG shape and $0.055$ for oracle-scaled CG. We therefore do not yet claim a Darcy solver win. The next step is to recover the exact dataset forcing/normalization convention from the Plaid metadata or learn in constrained coefficient/operator space rather than using an unconstrained image residual network.
The direct operator is nevertheless already a strong solver in a clearly defined regime. In \texttt{apps\_industrial\_breakthrough/pdebench\_darcy\_operator\_regime\_audit.py}, we evaluate $1000$ real coefficient fields from the same external shard and stratify by the fraction of low-conductivity cells. Table~\ref{tab:pdebench-darcy-operator-regime} reports the result. When low-conductivity inclusions occupy less than $25\%$ of the domain, the raw finite-volume CG solve reaches mean nRMSE below $6.2\times10^{-4}$. For the intermediate $25$--$50\%$ regime, it remains useful with mean nRMSE $0.0654$ and median $0.0236$. The failure begins once low-conductivity cells dominate the grid, which is exactly where the simple arithmetic-face cell-centered stencil diverges from the external generator's apparent discretization.
We then stress-tested that failure with \texttt{apps\_industrial\_breakthrough/pdebench\_darcy\_operator\_stress\_audit.py}. The goal was to determine whether the high-inclusion collapse was a shallow implementation issue. A generalized face-transmissibility sweep over power means $p\in\{-4,-2,-1,-0.5,0,0.5,1,2,4\}$ and boundary scales $\{0.5,1,2\}$ did not fix the regime: the best training-screen mean was $0.3300$ and the high-inclusion mean remained $0.5664$. A held-out jump-aware residual correction using local features $(u,|\nabla a|u,|\nabla a|,\Delta u,au,\mathbf{1}_{a<0.5}u,1)$ was unstable, increasing held-out mean error from $0.2931$ to between $4.525$ and $7.259$ depending on ridge strength. Finally, high-contrast coefficient convention tests rejected simple metadata mistakes: the original coefficient field gave oracle-scaled high-contrast error $0.1327$, while inverse coefficients, vertical flips, horizontal flips, and rotations worsened to $0.4069$, $0.2717$, $0.3174$, and $0.3096$, respectively. This negative result is useful. It says that the next Darcy improvement should not be another scalar transmissibility tweak or unconstrained residual network; it should recover the exact generator discretization or move to a constrained multiscale/interface operator that preserves ellipticity.
We also tested the more realistic hybrid idea: keep the direct operator solve, but train a small neural module only for the unknown correction. In \texttt{apps\_industrial\_breakthrough/darcy\_neural\_osnr\_pde\_layer.py}, a compact encoder observes the coefficient field and the direct CG solution. Three heads are compared on held-out Darcy fields: a direct low-resolution residual decoder, an OSNR/PDE-layer residual decoder that passes the predicted source through a screened Poisson inverse before upsampling, and a scalar normalization head. The correction is gated by the coefficient regime: low/mid-inclusion samples keep the direct operator output, while high-inclusion samples use the learned correction. On a $500$-sample smoke split with $350$ training samples and high-regime specialization, the base held-out mean nRMSE is $0.2876$. The gated OSNR/PDE residual reduces this to $0.2516$, the gated direct residual to $0.2536$, and the gated scale head to $0.2332$. In the hardest low-conductivity-dominant bin, the base mean error $0.8923$ drops to $0.7202$ with the scale head. This is not a final SOTA solver, but it is a real hybrid lesson: the current external Darcy gap is mostly a hidden sample-dependent normalization/interface convention, and a physics-gated neural correction is useful only when it respects the regimes already solved by the operator.
We then pushed directly on that convention gap with two additional real-data calibration experiments. First, \texttt{apps\_industrial\_breakthrough/pdebench\_darcy\_interface\_calibration\_search.py} evaluated $21$ positive symmetric face-transmissibility variants on a $360/240/120$ train/test split. The best learned calibration used a maximum-face law and reduced the held-out mean from $0.3715$ to $0.3416$, with the high-inclusion bins improving from $0.5535$ to $0.4752$ and from $0.8539$ to $0.7386$. More importantly, its oracle-scaled held-out error was only $0.0822$, showing that the operator family can produce a much better shape than our deployable scalar calibration recovers. Second, we tested whether the missing scalar is easily recoverable from coefficient geometry. A hand-feature kNN calibration in \texttt{apps\_industrial\_breakthrough/pdebench\_darcy\_geometry\_scale\_knn.py} failed, increasing gated mean error from $0.2876$ to $0.3678$. Supervised scalar-head variants that explicitly regress the oracle alignment scalar also failed to beat the earlier $0.2332$ gated scale result. These negative results narrow the remaining problem: the bottleneck is not generic network capacity or a simple geometry-to-scale map, but recovery of the external generator's discretization/normalization law or a richer constrained multiscale elliptic operator.
The same diagnosis suggests a different, stronger problem formulation: sparse-sensor PDE assimilation. In real deployments one often has a small number of pressure/head/flow probes but cannot afford dense field acquisition. In \texttt{apps\_industrial\_breakthrough/pdebench\_darcy\_sparse\_sensor\_assimilation.py}, the coefficient field defines the elliptic OSNR/PDE shape, and $m$ point observations determine the remaining amplitude by the closed-form least-squares scalar
\begin{equation}
\widehat{s}_m = \frac{\sum_{q\in\Omega_m} u_{\mathrm{CG}}(q)u_{\mathrm{obs}}(q)}{\sum_{q\in\Omega_m} u_{\mathrm{CG}}(q)^2+\epsilon}.
\end{equation}
On a larger $1000/700/300$ real PDEBench Darcy split, the blind direct operator has mean nRMSE $0.3262$. A single interior sensor reduces the error to $0.0889$, four sensors reduce it to $0.0681$, and $32$ sensors reach $0.0597$, close to the full-field oracle scalar ceiling $0.0585$. In the hardest low-conductivity-dominant bin, the same four-sensor assimilation reduces mean error from $0.8642$ to $0.1460$, while $32$ sensors reach $0.1304$ against an oracle of $0.1288$. This result is substantially stronger than the learned residual attempts: it uses no neural training, preserves the elliptic operator, and converts a failed blind coefficient-to-solution setting into a practical sparse-observation reconstruction problem.
\begin{table}[h]
\centering
\small
\begin{tabular}{lcc}
\toprule
Method & Mean held-out nRMSE & Hard-bin nRMSE \\
\midrule
Blind direct CG operator & $0.3262$ & $0.8642$ \\
1 sparse sensor & $0.0889$ & $0.2030$ \\
4 sparse sensors & $0.0681$ & $0.1460$ \\
16 sparse sensors & $0.0606$ & $0.1327$ \\
32 sparse sensors & $0.0597$ & $0.1304$ \\
Full-field oracle scalar & $0.0585$ & $0.1288$ \\
\bottomrule
\end{tabular}
\caption{Real PDEBench Darcy sparse-sensor assimilation on $300$ held-out coefficient fields. A few point observations close most of the blind-operator amplitude gap without neural training.}
\label{tab:pdebench-darcy-sparse-sensor}
\end{table}
\begin{table}[h]
\centering
\small
\begin{tabular}{lccc}
\toprule
Low-conductivity fraction & Samples & Raw CG nRMSE & Oracle-scaled nRMSE \\
\midrule
$0$--$0.10$ & $11$ & $0.000366$ & $0.000054$ \\
$0.10$--$0.25$ & $86$ & $0.000615$ & $0.000271$ \\
$0.25$--$0.50$ & $411$ & $0.065379$ & $0.009759$ \\
$0.50$--$0.75$ & $351$ & $0.471048$ & $0.074186$ \\
$0.75$--$1.00$ & $140$ & $0.876031$ & $0.114951$ \\
\bottomrule
\end{tabular}
\caption{External PhysArena/PDEBench Darcy coefficient-to-solution operator regime audit on $1000$ real coefficient fields. This is a true solver experiment, not target compression. The direct operator is highly accurate in low/mid contrast regimes and fails when low-conductivity inclusions dominate.}
\label{tab:pdebench-darcy-operator-regime}
\end{table}
\begin{figure}[h]
\centering
\includegraphics[width=0.92\linewidth]{../apps_industrial_breakthrough/pdebench_darcy_operator_regime_outputs/darcy_operator_easy_panel.png}
\caption{External Darcy coefficient-to-solution example in the low-inclusion regime. The direct operator solve reproduces the target flow without neural training.}
\label{fig:pdebench-darcy-operator-easy}
\end{figure}
\subsection{Weather-core validation: coupled rotating shallow water}
The next weather-facing benchmark is \texttt{apps\_industrial\_breakthrough/shallow\_water\_operator\_spline\_benchmark.py}. This moves beyond scalar advection--diffusion and Burgers equations to a coupled three-component linearized rotating shallow-water core. The unknown state is
\[
\mathbf{q}(t,x,y)=(\eta,u,v)^\top,
\]
where $\eta$ is the height anomaly and $(u,v)$ are horizontal velocities. The periodic operator is
\begin{align}
\eta_t + \mu\eta + H(u_x+v_y) &= s_\eta,\\
u_t + r u - f v + g\eta_x &= s_u,\\
v_t + r v + f u + g\eta_y &= s_v.
\end{align}
Here $H$ is mean depth, $g$ is gravity, $f$ is the Coriolis parameter, $r$ is velocity damping, and $\mu$ is a small height-relaxation gauge that removes the resonant zero-frequency mass mode. In Fourier space, every $(\omega,k_x,k_y)$ bin is a dense $3\times3$ complex linear system. OSNR solves the full coupled block by batched frequency-bin inversion, not by fitting a neural coordinate model or by stepping a recurrent simulator.
The forcing combines smooth planetary-wave structure with sparse localized height/vorticity impulses. The sparse impulses are treated as storm/front innovations in the operator domain. A low-pass spectral reconstruction is included as a smooth surrogate baseline. Table~\ref{tab:shallow-water-weather-core} reports the current results.
\begin{table}[h]
\centering
\small
\begin{tabular}{lccccc}
\toprule
Profile & Grid & OSNR PSNR & Low-pass PSNR & Atom error & Solve time \\
\midrule
Main, $24$ events & $32\times64^2$ & $33.2500$ dB & $29.0210$ dB & $0.0000$ px & $13.09$ ms \\
Dense, $48$ events & $32\times64^2$ & $34.9617$ dB & $27.1203$ dB & $0.3363$ px & $13.08$ ms \\
Heavy, $96$ events & $32\times64^2$ & $34.1593$ dB & $26.2245$ dB & $1.2163$ px & $12.85$ ms \\
$0.5\%$ noise, raw & $32\times64^2$ & $33.0138$ dB & $29.0210$ dB & $0.0000$ px & $13.08$ ms \\
$1.0\%$ noise, smoothed & $32\times64^2$ & $31.8288$ dB & $29.0210$ dB & $0.0000$ px & $13.32$ ms \\
$2.0\%$ noise, smoothed & $32\times64^2$ & $31.8257$ dB & $29.0210$ dB & $0.0000$ px & $13.84$ ms \\
Scale stress & $48\times96^2$ & $34.5164$ dB & $28.9019$ dB & $0.0711$ px & $40.55$ ms \\
Scale stress, $96$ events, $0.5\%$ noise & $48\times96^2$ & $32.3981$ dB & $27.7452$ dB & $0.4625$ px & $37.83$ ms \\
\bottomrule
\end{tabular}
\caption{Linearized rotating shallow-water weather-core benchmark. OSNR solves the coupled height/velocity operator by batched $3\times3$ frequency-bin inversions. The low-pass row is a smooth spectral-bias baseline. Atom error measures sparse storm/front innovation recovery from the height forcing channel after operator-domain residual extraction.}
\label{tab:shallow-water-weather-core}
\end{table}
\begin{figure}[h]
\centering
\includegraphics[width=0.92\linewidth]{../apps_industrial_breakthrough/shallow_water_operator_spline_outputs_48x96x96/shallow_water_operator_spline.png}
\caption{Coupled shallow-water OSNR benchmark at $48\times96^2$. The figure shows target height, OSNR reconstruction, low-pass baseline, recovered height innovation, sparse event atoms, and velocity magnitude. Unlike the scalar tests, the solve couples height and both velocity components through Coriolis and pressure-gradient terms.}
\label{fig:shallow-water-weather-core}
\end{figure}
This is the first result in the manuscript that begins to resemble a real weather core. It is still not a direct GraphCast/GenCast/WeatherNext comparison: those systems operate on global ERA5-scale atmospheric states and are trained on decades of data. The scientific significance is narrower but important. OSNR can invert a physically coupled, multi-variable periodic atmospheric operator in milliseconds, preserve sparse front/storm innovations, and outperform a smooth low-pass surrogate under noise. The degradation is graceful: even the $96$-event, $0.5\%$ noisy $48\times96^2$ stress case remains above $32$ dB with sub-pixel atom error. The next hard step is to leave the linearized core and introduce nonlinear advection, partial observations, and data assimilation windows while preserving this operator-domain sparse innovation advantage.
We also tested a first partial-observation assimilation variant in which only $\eta$ is treated as observed. A naive geostrophic lift fails badly because the synthetic state contains wave and forced components outside static balance. A dynamic momentum lift performs much better: given the observed $\eta$, it solves the two Fourier-domain momentum equations for $(u,v)$ while assuming small direct velocity forcing. This height-only lift reaches $33.2250$ dB on the $32\times64^2$ case and $35.1743$ dB on the $48\times96^2$ scale case, so balanced field reconstruction remains plausible from partial observations. However, sparse event localization degrades to $5.8550$ px and $10.9152$ px, respectively. The lesson is precise: partial state assimilation can recover smooth balanced dynamics, but front/storm innovation recovery needs an explicit sparse assimilation stage rather than a purely balanced velocity closure.
\subsection{Coupled weather-core operator identification}
The scalar operator-identification experiments show that field-only discovery is underdetermined, while sparse physical anchors make the problem well posed. We repeated the same idea on the coupled shallow-water core in \texttt{apps\_industrial\_breakthrough/shallow\_water\_operator\_identification.py}. The unknown parameter vector is now
\[
\theta=(\mu,H,r,f,g),
\]
corresponding to height relaxation, mean depth, velocity damping, Coriolis coupling, and gravity. Given observed $(\eta,u,v)$ and sparse samples of the forcing channels, the pointwise equations are linear in $\theta$:
\begin{align}
s_\eta-\eta_t &= \mu\eta + H(u_x+v_y),\\
s_u-u_t &= r u - f v + g\eta_x,\\
s_v-v_t &= r v + f u + g\eta_y.
\end{align}
Thus the multi-channel operator is recovered by one real ridge least-squares solve over analytic derivative columns. The important correction is the continuity-column term $u_x+v_y$; omitting $v_y$ makes $H$ unidentifiable.
\begin{table}[h]
\centering
\small
\begin{tabular}{lccccccc}
\toprule
Setting & Anchors & $\hat\mu$ & $\hat H$ & $\hat r$ & $\hat f$ & $\hat g$ & PSNR \\
\midrule
Clean, $1.0\%$ anchors & $3933$ & $0.1817$ & $0.9689$ & $0.0822$ & $0.7988$ & $0.9999$ & $25.7477$ dB \\
$0.1\%$ noise, $\sigma=0.5$, $0.5\%$ anchors & $1965$ & $0.2108$ & $0.9924$ & $0.0865$ & $0.8036$ & $0.9999$ & $31.3868$ dB \\
$0.2\%$ noise, $\sigma=0.75$, $0.5\%$ anchors & $1965$ & $0.2435$ & $1.0068$ & $0.0929$ & $0.8017$ & $1.0001$ & $32.5423$ dB \\
$0.2\%$ noise, $\sigma=0.75$, $1.0\%$ anchors & $3933$ & $0.2162$ & $1.0086$ & $0.0886$ & $0.7973$ & $1.0002$ & $31.4024$ dB \\
$0.2\%$ noise, $\sigma=0.75$, $2.0\%$ anchors & $7863$ & $0.2179$ & $1.0048$ & $0.0933$ & $0.7966$ & $1.0000$ & $34.2686$ dB \\
\bottomrule
\end{tabular}
\caption{Coupled shallow-water operator identification with true parameters $(\mu,H,r,f,g)=(0.2,1.0,0.08,0.8,1.0)$. Multi-channel derivative columns identify the physical operator from sparse forcing anchors.}
\label{tab:shallow-water-operator-identification}
\end{table}
\begin{figure}[h]
\centering
\includegraphics[width=0.92\linewidth]{../apps_industrial_breakthrough/shallow_water_operator_identification_outputs_noise002_denoise075_fixed/shallow_water_operator_identification.png}
\caption{Coupled shallow-water operator identification at $0.2\%$ observation noise with $\sigma=0.75$ pre-smoothing. Multi-channel physics anchors recover the operator and reconstruct the weather-core state without training a coordinate network.}
\label{fig:shallow-water-operator-identification}
\end{figure}
This result is stronger than the scalar anchor test in two ways. First, the multi-channel structure anchors the coupling constants $f$ and $g$ very tightly. Second, even with observation noise, small forcing-anchor fractions recover the coupled state above $31$ dB. The remaining weak parameter is $\mu$, the artificial height-relaxation gauge, because it is weakly excited relative to the wave and forcing terms. This suggests a practical design rule for weather-grade OSNR: learn physically meaningful coupling symbols from multi-channel states, and treat gauge/damping terms with explicit priors or assimilation-window constraints.
\section{FRI-Guided Adaptive Sparse Tier}
Uniform sparse dictionaries fail at sub-pixel discontinuities because a step located at $\tau \notin T\Z$ cannot be represented by a finite block of rigid grid atoms without tail error or leakage. FRI theory instead recovers the innovation coordinate first \cite{vetterli2002fri,dragotti2007moments}.
For a stream of $K$ weighted Diracs,
\[
w(t) = \sum_{k=1}^K a_k \delta(t-\tau_k),
\]
moments satisfy
\[
m_\ell = \int t^\ell w(t)\,\dd t
= \sum_{k=1}^K a_k \tau_k^\ell.
\]
The annihilating filter or matrix-pencil method recovers the roots $\tau_k$. Once $\tau_k$ are known, sparse step atoms are snapped exactly to those locations. This changes the sparse tier from an approximation grid into an adaptive representation of the true innovation geometry.
\section{Hybrid Sparse-Plus-Smooth Decomposition}
Let
\[
\A_s = \A_{\mathrm{smooth}},
\qquad
\A_x = \A_{\mathrm{sparse}}.
\]
The naive alternating update
\[
\cvec_s
=
\argmin_{\cvec}
\|\yvec - \A_x \zvec - \A_s \cvec\|_2^2
\]
is not wrong by itself, but it is incomplete if implemented as if the two bases are orthogonal. The full block normal equations contain cross terms:
\[
\begin{bmatrix}
\A_s^\top \A_s & \A_s^\top \A_x \\
\A_x^\top \A_s & \A_x^\top \A_x
\end{bmatrix}
\begin{bmatrix}
\cvec_s \\ \cvec_x
\end{bmatrix}
=
\begin{bmatrix}
\A_s^\top \yvec \\ \A_x^\top \yvec
\end{bmatrix}.
\]
The cross-Gram matrix
\[
\A_{\mathrm{cross}} = \A_s^\top \A_x
\]
must appear directly in the right-hand side of block updates:
\[
(\A_s^\top \A_s)\cvec_s
=
\A_s^\top \yvec - \A_{\mathrm{cross}}\zvec,
\]
\[
(\A_x^\top \A_x+\rho\I)\cvec_x
=
\A_x^\top \yvec - \A_{\mathrm{cross}}^\top \cvec_s
+ \rho \zvec - \uvec.
\]
This prevents smooth atoms from absorbing sparse shocks and prevents sparse atoms from chasing smooth energy. It is the finite-dimensional expression of the hybrid-spline coupling described by Debarre, Aziznejad, and Unser \cite{debarre2019hybrid,debarre2021composite}.
\section{Failure Modes and Corrections}
\subsection{Uncalibrated knot grids}
Defect: using weight $1$ and spreading biases over a fixed interval. Correction: use $v_k(x)=x/T-k$.
\subsection{Sequential smooth-first fitting}
Defect: solving the smooth component first lets the smooth basis approximate discontinuities through oscillatory combinations, producing Gibbs residuals. Correction: use joint or cross-Gram-shielded updates.
\subsection{Joint coherent dictionaries}
Defect: concatenating coherent dictionaries and applying naive ADMM can allocate smooth energy into sparse atoms and vice versa. Correction: explicitly include cross-Gram blocks and stabilize each block solve.
\subsection{Partition-of-unity trap}
Spline systems reproduce constants. A contiguous block of step atoms also reproduces a constant over an interval. This is not merely high coherence; it is an identifiability collision. Let
\[
\chi_m(x)=\mathbf{1}_{x\geq \tau_m}
\]
be step atoms sorted by their knot locations. On any interval tiled by an active adjacent block, a difference or finite linear combination of these atoms can reproduce an indicator plateau
\[
\mathbf{1}_{[\tau_a,\tau_b)}(x)
=
\chi_a(x)-\chi_b(x).
\]
Inside the plateau support, this function is exactly constant. At the same time, valid cardinal and exponential spline spaces satisfy partition-of-unity conditions and reproduce the global constant mode. Hence, after restriction to a local active support, the sparse step block and the smooth E-spline block contain indistinguishable constant directions.
During active-support debiasing,
\[
\A_f = [\A_s \mid \A_{x,\mathrm{active}}],
\]
the columns can become locally indistinguishable. Then $\A_f^\top \A_f$ is singular or nearly singular. This is the numerical origin of the observed coefficient explosions when unregularized \texttt{lstsq} was applied to a joint smooth-plus-step active set.
Correction: use a scale-invariant Tikhonov solve, consistent with ADMM/proximal regularization views of Bayesian denoising \cite{nguyen2018regularizers},
\[
\cvec_f =
(\A_f^\top \A_f + \gamma \I)^{-1}\A_f^\top \yvec,
\]
with
\[
\gamma =
\epsilon \, \mathrm{mean}(\mathrm{diag}(\A_f^\top \A_f)).
\]
In the production sparse core, $\epsilon=10^{-6}$. This makes the ridge invariant to the absolute scaling of the dictionary and shifts the zero singular directions by an amount proportional to the local Gram energy. The FRI stage further reduces the degeneracy by snapping sparse knots to physical innovation coordinates before debiasing, so the active sparse columns describe true shock interfaces rather than a diffuse uniform-grid approximation. In the batched Sprint 3 validation, the resulting ridge-stabilized debiasing matrices remain bounded with maximum condition number $18.1639<19.0$ while preserving $118.78$ dB mean reconstruction precision and $97.7\%$ hard-zero sparse parameters.
\section{Benchmarks}
The current prototype benchmark scripts demonstrate:
\begin{itemize}[leftmargin=2em]
\item calibrated Helmholtz coefficient recovery with zero autograd graph construction;
\item cascading derivative evaluation with machine-precision PDE residuals;
\item FFT/circulant inversion at $O(M\log M)$ complexity;
\item adaptive FRI sparse shock recovery with snapped knots and ridge-stabilized debiasing;
\item batched multi-edge sparse core execution;
\item second-order Hermite neural-operator coefficient recovery through block-circulant $3\times3$ Fourier solves;
\item 2D tensor-product Hermite fluid simulation through block-circulant $9\times9$ Fourier solves;
\item first SOTA-style SIREN comparison on a multi-edge non-bandlimited silhouette.
\end{itemize}
\subsection{Benchmark environment}
All timings in this draft were measured on a local Apple Silicon workstation. The benchmark scripts were executed as standalone Python processes from the repository virtual environment using \texttt{torch.no\_grad()} for every numerical solve. No PyTorch backward pass was constructed in any benchmark.
\begin{table}[h]
\centering
\begin{tabular}{@{}ll@{}}
\toprule
Component & Configuration \\
\midrule
CPU / SoC & Apple M4 Max \\
Memory & 128 GiB unified memory \\
Architecture & arm64 \\
Operating system & macOS 15.7.4, build 24G517 \\
Python & 3.14.3 \\
PyTorch & 2.12.0 \\
NumPy & 2.4.6 \\
LaTeX compiler & Tectonic 0.16.9 \\
\bottomrule
\end{tabular}
\caption{Hardware and software environment used for the current OSNR benchmark run.}
\end{table}
\subsection{Numerical results}
Table~\ref{tab:benchmarks} records the current measured outputs of the repository benchmark scripts after the calibrated-grid, cross-Gram, FFT, FRI, Hermite block-Gram, and ridge-stabilization corrections. The scripts live under the repository's benchmark and source directories. The timings are wall-clock processing times reported by the scripts and should be interpreted as prototype measurements rather than final library-level performance claims.
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{@{}p{0.28\linewidth}p{0.32\linewidth}p{0.32\linewidth}@{}}
\toprule
Tier & Mechanism & Official package validation profile \\
\midrule
Tier 1 steady-state & Pole-locked trigonometric E-splines plus circulant Fourier division & $314.86$ dB precision at $4.21$ ms latency \\
Tier 2 adaptive sparse & TLS matrix-pencil FRI, snapped step knots, cross-Gram ADMM, ridge debiasing & $118.78$ dB precision; $97.7\%$ hard parameter zeros; max edge error $7.627232\mathrm{e}{-14}$ \\
Tier 3 Hermite neural operator & Multistream Hermite Gram tensors and parallel DFT block solves & $183.77$ dB coefficient precision; $1.28$ ms 2D CFD frame latency; boundary residual $0.000000\mathrm{e}{+00}$ \\
\bottomrule
\end{tabular}
\caption{Current official validation profile of the compiled OSNR package cores.}
\label{tab:portfolio}
\end{table}
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{@{}p{0.24\linewidth}p{0.11\linewidth}p{0.14\linewidth}p{0.39\linewidth}p{0.06\linewidth}@{}}
\toprule
Benchmark & Time (ms) & PSNR (dB) & Residual / condition / sparsity & Autograd \\
\midrule
Helmholtz OSNR & 12.5436 & 294.38 & PDE residual $0.000000\mathrm{e}{+00}$ & 0 B \\
Cascading derivative bank & 14.7091 & 303.05 & derivative condition $1.0000$; PDE residual $2.546653\mathrm{e}{-14}$ & 0 B \\
Circulant FFT solver & 3.3497 & 316.43 & $O(M\log M)$; approx. $72192$ flops & 0 B \\
FRI shock tracker & 13.1712 & 124.84 & ridge condition $1.2847$; sparsity $99.2\%$ & 0 B \\
Batched adaptive sparse core & 9.9200 & mean $118.78$, min $116.73$ & max condition $18.1639$; edge error $7.627232\mathrm{e}{-14}$; sparsity $97.7\%$ & 0 B \\
Hermite operator core & 5.1442 & mean $157.60$ & boundary residual $0.000000\mathrm{e}{+00}$; max $3\times3$ condition $43199.1974$ & 0 B \\
2D tensor Hermite CFD & 1.2850 / frame & n/a & boundary leakage $0.000000\mathrm{e}{+00}$; clamp residual $0.000000\mathrm{e}{+00}$; final momentum residual $1.876066$ & 0 B \\
Multi-obstacle CFD cinema & 13.6396 / frame & n/a & $100$ frames; $20$ PNG exports; boundary leakage $0.000000\mathrm{e}{+00}$; clamp residual $0.000000\mathrm{e}{+00}$; final momentum residual $8.212812$ & 0 B \\
Graphics super-resolution & 110.5804 & 101.61 & $512\times512$ canvas; $K=6$ edges/scanline; edge error $1.674321\mathrm{e}{-08}$; sparsity $96.2\%$ & 0 B \\
Aerodynamic wind tunnel & 8.6717 / frame & n/a & $\operatorname{Re}=50{,}000$; $50$ frames; $5$ state exports; boundary leakage $0.000000\mathrm{e}{+00}$; clamp residual $0.000000\mathrm{e}{+00}$ & 0 B \\
HF video challenger, small high-quality profile & 26.4591 / frame & 60.2304 & SSIM $0.999686$; LPIPS $0.000004$; combined sparsity $80.21\%$; sparse-tier sparsity $95.00\%$ & 0 B \\
HF video challenger, full strict-sparsity profile & 167.0024 / frame & 29.1223 & SSIM $0.805561$; LPIPS $0.243320$; combined sparsity $95.31\%$; sparse-tier sparsity $96.88\%$ & 0 B \\
HF video challenger, full quality-prioritized profile & 167.8295 / frame & 35.0321 & SSIM $0.958564$; LPIPS $0.033300$; combined sparsity $86.81\%$; sparse-tier sparsity $96.88\%$ & 0 B \\
HF video challenger, 2D residual practical cinema & 170.7790 / frame & 41.3874 & SSIM $0.950734$; LPIPS $0.021302$; combined sparsity $87.50\%$; sparse-tier sparsity $96.88\%$ & 0 B \\
HF video challenger, 2D residual ceiling cinema & 169.8067 / frame & 142.7187 & SSIM $1.000000$; LPIPS $0.000000$; combined sparsity $77.50\%$; sparse-tier sparsity $96.88\%$ & 0 B \\
DL3DV multi-view stack, sparse residual profile & 140.1238 / view & 20.4800 & $30\times256^2$ RGB views; SSIM $0.448128$; LPIPS $0.718625$; combined sparsity $96.92\%$; peak footprint $193.30$ MB & 0 B \\
DL3DV multi-view stack, quality ceiling & 134.5901 / view & 117.2378 & $30\times256^2$ RGB views; SSIM $1.000000$; LPIPS $0.000000$; combined sparsity $77.50\%$; peak footprint $215.81$ MB & 0 B \\
DL3DV Pareto, $50\%$ 3D-DCT support & 136.9464 / view & 42.7562 & SSIM $0.981283$; LPIPS $0.002584$; combined sparsity $87.50\%$; max edge error $0.958244$ & 0 B \\
DL3DV Pareto, $25\%$ 3D-DCT support & 136.9298 / view & 35.5073 & SSIM $0.927023$; LPIPS $0.049832$; combined sparsity $92.50\%$; max edge error $0.958244$ & 0 B \\
Viscous Burgers, OSNR $+$ liquid residual & 1047.2 total & 29.4211 future & $128\times101$ periodic trajectory; future RMSE $3.085742\mathrm{e}{-02}$; full PSNR $31.8881$ dB; $32$ liquid states & small training graph \\
Biharmonic clamped plate & 42.9853 & n/a & boundary residual $0.000000\mathrm{e}{+00}$; relative operator residual $1.065298\mathrm{e}{-04}$; deflection RMS $4.697917\mathrm{e}{-08}$ & 0 B \\
\bottomrule
\end{tabular}
\caption{Measured benchmark results for the current OSNR prototype scripts.}
\label{tab:benchmarks}
\end{table}
\subsection{Interpretation}
The Helmholtz and cascading-derivative experiments isolate the Tier 1 deterministic case. Their high PSNR values and machine-precision PDE residuals confirm that, once the cardinal grid is calibrated, coefficient recovery in an operator-matched basis can replace iterative PINN-style optimization for this synthetic null-space task. The circulant experiment verifies the expected $O(M\log M)$ path when periodized shift-invariant structure is available.
The sparse experiments test the Tier 2 adaptive case. The single-edge FRI shock tracker recovers the discontinuity at $x=0.834200$, snaps the sparse atom to that location, and obtains $124.84$ dB PSNR with $99.2\%$ hard-zeroed sparse coefficients. The batched adaptive sparse core extends this to batch size $B=3$ with $K=3$ innovations per signal, achieving mean PSNR $118.78$ dB, minimum PSNR $116.73$ dB, maximum edge localization error $7.627232\mathrm{e}{-14}$, and ridge-stabilized condition number below $19$.
The Hermite operator experiment tests the Tier 3 coefficient-space trunk. A batched value/slope/curvature payload is passed through the exact Hermite block-circulant Gram system and recovered by independent $3\times3$ Fourier-domain solves. The boundary knots are overwritten with clamped value, slope, and curvature vectors, giving a structural boundary residual of exactly zero without a boundary loss term. The measured PSNR of $157.60$ dB confirms that the Hermite trunk can act as a deterministic neural-operator synthesis layer under the matched periodic benchmark model.
\subsection{Nonlinear PDE validation: Burgers PINN versus OSNR-liquid rollout}
The first nonlinear PINN-facing CFD experiment is \texttt{apps\_industrial\_breakthrough/burgers\_liquid\_pinn\_challenger.py}. It solves the viscous Burgers equation
\[
u_t + u u_x = \nu u_{xx},
\qquad x\in[-1,1),
\qquad \nu = 0.01/\pi,
\]
with periodic boundary conditions and initial condition $u(x,0)=-\sin(\pi x)$. A high-substep pseudospectral rollout provides the reference trajectory on a $128\times101$ space-time grid. The evaluation is intentionally split in time: the first $70\%$ of frames are available for fitting or residual correction, while the final $30$ frames are held out as a future prediction window.
The comparison has three profiles. The first is a compact SIREN-style coordinate PINN trained with Adam on data samples, initial-condition samples, periodic boundary consistency, and the Burgers residual computed by backward-mode automatic differentiation. The second is a truncated OSNR spectral rollout that keeps only a fixed number of Fourier/operator modes and evolves the known PDE directly without training. The third adds a tiny exact liquid residual in coefficient space. The liquid cell is trained only on the residual Fourier coefficients in the training window and uses the closed-form update
\[
x_{k+1}
=
\gamma_k\left(x_k-\frac{\sum_s A_s f_s(\mathbf{c}_k)}{\omega+\sum_s f_s(\mathbf{c}_k)}\right)
+
\frac{\sum_s A_s f_s(\mathbf{c}_k)}{\omega+\sum_s f_s(\mathbf{c}_k)}.
\]
Here $\mathbf{c}_k$ contains the truncated spectral coefficients and normalized time. This is a deliberately hybrid experiment: the spectral rollout remains no-autograd, while the liquid residual uses a small training graph to learn truncation-error compensation.
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{@{}p{0.28\linewidth}p{0.13\linewidth}p{0.13\linewidth}p{0.12\linewidth}p{0.14\linewidth}p{0.14\linewidth}@{}}
\toprule
Profile & Future PSNR & Future RMSE & Full PSNR & Latency & Autograd memory \\
\midrule
SIREN PINN, $1200$ epochs & $8.2031$ dB & $3.550245\mathrm{e}{-01}$ & $14.4821$ dB & $30229.3$ ms & est. $25{,}165{,}824$ B \\
OSNR spectral, $16$ modes & $19.1401$ dB & $1.007881\mathrm{e}{-01}$ & $21.6879$ dB & $7.0$ ms & $0$ B \\
OSNR $+$ liquid, $16$ modes & $19.8356$ dB & $9.303265\mathrm{e}{-02}$ & $23.5777$ dB & $852.3$ ms & small training graph \\
SIREN PINN, $800$ epochs & $7.6902$ dB & $3.766208\mathrm{e}{-01}$ & $13.9832$ dB & $19684.4$ ms & est. $25{,}165{,}824$ B \\
OSNR spectral, $32$ modes & $29.0669$ dB & $3.214163\mathrm{e}{-02}$ & $31.0045$ dB & $7.1$ ms & $0$ B \\
OSNR $+$ liquid, $32$ modes & $29.4211$ dB & $3.085742\mathrm{e}{-02}$ & $31.8881$ dB & $1047.2$ ms & small training graph \\
\bottomrule
\end{tabular}
\caption{First viscous Burgers CFD/PINN challenger. Metrics are computed on the held-out future window for the first two columns. The liquid rows train a tiny coefficient-space residual corrector; they are reported as hybrid OSNR-liquid profiles rather than zero-autograd deterministic solves.}
\label{tab:burgers-liquid-pinn}
\end{table}
\begin{figure}[h]
\centering
\includegraphics[width=0.24\linewidth]{../apps_industrial_breakthrough/burgers_liquid_pinn_outputs_modes32/burgers_target.png}
\includegraphics[width=0.24\linewidth]{../apps_industrial_breakthrough/burgers_liquid_pinn_outputs_modes32/burgers_osnr_coarse.png}
\includegraphics[width=0.24\linewidth]{../apps_industrial_breakthrough/burgers_liquid_pinn_outputs_modes32/burgers_liquid.png}
\includegraphics[width=0.24\linewidth]{../apps_industrial_breakthrough/burgers_liquid_pinn_outputs_modes32/burgers_pinn.png}
\caption{Burgers space-time heatmaps for the quality-first $32$-mode run. Left to right: reference trajectory, truncated OSNR spectral rollout, OSNR plus exact liquid coefficient residual, and SIREN PINN.}
\label{fig:burgers-liquid-pinn}
\end{figure}
The follow-up capacity sweep, \texttt{apps\_industrial\_breakthrough/burgers\_liquid\_capacity\_sweep.py}, separates three effects: retained operator modes, liquid hidden-state count, and forecast horizon. The result is not that arbitrarily larger liquid networks replace resolution. Instead, the dominant lever is still the operator basis. Increasing the spectral support from $8$ to $48$ retained modes raises the zero-training future PSNR from $11.6562$ to $38.2908$ dB on the $70\%/30\%$ train/future split. The liquid residual is most valuable when the basis is deliberately compressed: at $16$ modes and a $50\%/50\%$ split, it raises future PSNR from $19.5291$ dB to $22.2635$ dB with $128$ liquid states. At $48$ modes the same scaling gives only a sub-dB correction because little truncation error remains.
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{@{}p{0.14\linewidth}p{0.14\linewidth}p{0.16\linewidth}p{0.16\linewidth}p{0.15\linewidth}p{0.15\linewidth}@{}}
\toprule
Train fraction & Modes & Best hidden states & OSNR future PSNR & Liquid future PSNR & Gain \\
\midrule
$0.70$ & $8$ & $128$ & $11.6562$ dB & $14.3001$ dB & $+2.6439$ dB \\
$0.50$ & $8$ & $128$ & $12.7850$ dB & $15.8716$ dB & $+3.0866$ dB \\
$0.70$ & $16$ & $128$ & $19.1401$ dB & $20.6057$ dB & $+1.4656$ dB \\
$0.50$ & $16$ & $128$ & $19.5291$ dB & $22.2635$ dB & $+2.7344$ dB \\
$0.70$ & $32$ & $16$ & $29.0669$ dB & $29.4560$ dB & $+0.3891$ dB \\
$0.50$ & $32$ & $32$ & $29.0679$ dB & $30.3599$ dB & $+1.2920$ dB \\
$0.70$ & $48$ & $128$ & $38.2908$ dB & $38.8746$ dB & $+0.5838$ dB \\
$0.50$ & $48$ & $128$ & $37.9806$ dB & $38.2815$ dB & $+0.3009$ dB \\
\bottomrule
\end{tabular}
\caption{Burgers OSNR-liquid capacity sweep. The liquid cell is useful as a compact coefficient-space truncation-error corrector, especially under aggressive mode budgets. Once the operator basis is sufficiently resolved, additional liquid capacity yields diminishing returns.}
\label{tab:burgers-liquid-capacity}
\end{table}
\begin{figure}[h]
\centering
\includegraphics[width=0.45\linewidth]{../apps_industrial_breakthrough/burgers_liquid_capacity_sweep_outputs/burgers_capacity_target.png}
\includegraphics[width=0.45\linewidth]{../apps_industrial_breakthrough/burgers_liquid_capacity_sweep_outputs/burgers_capacity_best.png}
\caption{Best capacity-sweep Burgers profile. Left: reference trajectory. Right: $48$-mode OSNR rollout with $128$ liquid states on the $70\%/30\%$ train/future split, reaching $38.8746$ dB future PSNR.}
\label{fig:burgers-liquid-capacity}
\end{figure}
The spline-native liquid implementation \texttt{apps\_industrial\_breakthrough/burgers\_operator\_spline\_liquid.py} then replaces the generic sigmoid gates by compact cubic B-spline conductance banks and solves the liquid neuron ODE with an exponential Green update over Gauss--Legendre nodes inside each time cell. This is now treated as a rejected prototype rather than the final architecture: cubic B-spline gates are not matched to the liquid neuron operator, and numerical quadrature reintroduces the approximate integration step that OSNR is meant to remove. The ablation is still useful because it shows that merely making the gates ``spline-shaped'' is insufficient. A correct operator-spline liquid layer must derive its basis from the neuron ODE itself.
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{@{}p{0.27\linewidth}p{0.14\linewidth}p{0.14\linewidth}p{0.14\linewidth}p{0.16\linewidth}@{}}
\toprule
Profile & Future PSNR & Future RMSE & Full PSNR & Latency \\
\midrule
OSNR spectral, $32$ modes, $70\%/30\%$ split & $29.0669$ dB & $3.214163\mathrm{e}{-02}$ & $31.0045$ dB & $7.2$ ms \\
Generic sigmoid liquid, $64$ states & $29.3919$ dB & $3.096127\mathrm{e}{-02}$ & $31.8731$ dB & $1718.2$ ms \\
Operator-spline liquid, $64$ states, $13$ knots & $29.3667$ dB & $3.105131\mathrm{e}{-02}$ & $31.9368$ dB & $141870.5$ ms \\
OSNR spectral, $16$ modes, $50\%/50\%$ split & $19.5291$ dB & $1.086634\mathrm{e}{-01}$ & $21.6879$ dB & $7.2$ ms \\
Generic sigmoid liquid, $128$ states & $22.2946$ dB & $7.903278\mathrm{e}{-02}$ & $24.2495$ dB & $1232.0$ ms \\
Operator-spline liquid, $128$ states, $17$ knots & $20.7500$ dB & $9.441387\mathrm{e}{-02}$ & $23.0665$ dB & $125668.0$ ms \\
\bottomrule
\end{tabular}
\caption{First operator-spline liquid Green-solver ablation. The implementation uses spline-parametric conductance and forcing fields and solves the neuron ODE by exponential Green steps, but this initial parameterization is slower and less accurate than the generic sigmoid liquid control. This table identifies the next mathematical bottleneck: the spline-liquid state must be tied more directly to modal residual coefficients or initialized by a coefficient-space linear solve.}
\label{tab:operator-spline-liquid-ablation}
\end{table}
\section{Composite Operator-Spline Liquid Networks}
\subsection{Motivation: liquid dynamics as an operator equation}
Liquid time-constant networks and closed-form continuous-time networks model hidden states as continuous-time ODEs whose coefficients are modulated by the input and by the state itself \cite{hasani2020ltc,hasani2022cfc,cantini2025exact}. In scalar form, a liquid neuron can be written as
\[
\dot{x}(t)
=
-
\left[
w_{\mathrm{leak}} + f(I(t),x(t);\theta)
\right]x(t)
+
f(I(t),x(t);\theta)A.
\]
The productive OSNR interpretation is not to regard this as a black-box recurrent layer. It is a first-order operator equation with a known leak component and an innovation term. Splitting the deterministic leak from the nonlinear synaptic feedback gives
\[
\mathcal{L}x(t)
=
\left(D+w_{\mathrm{leak}}\right)x(t)
=
s(t),
\qquad
s(t)=f(I(t),x(t);\theta)\left(A-x(t)\right).
\]
The operator $\mathcal{L}=D+w_{\mathrm{leak}}$ has Green function
\[
g(t)=e^{-w_{\mathrm{leak}}t}\mathcal{H}(t),
\]
so the isolated liquid neuron is exactly an exponential-memory system. The correct OSNR liquid layer should therefore use exponential/operator splines matched to $D+w_{\mathrm{leak}}$, not polynomial splines inserted as generic nonlinear activations.
This viewpoint changes the computational target. Standard neural ODE, ODE-RNN, and latent-ODE implementations propagate states through numerical solvers and differentiate through either the solver trace or an adjoint system \cite{chen2018neuralode,rubanova2019latentode}. PINNs impose related continuous constraints by adding automatic-differentiation residuals at collocation points \cite{raissi2019pinn}. Standard numerical LNN implementations likewise propagate states sequentially by an ODE solver or by a constrained closed-form approximation. OSNR instead asks whether the entire hidden trajectory can be represented in an operator-matched spline space
\[
x(t)=\sum_k c[k]\beta_{\mathcal{L}}(t-k),
\]
where $\beta_{\mathcal{L}}$ is generated by the liquid leak operator. In the shift-invariant, fixed-conductance case, the coefficient-domain normal equations inherit a Toeplitz/circulant temporal structure and can be diagonalized by a one-dimensional FFT. In the state-dependent case, the exact global system is nonlinear; the mathematically controlled route is to isolate the nonlinear part as an innovation process $s(t)$ and solve the leak-filtered liquid trajectory exactly for a proposed innovation representation.
\subsection{Transfer-function bridge to state-free sequence models}
The strongest modern precedent for the OSNR liquid formulation is the transfer-function view of state-space sequence models. A linear time-invariant state-space model
\[
\dot{\mathbf{x}}(t)
=
\mathbf{A}\mathbf{x}(t)+\mathbf{B}u(t),
\qquad
y(t)=\mathbf{C}\mathbf{x}(t)+\mathbf{D}u(t)
\]
has Laplace-domain response
\[
Y(s)
=
\left[
\mathbf{C}(s\mathbf{I}-\mathbf{A})^{-1}\mathbf{B}
+
\mathbf{D}
\right]U(s)
=
H(s)U(s).
\]
Parnichkun et al. parameterize this dual representation directly as a rational transfer function and evaluate sequence blocks by FFT-based state-free inference, avoiding materialization of a hidden state tensor across the whole sequence \cite{parnichkun2024statefree}. This is conceptually aligned with OSNR's block-circulant operator calculus: when the governing operator is shift invariant, the recurrent scan can be replaced by a frequency-domain multiplication or division.
For the scalar liquid leak operator, the transfer function is the first-order rational filter
\[
H_{\mathrm{leak}}(s)
=
\frac{1}{s+w_{\mathrm{leak}}}.
\]
Thus, once an innovation trajectory $s(t)$ has been specified, the liquid state satisfies
\[
X(s)=H_{\mathrm{leak}}(s)S(s),
\qquad
x(t)=g*s,
\qquad
g(t)=e^{-w_{\mathrm{leak}}t}\mathcal{H}(t).
\]
On a uniform periodized grid this becomes a single FFT solve, while on nonuniform intervals it becomes the closed-form Green update derived below. The distinction is essential: Parnichkun's state-free result applies to linear transfer operators; OSNR does not claim that the original nonlinear liquid conductance is globally diagonalized. The OSNR claim is that the nonlinear term can be represented as a structured continuous innovation field and that the leak-filtered liquid trajectory can then be recovered by the exact transfer operator.
This reframes the proposed architecture as a Liquid State-Space Spline Operator. Compared with pure rational transfer-function layers, the OSNR addition is the continuous sparse-plus-smooth innovation sieve: FRI atoms capture non-bandlimited temporal events, and smooth exponential/DCT modes capture low-frequency drift. This is the part needed for impact-like events, irregular measurements, and causal regime changes emphasized in continuous-time sequence papers \cite{rubanova2019latentode,lechner2020odelstm,vorbach2021causal}. Compared with CfC and exact recursive LTC formulas, the OSNR contribution is not merely another cell update; it is a block trajectory representation that can use transfer-function diagonalization for uniform components and exact exponential Green kernels for local nonuniform components.
\subsection{Dual-continuum liquid innovation model}
The liquid innovation $s(t)$ need not be dense. Sequential signals often combine smooth trends with abrupt events: contact impacts, gait transitions, sensor dropouts, arrhythmia-like spikes, or regime switches. Continuous-time circuit policies and causal navigation experiments show that structured continuous dynamics can improve robustness and interpretability, but they still train recurrent neural circuits by gradient-based rollout rather than solving the operator algebraically \cite{lechner2020ncp,vorbach2021causal}. OSNR therefore models the innovation side as a sparse-plus-smooth continuum,
\[
\mathcal{L}x(t)
=
s_{\mathrm{sparse}}(t)+s_{\mathrm{smooth}}(t).
\]
The sparse tier is a temporal FRI model,
\[
s_{\mathrm{sparse}}(t)
=
\sum_{r=1}^{R} a_r \varphi(t-\tau_r),
\]
where matrix-pencil/TLS moment recovery estimates the nonuniform event times $\tau_r$. The smooth tier is represented by a low-frequency orthonormal dictionary such as a DCT or, more strictly, by an exponential-spline residual dictionary matched to the leak-filtered temporal statistics. A cross-Gram shielding step is required exactly as in the image/video model:
\[
\mathbf{A}_{\mathrm{cross}}
=
\mathbf{D}_{\mathrm{smooth}}^\top
\mathbf{A}_{\mathrm{sparse}},
\]
so smooth coefficients do not absorb sharp liquid events and sparse atoms do not duplicate slow drift. This is the liquid-network analogue of the OSNR sparse-plus-smooth decomposition.
\subsection{Global FFT solve under fixed leak}
For a uniform temporal grid and fixed leak $w_{\mathrm{leak}}$, the sampled operator
\[
\mathcal{L}=D+w_{\mathrm{leak}}
\]
is shift invariant under periodic or circulant boundary closure. Let $\mathbf{c}$ be the coefficient vector of the hidden trajectory and let $\mathbf{s}$ be the sampled innovation coefficients. The discrete operator relation has the form
\[
\mathbf{L}\mathbf{c}=\mathbf{s},
\]
where $\mathbf{L}$ is Toeplitz/circulant up to boundary treatment. With circulant closure,
\[
\widehat{\mathbf{c}}[\omega]
=
\frac{\widehat{\mathbf{s}}[\omega]}
{\widehat{D}[\omega]+w_{\mathrm{leak}}},
\]
and all frequency bins are solved concurrently by \texttt{torch.fft.fft}. For multiple liquid channels, this becomes either independent scalar divisions or small block solves when channels are coupled. If the coupling is constant, the frequency-bin update is
\[
\widehat{\mathbf{c}}[\omega]
=
\left(
\widehat{D}[\omega]\mathbf{I}
+
\mathbf{\Lambda}
-
\mathbf{W}
\right)^{-1}
\widehat{\mathbf{s}}[\omega],
\]
which is the spline-operator analogue of a rational transfer-function SSM. This is the non-iterative OSNR alternative to sequential liquid rollout, but it is exact only for the linear leak-filtered solve once the innovation sequence has been specified or estimated.
\subsection{Closed-form operator-spline liquid cell}
The liquid residual layer should be formulated as an operator-spline ODE solver, not as a generic recurrent neural network with spline activations. For a scalar liquid state,
\[
\dot{x}_i(t)
=
-\lambda_i(t)x_i(t)+b_i(t),
\qquad
\lambda_i(t)>0,
\]
the governing operator on a local time interval is
\[
L_{i,n}=D+\lambda_{i,n},
\qquad t\in[t_n,t_{n+1}],
\]
after freezing or spline-predicting the conductance rate $\lambda_i(t)$ on that cell. The correct basis is therefore the Green/operator spline of $D+\lambda_{i,n}$, not a polynomial cubic spline. Let $h=t_{n+1}-t_n$ and write local time as $\tau=t-t_n\in[0,h]$. The exact variation-of-constants formula is
\[
x_{i,n+1}
=
e^{-\lambda_{i,n}h}x_{i,n}
+
\int_0^h e^{-\lambda_{i,n}(h-\tau)}b_{i,n}(\tau)\,\dd\tau.
\]
To make this integral algebraic, the forcing is represented in an exponential-polynomial spline space on the same interval,
\[
b_{i,n}(\tau)
=
\sum_{m=1}^{M_\rho} q_{i,n,m}e^{\rho_m\tau},
\]
where the poles $\rho_m$ are chosen from the residual dynamics to be corrected: $\rho_0=0$ for constant forcing, real negative poles for dissipative memory, imaginary pairs $\pm j\omega$ for oscillatory modes, and repeated poles when polynomial-exponential terms are required. Substitution gives the closed-form kernel
\[
K_{\lambda,\rho}(h)
=
\int_0^h e^{-\lambda(h-\tau)}e^{\rho\tau}\,\dd\tau
=
\frac{e^{\rho h}-e^{-\lambda h}}{\lambda+\rho},
\qquad \lambda+\rho\neq 0.
\]
The removable singular case is handled by the analytic limit
\[
K_{\lambda,-\lambda}(h)=h\,e^{-\lambda h}.
\]
Thus the exact operator-spline liquid update is
\[
x_{i,n+1}
=
e^{-\lambda_{i,n}h}x_{i,n}
+
\sum_{m=1}^{M_\rho}q_{i,n,m}K_{\lambda_{i,n},\rho_m}(h).
\]
This is the central closed-form expression for the OSNR liquid cell.
The conductance and forcing coefficients must also live in coefficient space. Let $\mathbf{c}_n$ denote the OSNR PDE coefficients after the deterministic operator step, and let $\mathbf{r}_n$ denote the unresolved innovation or truncation residual to be modeled. Define a small set of operator-aligned sensors
\[
\mathbf{z}_n
=
\left[
\langle \phi_1,\mathbf{c}_n\rangle,\ldots,
\langle \phi_S,\mathbf{c}_n\rangle,
\langle \chi_1,\mathbf{r}_n\rangle,\ldots,
\langle \chi_R,\mathbf{r}_n\rangle
\right],
\]
where $\phi_s$ and $\chi_r$ are modal or Hermite coefficient probes, not raw coordinate samples. The positive liquid rate is then
\[
\lambda_{i,n}
=
\omega_i
+
\operatorname{softplus}
\left(
\sum_{\ell} a_{i,\ell}\,\eta_\ell(\mathbf{z}_n)
\right),
\qquad \omega_i>0.
\]
Here $\eta_\ell$ should be an exponential-spline dictionary matched to the coefficient process, for example modes generated by poles $\mu_\ell$ of an AR/CAR residual model. The interval forcing coefficients are
\[
q_{i,n,m}
=
\sum_{\ell} b_{i,m,\ell}\,\eta_\ell(\mathbf{z}_n).
\]
The liquid state contributes back to the PDE only through residual coefficient channels,
\[
\mathbf{c}_{n+1}
=
\Phi_{\mathrm{OSNR}}(\mathbf{c}_n)
+
\mathbf{B}\mathbf{x}_{n+1},
\]
where $\Phi_{\mathrm{OSNR}}$ is the deterministic operator-spline PDE step and $\mathbf{B}$ maps liquid states into the truncated/high-frequency innovation subspace. This prevents the learned liquid system from overwriting coefficients already explained by the physical operator.
\paragraph{Stability.}
The closed-form update is contractive in the homogeneous part whenever $\lambda_{i,n}>0$:
\[
|e^{-\lambda_{i,n}h}|<1.
\]
If $\lambda_{i,n}\ge\lambda_{\min}>0$ and $|q_{i,n,m}|\le Q_m$, then
\[
|x_{i,n+1}|
\le
e^{-\lambda_{\min}h}|x_{i,n}|
+
\sum_m Q_m |K_{\lambda_{i,n},\rho_m}(h)|.
\]
For residual poles with $\mathrm{Re}(\rho_m)\le0$, the kernel is uniformly bounded over finite $h$. This gives a direct route to stable long-horizon rollout: enforce positive rates, bound the forcing coefficient functionals, and restrict the liquid-to-PDE map $\mathbf{B}$ to residual subspaces.
\paragraph{Algorithm.}
The resulting implementation should follow this sequence.
\begin{enumerate}[leftmargin=*]
\item Advance the physical field coefficients by the deterministic OSNR operator step: $\widehat{\mathbf{c}}_{n+1}=\Phi_{\mathrm{OSNR}}(\mathbf{c}_n)$.
\item Project the unresolved defect or modal state into operator-aligned sensors $\mathbf{z}_n$.
\item Evaluate exponential-spline sensor dictionaries $\eta_\ell(\mathbf{z}_n)$.
\item Compute positive rates $\lambda_{i,n}$ and exponential forcing coefficients $q_{i,n,m}$.
\item Update each liquid state with the closed-form kernel $K_{\lambda,\rho}(h)$, using the analytic limit for $\lambda+\rho=0$.
\item Inject $\mathbf{B}\mathbf{x}_{n+1}$ only into the residual coefficient band, producing $\mathbf{c}_{n+1}$.
\end{enumerate}
The learnable objects are therefore not generic recurrent weights: they are the sensor dictionary coefficients, the forcing coefficients $q_{i,n,m}$, the positive rate functionals, and the residual injection map $\mathbf{B}$. This is the mathematically defensible Operator-Spline Liquid Network target for the next implementation.
\subsection{Repository algorithm target}
The first principled implementation should be a separate module rather than another residual experiment. The intended script, \texttt{apps\_industrial\_breakthrough/liquid\_transfer\_operator.py}, should implement the following deterministic sequence.
\begin{enumerate}[leftmargin=*]
\item Generate or ingest a continuous-time sequence tensor with explicit sample times, initially a Walker2d-like kinematic track with shape $B\times C\times T$.
\item Estimate sparse innovation events by temporal FRI/matrix-pencil recovery, yielding nonuniform knots $\tau_r$.
\item Fit the smooth innovation tier in an orthonormal or exponential-spline residual dictionary, with cross-Gram shielding against the sparse tier.
\item Solve the leak-filtered hidden trajectory by the exponential transfer relation $(D+w_{\mathrm{leak}})x=s$, using FFT diagonalization for the uniform component and closed-form Green updates for local nonuniform event cells.
\item Report trajectory RMSE, innovation sparsity, solve latency, and autograd allocation. The deterministic solver path should run under \texttt{torch.no\_grad()}.
\end{enumerate}
The benchmark comparison should be staged. First, compare against the existing generic sigmoid-liquid and rejected cubic-spline-liquid controls on synthetic sequences where the exact innovation structure is known. Second, compare the state-free transfer solve against recurrent CfC/LTC/Cantini-style exact cells on irregular synthetic sequences, separating numerical exactness from recurrent scan cost. Third, move to public liquid-network sequence benchmarks and compare against CfC, LTC, Neural ODE, ODE-RNN, and transfer-function SSM baselines under identical train/test splits. Only the first stage supports exact mathematical claims; the second and third stages are external efficiency and SOTA validation.
\subsection{Controlled validation: state-free liquid transfer solve}
The first controlled validation of this direction is \texttt{apps\_industrial\_breakthrough/liquid\_transfer\_operator.py}. The benchmark isolates the linear transfer claim from dataset, training, and irregular-sampling confounders. It evaluates
\[
(D+w_{\mathrm{leak}})x=s
\]
on a $512$-sample periodized grid with $w_{\mathrm{leak}}=0.5$ and \texttt{float32} tensors. The state-free solver applies the rational transfer function
\[
H(i\omega)=\frac{1}{i\omega+w_{\mathrm{leak}}}
\]
by one-dimensional FFT division. This validates the same linear state-free mechanism used by transfer-function state-space layers, but with the innovation $s$ represented by OSNR's sparse-plus-smooth model.
The validation uses three profiles. The first is a smooth periodic innovation made from a small number of Fourier modes, for which the analytic periodic Green response is known. The second is a sparse impulse innovation
\[
s(t)=\sum_r a_r\delta(t-\tau_r),
\]
where the event locations $\tau_r$ are recovered from Fourier moments by the regularized TLS matrix-pencil method. The third combines the sparse impulses with the smooth periodic component and estimates a joint sparse-plus-smooth frequency model. This last profile is a simple least-squares oblique separation, not yet the full ADMM cross-Gram solver.
\begin{table}[h]
\centering
\scriptsize
\setlength{\tabcolsep}{3pt}
\begin{tabular}{@{}p{0.24\linewidth}rrrrrr@{}}
\toprule
Profile & Traj. RMSE & Innov. RMSE & Event err. & OSNR ms & Rec. ms & Sparse \\
\midrule
\texttt{smooth\_periodic} & $1.1000\mathrm{e}{-07}$ & $3.2166\mathrm{e}{-07}$ & n/a & $0.2142$ & $19.4441$ & $98.83\%$ \\
\texttt{sparse\_impulse} & $6.4509\mathrm{e}{-07}$ & $1.4450\mathrm{e}{-05}$ & $3.7253\mathrm{e}{-09}$ & $0.4932$ & $21.1861$ & $99.22\%$ \\
\texttt{mixed\_sparse\_smooth} & $1.1767\mathrm{e}{-05}$ & $2.8567\mathrm{e}{-03}$ & $4.0978\mathrm{e}{-08}$ & $0.3237$ & $22.2715$ & $93.30\%$ \\
\bottomrule
\end{tabular}
\caption{Controlled state-free liquid transfer validation on a $512$-sample periodized sequence. The recurrent row is a periodic recurrent Green rollout reference, not a trained nonlinear LTC/CfC model. Event errors are measured in normalized time units.}
\label{tab:liquid-transfer-validation}
\end{table}
Table~\ref{tab:liquid-transfer-validation} shows that the FFT transfer solve recovers the leak-filtered state at near-\texttt{float32} accuracy for smooth and sparse controlled inputs, while reducing the measured block latency from roughly $19$--$22$ ms for the recurrent Green reference to less than $0.5$ ms. The mixed profile has a larger innovation error because sparse impulses have broadband Fourier support and the smooth dictionary is intentionally truncated; nevertheless, the leak-filtered trajectory error remains $1.1767\mathrm{e}{-05}$ with $93.30\%$ structural sparsity. These results validate the state-free linear liquid transfer layer and the sparse-plus-smooth innovation separation under controlled periodic assumptions. They do not yet claim superiority over trained nonlinear LTC/CfC models on external datasets; that comparison is the next benchmark stage.
The mixed profile is intentionally harder than the separated profiles because Dirac atoms occupy all Fourier frequencies. After the event times are localized, the script solves a joint sparse-plus-smooth least-squares system over sparse atoms and retained low-frequency smooth modes. The remaining innovation error therefore measures dictionary cross-talk and smooth-mode truncation rather than failure of the transfer solve itself. The much smaller trajectory error indicates that the stable leak transfer function attenuates part of this residual mismatch before it reaches the liquid state trajectory.
\subsection{Controlled validation: irregular causal Green evaluation}
The second validation script, \texttt{apps\_industrial\_breakthrough/liquid\_transfer\_irregular.py}, removes the periodized uniform-grid assumption from the state evaluation stage. It samples a $512$-point nonuniform time grid on $[0,1]$, keeps $w_{\mathrm{leak}}=0.5$, and evaluates the causal Green response
\[
x(t_i)
=
\int_0^{t_i}e^{-w_{\mathrm{leak}}(t_i-u)}s(u)\,\dd u
+
\sum_{\tau_r\le t_i}a_r e^{-w_{\mathrm{leak}}(t_i-\tau_r)}
\]
directly at all irregular sample locations. Smooth forcing terms are integrated analytically on each interval; sparse events remain off-grid. The recurrent reference is an exact causal Green rollout over the same nonuniform intervals, while the approximate baseline is a zero-order-hold recurrent update that represents the kind of local forcing approximation used by simple closed-form recurrent cells. Event times are again recovered from exact sparse Fourier moments, so this remains a controlled operator validation rather than a noisy inverse problem.
\begin{table}[h]
\centering
\scriptsize
\setlength{\tabcolsep}{3pt}
\begin{tabular}{@{}p{0.23\linewidth}rrrrrrr@{}}
\toprule
Profile & OSNR RMSE & Rec. RMSE & ZOH RMSE & Event err. & OSNR ms & Rec. ms & Sparse \\
\midrule
\texttt{smooth\_irregular} & $0.0000\mathrm{e}{+00}$ & $3.2506\mathrm{e}{-08}$ & $1.2631\mathrm{e}{-03}$ & n/a & $0.2467$ & $27.6731$ & $100.00\%$ \\
\texttt{sparse\_offgrid} & $3.6154\mathrm{e}{-07}$ & $4.7343\mathrm{e}{-07}$ & $4.7343\mathrm{e}{-07}$ & $6.1467\mathrm{e}{-08}$ & $0.6216$ & $3.0790$ & $99.22\%$ \\
\texttt{mixed\_irregular} & $3.6154\mathrm{e}{-07}$ & $1.0355\mathrm{e}{-06}$ & $1.2634\mathrm{e}{-03}$ & $6.1467\mathrm{e}{-08}$ & $0.2482$ & $27.3737$ & $99.22\%$ \\
\bottomrule
\end{tabular}
\caption{Controlled irregular causal liquid transfer validation. The OSNR column evaluates the closed-form Green response directly at nonuniform sample times. The recurrent column is an exact causal Green rollout; the ZOH column is a local zero-order forcing approximation.}
\label{tab:liquid-irregular-validation}
\end{table}
Table~\ref{tab:liquid-irregular-validation} shows that the operator-spline Green evaluation preserves near-\texttt{float32} agreement with the exact recurrent causal reference while avoiding the sequential scan over the full history. The zero-order recurrent approximation is accurate for pure off-grid impulses but loses roughly $10^{-3}$ RMSE on smooth forcing because it freezes the drive inside each irregular interval. This is the next step toward LNN relevance: the liquid response can be evaluated at irregular times by closed-form operator kernels, not only by a periodic FFT block. The remaining open problem is the harder one: estimating sparse-plus-smooth innovations from noisy irregular observations rather than from controlled moment access.
\subsection{Controlled validation: inverse innovation recovery}
The third liquid validation script, \texttt{apps\_industrial\_breakthrough/liquid\_innovation\_inverse.py}, begins to address the inverse problem. Instead of giving the solver exact innovation moments, it observes noisy irregular samples of several leak-filtered liquid traces driven by the same hidden innovation. This multi-leak setting is intentional: a single scalar trace only identifies the interval containing an off-grid impulse from adjacent samples, while two or more leak rates identify the event location inside the interval through cross-leak residual ratios.
For leak rates $\{0.35,0.70,1.25\}$ and a shared sparse-plus-smooth innovation, the script computes interval residuals
\[
r_{j,i}=x_j(t_{i+1})-e^{-w_j(t_{i+1}-t_i)}x_j(t_i).
\]
For an event $\tau\in(t_i,t_{i+1}]$, the sparse contribution obeys
\[
r_{j,i}^{\mathrm{event}}
=
a\,e^{-w_j(t_{i+1}-\tau)}.
\]
Thus ratios across leak channels localize $\tau$, after which a joint least-squares solve estimates event amplitudes and smooth forcing coefficients. This is still a controlled inverse problem: the smooth forcing dictionary and leak rates are known, and the detector is not yet a robust noisy-data estimator.
\begin{table}[h]
\centering
\scriptsize
\setlength{\tabcolsep}{4pt}
\begin{tabular}{@{}rrrrrrr@{}}
\toprule
Noise & Traj. RMSE & Innov. RMSE & FD innov. & Event err. & Events & Sparse \\
\midrule
$0.0\mathrm{e}{+00}$ & $1.4691\mathrm{e}{-07}$ & $8.8908\mathrm{e}{-07}$ & $4.7476\mathrm{e}{+01}$ & $1.7136\mathrm{e}{-07}$ & $4$ & $99.54\%$ \\
$1.0\mathrm{e}{-05}$ & $2.1515\mathrm{e}{-05}$ & $6.7998\mathrm{e}{-05}$ & $4.7477\mathrm{e}{+01}$ & $4.7682\mathrm{e}{-05}$ & $4$ & $99.54\%$ \\
$5.0\mathrm{e}{-05}$ & $6.4065\mathrm{e}{-05}$ & $1.8110\mathrm{e}{-04}$ & $4.7474\mathrm{e}{+01}$ & $1.9553\mathrm{e}{-04}$ & $4$ & $99.54\%$ \\
\bottomrule
\end{tabular}
\caption{Controlled inverse liquid innovation recovery from irregular multi-leak observations. The finite-difference baseline estimates $(D+w)x$ locally from one channel and is dominated by off-grid impulse discontinuities.}
\label{tab:liquid-inverse-validation}
\end{table}
Table~\ref{tab:liquid-inverse-validation} shows that, in the noise-free setting, the multi-leak inverse recovers both the trajectory and the smooth innovation near \texttt{float32} precision while localizing off-grid events to $1.7136\mathrm{e}{-07}$ normalized time units. At modest observation noise, the trajectory remains in the $10^{-5}$--$10^{-4}$ RMSE range. The large finite-difference innovation errors confirm the expected failure mode of local derivative estimates on discontinuous off-grid events. The next algorithmic requirement is a noise-robust event detector and regularized sparse-plus-smooth inverse solve; without that layer, this result should be read as an identifiability and controlled recovery validation rather than a full real-world sequence benchmark.
\subsection{Controlled validation: robust inverse noise sweep}
The robustness follow-up, \texttt{apps\_industrial\_breakthrough/liquid\_inverse\_robust\_sweep.py}, compares the residual-ratio inverse against a matched-dictionary OMP variant and a practical local refinement. The OMP solver precomputes a candidate off-grid event dictionary over each irregular interval, alternates event selection with a joint sparse-plus-smooth refit, and reports the pursuit/refit latency after dictionary setup. This is a more robust but grid-quantized detector: it gives up some low-noise event precision in exchange for stability under larger observation noise. The local refinement keeps OMP's selected support but re-estimates each event time by a small closed-form multi-leak search inside the selected interval, followed by one global coefficient refit. This removes most of the useful quantization error without the multi-second cost of full variable projection. A ratio-after-OMP refinement is also implemented as an optional ablation, but it inherits the high-noise collapse of the ratio estimator and is not used as the default path.
\begin{table}[h]
\centering
\scriptsize
\setlength{\tabcolsep}{4pt}
\begin{tabular}{@{}rlrrrrr@{}}
\toprule
Noise & Method & Traj. RMSE & Innov. RMSE & Event err. & Latency ms & Sparse \\
\midrule
$0.0\mathrm{e}{+00}$ & ratio & $1.4691\mathrm{e}{-07}$ & $8.8908\mathrm{e}{-07}$ & $1.7136\mathrm{e}{-07}$ & $6.9665$ & $99.54\%$ \\
$0.0\mathrm{e}{+00}$ & OMP & $9.0559\mathrm{e}{-05}$ & $5.6460\mathrm{e}{-07}$ & $4.1162\mathrm{e}{-04}$ & $5.4126$ & $99.54\%$ \\
$0.0\mathrm{e}{+00}$ & OMP-local & $1.4395\mathrm{e}{-05}$ & $4.4244\mathrm{e}{-07}$ & $2.4807\mathrm{e}{-05}$ & $10.8437$ & $99.54\%$ \\
$1.0\mathrm{e}{-04}$ & ratio & $2.0550\mathrm{e}{-04}$ & $6.0252\mathrm{e}{-04}$ & $5.3670\mathrm{e}{-04}$ & $5.6030$ & $99.54\%$ \\
$1.0\mathrm{e}{-04}$ & OMP & $1.7557\mathrm{e}{-04}$ & $6.0251\mathrm{e}{-04}$ & $4.7795\mathrm{e}{-04}$ & $5.3485$ & $99.54\%$ \\
$1.0\mathrm{e}{-04}$ & OMP-local & $8.9944\mathrm{e}{-05}$ & $6.0246\mathrm{e}{-04}$ & $2.1823\mathrm{e}{-04}$ & $11.3056$ & $99.54\%$ \\
$5.0\mathrm{e}{-04}$ & ratio & $5.2915\mathrm{e}{-01}$ & $2.5590\mathrm{e}{+00}$ & $7.8826\mathrm{e}{-04}$ & $5.4806$ & $99.54\%$ \\
$5.0\mathrm{e}{-04}$ & OMP & $2.6358\mathrm{e}{-04}$ & $3.0507\mathrm{e}{-03}$ & $5.3227\mathrm{e}{-04}$ & $5.3720$ & $99.54\%$ \\
$5.0\mathrm{e}{-04}$ & OMP-local & $1.7237\mathrm{e}{-04}$ & $3.0505\mathrm{e}{-03}$ & $3.0643\mathrm{e}{-04}$ & $10.9159$ & $99.54\%$ \\
$1.0\mathrm{e}{-03}$ & ratio & $5.4091\mathrm{e}{-01}$ & $3.5202\mathrm{e}{+00}$ & $8.8251\mathrm{e}{-04}$ & $5.6368$ & $99.54\%$ \\
$1.0\mathrm{e}{-03}$ & OMP & $7.2456\mathrm{e}{-04}$ & $4.5331\mathrm{e}{-03}$ & $6.3139\mathrm{e}{-04}$ & $5.4304$ & $99.54\%$ \\
$1.0\mathrm{e}{-03}$ & OMP-local & $7.8459\mathrm{e}{-04}$ & $4.5333\mathrm{e}{-03}$ & $7.6676\mathrm{e}{-04}$ & $10.5497$ & $99.54\%$ \\
\bottomrule
\end{tabular}
\caption{Robust inverse noise sweep on the controlled multi-leak liquid problem after vectorized event-design construction. The ratio method is more accurate at very low noise but fails catastrophically at higher noise. OMP remains stable under large noise; OMP-local reduces candidate-grid quantization at low and mid noise with roughly doubled millisecond-scale latency.}
\label{tab:liquid-robust-inverse-sweep}
\end{table}
Table~\ref{tab:liquid-robust-inverse-sweep} identifies the next engineering boundary. The state recovery problem is no longer limited by the Green transfer operator; it is limited by sparse event detection under noisy interval residuals. The matched-dictionary OMP variant removes the catastrophic high-noise failures of the residual-ratio method, and the OMP-local variant recovers much of the lost continuous timing precision at low and mid noise while remaining in the $10$--$11$ ms range for the full $512$-step controlled problem. At the largest tested noise, local refinement can overfit the noisy leak residuals, so the toolbox should expose both OMP and OMP-local as selectable estimators rather than treating refinement as uniformly dominant.
\subsection{Operator-matched innovation routing with an inferred operator}
The preceding experiments either specify the leak operator or fit a complete
trajectory. We next ask a different algorithmic question: can the operator be
inferred from contaminated observations and then used as an analytic
conditional-computation gate? The mechanical benchmark runner
generates separate training and test trajectories with an independent RK4
simulator. Each trajectory is a damped oscillator with sparse impulses at
unknown off-grid times. A robust iteratively reweighted fit estimates the
two-state flow map from noisy value/velocity jets. The interval innovation is
then
\[
\mathbf r_n=\mathbf x_{n+1}-\widehat{\mathbf F}\mathbf x_n,
\]
and a fixed median/MAD threshold fitted on the unlabelled training residuals
routes event intervals. Given a routed interval, the continuous event offset
$\delta\in[0,T]$ is estimated by projecting $\mathbf r_n$ onto
$\exp(\widehat{\mathbf A}\delta)\mathbf b$, where
$\widehat{\mathbf A}=T^{-1}\log\widehat{\mathbf F}$ and
$\mathbf b=(0,1)^\top$. No event label, event count, exact moment, pole, or
location is supplied to the estimator.
\begin{table}[H]
\centering
\small
\setlength{\tabcolsep}{4pt}
\begin{tabular}{@{}llrrrr@{}}
\toprule
Setting & Method & AP & Unsupervised F1 & Timing MAE$/T$ & Flow error \\
\midrule
linear, 40 dB & finite difference & $0.7800$ & $0.1996$ & $0.3673$ & $0.8188$ \\
linear, 40 dB & position-only AR(2) & $0.4758$ & $0.5245$ & n/a & n/a \\
linear, 40 dB & ordinary two-state & $1.0000$ & $0.9983$ & $0.0426$ & $0.0130$ \\
linear, 40 dB & robust inferred operator & $1.0000$ & $0.9994$ & $0.0425$ & $0.0005$ \\
linear, 30 dB & robust inferred operator & $1.0000$ & $0.9994$ & $0.1249$ & $0.0020$ \\
linear, 20 dB & robust inferred operator & $0.9907$ & $0.8553$ & $0.2830$ & $0.0115$ \\
strong Duffing, 30 dB & hybrid operator & $1.0000$ & $0.9963$ & $0.1222$ & $0.0800$ \\
\bottomrule
\end{tabular}
\caption{Blind operator and off-grid innovation routing over 12 seeds. The
event gate is unsupervised; AP is threshold-free. Roughly $5.5\%$ of intervals
contain events.}
\label{tab:blind-operator-innovation-routing}
\end{table}
Table~\ref{tab:blind-operator-innovation-routing} supports the routing
mechanism while isolating its limits. The inferred Hermite-jet residual gives
essentially perfect non-oracle event separation at 30--40 dB and estimates the
clean flow much more accurately than ordinary least squares. However, event
detection is easy enough in this benchmark that ordinary two-state least
squares also detects almost every event. More importantly, at 20 dB the
timing error is $0.2830T$, worse than the $0.25T$ expected from always choosing
the interval midpoint. Thus interval detection and sub-sample timing are
distinct claims: the latter must be confidence-gated under noise. The
position-only AR control is substantially weaker and cannot determine timing;
the derivative channel in the Hermite jet supplies genuine off-grid
information.
The field-level benchmark
tests whether the same principle survives a PDE and model mismatch. A
$1024$-point pseudo-spectral Strang-splitting solver generates periodic
advection--diffusion--reaction trajectories with compact cubic B-spline
sources at off-grid locations. Estimation sees only $256$-point block
averages. A trimmed unlabelled fit identifies transport parameters on a
separate trajectory, after which only the largest $1\%$ of test residuals are
retained as conditional corrections. The true injected support occupies
$0.283\%$ of space--time points.
\begin{table}[H]
\centering
\small
\setlength{\tabcolsep}{5pt}
\begin{tabular}{@{}llrrr@{}}
\toprule
Setting & Predictor & Source AP & Base nRMSE & nRMSE after $1\%$ \\
\midrule
linear & identity & $0.4323$ & $0.1077$ & $0.0813$ \\
linear & learned spectral map & $0.8604$ & $0.0506$ & $0.0036$ \\
linear & blind inferred operator & $0.8906$ & $0.0505$ & $0.0011$ \\
mild cubic & blind inferred operator & $0.8866$ & $0.0515$ & $0.0011$ \\
strong cubic & learned spectral map & $0.8586$ & $0.0673$ & $0.0057$ \\
strong cubic & blind inferred operator & $0.8630$ & $0.0672$ & $0.0045$ \\
strong cubic & hybrid operator & $0.8794$ & $0.0671$ & $0.0017$ \\
\bottomrule
\end{tabular}
\caption{Eight-seed field innovation routing. The correction budget retains
the largest $1\%$ of each one-step residual.}
\label{tab:pde-innovation-routing}
\end{table}
In the linear case, the blind fit recovers speed $0.72000001$, diffusivity
$0.00180000$, and decay $0.07999999$, from respective true values $0.72$,
$0.0018$, and $0.08$. Table~\ref{tab:pde-innovation-routing} shows that an
operator residual is substantially more compressible than a raw temporal
difference and also improves over an unconstrained learned spectral
transition. Under strong cubic mismatch the pure linear operator degrades,
but a four-feature closed-form local residual model lowers the $1\%$-budget
error from $0.0045$ to $0.0017$. This is evidence for an
\emph{analyze--\allowbreak annihilate--\allowbreak route--\allowbreak
reconstruct} algorithm: preserve the inferred
transport operator and spend flexible capacity on its localized mismatch. It
is not yet a SOTA claim; both studies are controlled, the mechanical study
observes the full value/velocity jet, and the PDE metric is one-step
correction rather than autonomous rollout.
The next test closes the loop and isolates the spline contribution. A sender
observes each new $256$-point field while the receiver retains only its previous
reconstruction. Both apply the same predictor; a median-plus-six-MAD gate,
fitted without labels on a separate trajectory, either sends no update or one
amplitude/location packet. The coordinate control sends a Kronecker impulse.
The spline codec sends a block-averaged cardinal cubic B-spline at one of four
sub-cell phases. The corrected receiver state is fed into the next prediction
for all $139$ transitions, so errors are allowed to accumulate.
\begin{table}[H]
\centering
\footnotesize
\setlength{\tabcolsep}{4pt}
\begin{tabular}{@{}llrrr@{}}
\toprule
Setting & Predictor / packet & Trajectory nRMSE & Terminal nRMSE & Payload \\
\midrule
linear & learned spectral / point & $0.03971$ & $0.06479$ & $0.199\%$ \\
linear & blind operator / point & $0.01713$ & $0.02242$ & $0.202\%$ \\
linear & blind operator / cardinal & $\mathbf{0.00965}$ & $\mathbf{0.01100}$ & $0.209\%$ \\
mild cubic & hybrid operator / point & $0.01731$ & $0.02299$ & $0.202\%$ \\
mild cubic & hybrid operator / cardinal & $\mathbf{0.00976}$ & $\mathbf{0.01119}$ & $0.210\%$ \\
strong cubic & learned spectral / point & $0.09334$ & $0.12578$ & $0.200\%$ \\
strong cubic & blind operator / cardinal & $0.12662$ & $0.11532$ & $0.298\%$ \\
strong cubic & hybrid operator / point & $0.02882$ & $0.03788$ & $0.201\%$ \\
strong cubic & hybrid operator / cardinal & $\mathbf{0.02286}$ & $\mathbf{0.02240}$ & $0.209\%$ \\
\bottomrule
\end{tabular}
\caption{Eight-seed closed-loop innovation codec. Payload includes a float32
amplitude and the location bits, and is normalized by one dense float32 field.}
\label{tab:operator-spline-streaming}
\end{table}
At matched predictor and nearly matched payload, the cardinal packet reduces
trajectory error relative to a point packet by $43.1\%$ in the linear case,
$43.1\%$ for the hybrid under mild nonlinearity, and $21.3\%$ under strong
nonlinearity, winning all eight paired seeds in each comparison. This is the
specific value of compact cardinal reproduction: one coefficient reconstructs
the off-grid event footprint rather than one sampled coordinate. The
factorization is necessary as well as the spline. A temporal-difference gate
fails because smooth transport dominates its robust scale estimate. Under
strong cubic feedback, the pure inferred linear operator false-triggers and
loses to the learned spectral control; the small local mismatch model is what
restores sparse routing.
A public-data follow-up uses the released PDEBench Test-17 FNO predictions.
The frozen FNO is the neural prior; samples $900$--$999$ and forecast frames
$8$--$20$ yield $3900$ held-out, three-channel, $256$-point residual fields.
Equal-accounting packets encode the target-time FNO residual, with location and
scale bits included. This is a residual codec/assimilation test, not a blind
forecast improvement.
\begin{table}[H]
\centering
\footnotesize
\setlength{\tabcolsep}{4pt}
\begin{tabular}{@{}lrrrrr@{}}
\toprule
Codec & Packets & Payload & Field nRMSE & Frobenius nRMSE & Residual left \\
\midrule
point & $9$ & $4.395\%$ & $0.004064$ & $0.001242$ & $69.80\%$ \\
DCT & $9$ & $4.395\%$ & $0.001941$ & $0.000500$ & $36.18\%$ \\
one-scale cardinal & $9$ & $4.395\%$ & $0.003810$ & $0.001169$ & $65.52\%$ \\
multiscale cardinal OMP & $8$ & $\mathbf{4.199\%}$ & $\mathbf{0.001460}$ & $\mathbf{0.000390}$ & $\mathbf{27.36\%}$ \\
\midrule
Test 19: point & $9$ & $4.395\%$ & $0.018400$ & $0.006024$ & $59.65\%$ \\
Test 19: DCT & $9$ & $4.395\%$ & $0.018944$ & $0.006492$ & $62.81\%$ \\
Test 19: one-scale cardinal & $9$ & $4.395\%$ & $0.016993$ & $0.005611$ & $55.21\%$ \\
Test 19: multiscale OMP & $8$ & $\mathbf{4.199\%}$ & $\mathbf{0.012304}$ & $\mathbf{0.004138}$ & $\mathbf{42.05\%}$ \\
\bottomrule
\end{tabular}
\caption{Public PDEBench FNO residual coding. The uncorrected FNO has field
nRMSE $0.004906$ and sample-wise Frobenius nRMSE $0.001431$.}
\label{tab:pdebench-fno-spline-codec}
\end{table}
Eight multiscale packets use fewer bits than nine DCT packets yet reduce paired
sample error by $10.1\%$ (bootstrap $95\%$ interval $4.3$--$15.6\%$) and win on
$77\%$ of samples. The one-scale cubic control is weak and DCT wins at the
smallest budgets: hierarchical scale selection is essential. A direct Python
OMP costs approximately $477\,\mu$s per field, versus $8.5\,\mu$s for DCT.
Replacing the per-field loop by batched FFT correlations and batched Gram
solves reproduces its errors within $4.44\times10^{-10}$ at $138\,\mu$s per field, a
$3.5\times$ speedup. A refit-free batched pursuit reaches $36\,\mu$s with a
small accuracy loss. The remaining latency gap is an explicit engineering
boundary.
The same scales and budgets transfer without tuning to PDEBench Test 19.
Eight multiscale packets again use fewer bits than nine DCT packets, but now
reduce paired sample error by $32.5\%$ (bootstrap $95\%$ interval
$30.3$--$34.8\%$) and win on all $100$ held-out samples. Per-field/Frobenius
nRMSE is $0.012304/0.004138$, versus $0.018944/0.006492$ for DCT. Even the
single-scale cardinal control beats DCT on Test 19, while the multiscale
dictionary remains decisively better. The fast refit-free pursuit retains a
$31.5\%$ paired reduction at approximately $35\,\mu$s per field.
Two stronger dictionary controls sharpen this result. We add an exact
orthonormal Haar transform and a channel-specific PCA/KLT basis fit only on
samples $0$--$899$, then frozen for samples $900$--$999$. With eight spline
versus nine control packets, the Test-17 paired reductions are $14.0\%$
(95\% interval $8.4$--$18.9\%$) against Haar and $6.8\%$
($1.1$--$12.3\%$) against learned PCA. On Test 19 they are $17.9\%$ and
$32.3\%$, respectively, with $100/100$ wins. The effect therefore survives
both a localized multiscale control and a training-only learned residual basis.
A separate channel audit gives each control ten packets against eight spline
packets and includes a shared 16-bit block scale, amplitude quantization,
coefficient noise, packet loss, and noisy sender residuals. Eight-bit payload
is $2.051\%$ of dense float32 for the spline versus $2.148\%$ for controls; at
four bits all methods use exactly $1.660\%$.
\begin{table}[H]
\centering
\footnotesize
\setlength{\tabcolsep}{3.5pt}
\begin{tabular}{@{}llrrr@{}}
\toprule
Set & Channel & Spline & learned PCA & paired reduction (95\% CI) \\
\midrule
17 & float32 clean & $0.001460$ & $0.001794$ & $1.4\%$ ($-4.9$--$7.3\%$) \\
17 & int4 clean & $\mathbf{0.001496}$ & $0.001878$ & $\mathbf{7.4\%}$ ($3.0$--$11.5\%$) \\
17 & int2 clean & $0.002837$ & $0.003145$ & $1.0\%$ ($-2.1$--$3.9\%$) \\
19 & float32 clean & $\mathbf{0.012304}$ & $0.018570$ & $\mathbf{30.7\%}$ ($28.5$--$32.9\%$) \\
19 & int4 clean & $\mathbf{0.012407}$ & $0.018601$ & $\mathbf{30.3\%}$ ($28.2$--$32.5\%$) \\
19 & int2 clean & $\mathbf{0.017709}$ & $0.020802$ & $\mathbf{13.4\%}$ ($12.5$--$14.3\%$) \\
19 & 10-dB coefficient SNR & $\mathbf{0.014436}$ & $0.019466$ & $\mathbf{23.6\%}$ ($22.0$--$25.3\%$) \\
19 & 20\% packet loss & $\mathbf{0.015939}$ & $0.020264$ & $\mathbf{18.7\%}$ ($17.4$--$20.0\%$) \\
\bottomrule
\end{tabular}
\caption{Quantized/impaired public residual packets. Entries are mean
per-field nRMSE; paired intervals bootstrap the 100 held-out samples.}
\label{tab:pdebench-packet-robustness}
\end{table}
The boundary is dataset-dependent rather than cosmetic. Test 19 wins all 100
paired samples against learned PCA in every tested clean, quantized, noisy, and
loss-impaired condition. On heterogeneous Test 17, however, the lower clean
aggregate nRMSE does not imply a resolved paired advantage over the stronger
ten-packet PCA control: the clean, eight-bit, 5\% loss, and 20-dB sender-noise
intervals cross zero. Four bits is positive at exactly matched payload, while
at two bits all Test-17 transform-comparison intervals cross zero and the
spline payload is slightly larger. This establishes a low-rate boundary and
supports a regime-dependent matched-residual claim, not universal codec
dominance. Moreover all public experiments still encode the true target-time
residual; predicted-residual or two-dimensional tests are the next gate.
We tested the predicted-residual gate directly. At time $t$ a causal codec may
use only previously revealed residuals. A cross-channel spectral AR selects
order and ridge on samples $800$--$899$ after candidate fits on $0$--$799$,
then refits on $0$--$899$. Test 17 selects order four: eight delayed spline
packets improve paired sample error by $18.25\%$ (95\% interval
$15.47$--$21.10\%$) over no correction. This is a causal correction win but
not a spline win: DCT and learned PCA are slightly better (spline reductions
$-0.51\%$ and $-0.64\%$), while spline beats Haar by only $0.63\%$. Test 19
selects order one; spline packets worsen no correction by $0.30\%$ (interval
$-0.52$--$-0.08\%$) and lose to all three transforms. Thus spatial residual
compressibility does not imply temporal predictability, and an assimilation
result cannot be relabelled as a forecast result. A structurally different
causal state, or genuine target-time sensor assimilation, is required.
The two-dimensional gate is more encouraging but remains regime-dependent.
On public $128\times128$ PDEBench FNO residuals, we compare eight
tensor-product multiscale cubic-cardinal packets with nine point, 2D-DCT,
exact 2D-Haar, single-scale cardinal, and separable-KLT packets. The KLT row
and column bases are fit per channel on samples $0$--$79$ and frozen on held-out
samples $80$--$99$. Including two-dimensional locations and scale indices,
spline payload is $0.0748\%$ of a dense float32 field versus $0.0790\%$ for
controls.
\begin{table}[H]
\centering
\footnotesize
\begin{tabular}{@{}llrr@{}}
\toprule
Set & Codec & Sample nRMSE & spline reduction (95\% CI) \\
\midrule
26 & point / DCT / Haar & $0.002315/0.002334/0.002245$ & $19.1/18.7/16.7\%$ \\
26 & one-scale cardinal & $0.002187$ & $14.9\%$ \\
26 & learned separable KLT & $0.001889$ & $\mathbf{2.08\%}$ ($0.55$--$3.60\%$) \\
26 & multiscale cardinal & $\mathbf{0.001850}$ & -- \\
26 & adaptive spline/KLT atlas & $\mathbf{0.001828}$ & $1.12\%$ vs. spline \\
26 & cheap preselector atlas & $0.001848$ & $2.19\%$ vs. KLT \\
27 & learned separable KLT & $\mathbf{0.001462}$ & $-6.48\%$ ($-8.46$--$-4.51\%$) \\
27 & multiscale cardinal & $0.001562$ & -- \\
27 & adaptive spline/KLT atlas & $\mathbf{0.001449}$ & $1.01\%$ vs. KLT \\
27 & cheap preselector atlas & $0.001451$ & $0.88\%$ vs. KLT \\
\bottomrule
\end{tabular}
\caption{Public 2D target-time FNO residual coding on 20 held-out samples.
Eight multiscale packets use fewer bits than nine controls.}
\label{tab:pdebench-2d-spline-codec}
\end{table}
On Test 26 the spline beats all fixed controls on $20/20$ samples and learned
KLT on $15/20$, so the positive paired interval establishes a small but real
2D matched-dictionary gain. Untuned Test 27 supplies the counterexample: the
spline still beats point, DCT, Haar, and one-scale cardinal atoms, but learned
KLT wins decisively. Tensor-product cardinal atoms are therefore a strong
compact prior for some 2D residual geometries, not a universal replacement for
a learned covariance basis.
The fixed-basis boundary suggests an adaptive atlas. Since an assimilation
encoder observes the residual being coded, it may evaluate both reconstructions
and send one mode bit choosing the lower residual norm. Including this bit,
payload is $0.0792\%$. The atlas selects spline on $70.0\%$ of Test-26 fields
but only $12.3\%$ of Test-27 fields. It significantly improves both experts:
on Test 26, $1.12\%$ versus spline (interval $0.46$--$1.89\%$) and $3.22\%$
versus KLT; on Test 27, $6.91\%$ versus spline and $1.01\%$ versus KLT
(interval $0.46$--$1.70\%$). This converts the transfer failure into a useful
design principle: route innovations among compact analytic and learned basis
experts, treating operator splines as a specialized expert rather than a
universal representation. The remaining systems question is whether a cheap
preselector can preserve the gain without evaluating every encoder.
A training-only closed-form preselector answers that systems question
positively in this audit. Thirteen cheap residual statistics---energy and
tail ratios, periodic derivative energies, radial spectral fractions, and
top-nine DCT/Haar energy---feed a ridge predictor of the spline/KLT error
ratio. On Test 26 it ties the pure spline statistically and beats KLT by
$2.19\%$ (interval $0.91$--$3.45\%$); on Test 27 it beats spline by $6.79\%$
and KLT by $0.88\%$ (interval $0.40$--$1.52\%$). It selects spline on
$84.3\%$ and $8.6\%$ of the respective fields. An actual conditional CPU path
computes features and runs only the selected encoder, reproducing the reference
reconstruction exactly. It costs $5.61$ versus $6.13$ ms/field for exhaustive
selection on Test 26 and $1.09$ versus $5.75$ ms/field on Test 27, measured
$1.09\times$ and $5.26\times$ speedups. Exact dual-mode search remains $1.03\%$
and $0.13\%$ more accurate, quantifying the price of preselection.
The atlas also survives coefficient quantization. With a shared 16-bit block
scale and ten KLT packets against eight spline packets, exact-atlas sample
nRMSE at int8/int4 is $0.001819/0.001828$ on Test 26, versus
$0.001868/0.001877$ for KLT (both $2.57\%$ paired reductions), and
$0.001429/0.001431$ on Test 27, versus $0.001439/0.001441$ for KLT
($0.81\%/0.80\%$). Cheap preselection also beats KLT with positive intervals
in all four cases. Atlas payload is $0.0452\%$ at int8 and $0.0376\%$ at int4,
only one mode bit above KLT. The fixed Test-26 spline/KLT interval crosses
zero after quantization; adaptive routing, rather than the spline alone, is the
robust contribution. An optimized accelerator implementation remains open.
\paragraph{Rate--distortion atlas guarantee.}
Let $E_m(r)$ denote the reconstruction produced by packet codec $m$ for a
residual field $r$. With equal or padded packet rates, the encoder-side mode
decision
\[
m^*(r)=\arg\min_{m\in\{1,\ldots,M\}}\|r-E_m(r)\|_2^2
\]
requires only $\lceil\log_2 M\rceil$ mode bits and satisfies
$\|r-E_{m^*}(r)\|_2^2\leq\min_m\|r-E_m(r)\|_2^2$ field by field. Consequently
the summed squared error, and hence sample Frobenius error, cannot exceed any
constituent codec. Unequal-rate operation replaces the objective by
$D_m+\lambda R_m$. The guarantee is elementary but important: complementing
an operator-matched spline expert with a learned covariance expert is safe at
negligible rate, and empirical gains quantify whether their errors are truly
complementary rather than redundant.
The complementarity transfers across all five public static $128\times128$
Darcy residual sets, using samples $0$--$899$ for training-only construction
and $900$--$999$ for evaluation.
\begin{table}[H]
\centering
\footnotesize
\begin{tabular}{@{}rrrrr@{}}
\toprule
Test & Spline mode & KLT nRMSE & Atlas nRMSE & reduction (95\% CI) \\
\midrule
21 & $18\%$ & $0.041215$ & $\mathbf{0.040302}$ & $2.65\%$ ($1.48$--$4.01\%$) \\
22 & $22\%$ & $0.025046$ & $\mathbf{0.024501}$ & $2.01\%$ ($1.06$--$3.08\%$) \\
23 & $42\%$ & $0.008331$ & $\mathbf{0.007922}$ & $4.02\%$ ($2.68$--$5.66\%$) \\
24 & $55\%$ & $0.005932$ & $\mathbf{0.005487}$ & $5.97\%$ ($4.21$--$7.93\%$) \\
25 & $62\%$ & $0.005612$ & $\mathbf{0.005193}$ & $5.68\%$ ($3.99$--$7.58\%$) \\
\bottomrule
\end{tabular}
\caption{One-bit spline/KLT atlas on held-out public static 2D residuals.}
\label{tab:pdebench-2d-atlas-transfer}
\end{table}
Together with dynamic Tests 26--27, the exact atlas beats KLT on all seven
public 2D sets, with every paired interval positive. Spline mode use spans
$12.3\%$--$70.0\%$, direct evidence of regime-dependent complementarity.
Cheap preselection retains a significant KLT gain on six of seven sets; Test
22 is unresolved ($0.81\%$, interval $-0.16$--$1.81\%$).
The two-expert atlas is nevertheless incomplete: on two Test-29 CFD regimes,
DCT is substantially better than both spline and cross-regime KLT. We
therefore apply the same rate--distortion rule to four experts---multiscale
cardinal, training-only separable KLT, DCT, and Haar---at the cost of two mode
bits. Rerunning Tests 21--27, the exact four-mode atlas improves on the
strongest constituent by $5.85\%$, $3.03\%$, $4.44\%$, $6.41\%$, $4.40\%$,
$1.13\%$, and $1.01\%$, respectively; every paired interval is positive.
The harder transfer protocol leaves out each of the three Test-29
four-channel CFD configurations in turn. KLT bases and the closed-form
selector use only the other two configurations, while all ten samples of the
third are held out.
\begin{table}[H]
\centering
\footnotesize
\begin{tabular}{@{}lrrr@{}}
\toprule
Held-out regime & Best fixed expert & Four-mode atlas & gain (95\% CI) \\
\midrule
\texttt{M01\_Eta01} & DCT $0.001239$ & $\mathbf{0.001157}$ & $5.17\%$ ($2.47$--$8.51\%$) \\
\texttt{M10\_Eta001} & spline $0.010867$ & $\mathbf{0.010846}$ & $0.28\%$ ($0.11$--$0.49\%$) \\
\texttt{M10\_Eta01} & DCT $0.002264$ & $\mathbf{0.002085}$ & $5.05\%$ ($2.13$--$9.68\%$) \\
\bottomrule
\end{tabular}
\caption{Leave-one-regime-out Test-29 residual coding. Training-only
statistics come from the other two CFD configurations; the encoder still
observes each target-time assimilation residual.}
\label{tab:pdebench-four-expert-cross-regime}
\end{table}
Thus exact routing beats the strongest included expert with a positive paired
interval on all ten public 2D datasets/regimes. This does not follow merely
from averaging: field choices change sharply with physics. Spline accounts
for $24.8\%$, $87.3\%$, and $22.2\%$ of the three Test-29 regimes, while DCT
accounts for $60.0\%$, $2.3\%$, and $69.1\%$. A multi-output version of the
13-feature selector runs only its predicted encoder and is $1.58\times$--
$4.49\times$ faster than exhaustive four-mode evaluation. Its boundary is
equally clear: it is significantly better than the strongest fixed expert on
six of ten public sets, unresolved on three, and $1.29\%$ worse than spline on
held-out \texttt{M10\_Eta001}. Exact encoder-side selection is robust;
low-cost out-of-regime routing remains an open learning problem. Most
importantly, the result rejects universal spline dominance: operator splines
are a complementary analytic expert inside a compact adaptive atlas.
The cheap-router failure can be reduced by enforcing more of the transform
calculus. If $U_m$ is an orthonormal codec and $I_{m,K}(r)$ indexes its $K$
largest coefficients, Parseval gives
\[
\|r-E_m(r)\|_2^2
=\|r\|_2^2-\sum_{k\in I_{m,K}(r)}|\langle r,u_{m,k}\rangle|^2.
\]
Thus KLT, DCT, and Haar need no learned error predictor: the encoder chooses
their minimum-distortion member exactly from retained energy. Only the
nonorthogonal spline comparison remains unknown. Our guarded pilot computes
one residual FFT and one matched-filter inverse FFT per cardinal scale,
providing first-step OMP capture and top-correlation energies without the eight
OMP iterations or least-squares refits. A binary ridge gate then chooses
between the spline and the Parseval-best orthogonal transform using these
features, residual morphology, and channel type.
Across Tests 21--27 and the three leave-one-regime-out Test-29 configurations,
the guarded pilot is significantly better than the strongest fixed expert on
nine of ten sets and statistically tied on the tenth. In the three OOD CFD
rows it improves on DCT by $3.08\%$ and $4.13\%$ in the two DCT-dominant
regimes, and ties spline at $-0.004\%$ (interval crosses zero) in the
spline-dominant regime. This removes the old cheap router's significant
$1.29\%$ OOD loss. The guarded pilot improves the old router on nine of ten
sets, with Test 27 unresolved, while measuring $1.13\times$--$3.28\times$
faster than exhaustive encoding. Exact search remains $0.04\%$--$2.20\%$
better. A feature-only four-output regressor is a decisive ablation: despite
receiving the same 28 pilot features, it still loses $0.54\%$ to spline in the
hard regime. The robust gain therefore comes from decomposing the decision by
Parseval and learning only the nonorthogonal comparison, not from adding
features indiscriminately. Pilot transform coefficients are cached for the
selected synthesis. Quantization also preserves the analytic decision:
for retained coefficients $c$ and quantized coefficients $q$, distortion is
$\|r\|^2-2\langle c,q\rangle+\|q\|^2$. This value matches explicit KLT/DCT/
Haar reconstruction to $3.9\times10^{-15}$ and selects the true best transform
on every audited field. At int8/int4, guarded-pilot gains against the strongest
fixed expert are $1.23\%/1.06\%$ on Test 26 and $0.79\%/0.78\%$ on Test 27,
all with positive intervals, at $0.0454\%/0.0378\%$ dense payload.
\paragraph{Sparse-station operator-kernel atlas.}
We next remove the encoder's full-field residual access. On the last target
frame of each public Test-29 forecast, the method receives contemporary
residual values at only $S$ random pixels and reconstructs the complete
$128\times128$ four-channel residual. This is sparse-observation
assimilation, not blind forecasting. Periodic IDW is the equal-observation
baseline. The spline expert uses the sampled columns of a periodized cardinal
cubic interpolation operator; a Gaussian kernel supplies a smooth radial
control. Kernel scales and ridges are selected on balanced fields from the
other two CFD regimes. Within a field, $75\%$ of stations fit each expert and
$25\%$ validate it. Training-only routing/audit splits select a switching
margin by minimizing the worst source-regime squared-error ratio to IDW; the
chosen expert is then refit on all $S$ observations.
\begin{table}[H]
\centering
\footnotesize
\begin{tabular}{@{}lrrrr@{}}
\toprule
Held-out regime & stations & IDW nRMSE & guarded atlas & gain (95\% CI) \\
\midrule
\texttt{M01\_Eta01} & 512 & $0.001023$ & $\mathbf{0.000972}$ & $3.12\%$ ($0.00$--$7.41\%$) \\
\texttt{M10\_Eta001} & 512 & $0.004882$ & $\mathbf{0.004686}$ & $3.48\%$ ($0.34$--$8.09\%$) \\
\texttt{M10\_Eta01} & 512 & $0.001593$ & $\mathbf{0.001572}$ & $0.67\%$ ($0.01$--$2.00\%$) \\
\bottomrule
\end{tabular}
\caption{Leave-one-regime-out sparse-station Test-29 assimilation. All
hyperparameters and routing margins use only the other two CFD regimes; 512
stations are $3.125\%$ of the spatial grid.}
\label{tab:pdebench-sparse-station-atlas}
\end{table}
The guarded atlas therefore improves periodic IDW point estimates at the
predeclared 512-station operating point in all three unseen regimes. The last
two gains have strictly positive bootstrap intervals; the first interval
touches zero and is unresolved. It routes $12.5\%$,
$50.0\%$, and $15.0\%$ of fields, respectively, to a cardinal or Gaussian
expert, so the improvement is not a renamed IDW result. The wider density
sweep is deliberately mixed: seven of nine point estimates at
$S\in\{256,512,1024\}$ favor the atlas, three have strictly positive bootstrap
lower bounds, five are unresolved, and \texttt{M10\_Eta01} at 256 stations
loses $1.66\%$ (interval $-4.08$--$-0.00\%$). Thus current evidence supports
moderate-density guarded interpolation, not uniform safety under extreme
sparsity.
A compact-packet ablation is negative. Sensor-only OMP with eight cardinal
or nine KLT/DCT/Haar coefficients loses substantially to IDW, which retains
all station values. At 512 stations, cardinal nRMSE is $0.003960$ versus
$0.001023$ on \texttt{M01\_Eta01}, and $0.006970$ versus $0.004882$ on
\texttt{M10\_Eta001}. Sparse local support alone cannot identify atoms that
receive no measurements. The successful spline mechanism is therefore a
full-capacity periodized interpolation operator used behind a conservative
router, rather than aggressive coefficient packetization.
The resolved low-density loss suggests that routing uncertainty, rather than
expert capacity alone, is the immediate failure. We therefore add a
cross-fitted agreement gate. Two disjoint validation-station folds must
independently select the same kernel, and on both folds it must improve on IDW
by a nonzero margin. Otherwise the gate abstains to IDW. The margin is chosen
from $\{0.10,0.20,0.30\}$ on the other two regimes by the same worst-group
training audit; the accepted expert is finally refit on all stations.
Three independent station permutations, three held-out regimes, and three
station densities yield 27 comparisons. The cross-fitted gate has 19 positive
point estimates, six changes within $0.0001\%$ of an exact tie, and two
unresolved negative estimates. Eleven bootstrap lower bounds are strictly
positive and none is a resolved loss. By comparison, the original gate has
nine resolved gains and one resolved loss. Cross-fitting changes the worst
point result from $-1.659\%$ to an unresolved $-0.496\%$, while average gain
decreases from $1.743\%$ to $1.372\%$. At 512 stations, mean gains across the
three layouts are $2.52\%$, $2.78\%$, and $0.23\%$ for the three regimes.
Thus fold agreement plus exact IDW abstention is an effective empirical safety
device, but not a formal no-harm certificate: the remaining two negative point
estimates, although unresolved, prevent that stronger claim.
\newpage
We next strengthen the spline expert without increasing its coefficient count.
Let $K_{h_1}$ and $K_{h_2}$ be unit-diagonal, periodized tensor-product
cardinal cubic kernels at two scales. Their direct-sum RKHS kernel is
\[
K_{\mathrm{multi}}=\frac{K_{h_1}+\eta K_{h_2}}{1+\eta},\qquad
c=(K_{\mathrm{multi}}(X,X)+\lambda I)^{-1}y,
\quad \eta>0.
\]
The construction remains positive and has one coefficient per station, equal
to a single-kernel interpolant. Scale pair, mixture weight, and ridge are
minimax-selected on balanced fields from the other two regimes.
Across three station permutations, three densities, and three held-out
regimes, the fixed pyramid improves the single cardinal kernel in 24 of 27
point comparisons, with 22 resolved gains. Hierarchy therefore improves the
representation itself, but is unsafe alone: two low-density comparisons are
resolved losses. On the base layout, the pyramid beats IDW by $4.56\%$ at 512
stations and $14.43\%$ at 1024 in \texttt{M10\_Eta001}, but loses $37.90\%$
and $24.37\%$ to IDW at 256 stations in \texttt{M01\_Eta01} and
\texttt{M10\_Eta01}, respectively.
Applying the same two-fold nonzero-margin abstention to this stronger expert
gives the following means over three station layouts.
\begin{table}[H]
\centering
\footnotesize
\begin{tabular}{@{}lrrr@{}}
\toprule
Held-out regime & 256 stations & 512 stations & 1024 stations \\
\midrule
\texttt{M01\_Eta01} & $0.00\%$ & $1.49\%$ & $1.95\%$ \\
\texttt{M10\_Eta001} & $0.26\%$ & $2.08\%$ & $8.97\%$ \\
\texttt{M10\_Eta01} & $0.17\%$ & $0.18\%$ & $4.11\%$ \\
\bottomrule
\end{tabular}
\caption{Mean paired reduction versus periodic IDW for the cross-fitted
multiscale-cardinal pyramid, over three independent random station layouts.}
\label{tab:pdebench-cardinal-pyramid}
\end{table}
Across all 27 comparisons, the routed pyramid has 18 positive estimates,
seven numerical ties, two unresolved negatives, 12 resolved gains, and no
resolved loss. Its mean gain is $2.134\%$ and worst point estimate is an
unresolved $-0.082\%$. It lowers nRMSE relative to the earlier stable
single-scale atlas in 16 of 27 cases and improves that method by $1.00\%$ on
average. The central result is consequently not ``more scales always win.''
A capacity-matched cardinal hierarchy creates a stronger analytic specialist;
cross-fit abstention supplies its empirical robustness.
Sensor geometry is itself part of the sampling operator. We therefore repeat
the 27-case audit with shifted periodic grids and with stratified layouts that
place one sensor at a random subcell location in each grid cell. Relative to
random stations, grid IDW reduces mean nRMSE by $11.2\%$, $12.1\%$, and
$21.3\%$ at 256, 512, and 1024 stations. Stratified IDW retains reductions of
$7.6\%$, $8.2\%$, and $13.2\%$. The gain follows fill distance: averaged
over three seeds, random fill radii are $13.54$, $9.39$, and $7.00$ pixels;
grid radii are $5.66$, $4.47$, and $2.83$; stratified radii are $8.19$,
$6.63$, and $4.24$.
Coverage alone is insufficient. The normalized sampling mask of every grid
has maximum non-DC Fourier magnitude one, the signature of exact
reciprocal-lattice replicas. Random and stratified masks have maxima only
$0.09$--$0.19$. Consistent with this alias nullspace, the grid pyramid router
has eight resolved gains but one resolved loss among 27 comparisons, reaching
$-2.06\%$ in its worst case: validation on the same lattice cannot observe an
off-lattice component. The stratified router has 19 positive estimates,
eight ties, no negative estimate, nine resolved gains, and no resolved loss;
its mean gain over the already stronger stratified IDW is $1.58\%$.
Consequently the practical cardinal acquisition rule is to allocate one
sensor per spline-scale cell to bound holes, then dither within cells to break
coherent aliases. This is an empirical design rule rather than a universal
optimality theorem.
The next experiment adapts the operator spectrum rather than the sampling
grid. For channel $c$, we augment the polynomial pyramid by
\[
K_{\mathrm{exp},c}(\mathbf{x},\mathbf{y})=
\frac{K_{\mathrm{multi}}(\mathbf{x},\mathbf{y})+
\gamma_c K_{h_c}(\mathbf{x},\mathbf{y})
\cos\!\left(\boldsymbol{\omega}_c^\top
(\mathbf{x}-\mathbf{y})\right)}{1+\gamma_c}.
\]
The modulated term is the real sum of the two spectral shifts induced by the
conjugate poles $\pm i\boldsymbol{\omega}_c$. Since both the cardinal kernel
and the stationary cosine kernel are positive semidefinite, their pointwise
product is positive semidefinite by the Schur product theorem; the positive
direct sum remains a valid kernel. It also retains one coefficient per
station. Scale, frequency, weight, and ridge are selected separately for the
four physical channels on balanced fields from the other two CFD regimes.
This gives an unusually clean positive/negative boundary. Used everywhere,
the channelwise exponential pyramid beats its polynomial parent in only nine
of 27 stratified held-out comparisons and loses in 18; it has two resolved
gains, 13 resolved losses, and a mean relative change of $-2.438\%$. Yet on
the difficult \texttt{M10\_Eta001} regime its mean gains over the polynomial
pyramid are $0.830\%$, $0.596\%$, and $0.306\%$ at 256, 512, and 1024
stations. Learned poles are therefore a specialized operator hypothesis, not
a universally better spline degree.
We consequently place IDW, the polynomial pyramid, and the channelwise
exponential pyramid in a cross-fitted operator atlas. Both station folds must
clear a training-selected margin over IDW; exponential selection must also
dominate polynomial selection on every fold, and routing below $7.5\%$ of the
field population triggers exact abstention. Over three regimes, three
stratified layouts, and three densities, this guarded atlas has 19 positive
results and eight exact abstentions, with no negative result, 12 resolved
gains, and no resolved loss. Mean gain over stratified IDW is $1.829\%$ and
the best case reaches $9.231\%$. Its aggregate nRMSE is $0.429\%$ lower than
the polynomial-only router on average (16 wins, six ties, five losses), with a
$1.016\%$ average improvement in \texttt{M10\_Eta001}. Thus operator-pole
adaptation is useful here only when evidence-gated; the negative fixed-basis
result is as important as the atlas gain.
The fieldwise audit localizes the mechanism: in \texttt{M10\_Eta001}, channels
1 and 2 average $18.78\%$ and $21.49\%$ reductions relative to IDW and route to
the modulated expert on $52.2\%$ and $53.3\%$ of fields. Channel 3 selects
zero modulation in all nine cases, so its nominal exponential routes are
regularization-only. The gain is therefore specific to two channel dynamics,
not generic added flexibility.
Cardinality also resolves the regular-grid hardware bottleneck exactly. On a
rectangular periodic station lattice the pyramid Gram matrix is block
circulant with circulant blocks. With the two-dimensional lattice DFT $F$,
its coefficient solve is
\[
\mathbf{c}=F^*\frac{F\mathbf{y}}
{\widehat{K}_{\mathrm{grid}}+\lambda},
\]
and full-field synthesis is one FFT convolution after scattering
$\mathbf{c}$ to the station lattice. No dense station-to-field matrix is
formed. Across three regimes, three layouts, and three densities, the FFT
implementation agrees with the dense solution to at worst
$4.57\times10^{-15}$. Median CPU speedups are $4.78\times$, $12.92\times$,
and $19.72\times$ at 256, 512, and 1024 stations. At 1024
stations the explicit Gram plus synthesis operators occupy about $136$ MiB,
versus $0.25$ MiB for the kernel and its spectrum.
Dither makes the station Gram noncirculant, but does not destroy translation
invariance of the much larger station-to-field synthesis map. We therefore
retain the exact dense $n\times n$ irregular Gram solve, scatter its
coefficients at the true station locations, and FFT-convolve on the output
grid. This split diagonalization agrees with the original dense dithered
implementation to at worst $3.42\times10^{-15}$ across all 27 cases. Median
speedups are $4.78\times$, $7.49\times$, and $6.54\times$, with runtimes
$8.9$, $11.7$, and $28.3$ ms. At 1024 stations it stores an $8$ MiB Gram plus
a $0.25$ MiB spectrum rather than the $136$ MiB Gram--synthesis pair, a
$16.5\times$ operator-memory reduction. The key is to preserve irregularity
only where geometry requires it and diagonalize the globally stationary map.
This accelerator exposes the price of the anti-aliasing geometry above.
Snapping stratified measurements to cell centers restores circulant structure
but loses $4.20\%$, $10.86\%$, and $21.00\%$ relative to the true dithered
solve, with resolved losses in four of nine, nine of nine, and nine of nine
cases. An exact irregular matvec can still scatter, FFT-convolve, and sample
at the true stations. Preconditioned by the lattice inverse, it recovers the
dense field to within $6.11\times10^{-9}$, but needs 21--100 iterations and is
$4$--$7\times$ slower than optimized dense algebra at this scale. At 1024
stations, eight truncated iterations are $1.92\times$ faster but lose
$5.29\%$ to dense and tie IDW; 16 iterations retain a $1.06\times$ speedup and
beat IDW by $4.89\%$ on average, but still have five of nine resolved losses
to dense. Snapping and iterative replacement of the Gram are therefore
negative controls; the exact irregular solution is the dense-Gram/FFT-
synthesis split above.
Continuously located sensors admit a further operator-derived bridge without
a generic NUFFT. Write an off-pixel center as
$\mathbf{x}_j=\mathbf{n}_j+\boldsymbol{\delta}_j$, where $\mathbf{n}_j$ is its
nearest grid point. For the cardinal pyramid $K$, full-grid synthesis becomes
the truncated Hermite moment expansion
\[
\sum_j c_j K(\mathbf{q}-\mathbf{x}_j)
\simeq
\sum_{|\boldsymbol{\alpha}|\le p}
\left[D^{\boldsymbol{\alpha}}K *
\sum_j \frac{c_j(-\boldsymbol{\delta}_j)^{\boldsymbol{\alpha}}}
{\boldsymbol{\alpha}!}\,\delta_{\mathbf{n}_j}\right](\mathbf{q}).
\]
The continuous-coordinate station Gram is still assembled and solved exactly;
only the much larger station-to-grid map is replaced. In two dimensions the
order-$p$ expansion needs $(p+1)(p+2)/2$ FFT convolutions, independent of the
number of sensors.
We evaluate bilinearly sampled PDEBench residual measurements with random
offsets up to $0.49$ pixel over the same three regimes, three seeds, and three
densities. The six-channel, second-order expansion has median relative
synthesis error $3.40\times10^{-4}$ and worst error $2.22\times10^{-3}$ over
27 cases; its maximum absolute change in normalized reconstruction error is
$3.71\times10^{-7}$. Median synthesis speedups are
$11.09\times$, $23.19\times$, and $45.13\times$ at 256, 512, and 1024
sensors. Including the shared exact Gram solve, median speedups are
$9.53\times$, $13.55\times$, and $10.92\times$. At 1024 sensors, six complex
derivative spectra require $1.5$ MiB rather than a $128$ MiB dense synthesis
matrix; including the common Gram gives approximately $9.5$ versus $136$ MiB
of operator storage.
The controls expose the approximation mechanism. Nearest and bilinear
coefficient gridding reach worst relative synthesis errors $0.357$ and
$1.992$. A displacement-radius sweep gives error exponents $1.99$ for first
order and $3.01$ for second order, as predicted by the Taylor remainder.
Third order lowers the constant but its exponent remains $3.03$, because the
cubic B-spline is globally only $C^2$ and knot crossings preclude a uniform
fourth-order remainder. It therefore adds four FFT channels without material
reconstruction benefit. Second-order Hermite moment gridding is the practical
Pareto point. This closes the simulated off-pixel synthesis gate; validation
on physical continuous-coordinate stations, rather than bilinearly sampled
gridded fields, remains open.
The result is not specific to bilinear measurement formation. Replacing it
by periodic Fourier upsampling and continuous sampling over the same 27 cases
gives median/worst second-order synthesis errors
$3.60\times10^{-4}/2.30\times10^{-3}$ and a maximum absolute nRMSE change of
$3.93\times10^{-7}$. The near-identical envelope supports the translated-
spline approximation mechanism rather than an accidental match to the pixel
sampler.
Compact support also removes the dense storage assumption from the remaining
continuous Gram. A periodic neighbor search assembles only nonzero cubic-
cardinal interactions, after which one sparse factorization serves every
field. Across 27 cases the sparse and dense matrices agree to at worst
$3.68\times10^{-16}$ and their coefficients to $9.83\times10^{-14}$. On the
fixed $128^2$ domain the Gram is $25\%$ dense, uses $2.66\times$ less raw
storage, and gives median Gram-stage speedups $1.43\times$, $1.40\times$, and
$1.19\times$ at 256, 512, and 1024 sensors; one 512-sensor case is a
$0.98\times$ tie. Combined with second-order Hermite synthesis, median
end-to-end speedups become $9.48\times$, $14.54\times$, and $12.74\times$.
This is an exact storage result but not yet an asymptotically fast sparse
solver. On a $256^2$ diagnostic with 4096 continuous sensors, density falls
to $6.25\%$ and raw Gram storage improves $10.65\times$, yet generic sparse LU
reaches only $1.03\times$ dense parity because of fill-in. Unpreconditioned
CG takes 466 iterations to reach relative residual $10^{-8}$. Thus the full
operator need not be dense, but larger deployments require a cardinal
multilevel or lattice-corrected preconditioner rather than generic sparse
algebra.
The lattice-corrected preconditioner makes the scaling boundary constructive.
Using the inverse BCCB operator of the underlying one-sensor-per-cell lattice
reduces batched PCG from 466 to 100 iterations, although exact convergence is
still slower than dense. At 4096 sensors, 32 iterations give
$2.21\times$ end-to-end speed with relative field error
$4.998\times10^{-3}$, while 64 iterations give $1.30\times$ speed with error
$9.99\times10^{-6}$. Sixteen iterations are rejected: their
$3.41\times$ speed costs $11.3\%$ field error. At 256--1024 sensors sparse
direct factorization remains faster than the accurate truncated variants.
The resulting solver policy is size dependent: sparse direct below the
factorization crossover, 64-step lattice-PCG above it for high fidelity, and
32 steps only under an explicit $0.5\%$ operator-error budget.
At this stage the liquid module is mature enough to be used as an OSNR toolbox component under controlled assumptions: exact Green/transfer evaluation is solved, irregular timing is supported, and sparse event recovery has both a fast robust estimator and a more precise local estimator. It is not yet a stand-alone SOTA learning claim against trained LTC/CfC/ODE-RNN models. The next learning benchmark should therefore use this module as a structured layer inside a small trainable hybrid system, while the main hard-core validation remains the CFD/PINN setting where analytic derivatives, hard boundary constraints, and FFT operator diagonalization are the central advantage.
\subsection{Controlled validation: hybrid impact sequence learning}
The first learning-facing liquid benchmark tests whether the liquid toolbox is useful beyond deterministic reconstruction by constructing a train/test family of damped hybrid impact sequences. Each trajectory combines a smooth damped oscillatory component with sparse exponential impact responses. The model observes the first $60\%$ of each noisy sequence and predicts the held-out future. This task is intentionally structured: the future is predictable when the prefix identifies the continuous operator, the smooth modes, and the sparse impact responses, but a generic neural sequence model must learn this structure from examples.
The benchmark uses $96$ training sequences, $32$ held-out test sequences, $128$ time samples, three sparse impacts per sequence, and observation noise $10^{-3}$. Two gradient-trained baselines are included: a prefix MLP and a GRU encoder that maps the observed prefix directly to the future suffix. Three OSNR profiles are included. OSNR-head is a small neural head that maps the observed prefix to coefficients over a fixed normalized liquid dictionary and then renders the trajectory through the analytic basis. OSNR-support trains a classifier to imitate the OMP event support, selects a diverse top-$K$ set of event atoms, and then solves amplitudes and smooth coefficients by ridge least squares on the prefix. OSNR-liquid performs no gradient training on this dataset. It fits each test prefix by a sparse-plus-smooth operator dictionary: damped Fourier/exponential smooth atoms plus a causal exponential event dictionary, selected by OMP and refit by ridge least squares, then extrapolated through the same closed-form liquid response.
\begin{table}[h]
\centering
\small
\begin{tabular}{@{}lrrrr@{}}
\toprule
Profile & Future RMSE & Future PSNR & Train ms & Infer/Fit ms \\
\midrule
ZOH hold & $9.7691\mathrm{e}{-01}$ & $4.8727$ & $0.0$ & $0.000$ \\
MLP prefix & $3.3785\mathrm{e}{-01}$ & $14.0954$ & $241.4$ & $0.044$ \\
GRU prefix & $2.5258\mathrm{e}{-01}$ & $16.6219$ & $10728.8$ & $1.807$ \\
OSNR-head & $3.2231\mathrm{e}{-01}$ & $14.5044$ & $1010.7$ & $0.112$ \\
OSNR-support & $9.0795\mathrm{e}{+00}$ & $-14.4914$ & $1121.2$ & $1.075$ \\
OSNR-liquid & $1.9521\mathrm{e}{-01}$ & $18.8600$ & $0.0$ & $61.654$ \\
\bottomrule
\end{tabular}
\caption{Controlled hybrid impact sequence learning benchmark. Metrics are computed on the held-out future suffix across $32$ test sequences. OSNR-head is a trained coefficient predictor over a fixed analytic liquid dictionary. OSNR-support is a trained event-support classifier followed by analytic coefficient refitting. OSNR-liquid is a per-sequence structured sparse-plus-smooth fit rather than a trained neural baseline.}
\label{tab:liquid-hybrid-learning}
\end{table}
Table~\ref{tab:liquid-hybrid-learning} gives three distinct lessons. First, on this operator-matched hybrid family, the per-sequence OSNR-liquid estimator improves future RMSE over the trained GRU by roughly $22.7\%$ and improves future PSNR by $2.2381$ dB, without gradient training. Second, the naive trainable coefficient head is not yet competitive with the GRU: it improves over the direct MLP but underperforms the recurrent encoder. Third, support imitation alone fails catastrophically. Even with diversity-constrained top-$K$ selection, small support mistakes produce an ill-conditioned analytic extrapolation. This is an important boundary condition: the trainable hybrid should not try to classify sparse support independently from amplitude and trajectory fit. The next viable learned version should be a distillation/correction model around the OMP-selected OSNR fit, or a differentiable sparse solver layer with the support decision coupled to the reconstruction loss.
A sample-efficiency stress run reduces the gradient-trained training set while leaving the per-sequence OSNR-liquid fit unchanged. With only $8$ training trajectories and $64$ held-out test trajectories, the trained GRU reaches future RMSE $4.2430\mathrm{e}{-01}$ while OSNR-liquid reaches $2.1714\mathrm{e}{-01}$, a $48.8\%$ reduction without dataset-level backpropagation. With only $4$ training trajectories, the GRU reaches $3.9635\mathrm{e}{-01}$ and OSNR-liquid remains $2.1714\mathrm{e}{-01}$, a $45.2\%$ reduction. This is the strongest liquid-network learning signal so far: not a public LTC/CfC benchmark victory, but a clean demonstration that a biologically inspired leak-filtered operator dictionary can replace gradient training when the sequence family is sparse-plus-smooth and operator matched. A multi-leak dictionary ablation was added to the runner, but the naive wider leak bank overfit the prefix and underperformed the fixed leak; future learned liquid hybrids should regularize leak selection or validate it inside the observed prefix rather than simply expanding the atom bank.
\subsection{Controlled validation: dendritic cellular operator learning}
The liquid experiments above still treat the neuron mostly as a scalar leak-filtered state. The more radical cellular benchmark makes each artificial cell a multicompartment dendritic operator followed by a soma leak, and learning is a local sparse inverse problem rather than reverse-mode differentiation through a network.
For branch $b$ with dendritic coordinate $s\in[0,1]$, the controlled model starts from the cable-like PDE
\[
\partial_t v_b(s,t)
=
D_b\partial_{ss}v_b(s,t)
-
\lambda_b v_b(s,t)
+
\sum_j \theta_{bj} r_j(t)\delta(s-s_{bj}),
\]
with cosine eigenmodes used as the exact reduced basis. The modal state obeys
\[
\dot z_{bm}(t)
=
-
\left(\lambda_b+D_b\pi^2m^2\right)z_{bm}(t)
+
\sum_j\theta_{bj}\cos(\pi m s_{bj})r_j(t),
\]
so every synapse generates an analytically known branch atom. The soma then integrates the branch root voltages by
\[
\dot V(t)
=
-\lambda_s V(t)
+
\sum_b a_b v_b(0,t).
\]
This gives a cellular analogue of OSNR: the dictionary atoms are not generic lags or learned hidden units; they are Green responses of dendritic diffusion plus soma leakage.
The learning rule is local. For each branch, dendritic probe traces at $s\in\{0,0.43,0.86\}$ provide the local voltage/calcium-style observables. The branch solves
\[
\min_{\theta_b}
\left\|
\mathbf{P}_b\sum_j\theta_{bj}\mathbf{a}_{bj}
-
\mathbf{y}_b
\right\|_2^2
+
\alpha\|\theta_b\|_2^2,
\qquad \|\theta_b\|_0\le K_b,
\]
by OMP plus ridge refitting. No global reverse pass, adjoint, or dataset-level backpropagation is used for the cellular row. A soma-only OMP ablation is also included: it sees only $V(t)$, not the local branch probes, and therefore tests whether exact synapse placement is identifiable from the soma alone.
\begin{table}[h]
\centering
\scriptsize
\setlength{\tabcolsep}{3pt}
\begin{tabular}{@{}llrrrrr@{}}
\toprule
Train seq. & Profile & Test RMSE & PSNR & Train ms & Infer ms & Supp. F1 \\
\midrule
$8$ & Delay-ridge local & $5.7633\mathrm{e}{-04}$ & $29.0007$ & $831.1$ & $0.024$ & n/a \\
$8$ & Dense cable ridge & $4.9364\mathrm{e}{-04}$ & $30.3459$ & $4106.4$ & $0.043$ & n/a \\
$8$ & GRU backprop & $6.1679\mathrm{e}{-03}$ & $8.4113$ & $7025.4$ & $9.963$ & n/a \\
$8$ & DOS-NC soma OMP & $3.8017\mathrm{e}{-04}$ & $32.6145$ & $4107.6$ & $0.087$ & $0.100$ \\
$8$ & DOS-NC branch OMP & $2.1202\mathrm{e}{-05}$ & $57.6867$ & $5324.2$ & $91.379$ & $1.000$ \\
\midrule
$2$ & Dense cable ridge & $4.5093\mathrm{e}{-03}$ & $11.1320$ & $4151.8$ & $0.025$ & n/a \\
$2$ & GRU backprop & $6.3951\mathrm{e}{-03}$ & $8.0972$ & $6274.8$ & $9.341$ & n/a \\
$2$ & DOS-NC soma OMP & $1.8702\mathrm{e}{-03}$ & $18.7764$ & $4152.3$ & $0.112$ & $0.000$ \\
$2$ & DOS-NC branch OMP & $8.6729\mathrm{e}{-06}$ & $65.4508$ & $5356.5$ & $92.441$ & $1.000$ \\
\bottomrule
\end{tabular}
\caption{Controlled dendritic cellular operator learning. The teacher is a $4$-branch dendritic cable cell with $6$ cosine modes, $8$ presynaptic traces, $10$ sparse synapses, $160$ time samples, $48$ held-out test sequences, and observation noise $2\mathrm{e}{-3}$. DOS-NC branch OMP learns from local dendritic probe traces by sparse inverse solving; the GRU baseline uses ordinary backpropagation.}
\label{tab:dendritic-cellular-operator-learning}
\end{table}
Table~\ref{tab:dendritic-cellular-operator-learning} is the first controlled validation of the proposed biological learning thesis. With only $8$ training sequences, the branch-local dendritic operator learner reaches RMSE $2.1202\mathrm{e}{-05}$ on held-out soma voltage, about $291\times$ lower than the backprop-trained GRU on the same split, and exactly recovers the teacher support. With only $2$ training sequences, it remains at $8.6729\mathrm{e}{-06}$ RMSE with support F1 $1.000$. The soma-only OMP row is the critical ablation: it improves over generic temporal baselines but fails to identify the true synapses. This matches the biological premise. Local dendritic observables are not an implementation detail; they are the information channel that makes no-backprop synaptic learning identifiable.
\begin{figure}[h]
\centering
\includegraphics[width=0.85\linewidth]{../apps_industrial_breakthrough/dendritic_cellular_operator_outputs_main/dendritic_cellular_operator_comparison.png}
\caption{Held-out dendritic cell trace for the $8$-sequence benchmark. The branch-local DOS-NC curve is visually indistinguishable from the teacher soma, while the GRU and generic delay ridge baselines miss the operator-matched cellular dynamics.}
\label{fig:dendritic-cellular-operator-learning}
\end{figure}
This is still a controlled cellular-identifiability experiment, not a public LNN benchmark or a claim that arbitrary supervised learning can be replaced by local rules. Its significance is narrower and stronger: once the cell is modeled as a dendritic diffusion operator, local branch traces plus sparse inverse solving can learn the synaptic operator dramatically faster and more accurately than a small backprop-trained recurrent network in the matched regime. The next step is to stack these cells into layers where each branch receives its own local predictive or modulatory innovation, so the network-level credit signal is carried by local residual fields instead of exact reverse-mode gradients.
\subsection{Real-data stress: no-backprop CNN/RNN/attention/GNN views}
The next script, \texttt{apps\_industrial\_breakthrough/dendritic\_cross\_arch\_benchmark.py}, deliberately tests whether the no-backprop hypothesis survives contact with real benchmark data. It uses torchvision MNIST and FashionMNIST, keeps all feature extractors fixed after random or analytic wiring, and learns only closed-form ridge readouts. The profiles are architectural analogues rather than trained networks: DCT operator coefficients, fixed local convolutional banks, liquid row/column scans, fixed patch attention, grid-graph diffusion, class-balanced dendritic RBF memory cells, and a fused ridge readout. The comparison baselines are small MLP/CNN models trained by ordinary backprop for one epoch under the same CPU-only runner. This is not a SOTA protocol; it is a fast falsification test for whether local operator learning has benchmark-scale signal beyond the controlled dendritic-cell setting.
\begin{table}[h]
\centering
\scriptsize
\setlength{\tabcolsep}{3pt}
\begin{tabular}{@{}llrrrr@{}}
\toprule
Dataset/split & Profile & Accuracy & Train ms & Est. MB & Backprop \\
\midrule
MNIST $10\mathrm{k}/10\mathrm{k}$ & CNN local ridge & $97.44\%$ & $1082.0$ & $289.8$ & no \\
MNIST $10\mathrm{k}/10\mathrm{k}$ & CNN local ensemble $\times4$ & $97.85\%$ & $4065.0$ & $289.8$ & no \\
MNIST $10\mathrm{k}/10\mathrm{k}$ & DOS-NC fused ridge & $97.58\%$ & $7102.1$ & $515.1$ & no \\
MNIST $10\mathrm{k}/10\mathrm{k}$ & CNN, one epoch & $92.79\%$ & $2083.6$ & $4.4$ & yes \\
\midrule
MNIST full & CNN local ridge & $97.97\%$ & $4247.8$ & $912.5$ & no \\
MNIST full & CNN local ensemble $\times4$ & $98.03\%$ & $16325.2$ & $912.5$ & no \\
MNIST full & DOS-NC fused ridge & $98.07\%$ & $26602.2$ & $1531.9$ & no \\
MNIST full & CNN, one epoch & $98.25\%$ & $12278.1$ & $4.4$ & yes \\
\midrule
Fashion $10\mathrm{k}/10\mathrm{k}$ & CNN local ridge & $88.06\%$ & $1094.2$ & $289.8$ & no \\
Fashion $10\mathrm{k}/10\mathrm{k}$ & CNN local ensemble $\times4$ & $89.10\%$ & $4072.5$ & $289.8$ & no \\
Fashion $10\mathrm{k}/10\mathrm{k}$ & DOS-NC fused ridge & $88.30\%$ & $7104.4$ & $515.1$ & no \\
Fashion $10\mathrm{k}/10\mathrm{k}$ & CNN, one epoch & $77.60\%$ & $2079.6$ & $4.4$ & yes \\
\midrule
Fashion full & CNN local ridge & $89.48\%$ & $4264.5$ & $912.5$ & no \\
Fashion full & CNN local ensemble $\times4$ & $89.53\%$ & $16477.9$ & $912.5$ & no \\
Fashion full & DOS-NC fused ridge & $89.56\%$ & $26816.8$ & $1531.9$ & no \\
Fashion full & CNN, one epoch & $86.36\%$ & $12664.3$ & $4.4$ & yes \\
\bottomrule
\end{tabular}
\caption{Real-data no-backprop cross-architecture stress test. The no-backprop rows use fixed local operator features and closed-form ridge readouts; the CNN baseline is a small gradient-trained model, not a tuned SOTA model. Full MNIST/Fashion use the standard $60{,}000/10{,}000$ train/test split.}
\label{tab:dendritic-cross-arch-realdata}
\end{table}
Table~\ref{tab:dendritic-cross-arch-realdata} is encouraging but not a moonshot. On the low-data MNIST split, the local CNN ensemble reaches $97.85\%$ and beats the one-epoch backprop CNN by $5.06$ percentage points. On FashionMNIST, the no-backprop local ensemble reaches $89.10\%$ with $10{,}000$ training examples and $89.53\%$ on the full split, beating the bounded one-epoch CNN by $11.50$ and $3.17$ percentage points respectively. However, full MNIST remains below the same one-epoch CNN ($98.07\%$ fused no-backprop versus $98.25\%$), the fixed-attention and grid-GNN views are weak, and the memory estimate for the closed-form full-feature solves is much larger than the small CNN. The correct interpretation is therefore not ``SOTA without backprop.'' The result is a fast, CPU-only sample-efficiency signal for local operator features plus algebraic readouts, and a clear boundary: generic CNN/RNN/attention/GNN replacement will require true stacked local learning and streaming/local normal equations, not merely wider fixed feature banks.
\subsection{No-backprop optimization: neuromodulated local control}
The next runner, \texttt{neuromodulated\_local\_learning\_benchmark.py}, moves from fixed features toward a biologically motivated optimization loop. The intended replacement for global reverse-mode differentiation is not ``no loss.'' It is a different decomposition of the loss. Each cell or local branch receives a local state, a local eligibility trace, and a low-dimensional modulatory innovation. For a branch state $v_{ib}(s,t)$,
\[
\begin{aligned}
\partial_t v_{ib}
&=
D_{ib}\partial_{ss}v_{ib}
-
\lambda_{ib}v_{ib}
+
\sum_j \theta_{ijb}r_j(t)\delta(s-s_{ijb}),\\
\tau_i\dot V_i
&=
-V_i+\sum_b a_{ib}v_{ib}(0,t),
\end{aligned}
\]
the branch-level update should be a three-factor control rule,
\[
\Delta\theta_{ijb}
=
\eta\,m_i(t)e_{ijb}(t)
-
\eta_h\,\partial_{\theta_{ijb}}\mathcal{H}_{ijb},
\qquad
e_{ijb}(t)
=
\int r_j(\tau)G_{ib}(t-\tau)\chi_{ib}(\tau)\,d\tau .
\]
Here $e_{ijb}$ is the local eligibility trace induced by the dendritic Green function, $m_i$ is a reward, dopamine, prediction-error, or observation-innovation field available to the cell or region, and $\mathcal{H}$ is a homeostatic stability cost. This is closer to feedback control than to backpropagation. The global task loss is allowed to create a modulatory signal, but it is not differentiated through every downstream operation to produce an exact adjoint for every upstream synapse.
This leads to a different architecture search space. A layer should be a population of multicompartment cells with lateral competition and residual identity highways,
\[
x_{\ell+1}
=
x_\ell
+
P_\ell\,\Gamma_\ell(\{V_{\ell i}\}_i),
\]
where $\Gamma_\ell$ can include soma thresholds, local winner-take-all inhibition, liquid leak filters, and branch-local sparse solves. The residual path is not a stylistic copy of ResNets; it is a stability/control channel that prevents local cell updates from having to preserve the whole signal while they learn a correction. Attention should also be reinterpreted. Instead of backpropagating through dense learned QKV matrices, a biological attention analogue stores local keys or prototypes, routes by kernel similarity and competition, and updates the keys only when a modulatory innovation indicates that a region was informative or surprising. Transformer-like selectivity is still needed, but the learning mechanism must be local memory deposition and residual routing rather than exact gradient transport.
The benchmark implements a first small version of this idea on MNIST and FashionMNIST. It learns patch filters by local competitive quantization, learns class-gated prototype keys by per-class local clustering, solves the readout by ridge normal equations, then performs a residual-memory correction:
\[
R_0 = Y-\widehat Y_0,\qquad
C_c=\operatorname{top}_{K}\{x_i:y_i=c,\|R_{0,i}\|_2\},\qquad
A^\star
=
\arg\min_A\|\Phi_C A-R_0\|_F^2+\lambda\|A\|_F^2,
\]
and predicts by $\widehat Y=\widehat Y_0+\Phi_C A^\star$. The residual centers $C_c$ are a simple dopamine analogue: high-innovation examples deposit class-local memory, and the correction controller is fitted algebraically. No representation layer is trained by reverse-mode AD.
\begin{table}[h]
\centering
\scriptsize
\setlength{\tabcolsep}{3pt}
\begin{tabular}{@{}llrrrr@{}}
\toprule
Dataset/split & Profile & Accuracy & Train ms & Est. MB & Backprop \\
\midrule
MNIST $10\mathrm{k}/10\mathrm{k}$ & Fixed local CNN ridge & $96.76\%$ & $2052.7$ & $660.7$ & no \\
MNIST $10\mathrm{k}/10\mathrm{k}$ & Hebbian CNN ridge & $86.76\%$ & $3206.0$ & $1113.0$ & no \\
MNIST $10\mathrm{k}/10\mathrm{k}$ & Dopamine attention ridge & $93.56\%$ & $3970.2$ & $52.1$ & no \\
MNIST $10\mathrm{k}/10\mathrm{k}$ & NML fused local ridge & $98.04\%$ & $2175.3$ & $1193.9$ & no \\
MNIST $10\mathrm{k}/10\mathrm{k}$ & Dopamine residual memory & $98.12\%$ & $2518.2$ & $1294.8$ & no \\
MNIST $10\mathrm{k}/10\mathrm{k}$ & CNN, one epoch & $92.79\%$ & $2076.1$ & $4.4$ & yes \\
\midrule
Fashion $10\mathrm{k}/10\mathrm{k}$ & Fixed local CNN ridge & $87.90\%$ & $2084.2$ & $660.7$ & no \\
Fashion $10\mathrm{k}/10\mathrm{k}$ & Hebbian CNN ridge & $88.28\%$ & $3166.6$ & $1113.0$ & no \\
Fashion $10\mathrm{k}/10\mathrm{k}$ & Dopamine attention ridge & $80.75\%$ & $4008.6$ & $52.1$ & no \\
Fashion $10\mathrm{k}/10\mathrm{k}$ & NML fused local ridge & $88.31\%$ & $1967.0$ & $1193.9$ & no \\
Fashion $10\mathrm{k}/10\mathrm{k}$ & Dopamine residual memory & $88.53\%$ & $2827.5$ & $1611.4$ & no \\
Fashion $10\mathrm{k}/10\mathrm{k}$ & CNN, one epoch & $77.60\%$ & $2075.5$ & $4.4$ & yes \\
\bottomrule
\end{tabular}
\caption{Neuromodulated local-learning stress test. Patch filters and prototype keys are learned by local clustering, readouts are closed-form ridge solves, and the residual-memory row uses high-innovation examples as class-local memory centers. The Fashion residual-memory row uses $256$ centers per class; the MNIST row uses $64$ centers per class.}
\label{tab:neuromodulated-local-learning}
\end{table}
Table~\ref{tab:neuromodulated-local-learning} gives a precise result rather than the desired universal breakthrough. On MNIST, the residual-memory controller improves the fused no-backprop model from $98.04\%$ to $98.12\%$, beating the one-epoch CNN control by $5.33$ points on the same split. On FashionMNIST, increasing residual memory from $64$ to $128$ and $256$ centers per class improves the residual row from $88.39\%$ to $88.51\%$ and $88.53\%$, but it still remains below the earlier four-bank fixed local ensemble at $89.10\%$. The attention-like prototype branch is also weak as a stand-alone model. The conclusion is important: a scalar/vector modulatory innovation can improve no-backprop local learning, but the current single-stage memory deposition is still too shallow and too memory-heavy to replace stacked backprop-trained architectures at SOTA scale. The next serious architecture must stack these local residual controllers, expose intermediate local targets or predictive residuals at each layer, and update the residual memory by streaming/local normal equations rather than by one global dense solve.
\subsection{Cell-operator mismatch audit: poles, conductance, and activation}
The criticism of the previous liquid/cellular experiments is correct: a hand-chosen linear cable basis with a sigmoid release transform is not yet the correct neuron model. Hasani's LTC formulation and the exact multi-synapse extension instead make the synapse the nonlinear operator \cite{hasani2020ltc,hasani2022cfc,cantini2025exact}. In scalar form,
\[
\dot x(t)
=
-\omega x(t)
+
\sum_{s=1}^{S} f_s(g_s(t);\theta_s)\bigl(A_s-x(t)\bigr),
\]
so the instantaneous pole is not fixed. It is
\[
p(t)=-\left(\omega+\sum_s f_s(g_s(t);\theta_s)\right),
\]
and the driving equilibrium is the conductance-weighted reversal potential. This means that the ``activation function'' is not a pointwise ReLU/SIREN-style nonlinearity after a linear map. It is a synaptic conductance field that simultaneously controls gain, sign, equilibrium, and time constant. A linear exponential-pole dictionary can approximate its traces, but it is structurally mismatched because it does not include the multiplicative feedback term $(A_s-x)$.
The diagnostic runner \texttt{cellular\_operator\_model\_audit.py} isolates this issue. It generates a teacher from the exact zero-order-hold multi-synapse LTC recurrence
\[
x_{k+1}
=
\gamma_k x_k
+
(1-\gamma_k)
\frac{\sum_s f_s(g_{s,k};\theta_s)A_s}{\omega+\sum_s f_s(g_{s,k};\theta_s)},
\qquad
\gamma_k=
\exp\left[-\Delta t_k\left(\omega+\sum_s f_s(g_{s,k};\theta_s)\right)\right],
\]
then compares four no-backprop identification families: fixed linear poles, conductance with wrong gates, conductance with oracle gates, and sparse search over an overcomplete conductance-gate dictionary. The conductance learners use the locally observed voltage and solve
\[
\dot x+\omega x
=
\sum_s w_s f_s(g_s;\theta_s)(A_s-x)
\]
by ridge or OMP/ridge, followed by exact rollout. This is a local operator-identification rule, not reverse-mode training through a network.
\begin{table}[h]
\centering
\scriptsize
\setlength{\tabcolsep}{3pt}
\begin{tabular}{@{}lrrrrr@{}}
\toprule
Profile & Test RMSE & PSNR & Terms & Train ms & Est. MB \\
\midrule
Linear raw multi-pole ridge & $2.8644\mathrm{e}{-02}$ & $16.947$ & $49$ & $157.6$ & $0.58$ \\
Linear sigmoid multi-pole ridge & $2.5585\mathrm{e}{-02}$ & $17.928$ & $49$ & $161.7$ & $0.58$ \\
Wrong-gate conductance ID & $8.2627\mathrm{e}{-03}$ & $27.745$ & $8$ & $12.0$ & $0.09$ \\
Oracle-gate conductance ID & $3.0990\mathrm{e}{-04}$ & $56.263$ & $8$ & $11.8$ & $0.09$ \\
Dense grid-gate conductance ID & $6.3447\mathrm{e}{-02}$ & $10.040$ & $384$ & $476.5$ & $5.04$ \\
Sparse grid-gate OMP ID & $2.1460\mathrm{e}{-03}$ & $39.455$ & $24$ & $39.6$ & $4.48$ \\
\bottomrule
\end{tabular}
\caption{Cell-operator mismatch audit on a synthetic exact multi-synapse LTC teacher with $16$ training sequences, $64$ test sequences, $192$ time steps, $8$ synapses, and observation noise $10^{-3}$. Linear pole dictionaries are the wrong operator family. Matched conductance identification is both more accurate and leaner. Blind dense gate expansion is ill-conditioned; sparse gate selection is the viable unknown-operator path.}
\label{tab:cellular-operator-mismatch-audit}
\end{table}
Table~\ref{tab:cellular-operator-mismatch-audit} identifies the current mistake sharply. The old fixed-pole view is not merely under-tuned; it is the wrong operator for an LTC-style cell. Matching the conductance law reduces test RMSE by roughly $83\times$ versus the best linear-pole row while using only $8$ terms and about $0.09$ MB in this audit. Even a wrong fixed gate improves substantially over linear poles, proving that the multiplicative reversal-potential structure matters. Conversely, a dense overcomplete gate dictionary fails, while sparse OMP over gate candidates recovers much of the gap. The next no-backprop architecture should therefore start with constrained local identification of $\omega$, $A_s$, $\theta_s$, and active synapses, under positivity/stability bounds, before any CNN/RNN/transformer-scale benchmark. Architecture comes after the cell operator is right.
\subsection{Grown-topology operator networks and sample-efficient closed-form identification}
\label{sec:grown-topology}
The cell-operator mismatch audit establishes that an LTC-style cell is governed by a
conductance operator, not a fixed-pole linear filter. This subsection develops the
learning-time consequence. Mainstream artificial networks fix the architecture in
advance and brute-force a generic function approximator by backpropagation. Biological
networks instead \emph{grow}: capacity is added developmentally while the system is
learning, and the learning rules and topology themselves were shaped by evolution
\cite{stanley2002neat}. A single biological neuron is correspondingly far richer than a
weighted-sum-plus-activation unit; a layer-five pyramidal cell requires a five-to-eight
layer temporal network to reproduce \cite{beniaguev2021single}. These two observations
motivate a different training regime, which we state as a falsifiable thesis.
\paragraph{Thesis.} For systems whose structure is \emph{specifiable} as a known
operator family --- a connectome-shaped dynamical system, a governing differential
operator --- one should not learn a generic function. One should parameterize the
operator and identify its few free parameters: solve everything that is linear in its
coefficients by an OSNR closed-form solve, and reserve a small gradient-free
evolutionary search for the nonlinear and structural parameters, growing the topology
with warm starts. The claim is that such a model matches or beats a backpropagation
network of equal budget on \emph{sample-efficiency, parameter count, and
out-of-distribution robustness}, not on raw task score.
\paragraph{Two-timescale decomposition.} The regime separates exactly along the
linear/nonlinear boundary already used throughout OSNR.
\begin{itemize}[leftmargin=1.4em]
\item \textbf{Inner (fast), linear.} Given a frozen nonlinearity shape and topology,
the conductance/readout coefficients enter the operator linearly. They are recovered
by a single scale-invariant ridge solve over an operator-matched feature dictionary,
with derivatives supplied by the autograd-free spline ladder of
Section~\ref{sec:autograd-free} rather than by automatic differentiation. No
iteration, no backward graph.
\item \textbf{Outer (slow), nonlinear and structural.} Time constants, synaptic
gains, gate shapes, and the topology itself are searched by an evolutionary method
(CMA-ES/PGPE \cite{sehnke2010pgpe,salimans2017es}) with NEAT-style complexification
\cite{stanley2002neat}. A neuron is added only when capacity saturates, and its
outgoing weights are initialized near zero so the addition is a function-preserving
no-op \cite{chen2016net2net}; the inner solve then re-fits in milliseconds and the new
unit is retained only if held-out fitness improves.
\end{itemize}
The mutual dependency is the point: NEAT-style growth is normally bottlenecked because
scoring each candidate topology needs a full training run, while a closed-form inner
solve makes candidate evaluation nearly free. Cheap convex identification and
evolutionary growth each enable the other. This is the operator-spline counterpart of
sparse governing-equation discovery \cite{brunton2016sindy} and universal differential
equations \cite{rackauckas2020universal}, specialized to grown conductance networks.
\paragraph{Rung~0: verifying the identification engine.} Before any closed-loop control
study, the premise must be checked in isolation: when the operator form is exactly
known, is the closed-form solve genuinely more sample-efficient and cheaper than
backpropagation on the \emph{same} model class? The runner
\texttt{bio\_growth/rung0\_osnr\_id\_verification.py} fixes a known LTC teacher with
$N{=}8$ neurons, $M{=}2$ inputs, a known sigmoidal synaptic feature bank, and unknown
$(\tau_i, w_{ij}, v_{ik}, A_i)$. The expanded right-hand side lies exactly in the
matched dictionary
\[
\Phi_i(\mathbf{x},\mathbf{I})
=
\bigl[\,x_i,\;\{\sigma_j(\mathbf{x})\}_j,\;\{x_i\sigma_j(\mathbf{x})\}_j,\;
\{I_k\}_k,\;\{x_iI_k\}_k\,\bigr],
\]
so identification reduces to per-neuron ridge regression of $\dot{x}_i$ onto $\Phi_i$.
Four methods of the identical model class are compared on $24$ held-out clean
trajectories (rollout normalized RMSE): \textbf{OSNR}, closed-form ridge on a
Tikhonov/curvature-smoothed derivative (the operator-spline derivative);
\textbf{FD}, the same closed-form ridge on a raw finite-difference derivative ---
i.e. the exact minimizer of the one-step linear least-squares objective;
\textbf{1-step SGD}, that same linear objective optimized by Adam for $4000$ epochs;
and \textbf{rollout BPTT}, the naive recurrent fit by backpropagation through a
$300$-step integrator ($300$ epochs). Observation noise is $10\%$ of per-state standard
deviation.
\begin{table}[H]
\centering
\begin{tabular}{rcccccc}
\toprule
trajectories & OSNR nRMSE & OSNR time & FD nRMSE & 1-step SGD nRMSE & rollout BPTT nRMSE & BPTT time \\
\midrule
$1$ & $9.146$ & $0.00$\,s & $10.136$ & $4.139$ & $0.134$ & $40.5$\,s \\
$2$ & $1.010$ & $0.01$\,s & $2.267$ & $1.603$ & $0.088$ & $45.5$\,s \\
$4$ & $0.0815$ & $0.01$\,s & $2.793$ & $1.314$ & $0.103$ & $45.3$\,s \\
$8$ & $0.0387$ & $0.03$\,s & $0.0771$ & $0.734$ & $0.109$ & $45.2$\,s \\
$16$ & $0.0318$ & $0.06$\,s & $0.0585$ & $0.842$ & $0.0494$ & $46.0$\,s \\
$32$ & $0.0276$ & $0.11$\,s & $0.0215$ & $0.189$ & $0.0481$ & $47.3$\,s \\
\bottomrule
\end{tabular}
\caption{Rung~0 LTC operator identification (held-out rollout nRMSE) versus number of
training trajectories at $10\%$ observation noise. The closed-form OSNR solve runs in
$0.01$--$0.11$\,s versus $\sim\!45$\,s for backpropagation-through-time --- a
$400$--$4000\times$ wall-clock reduction --- and from four trajectories upward it is
also more accurate than the trained recurrent BPTT fit. Finite-difference closed-form is
the exact one-step least-squares optimum; one-step SGD on the identical objective has
not reached it after $4000$ epochs, illustrating that the direct solve dominates
iterative optimization even on the linear sub-problem at fixed budget.}
\label{tab:rung0-sample-efficiency}
\end{table}
\begin{table}[H]
\centering
\begin{tabular}{rcc}
\toprule
observation noise & OSNR-spline nRMSE & finite-difference nRMSE \\
\midrule
$0\%$ & $0.0357$ & $0.0225$ \\
$5\%$ & $0.0361$ & $0.0303$ \\
$10\%$ & $0.0387$ & $0.0771$ \\
$20\%$ & $0.0626$ & $1.509$ \\
$40\%$ & $0.674$ & $5.662$ \\
\bottomrule
\end{tabular}
\caption{Rung~0 noise robustness at eight training trajectories. The operator-spline
(Tikhonov-curvature) derivative is the active ingredient: at low noise it is unnecessary
(finite difference is marginally better, since the smoother tends to the identity), but
finite-difference identification collapses as noise grows while spline-OSNR degrades
gracefully --- a $24\times$ advantage at $20\%$ noise.}
\label{tab:rung0-noise}
\end{table}
\paragraph{Interpretation and honest scope.} Three conclusions hold robustly.
First, the wall-clock advantage is unconditional: a deterministic closed-form solve in
tens of milliseconds replaces tens of seconds of backpropagation, which is precisely the
property that makes evolutionary topology growth affordable. Second, the operator-spline
derivative, not merely the closed form, is what buys sample-efficiency and noise
robustness: the finite-difference control collapses at four trajectories
(Table~\ref{tab:rung0-sample-efficiency}) and under noise
(Table~\ref{tab:rung0-noise}), whereas the spline-smoothed solve remains accurate. Third,
from four trajectories upward the closed-form identification is at least as accurate as a
fully trained recurrent backpropagation fit. The honest caveats are equally explicit. In
the data-starved regime ($\le 2$ trajectories) the closed-form solve is unstable while
rollout BPTT, which is implicitly regularized by having to produce a stable trajectory,
is more robust; a low-data identification therefore needs stronger rank-revealing
regularization. And the accuracy comparison is against a recurrent BPTT fit whose
difficulty is partly the long-horizon credit-assignment problem the closed form
sidesteps --- so the unconditional claim is wall-clock and compute, with the accuracy
advantage holding once a minimal data threshold is met. Rung~0 thus validates the
inner-solve premise and clears the path to the closed-loop control study (Rung~1), where
the evolutionary outer loop and warm-started growth are exercised directly.
\paragraph{Rung~1: closed-loop control, and an honest negative.} We exercised the full
regime on a 2D chemotaxis control task (a noisy gradient-climbing agent, the canonical
\emph{C.\ elegans} behaviour), with the policy a small liquid reservoir whose readout is the
closed-form solve and whose dynamics and topology are grown by evolution. Two findings, one
methodological and one sobering. First, behaviour cloning from a privileged teacher fails for
\emph{both} the structured network and a backpropagation baseline, because teacher-forced
training drifts off-distribution in closed loop; reframing the task as direct reward
optimisation (the evolutionary outer loop) fixes this and the structured policy solves the
task with $\sim\!50$ parameters. Second, and honestly, the architectural advantage on control
is \emph{modest}: at convergence a generic recurrent network nearly matches the structured
liquid network on task score and robustness, and the structured model's remaining edge is
roughly a factor of four in trained-parameter count, not a decisive win. The clean, decisive
advantages of the operator-matched approach are therefore in \emph{identification}, not
control, which the next sections quantify against the standard identification baselines.
On the related continual-learning axis---catastrophic forgetting, often cited as a place
where biological learning outperforms backpropagation---the operator-matched approach
admits a more ambitious construction that brings together the three ingredients of the
biological thesis: \emph{grow} structure on demand, learn each piece by a \emph{local
closed-form} solve rather than global backpropagation, and \emph{do not overwrite} what
was already learned. We test it in the hardest fair setting: a stream of dynamical regimes
arrives in blocks with \emph{no regime labels} and \emph{no replay}, and the learner must
itself detect when the dynamics have changed. The grown learner maintains a bank of
closed-form matched experts; each incoming window is routed to the expert that best explains
it, a \emph{new} expert is grown whenever the best residual exceeds a novelty threshold, and
a final consolidation pass merges experts that turn out to capture the same regime. Crucially,
the baselines are given the \emph{same} matched feature library, so the comparison isolates the
mechanism (grow-and-consolidate with closed-form solves) rather than the feature prior: a single
shared matched model updated online by stochastic gradient descent, a black-box neural vector
field trained online, the same field regularized by \emph{elastic weight consolidation} (EWC,
the canonical deep continual-learning method \cite{kirkpatrick2017overcoming}, given the task
boundaries and its best regularization strength---an advantage the grown learner is not given),
and---as an upper bound---the same black-box field trained jointly on all regimes with full replay.
The result is decisive and robust across three independent system sets (five regimes each;
Figure~\ref{fig:grown-continual}). The grown learner reaches mean forecast nRMSE $0.007$--$0.024$ with essentially no forgetting of the
first regime ($0.006$--$0.032$), while every backpropagation baseline exhibits textbook catastrophic
forgetting (mean $0.9$--$3.1$, retaining only the most recent regime). Critically, this includes
EWC---the method deep learning built specifically to prevent forgetting: even with task boundaries
handed to it and its regularization strength tuned to its best, EWC ($0.9$--$2.7$) is essentially
no better than naive online training, because the five regimes are genuinely distinct vector fields
and no single network can hold them all. The grown learner is roughly two orders of magnitude
better than EWC. The replay upper bound,
despite seeing every regime jointly, never falls below mean $\approx 1.2$---a single context-free
vector field cannot represent five distinct dynamics at once---so the grown bank is roughly two
orders of magnitude better than even the strongest backpropagation control. The growth mechanism
also recovers the latent structure: from eleven to thirteen experts grown online, consolidation
returns \emph{exactly} the five true regimes on every run. The memory is moreover persistent with
instant recognition: when the entire five-regime sequence is presented a second time, the learner
grows \emph{zero} new experts on the revisit---it routes each returning regime straight back to its
existing expert, with no relearning---whereas the eleven experts were all grown on the first pass. The single hyperparameter---the
novelty threshold for growing---does not need tuning: sweeping it over a $3.5\times$ range
($0.20$ to $0.70$) leaves the forecast accuracy unchanged ($0.0089$--$0.0092$) and the consolidated
regime count exactly five, even though the number of raw experts grown varies from twenty down to
seven; the consolidation pass absorbs the difference. Unlike the naive version, this is not
merely an argument from modularity handed regime labels for free; it is a label-free, online,
self-structuring learner whose only prior is the matched feature family, and on this home turf of
biological learning it decisively outperforms backpropagation.
The advantage is also a \emph{sample-efficiency} advantage, biology's other reputed strength. As
the data per regime is reduced (Figure~\ref{fig:sample-efficiency}), the grown learner degrades
gracefully---mean nRMSE
$0.013\to0.08\to0.21\to0.70$ at eight, four, two, and one trajectory per regime---and already beats
every backpropagation variant, including full replay, at a single trajectory per regime (by roughly
$4\times$, widening to $\sim$$170\times$ at eight). The backpropagation baselines never improve with
more per-regime data because forgetting, not data, is their binding constraint. The grown learner
thus wins on both axes at once: it needs little data per regime and it does not overwrite.
The separation widens as the stream lengthens (Figure~\ref{fig:continual-scaling}). Scaling the
number of sequential regimes from eight to twenty, the grown learner's mean forecast nRMSE stays
flat at $\approx 0.02$ and its forgetting of the first regime is \emph{constant} at $0.009$---adding
twenty regimes after the first degrades the first not at all---while every backpropagation baseline
remains pinned near nRMSE $1$, retaining essentially only the most recent regime. The structure
discovery also stays exact at scale: consolidation returns precisely the true number of regimes at
every $K$ tested (recovering eight, twelve, sixteen, and twenty experts from fourteen, twenty-six,
thirty-four, and forty-two grown online). The grown learner therefore scales the way biological
memory is supposed to: capacity is added as needed, old skills are untouched, and the cost of a
new regime is one new expert rather than interference with all the others.
\begin{figure}[H]
\centering
\includegraphics[width=0.92\linewidth]{figures/osnr_grown_continual.png}
\caption{Label-free, online, growing continual learning across five sequential dynamical regimes
(no labels, no replay), averaged over three independent system sets; bars show the mean and
whiskers the min--max across sets, on a log scale. Left: mean forecast nRMSE over all regimes after
the full stream. Right: forgetting, measured as the error on the first regime once the stream has
ended. The grown learner (grow a closed-form matched expert on novelty, then consolidate
duplicates) is roughly two orders of magnitude better than a single shared matched model updated
online, a black-box neural field trained online, EWC (the canonical deep continual-learning method,
given task boundaries and its best regularization strength), and even a black-box field trained
jointly with full replay---all of which see the same matched features. Growth recovers exactly the
five true regimes on every run. This is the biological recipe---grow, learn locally in closed form, do not
forget---winning on its home turf.}
\label{fig:grown-continual}
\end{figure}
\begin{figure}[H]
\centering
\includegraphics[width=0.56\linewidth]{figures/osnr_sample_efficiency.png}
\caption{Sample efficiency on the same five-regime continual stream: mean forecast nRMSE versus the
number of trajectories seen per regime (log--log). The grown learner improves steeply with data and
sits below both the online and full-replay backpropagation baselines at every data budget, beating
them even at a single trajectory per regime. The backpropagation curves stay flat because their
limiting factor is catastrophic forgetting, not the amount of per-regime data.}
\label{fig:sample-efficiency}
\end{figure}
\begin{figure}[H]
\centering
\includegraphics[width=0.92\linewidth]{figures/osnr_continual_scaling.png}
\caption{Scaling the continual stream from eight to twenty sequential regimes. Left: mean forecast
nRMSE stays flat at $\approx 0.02$ for the grown learner while the online and full-replay
backpropagation baselines remain near $1$ regardless of stream length. Right: the consolidated
expert count (solid) tracks the true number of regimes (dashed) exactly at every scale, while the
raw number of experts grown online (dotted) runs ahead before consolidation collapses it. Capacity
grows with the task; old regimes are not disturbed.}
\label{fig:continual-scaling}
\end{figure}
\paragraph{Generalization beyond dynamical systems: Permuted- and Rotated-MNIST.} The
grow-and-consolidate mechanism is not specific to differential equations. To test it on a benchmark the continual-learning
community actually tracks, we apply the identical recipe to Permuted-MNIST (ten tasks, each a fixed
random permutation of the $784$ pixels, presented sequentially without replay). The only change is
the per-task solver: a frozen bank of random ReLU features feeds a per-task closed-form ridge
classifier, grown on each task. Against the same EWC, online, and joint-replay deep baselines (an
MLP with a shared head), the grown bank attains $96.1\%$ average accuracy across the ten tasks when
the task is known at test time---exceeding EWC at its best ($87.6\%$) and even the joint upper bound
($95.8\%$), with no backpropagation and no forgetting (the online MLP collapses to $68.5\%$,
retaining $38\%$ on the first task). When the task is \emph{not} given it must be inferred from the
input. A confidence router (send each image to the most confident expert) recovers $81.6\%$; a
simple generative router---a per-task diagonal Gaussian over the same random features, choosing the
task of highest likelihood---routes \emph{perfectly} ($100\%$), so task-free accuracy equals the
task-known accuracy at $96.1\%$, again well above EWC. The earlier task-free shortfall was thus a
weak router, not a limitation of the modular learner: the per-task feature distributions are cleanly
separable. The same recipe on Rotated-MNIST (ten tasks, each a fixed rotation of the digits) is also
decisive: $96.1\%$ task-known and $86.9\%$ task-free, both far above EWC's $\sim$$64$--$71\%$ on this
harder shift. Interestingly the routers swap roles here---adjacent rotation angles overlap
distributionally, so the Gaussian router degrades, but the confidence router is rescued by
cross-generalization (routing to a neighboring-angle expert still classifies the digit). In every
case at least one simple router beats EWC. Across both standard benchmarks, then, the grown learner
beats the dedicated deep continual-learning method in all four settings (task-known and task-free,
permuted and rotated), with no backpropagation and no forgetting---and where the tasks are
distributionally distinct, task inference is essentially exact, so task-free operation is free.
The most stringent test is \emph{class-incremental} learning, where deep methods are weakest: on
Split-MNIST (five tasks of two digits each, no task label at test, classification over all ten
classes), EWC and naive online training both collapse to about $19$--$20\%$, the well-documented
class-IL failure of regularization-based methods, since a single shared head cannot keep ten
once-seen classes separable. The grown bank instead reaches $91.1\%$ with a shared-covariance Mahalanobis router (and
$99.6\%$ with an oracle task label, confirming the per-task experts are near-perfect and the only
loss is in routing), against a joint upper bound of $97.5\%$. That is roughly a four-and-a-half-fold
improvement over EWC in precisely the setting deep continual learning finds hardest---again with no
backpropagation, no replay, and no forgetting. Figure~\ref{fig:continual-vision} summarizes the
three vision settings.
\begin{figure}[H]
\centering
\includegraphics[width=0.78\linewidth]{figures/osnr_continual_vision.png}
\caption{Continual learning on standard vision benchmarks with no task label at test. Across
Permuted-MNIST and Rotated-MNIST (domain-incremental, task-free) and Split-MNIST
(class-incremental), the closed-form grown bank (blue) stays close to the joint upper bound (green)
and well above EWC (red), the deep continual-learning baseline---most dramatically in the
class-incremental setting, where EWC collapses to near chance. No backpropagation, no replay, no
forgetting.}
\label{fig:continual-vision}
\end{figure}
\paragraph{Scaling up, and a fair fight against strong replay.} Two caveats must be met for these
results to mean anything to the continual-learning community: EWC is by now a weak baseline, and
MNIST is a toy. We therefore move to a frozen ImageNet-pretrained ResNet-18 backbone (the
parameter-efficient protocol in which all methods share identical features, so only the
continual mechanism differs) and to Split-CIFAR, scored against \emph{dark experience replay}
(DER++ \cite{buzzega2020dark}), a strong rehearsal baseline rather than EWC. On Split-CIFAR-10
class-incremental, the closed-form bank reaches $77\%$ versus $49\%$ for EWC. On the standard
Split-CIFAR-100 (ten tasks of ten classes), the per-task-routed bank is held back by the harder
ten-way task inference ($48\%$ routing), but the \emph{class-incremental} instance of the same
closed-form idea---a per-class prototype classifier with a shared-covariance (Mahalanobis) metric,
which is simply a prototype grown for each class as it is seen---reaches $57.0\%$, ahead of DER++
at $47.6\%$ and essentially at the joint upper bound of $58.7\%$, while EWC and online training
collapse to $9\%$ (Figure~\ref{fig:cifar100-classil}). The point is not that the prototype
classifier is novel---nearest-class-mean on
frozen features is a known strong rehearsal-free baseline---but that the entire family is
\emph{closed-form}: it carries no replay buffer, takes no gradient step, and fits in under a second,
whereas DER++ requires backpropagation and a two-thousand-example buffer for a lower score. This is
exactly the efficiency axis the field has turned to.
The same picture holds, and sharpens, on the backbone the prompt-based literature actually uses. With
a frozen ImageNet-21k ViT-B/16 (features extracted once on a laptop GPU), Split-CIFAR-100
class-incremental accuracy rises to $88.0\%$ for the Mahalanobis prototype classifier and $89.8\%$
for a random-projection Gram-ridge variant (the closed-form core of RanPAC \cite{mcdonnell2023ranpac}),
versus $86.1\%$ for DER++ and a collapse to $16\%$ for EWC; the published numbers for prompt-tuning
on the identical backbone are roughly $83$--$84\%$ (L2P \cite{wang2022l2p}) and $84$--$86\%$
(DualPrompt \cite{wang2022dualprompt}). Our closed-form classifiers thus exceed DER++ and the prompt
methods while training in seconds with no buffer and no gradient step. We are explicit that the
prototype and random-projection classifiers are \emph{not} our invention---they are the established
strong rehearsal-free baselines of this regime---and that we did not re-run the prompt methods; the
contribution is the unification (the same closed-form, operator-/structure-matched principle that
identifies dynamical systems and PDEs also drives a competitive continual-vision learner) together
with the efficiency demonstration: in the frozen-backbone regime, closed-form modular learning
matches or beats strong replay and prompt-tuning at one to two orders of magnitude less compute and
no stored data.
We also report the boundary honestly. On Split-ImageNet-R---a deliberately harder benchmark whose
renditions shift \emph{away} from the pretraining distribution---the same closed-form classifiers
reach $67\%$ (random-projection) and $66\%$ (Mahalanobis), still above DER++ ($59\%$) and the
published L2P ($\sim$$61$--$65\%$), and on par with DualPrompt ($\sim$$66$--$69\%$), but now
\emph{below} CODA-Prompt ($\sim$$73$--$75\%$) and the full RanPAC ($\sim$$74$--$78\%$). The gap is
attributable to first-session backbone adaptation, a one-time gradient pass those methods include
and our purely closed-form variant omits; restoring it is possible but would forfeit the
zero-backpropagation property that is the point here. The honest summary across both benchmarks is
that closed-form modular learning is at or near the accuracy frontier while being categorically
cheaper---decisively so when the frozen features already suit the data, competitively so when they
do not.
Backbone quality, not the classifier, turns out to be the lever. Swapping the supervised ViT-B for a
self-supervised DINOv2 ViT-L/14 (still frozen, features extracted once on a laptop GPU) raises the
closed-form Split-ImageNet-R class-incremental accuracy to $89.6\%$ for the random-projection variant
and $84.1\%$ for the Mahalanobis prototype, above DER++ ($86.7\%$) on the same features and well
above the published prompt-tuning and RanPAC numbers reported on ViT-B ($\sim$$61$--$78\%$). We are
careful about what this does and does not show: it is a stronger-backbone result, not a
classifier-versus-classifier victory over RanPAC, whose closed-form core our random-projection
variant essentially is. The honest takeaway is that the entire closed-form family scales with
backbone quality at no training cost---a better frozen representation is free to adopt and turns a
covariance solve into a state-of-the-art-rivaling continual learner---so the practical frontier in
this regime is set by representation quality and a closed-form read-out, not by the expensive
prompt- or replay-based adaptation machinery.
Table~\ref{tab:ptmcil} places our method on the canonical rehearsal-free pretrained-model CIL
leaderboard \cite{mcdonnell2023ranpac} (ten tasks, final average accuracy, no rehearsal buffer). On
the identical ViT-B/16-in21k backbone our closed-form learner uses no gradient step at all, yet it
surpasses every prompt-based method---L2P, DualPrompt, CODA-Prompt---and the adapter method ADaM, and
it reproduces RanPAC's own no-adaptation ablation ($89.9\%$ measured versus their $89.0\%$ on
CIFAR-100). The one method above us is the full RanPAC, whose advantage is precisely its first-session
\emph{gradient} adaptation of the backbone; matching that with a purely closed-form routine is left
open, and we instead claim the strongest \emph{zero-backpropagation} position on the board. With a
stronger frozen backbone (DINOv2 ViT-L/14) the same closed-form method reaches $92.3\%$ on CIFAR-100
and $89.6\%$ on ImageNet-R, exceeding the best published numbers---though we note two caveats in
fairness: full RanPAC would also benefit from the stronger backbone, and our ImageNet-R figure uses an
$80/20$ split rather than the reference split.
\begin{table}[H]
\centering
\small
\begin{tabular}{lcc c}
\toprule
Method (ViT-B/16-in21k, final acc.) & CIFAR-100 & ImageNet-R & gradient? \\
\midrule
L2P \cite{wang2022l2p} & 84.6 & 72.5 & yes (prompts) \\
DualPrompt \cite{wang2022dualprompt} & 81.3 & 71.0 & yes (prompts) \\
CODA-Prompt & 86.3 & 75.5 & yes (prompts) \\
ADaM & 87.6 & 72.3 & yes (adapter) \\
RanPAC (full, with FSA) \cite{mcdonnell2023ranpac} & \textbf{92.2} & \textbf{78.1} & yes (FSA) \\
RanPAC, no FSA (= ours) & 89.0 & 71.8 & \textbf{no} \\
NCM only & 83.4 & 61.2 & \textbf{no} \\
\midrule
\textbf{Ours, closed-form (ViT-B/16-in21k)} & \textbf{89.9} & -- & \textbf{no} \\
\textbf{Ours, closed-form (DINOv2 ViT-L/14)} & \textbf{92.3} & \textbf{89.6} & \textbf{no} \\
\bottomrule
\end{tabular}
\caption{Rehearsal-free pretrained-model class-incremental learning (final average accuracy, ten
tasks; published numbers from \cite{mcdonnell2023ranpac}). Among methods that take \emph{no gradient
step}, ours is the strongest; the only entry above it on the shared backbone is full RanPAC, whose
edge is a gradient first-session adaptation. A stronger frozen backbone lifts the closed-form method
past the best published numbers (caveats in text).}
\label{tab:ptmcil}
\end{table}
\begin{figure}[H]
\centering
\includegraphics[width=0.6\linewidth]{figures/osnr_cifar100_classil.png}
\caption{Split-CIFAR-100 class-incremental on a frozen ResNet-18, against a strong replay baseline.
The closed-form per-class prototype classifier (blue) edges out DER++ (orange) and reaches the joint
upper bound (green), while EWC (red) collapses---and it does so with no replay buffer and no
gradient steps, in under a second versus several seconds plus a two-thousand-example buffer for
DER++.}
\label{fig:cifar100-classil}
\end{figure}
\subsection{The sparse-stochastic-process view: when does matched beat random?}
\label{sec:ssp-view}
The continual-vision results motivate a conditional modeling question:
when does a physically matched representation improve the task beyond a
generic feature map? The SSP framework \cite{unser2014sparse1,unser2014sparse2}
supplies the innovation model $Ls=w$, with an admissible inverse and stated
boundary conditions. The generalized white noise $w$ is defined through test
functions, not pointwise samples; $L$ whitens and $L^{-1}$ colors.
For a suitable localization filter $L_d$, its useful spline bridge is
$L_d s=\beta_L*w$, where $\beta_L=L_dL^{-1}\delta$.
Overlapping increment kernels can retain dependence. Non-Gaussian
compressibility is not necessarily finite-rate innovation; a Green-atom
expansion for atomic forcing is not a universal cardinal expansion of white noise.
\paragraph{Correction of an earlier overstatement.}
A random feature map is not itself the driving white noise, and a frozen
vision embedding has not thereby been shown to be white or Gaussian.
We withdraw the claimed if-and-only-if theorem that every non-scalar $L$
gives a strict matched-ridge risk advantage. For orthonormal design,
ridge has effective degrees of freedom $d/(1+\lambda)$ and risk
$(\lambda^2\|c\|^2+\sigma^2d)/(1+\lambda)^2$ under independent zero-mean
variance-$\sigma^2$ noise, irrespective of the unknown coefficient support.
An orthogonal feature rotation preserves all ridge predictions. A non-scalar
orthogonal $L$ can leave Gaussian covariance isotropic; a nonorthogonal
inverse can concentrate it despite dense Gaussian innovations.
Five executable checks in \texttt{tests/test\_ssp\_claim\_boundaries.py}
establish these elementary counterexamples. They are not a new SSP theorem.
The historical feature-swap results remain: PCA whitening $0.881$,
linear discriminant $0.849$, Nystr\"om-RBF $0.846$, covariance-shaped random
weights $0.894$, Student-$t$ $0.894$, Laplace $0.895$, and Gaussian random
projection $0.894$. In the reported ten-examples-per-class comparison,
the $\ell_1$ readout obtains $0.827$ versus ridge $0.832$.
These outcomes establish no advantage for those tested alternatives, not
Gaussian optimality, absence of all exploitable structure, or universal
optimality of random projections. An $\ell_1$ estimator is not generically
optimal for every sparse stochastic process.
The defensible interpretation is conditional: appropriate operator priors,
observation models, regularization and computational structure can help.
The cited dynamical-system/PDE gains are measured under their particular
information and baseline protocols. Improvement in those experiments is not
a universal ordering of representations, and changing coordinates within
one function space is distinct from changing that space or its prior.
\subsection{Closed-form meta-adaptation as a stability primitive for recursive self-improvement}
\label{sec:rsi}
Fixed-feature pooled estimation is useful in a \emph{recursive self-improvement}
loop because it retains earlier objective contributions without raw-data replay.
It is not structurally immune to catastrophic forgetting or self-label error
amplification. With conflicting labels the pooled optimum can worsen an old
task; repeated erroneous pseudo-labels can bias the sufficient statistics.
A frozen backbone and exact readout $W=(G+\lambda I)^{-1}C$ prevent backbone
updates and avoid iterative readout-solver error, but neither fact guarantees
correct self-generated supervision. The following low-drift comparisons are
empirical outcomes of their stated protocols, not consequences of universal
immunity.
We test this directly with a ``telephone-game'' self-training loop on a frozen backbone. Starting
from a tiny labelled seed (one or two examples per class on CIFAR-100 features), the agent repeatedly
pseudo-labels a fresh batch of unlabelled data with its \emph{own current model}, folds it into its
experience, and updates---for many generations, with no further ground truth. The only thing that
differs between the two learners is the update rule: an online gradient step on each self-labelled
batch (the dense-weight route), versus accumulation into the closed-form Gram memory. The outcome is
unambiguous. The closed-form learner \emph{bootstraps and stays stable}: from a single label per
class it climbs from $49\%$ to $62\%$ and holds there, its self-generated labels \emph{improving}
across generations ($0.42\to0.62$). The gradient learner \emph{collapses}: it drifts downward from
$40\%$ to $34\%$ as its pseudo-label accuracy decays generation over generation---the textbook
autoregressive failure---ending roughly $28$ points below the closed-form learner (the same pattern
holds at two labels per class, $73\%$ stable versus $44\%$ and falling). The closed-form update is
thus not merely a cheaper continual learner; it is a mechanism for self-improvement that accumulates
capability without forgetting or drift where the gradient loop degenerates. The pattern is not an artifact of one
backbone: across a supervised ViT-B, a self-supervised DINOv2 ViT-L, and a weaker ResNet-18, and on
both CIFAR-100 and ImageNet-R features, the closed-form learner stays stable or improves across
generations while the gradient learner drifts downward in every case. This connects the
operator-matched, closed-form philosophy of this paper to the stability of open-ended, self-improving
systems (Figure~\ref{fig:rsi}).
\begin{figure}[H]
\centering
\includegraphics[width=0.62\linewidth]{figures/osnr_rsi_selftrain.png}
\caption{Recursive self-training (``telephone game'') from a one-label-per-class seed on a frozen
backbone: each generation the agent pseudo-labels fresh data with its own current model and updates.
The closed-form Gram learner (blue) bootstraps and stays stable; the gradient learner (red) drifts
downward as its self-generated errors compound---the autoregressive collapse. Only the update rule
differs.}
\label{fig:rsi}
\end{figure}
\paragraph{Evolving architecture plus stable self-improvement: the full loop.} The same closed-form
memory composes with \emph{architecture growth} to give the open-ended picture in full. We stream a
hundred skills (CIFAR-100 classes) at ten new skills per generation; at each generation the agent
grows fresh closed-form capacity for the new skills and must retain all earlier ones, against a
gradient agent that expands its head and trains by stochastic gradient descent. The contrast is
categorical (Figure~\ref{fig:grow-rsi}): growing from a small labelled seed, the closed-form agent
holds $83\%$ accuracy over all hundred accumulated skills and answers the first generation's skills at
$89\%$, while the gradient agent collapses to $1\%$ overall with \emph{zero} retention of the first
skills. Capacity grows on demand and nothing already learned is disturbed.
Closing the loop requires self-improvement to \emph{add} capability rather than corrupt it, and this
is where the design matters. A naive self-labelling loop does corrupt even the closed-form learner---a
new skill's unlabelled data is confidently mislabelled as old skills before its prototype is
established, and a plain confidence gate only softens this ($52\%$). The fix uses the temporal
structure that is genuinely available without labels: a freshly arrived unlabelled batch belongs to
the \emph{new} skills, so its pseudo-labels are confined to the current generation's classes and
admitted under a confidence gate, letting the new prototypes bootstrap from unlabelled data instead of
being absorbed by old ones. With this, self-improvement \emph{exceeds} the labelled-seed-only model
($83.4\%\to88.5\%$ over all hundred skills) while retaining the first skills at $94\%$, approaching the
fully-supervised ceiling of $89.8\%$---all with growth on demand, no forgetting, and the gradient
agent still at $1\%$. The open-ended loop is therefore complete: grow new capacity, improve it from a
few labels plus unlabelled experience, and never forget or drift---a closed-form realisation of the
stability a self-improving system requires.
\begin{figure}[H]
\centering
\includegraphics[width=0.62\linewidth]{figures/osnr_grow_rsi.png}
\caption{Evolving architecture with stable self-improvement: a hundred skills streamed ten at a time,
the agent growing capacity for each and self-improving from a few labels plus unlabelled data
(temporally-restricted, confidence-gated). The closed-form agent (blue/green) reaches $88.5\%$ over
all skills and $94\%$ on the first generation after ten generations---above its own labelled-seed-only
model ($83.4\%$) and approaching the fully-supervised ceiling ($89.8\%$); the gradient agent
(red/orange) collapses to chance and forgets the first skills entirely. Same frozen backbone; only the
learning machinery differs.}
\label{fig:grow-rsi}
\end{figure}
\subsection{Operator-matched identification versus generic polynomial discovery}
\label{sec:matched-vs-sindy}
The grown-topology study (\S\ref{sec:grown-topology}) and the cell-operator mismatch audit
establish that the basis must match the operator. This subsection makes the consequence
quantitative against the standard equation-discovery baseline, sparse identification of
nonlinear dynamics (SINDy) \cite{brunton2016sindy}. SINDy regresses estimated state
derivatives onto a fixed feature library, typically polynomial, with sequential thresholded
least squares. The OSNR position is identical in algorithm but insists that the library be
the operator-matched dictionary rather than a generic polynomial one, and that derivatives
come from the autograd-free operator-spline ladder of \S\ref{sec:autograd-free}.
\paragraph{Polynomial systems: parity.} On the Lorenz system, which lies in a degree-two
polynomial span, the library is matched for every method and only the derivative estimator
differs. Across a noise sweep (five seeds) the operator-spline derivative ties a well-tuned
smoothed-SINDy and clearly beats the naive finite-difference SINDy default (coefficient
error at five percent noise: finite difference $0.046$, smoothed $0.015$, operator-spline
$0.012$); at the highest noise a hand-tuned Gaussian smoother is marginally better. The
honest conclusion is parity: when the operator is unknown or polynomial, OSNR is never worse
than SINDy but does not dominate it.
\paragraph{Non-polynomial systems: decisive separation.} The picture changes when the
operator is non-polynomial, which is the physically relevant regime for saturating,
conductance-like, or biological dynamics. Consider the coupled saturating system
\[
\dot{x}_i = \sum_j A_{ij}\tanh(x_j) - d_i x_i,
\]
which is linear in its coefficients under the matched library $[\,x_i,\ \tanh(x_j)\,]$.
Training uses multiple trajectories spanning the saturating regime $|x|\le 3$ so the
nonlinearity is identifiable, and models are tested both in-distribution and on
extrapolation to $|x|\le 6$. Table~\ref{tab:nonpoly-sysid} reports normalized RMSE over six
seeds.
\begin{table}[H]
\centering
\begin{tabular}{lccc}
\toprule
method & in-distribution nRMSE & extrapolation nRMSE & terms \\
\midrule
SINDy, degree-3 polynomial (standard) & $0.31$ & $12.5$ (diverges) & $12$ \\
SINDy, polynomial $+$ $\tanh$ features & $0.0066$ & $0.026$ & $11$ \\
OSNR, operator-matched $[x,\tanh x]$ & $\mathbf{0.0050}$ & $\mathbf{0.0065}$ & $11$ \\
\bottomrule
\end{tabular}
\caption{Identification of a non-polynomial saturating system. The standard polynomial
SINDy library is $40$--$60\times$ worse in-distribution and roughly $400$--$2000\times$
worse on extrapolation, where the polynomial approximation of $\tanh$ yields an unstable
identified model that diverges outside the training range. The matched basis extrapolates
exactly. The result is robust to observation noise (extrapolation nRMSE for the polynomial
library stays near $12.5$ at $0$, $5$, and $10$ percent noise, versus $0.006$, $0.007$,
$0.029$ for the matched basis).}
\label{tab:nonpoly-sysid}
\end{table}
Combined with the polynomial-system parity above, OSNR is never worse than generic sparse
discovery and is decisively better when the governing operator is known and non-polynomial,
the physics-informed and biological regime this work targets.
\paragraph{Generality across operator families.} The separation is not specific to the
saturating $\tanh$ nonlinearity. Repeating the experiment with an oscillatory family
($g=\sin$, matched library $[x,\sin x,\cos x]$) and a rational family ($g=1/(1+x^2)$,
matched library $[x,1/(1+x^2)]$) gives the same outcome: the standard polynomial library
diverges on extrapolation (nRMSE $13.3$ and $7.2$ respectively) while the matched basis
remains accurate ($0.0076$ and $0.0034$), a three-order-of-magnitude gap in every case.
\paragraph{Scaling with dimension.} The separation is not a small-system artifact; it widens
with dimension. For $N$-unit saturating networks $\dot x_i=\sum_j A_{ij}\tanh(x_j)-d_i x_i$,
the matched library has $2N$ features (linear in $N$), whereas a polynomial library grows
combinatorially. Across $N\in\{3,6,12,20\}$ the matched basis stays accurate (forecast nRMSE
$0.006$ to $0.028$) and fast, while a degree-two polynomial library degrades catastrophically
($0.10$ at $N{=}3$ to $65$ at $N{=}20$, already diverging by $N{=}6$): the matched
representation is roughly three orders of magnitude better at $N{=}20$ while using $40$
features against the polynomial's $231$.
\begin{figure}[H]
\centering
\includegraphics[width=0.6\linewidth]{figures/osnr_dim_scaling.png}
\caption{Extrapolation error versus system dimension. The operator-matched basis stays
accurate while a polynomial library degrades and diverges; the gap widens with dimension.}
\label{fig:dim-scaling}
\end{figure}
\paragraph{Physical example: the large-angle pendulum.} The effect is not an artifact of
synthetic systems. For the damped pendulum $\ddot\theta=-(g/L)\sin\theta-\gamma\dot\theta$,
physics supplies the matched feature $\sin\theta$, whereas the small-angle polynomial
approximation $\sin\theta\approx\theta-\theta^3/6$ is famously wrong at large amplitude.
Training at moderate amplitude and forecasting a near-inverted swing ($\theta_0\approx2.8$,
five seeds), the matched library reaches extrapolation nRMSE $0.30$, while a degree-five
polynomial SINDy model diverges ($7.2$) and a black-box neural ODE---which fits
in-distribution best---fails to extrapolate the physics ($1.3$).
\paragraph{Biological capstone: a conductance neuron.} The motivating case for this entire
program is a neuron, whose dynamics are conductance ODEs with sigmoidal gating. For a
Morris--Lecar-type model,
\[
\dot V = I - g_L(V-E_L) - g_K\,w\,(V-E_K) - g_{Na}\,m_\infty(V)\,(V-E_{Na}),
\qquad
\dot w = \frac{w_\infty(V)-w}{\tau},
\]
with $m_\infty,w_\infty$ sigmoidal ($\tanh$) gating, the model is linear in its conductances
under the matched library $[\,1,V,w,wV,m_\infty,m_\infty V,w_\infty\,]$ once the gating
midpoints and slopes are taken from biophysics. Identifying the model from noisy voltage
traces (five seeds), the matched basis reaches forecast nRMSE $0.0019$ in-distribution and
$0.0023$ extrapolated to a large voltage excursion, versus $0.061/0.080$ for a degree-five
polynomial SINDy model and $0.018/0.156$ for a neural ODE---a $30$--$70\times$ advantage,
largest in extrapolation. This is the thesis in its native setting: when the operator is the
biophysics, encoding it is decisively better than approximating it.
\paragraph{Comparison with neural ordinary differential equations.} The other deep-learning
approach to learning dynamics from data is the neural ODE, a black-box multilayer-perceptron
vector field $f_\theta$ trained by backpropagation through an ODE solver
\cite{chen2018neuralode}. Table~\ref{tab:osnr-vs-node} compares OSNR's closed-form matched
identification against a well-trained neural ODE (sixty-four hidden units, roughly nine
thousand parameters, fifteen hundred epochs of RK4 backpropagation) on the saturating system,
as a function of the number of training trajectories.
\begin{table}[H]
\centering
\begin{tabular}{rccccc}
\toprule
trajectories & OSNR in/extrap & OSNR time & neural ODE in/extrap & NODE time & speedup \\
\midrule
$2$ & $0.021 / 0.076$ & $0.002$\,s & $0.317 / 0.487$ & $11.1$\,s & $4621\times$ \\
$8$ & $0.010 / 0.008$ & $0.007$\,s & $0.118 / 0.451$ & $11.6$\,s & $1608\times$ \\
$32$ & $0.010 / 0.007$ & $0.027$\,s & $0.037 / 0.114$ & $11.7$\,s & $430\times$ \\
\bottomrule
\end{tabular}
\caption{OSNR closed-form matched identification versus a well-trained neural ODE on the
saturating system (forecast nRMSE, four seeds). OSNR is more accurate at every data budget
(its two-trajectory model already beats the neural ODE trained on thirty-two), extrapolates
far better, and is $430$--$4600\times$ faster to fit. The neural ODE is converged, not a
strawman; it loses because it is a structure-free black box. This is the physics-informed
regime: OSNR exploits the known operator family while the neural ODE learns it from scratch.}
\label{tab:osnr-vs-node}
\end{table}
\paragraph{Interpretation.} Two controls make the claim honest. First, the SINDy advantage is
not a derivative trick: both methods use the same derivative and sparse regression. Second, a
SINDy variant \emph{given} the matched features also succeeds, so the separation is entirely
about exploiting known operator structure---precisely the OSNR premise. The matched library
is additionally leaner and more noise-robust than the augmented one. Together with the
closed-form speed and data-efficiency over neural ODEs and the autograd-free identification
speed over backpropagation (\S\ref{sec:grown-topology}, Rung~0), the operator-matched
representation is Pareto-dominant for identifying known-structure dynamical systems.
\paragraph{Scope and a negative control.} The advantage is specific to the regime where the
operator family is known. When the basis must instead be \emph{discovered} from a large
overcomplete library deliberately populated with features collinear to the truth, sparse
thresholded regression is already strong: on such a coherent library SINDy attains
extrapolation nRMSE $0.015$, whereas an OSNR rank-revealing column selection followed by
sparse refitting reaches only $0.23$, despite yielding a leaner ($13$ versus $33$ terms),
far better-conditioned ($10^{8}$ versus $10^{13}$), and more seed-stable model. We therefore
do not claim that operator-matched machinery improves blind equation discovery; the claim is
narrower and the experiments support it: when the governing operator family is known---the
physics-informed setting---encoding it in the basis decisively beats both generic sparse
discovery and black-box neural learning, especially in extrapolation.
\subsection{Scaling to PDEs: operator-matched identification versus neural operators}
\label{sec:pde-vs-fno}
The dimension-scaling result of \S\ref{sec:matched-vs-sindy} predicts that the
operator-matched advantage should be largest for high-dimensional spatial operators, i.e.
partial differential equations. We test this against the neural-operator state of the art,
the Fourier neural operator (FNO) \cite{li2021fno}, on the one-dimensional viscous Burgers
equation $u_t=-u u_x + \nu u_{xx}$ on a periodic domain. The OSNR method estimates a small
matched spatial-operator library $[u_x,u_{xx},u u_x]$ from trajectory data, solves the
coefficients in closed form, and forecasts by integrating the identified equation
(method of lines, spectral spatial derivatives). The FNO is trained autoregressively.
\paragraph{Data-efficiency, extrapolation, speed.} Table~\ref{tab:pde-fno} forecasts fresh
initial conditions over a horizon twice as long as any seen in training, as a function of the
number of training trajectories. A well-trained FNO ($74$k parameters, GPU) is the baseline.
\begin{table}[H]
\centering
\begin{tabular}{rccccc}
\toprule
trajectories & OSNR in/extrap & OSNR time & FNO in/extrap & FNO time \\
\midrule
$2$ & $0.043 / 0.062$ & $0.002$\,s & $0.451 / 0.720$ & $8.1$\,s \\
$8$ & $0.029 / 0.041$ & $0.007$\,s & $0.171 / 0.433$ & $8.2$\,s \\
$32$ & $0.014 / 0.020$ & $0.029$\,s & $0.041 / 0.073$ & $8.3$\,s \\
\bottomrule
\end{tabular}
\caption{Burgers forecasting, OSNR operator-matched identification versus a Fourier neural
operator (nRMSE, in-horizon / twice-horizon extrapolation; twelve test trajectories). OSNR is
more accurate at every data budget---its two-trajectory model beats the FNO trained on
thirty-two---roughly three to four times better on long-horizon extrapolation because it
integrates the identified equation rather than rolling out a learned autoregressive map, and
is two-to-four orders of magnitude faster to fit.}
\label{tab:pde-fno}
\end{table}
The advantage is not specific to the FNO. Against the other neural-operator family, DeepONet
\cite{lu2021deeponet}, in its native operator-map mode (branch encoding the initial condition,
trunk the space-time query) and fully trained, the forecast nRMSE on the same problem is
$1.14$, $0.67$, and $0.70$ at two, sixteen, and thirty-two training trajectories---worse than
both the FNO and OSNR, partly because the forecast horizon extends beyond the training window
and the trunk does not extrapolate in time. OSNR therefore outperforms both standard
neural-operator baselines, with the FNO the stronger of the two.
\paragraph{Cross-regime adaptation.} The black-box weakness is generalization across physical
regimes. An FNO trained at one viscosity cannot forecast another; OSNR re-identifies the
viscosity from a short snippet of the new regime in closed form and forecasts any of them.
Training at $\nu_0=0.06$ and testing across $\nu\in[0.045,0.12]$ (Table~\ref{tab:pde-crossreg}),
the FNO degrades by up to an order of magnitude away from $\nu_0$, while OSNR re-identifies the
viscosity to within one percent from thirty frames and forecasts at sub-percent error
throughout. A frozen-coefficient OSNR control degrades just as the FNO does, confirming that
the advantage is the closed-form re-identification, not the representation alone.
\begin{table}[H]
\centering
\begin{tabular}{rcccc}
\toprule
test $\nu$ & FNO (trained at $0.06$) & OSNR frozen & OSNR adapt & $\nu$ recovered \\
\midrule
$0.045$ & $0.109$ & $0.046$ & $\mathbf{0.0059}$ & $0.0447$ \\
$0.060$ & $0.085$ & $0.006$ & $\mathbf{0.0017}$ & $0.0598$ \\
$0.100$ & $0.174$ & $0.143$ & $\mathbf{0.0013}$ & $0.0998$ \\
$0.120$ & $0.231$ & $0.210$ & $\mathbf{0.0012}$ & $0.1196$ \\
\bottomrule
\end{tabular}
\caption{Cross-regime forecasting (nRMSE). OSNR adapts to an unseen viscosity by closed-form
re-identification from a thirty-frame snippet and forecasts at sub-percent error across the
range; the FNO, trained at a single viscosity, cannot adapt and degrades away from it. The
frozen-coefficient OSNR control degrades like the FNO, isolating re-identification as the
mechanism.}
\label{tab:pde-crossreg}
\end{table}
\begin{figure}[H]
\centering
\includegraphics[width=0.6\linewidth]{figures/osnr_cross_regime.png}
\caption{Cross-regime forecasting versus test viscosity. The FNO (trained at one viscosity)
and frozen-coefficient OSNR both degrade away from the training regime; OSNR with closed-form
re-identification stays at sub-percent error across all regimes.}
\label{fig:cross-regime}
\end{figure}
The same adaptation holds in two dimensions: training the FNO at $\nu_0=0.07$ on a $64\times64$
grid and testing across $\nu\in[0.05,0.09]$, OSNR re-identifies the viscosity near-exactly and
forecasts at $\sim\!10^{-3}$ at every regime, while the FNO sits at $0.17$--$0.28$ throughout
(and a frozen-coefficient OSNR control again degrades off-regime).
\paragraph{Generality across PDEs.} The advantage is not specific to Burgers. On a
bistable reaction-diffusion (Allen--Cahn) equation $u_t=\nu u_{xx}+u-u^3$ with matched library
$[u_{xx},u,u^3]$, OSNR recovers the exact coefficients from two trajectories and forecasts at
nRMSE $\approx 10^{-4}$ at every data budget (the dynamics are shock-free, so integrating the
identified equation is near machine precision), whereas the same well-trained FNO ranges from
$0.40$ at two trajectories to $0.045$ at thirty-two---a three-to-four-order-of-magnitude gap.
\paragraph{Noise robustness.} Real data is noisy, and under noise both derivatives are
fragile: spectral spatial derivatives amplify high-wavenumber noise by $k^2$, and
finite-difference time derivatives amplify white noise. The OSNR remedy is operator-spline
smoothing before differentiating---a data-adaptive spatial low-pass (the cutoff set from the
high-wavenumber noise floor) plus per-point Tikhonov temporal smoothing. Identifying Burgers
from noisy trajectories and forecasting from a clean state (Table~\ref{tab:pde-noise}), the
smoothed solver degrades gracefully and beats the FNO at every noise level, while the
unsmoothed solver collapses---confirming that the smoothing, not merely the closed form, is
the essential ingredient---and the FNO's autoregressive rollout diverges at ten percent noise.
\begin{table}[H]
\centering
\begin{tabular}{rccc}
\toprule
observation noise & OSNR raw & OSNR spline & FNO \\
\midrule
$0\%$ & $0.007$ & $0.017$ & $0.082$ \\
$2\%$ & $0.495$ & $\mathbf{0.090}$ & $0.192$ \\
$5\%$ & $0.748$ & $\mathbf{0.173}$ & $0.281$ \\
$10\%$ & $0.913$ & $\mathbf{0.245}$ & diverged \\
\bottomrule
\end{tabular}
\caption{Burgers identification under observation noise (forecast nRMSE from a clean state).
Operator-spline smoothing (adaptive spatial low-pass + Tikhonov temporal) degrades gracefully
and beats the FNO at every level; the unsmoothed spectral solver collapses; the FNO rollout
diverges at $10\%$. At high noise the smoothed solver trades mild coefficient bias for
stability.}
\label{tab:pde-noise}
\end{table}
\begin{figure}[H]
\centering
\includegraphics[width=0.6\linewidth]{figures/osnr_noise.png}
\caption{Identification under observation noise. The operator-spline solve degrades
gracefully; the unsmoothed solve collapses and the FNO rollout diverges at high noise.}
\label{fig:noise}
\end{figure}
\paragraph{Chaotic fourth-order operators via the weak form.} The hardest case is a
chaotic, fourth-order operator, the Kuramoto--Sivashinsky equation
$u_t=-u u_x - u_{xx} - u_{xxxx}$. Here the strong form fails: the $u_{xxxx}$ feature amplifies
high-wavenumber content by $k^4$, so differentiating chaotic data directly gives unstable,
trajectory-dependent coefficients (recovered values scattered over $-0.5$ to $-1.0$, a
relative error of $0.32$). The remedy is the operator-spline weak form: integrate the equation
against smooth compactly-supported test functions and move every derivative analytically onto
the test function, so the data enters only as $u$ and $u^2$ and is never differentiated. With
this, the coefficients are recovered as $(-1.0002,-0.998,-0.998)$ against a true
$(-1,-1,-1)$---a relative error of $0.002$ with negligible seed-to-seed variance, two orders of
magnitude better than the strong form. The matched-operator approach thus extends, with the
weak form, even to chaotic high-order PDEs.
\paragraph{Two-dimensional capstone.} Neural operators are used above all in two and three
spatial dimensions, and the dimension-scaling argument predicts the matched-operator advantage
should be largest there. On a $64\times64$ periodic 2D reaction-diffusion
$u_t=\nu(u_{xx}+u_{yy})+u-u^3$ with matched library $[\nabla^2 u, u, u^3]$, against a
well-trained 2D FNO ($4.6\times10^5$ parameters, fifteen hundred epochs, $34$\,s per fit),
Table~\ref{tab:pde-2d} shows the largest separation of all: OSNR recovers near-exact
coefficients from two trajectories, forecasts at $\sim10^{-3}$ at every budget, and is about
sixty times more accurate than the FNO even at the FNO's best data budget, two to three orders
of magnitude better in extrapolation, and three to four orders of magnitude faster to fit.
\begin{table}[H]
\centering
\begin{tabular}{rcccc}
\toprule
trajectories & OSNR in/extrap & OSNR time & FNO in/extrap & FNO time \\
\midrule
$2$ & $0.0010 / 0.0012$ & $0.01$\,s & $0.524 / 0.926$ & $35$\,s \\
$8$ & $0.0009 / 0.0010$ & $0.05$\,s & $0.116 / 0.243$ & $34$\,s \\
$16$ & $0.0009 / 0.0011$ & $0.11$\,s & $0.060 / 0.105$ & $34$\,s \\
\bottomrule
\end{tabular}
\caption{2D reaction-diffusion forecasting (nRMSE, in-horizon / twice-horizon). In the
two-dimensional setting where neural operators are normally deployed, OSNR's two-trajectory
model is roughly sixty times more accurate than the FNO trained on sixteen, with two-to-three
orders of magnitude better extrapolation and three-to-four orders faster fitting. The gap is
not a single-seed artifact: over five FNO training seeds the in-horizon error is
$0.27\pm0.09$ at eight trajectories and $0.13\pm0.04$ at sixteen, far larger than OSNR's
$\sim\!10^{-3}$.}
\label{tab:pde-2d}
\end{table}
\begin{figure}[H]
\centering
\includegraphics[width=0.49\linewidth]{figures/osnr_vs_fno_2d.png}
\includegraphics[width=0.49\linewidth]{figures/osnr_vs_fno_3d.png}
\caption{OSNR versus FNO forecast error against the number of training trajectories, in 2D
(left) and 3D (right). OSNR is flat near machine precision; the FNO needs far more data and
never closes the gap.}
\label{fig:osnr-fno}
\end{figure}
The same separation holds in three dimensions, the most demanding neural-operator setting:
on a $32^3$ grid OSNR recovers near-exact coefficients from two trajectories and forecasts at
$\sim\!10^{-3}$, while a $3.7\times10^5$-parameter 3D FNO reaches only $0.25$--$0.32$
in-horizon and $0.61$--$0.71$ in extrapolation---roughly a $300\times$ and $500\times$ gap---at
two orders of magnitude more compute. The matched-operator advantage thus grows monotonically
across one, two, and three dimensions, exactly as the dimension-scaling argument predicts.
\paragraph{Partial knowledge: a hybrid of known core and learned residual.} The scope caveat
of the whole approach is that it presumes the operator family is known. Real physics is usually
only \emph{partially} known. The operator-matched representation extends naturally to this case
by the sparse-plus-smooth construction of \S\ref{sec:matched-vs-sindy}: keep the known operator
terms as a closed-form core and add a small learned residual dictionary for the unmodeled part.
On Burgers with an unknown non-polynomial reaction added,
$u_t=-u u_x+\nu u_{xx}+0.7\sin(2.5u)$, where the method is told only the advection-diffusion
core, Table~\ref{tab:pde-hybrid} compares a misspecified pure core, the hybrid (core plus a
fourteen-function radial-basis residual in $u$, still one closed-form ridge solve), and an FNO.
\begin{table}[H]
\centering
\begin{tabular}{rccc}
\toprule
trajectories & OSNR core only (misspecified) & OSNR hybrid & FNO \\
\midrule
$2$ & $0.446$ & $\mathbf{0.027}$ & $0.560$ \\
$8$ & $0.453$ & $\mathbf{0.024}$ & $0.264$ \\
$32$ & $0.464$ & $\mathbf{0.022}$ & $0.060$ \\
\bottomrule
\end{tabular}
\caption{Partial-knowledge regime (Burgers plus an unknown reaction). The misspecified core
cannot improve with data (a model-form error), and the black-box FNO needs many trajectories;
the hybrid---known core plus a small learned residual, solved in closed form---captures the
unknown reaction and is data-efficient (its two-trajectory model beats the FNO trained on
thirty-two). This extends the approach from a fully known operator to physics-plus-discrepancy.}
\label{tab:pde-hybrid}
\end{table}
\begin{figure}[H]
\centering
\includegraphics[width=0.6\linewidth]{figures/osnr_hybrid.png}
\caption{Partial-knowledge regime. The misspecified core is flat (model-form error); the FNO
improves slowly with data; the hybrid (known core plus a small learned residual) is accurate
and data-efficient.}
\label{fig:hybrid}
\end{figure}
The PDE results mirror the ODE ones (\S\ref{sec:matched-vs-sindy}) at higher dimension and
against a stronger black-box baseline: in the regime where the operator family is known,
encoding it yields a representation that is more data-efficient, more accurate in
extrapolation, far faster, and---unlike a learned operator---instantly adaptable to a new
physical regime. The same scope caveat applies: this is the physics-informed regime, and a
neural operator remains the tool of choice when the governing equations are unknown.
\subsection{Real measured data}
\label{sec:real-data}
All results so far use synthetic or simulated systems. We close the loop with three real measured
benchmarks from the nonlinear system-identification literature, scored by free-run simulation
error against published results, and summarized on a common axis in Figure~\ref{fig:real-data}.
\paragraph{Silverbox (a real Duffing oscillator).} The Silverbox is a measured electronic
circuit implementing a Duffing oscillator, $m\ddot y + c\dot y + ky + k_3 y^3 = u$. We identify
the discrete-time matched form (a NARX whose nonlinear term is the known cubic) by a closed-form
least-squares solve, then refine the coefficients by output-error (free-run) minimization from
that initialization. On the three official test sets, the linear model gives $9.2$, $14.9$,
$8.3$\,mV free-run RMSE; the closed-form matched model $8.3$, $9.5$, $7.1$\,mV; and the refined
model $2.0$, $3.4$, $1.8$\,mV. The matched cubic is necessary (it separates from the linear model
on the high-amplitude multisine), and the refined result is competitive with strong published
nonlinear-identification methods (which report roughly $0.2$--$1$\,mV for the very best and
$1$--$7$\,mV more typically); it is not the absolute state of the art on this much-studied
benchmark, but it validates the matched-operator-plus-closed-form-plus-refinement pipeline on
genuinely measured data.
\paragraph{Cascaded Tanks (an honest negative).} The Cascaded Tanks benchmark is a real
two-tank fluid system with only $1024$ training samples, a hidden upper-tank state, and an
unknown overflow saturation. Here the matched-physics advantage does \emph{not} materialize:
the closed-form matched and hybrid models ($0.87$ and $0.98$\,V) do not beat a linear model
($0.84$\,V), and only output-error refinement reaches the edge of the competitive range
($0.75$\,V versus a published $0.3$--$0.7$\,V). The cause is structural and worth stating: the
governing physics lives partly in the unobserved upper tank, so it cannot be expressed in lags
of the measured lower-tank level alone, and the data is scarce. This sharpens the scope of the
whole approach: the operator-matched advantage requires the relevant dynamics to be observable
(as in Silverbox), and degrades to parity with generic models when a dominant state is hidden.
\paragraph{The remedy the diagnosis prescribes (latent-augmented matched model).} If the failure
on Cascaded Tanks is caused by a hidden state, the principled fix is to restore that state
explicitly: a grey-box model in which the unobserved upper-tank level is a \emph{learned latent
variable} evolved by its own known-form dynamics (Bernoulli square-root outflow, linear pump
inflow, an overflow spill into the lower tank), with the lower tank as the observed output, fit
end-to-end by output-error (back-propagation through the two-state free-run rollout, all physical
parameters positive). This is the matched-operator principle carried into the partially-observed
regime: a known-physics core paired with the minimal latent state the system requires. It works.
The latent-augmented model reaches $0.55$\,V free-run RMSE on the test set, down from $0.75$\,V for
the observed-only model and now \emph{inside} the published competitive range of $0.3$--$0.7$\,V.
Restoring the hidden state turns the honest negative into a competitive result, which is the
strongest possible confirmation of the diagnosis: the obstacle was observability, not the
matched-operator idea, and the same latent-augmentation recipe is what a real conductance-neuron
recording (with its hidden gating variables) would require. The result is robust: across ten random
initializations the fit converges to the same input-output behavior ($0.55$\,V on every restart),
with only mild non-identifiability in the absolute scale of the latent state (which is itself
unobservable)---it is a stable basin, not a lucky seed.
\paragraph{EMPS (a friction-dominated positioning system).} The EMPS benchmark is a real
electro-mechanical positioning system, a double integrator dominated by friction:
$M\ddot q = u - F_v\dot q - F_c\,\mathrm{sign}(\dot q) - \tau_0$. The known nonlinearity is the
Coulomb friction term $\mathrm{sign}(\dot q)$, which we encode in the discrete-time matched NARX
(with a small Stribeck residual of velocity radial basis functions for the hybrid model). The
result isolates the value of the matched nonlinearity cleanly: the linear (no-friction) model
\emph{diverges in free-run} (it is numerically unstable on this marginally-stable plant), whereas
adding the known Coulomb term makes the simulation stable at $22$\,mm RMSE, and the hybrid Stribeck
residual reaches $14$\,mm (about $17\%$ of the output standard deviation), inside the published
range of roughly $3$--$15$\,mm. Two points are worth noting. First, the matched friction term is
\emph{decisive for stability}, not merely accuracy: without it the free-run model has no usable
prediction at all. Second, output-error refinement yields no improvement here---the closed-form
one-step fit is already at a free-run optimum---in contrast to Silverbox, where refinement was
essential. The closed-form solution is thus sometimes already output-error-optimal, and sometimes
only a good initialization; which case obtains depends on the conditioning of the simulated rollout.
We also tried, for completeness, the latent-state continuous grey-box that succeeds on Cascaded
Tanks (below)---treating velocity as an explicit hidden state and integrating the friction ODE---but
on EMPS it is markedly worse ($148$\,mm) because the plant is a pure double integrator: open-loop
integration of a slightly imperfect acceleration accumulates unbounded position drift over the long
free-run, whereas the position-feedback NARX form is anchored and stable. The right matched form
therefore depends on the stability character of the operator (dissipative versus integrating), not
only on observability---latent augmentation helps the bounded, dissipative tank dynamics and hurts
the marginally-stable integrator.
\begin{figure}[H]
\centering
\includegraphics[width=0.78\linewidth]{figures/osnr_real_data.png}
\caption{Real measured data, three benchmarks, on a common axis (free-run RMSE as a percentage of
the test-output standard deviation; log scale). Where the governing dynamics are observable in the
measured output (Silverbox, EMPS), the matched/hybrid OSNR model improves by a large factor over a
linear baseline and approaches the strong end of the published range; for EMPS the linear
no-friction model is numerically unstable in free-run (hatched bar, capped). When a dominant state
is hidden (Cascaded Tanks, unobserved upper tank), the matched model collapses to parity with the
linear baseline and stays far from the state of the art. Observability---not the presence of a
nonlinearity per se---governs whether the matched-operator advantage materializes. The black
diamond on the Cascaded Tanks group is the remedy: a latent-augmented matched model (the hidden
upper tank restored as a learned state) drops back into the published competitive band.}
\label{fig:real-data}
\end{figure}
\subsection{Bio-conductance vision: retina, V1, and predictive residual fields}
The next runner, \texttt{apps\_industrial\_breakthrough/bio\_conductance\_vision\_benchmark.py}, is the first real-data attempt to move away from a rigid fixed-feature interpretation of the liquid thesis. The architecture is deliberately cellular rather than MLP-like. A retinal front end performs local contrast normalization. A V1-like bank computes oriented even/odd quadrature energy, divisive normalization, and lateral inhibition. The inhibited visual field is pooled into $7\times7$ token maps and scanned in row, reverse-row, column, and center-out orders by exact conductance cells.
For a token $z_k$, each cell uses excitatory and inhibitory conductances
\[
g^+_{i,k}
=
\sigma\!\left(\gamma^+_i(w_i^{+\top}z_k+b_i^+)\right),
\qquad
g^-_{i,k}
=
\sigma\!\left(\gamma^-_i(w_i^{-\top}z_k+b_i^-)\right),
\]
and the exact zero-order-hold update
\[
x_{i,k+1}
=
\rho_{i,k}x_{i,k}
+
(1-\rho_{i,k})
\frac{g^+_{i,k}A_i^+ + g^-_{i,k}A_i^-}
{\lambda_i+g^+_{i,k}+g^-_{i,k}},
\qquad
\rho_{i,k}
=
\exp\!\left[-\frac{\lambda_i+g^+_{i,k}+g^-_{i,k}}{K}\right].
\]
A lateral competition step $x_i\leftarrow\tanh(x_i-\eta \bar x)$ follows each token update. This keeps the activation mechanism aligned with the LTC/conductance audit: the nonlinearity changes the pole and reversal equilibrium, not merely a pointwise post-activation.
The readout is still algebraic. The base controller fuses four streams: random/analytic local convolution responses, low-frequency DCT identity coefficients, V1 inhibited energy, and conductance-cell scan states. A dopamine-style residual memory then treats the output innovation as a local control signal. For stage $q$,
\[
R_q=Y-\widehat Y_q,\qquad
C_{c,q}=
\operatorname{top}_{K}
\{s_i:y_i=c,\|R_{q,i}\|_2\},
\]
where $s_i$ is a compact retinal state plus a fixed random feedback projection of the V1/conductance state. Gaussian class-local memory features $\Phi_{C_q}$ solve
\[
B_q^\star
=
\arg\min_B
\|\Phi_{C_q}B-R_q\|_F^2+\lambda\|B\|_F^2,
\qquad
\widehat Y_{q+1}
=
\widehat Y_q+\alpha\Phi_{C_q}B_q^\star .
\]
This is a small predictive-coding stack: each stage reselects high-innovation examples and applies a damped local correction. There is no reverse-mode differentiation through the visual front end, the conductance scan, or the residual stack.
\begin{table}[h]
\centering
\scriptsize
\setlength{\tabcolsep}{3pt}
\begin{tabular}{@{}llrrrr@{}}
\toprule
Dataset/split & Profile & Accuracy & Train ms & Est. MB & Backprop \\
\midrule
Fashion $10\mathrm{k}/10\mathrm{k}$ & Prior local ensemble $\times4$ & $89.10\%$ & $4072.5$ & $289.8$ & no \\
Fashion $10\mathrm{k}/10\mathrm{k}$ & Fixed local CNN ridge & $89.19\%$ & $2118.3$ & $660.7$ & no \\
Fashion $10\mathrm{k}/10\mathrm{k}$ & Retina/V1 inhibited ridge & $88.46\%$ & $3964.6$ & $134.8$ & no \\
Fashion $10\mathrm{k}/10\mathrm{k}$ & Conductance scan ridge & $84.54\%$ & $5.9^\dagger$ & $60.9$ & no \\
Fashion $10\mathrm{k}/10\mathrm{k}$ & Bio residual fused ridge & $89.10\%$ & $62.2^\dagger$ & $223.9$ & no \\
Fashion $10\mathrm{k}/10\mathrm{k}$ & Bio+local fused ridge & $90.35\%$ & $1053.8$ & $1013.8$ & no \\
Fashion $10\mathrm{k}/10\mathrm{k}$ & Bio dopamine residual & $90.41\%$ & $1139.7$ & $1217.1$ & no \\
Fashion $10\mathrm{k}/10\mathrm{k}$ & Predictive dopamine stack, $2$ stages & $90.44\%$ & $1177.8$ & $1420.4$ & no \\
Fashion $10\mathrm{k}/10\mathrm{k}$ & Predictive dopamine stack, $3$ stages & $90.36\%$ & $1274.1$ & $1623.6$ & no \\
\midrule
MNIST $10\mathrm{k}/10\mathrm{k}$ & Bio+local fused ridge & $98.14\%$ & $1012.7$ & $1013.8$ & no \\
MNIST $10\mathrm{k}/10\mathrm{k}$ & Bio dopamine residual & $98.23\%$ & $1139.0$ & $1217.1$ & no \\
MNIST $10\mathrm{k}/10\mathrm{k}$ & Predictive dopamine stack, $2$ stages & $98.22\%$ & $1198.0$ & $1420.4$ & no \\
Fashion $10\mathrm{k}/10\mathrm{k}$ & CNN, one epoch & $75.83\%$ & $2145.8$ & $4.4$ & yes \\
MNIST $10\mathrm{k}/10\mathrm{k}$ & CNN, one epoch & $91.43\%$ & $2108.8$ & $4.4$ & yes \\
\bottomrule
\end{tabular}
\caption{Bio-conductance vision benchmark. The daggered rows reuse previously computed shared V1/conductance or DCT features, so their train time should not be read as a standalone full pipeline cost. All no-backprop rows use closed-form ridge or local residual solves.}
\label{tab:bio-conductance-vision}
\end{table}
Table~\ref{tab:bio-conductance-vision} is a real improvement over the earlier no-backprop vision frontier, but not a SOTA claim. On FashionMNIST $10\mathrm{k}/10\mathrm{k}$, the best profile improves the previous local-ensemble mark from $89.10\%$ to $90.44\%$, a $+1.34$ point gain. The improvement does not come from the conductance scan alone; by itself that scan reaches only $84.54\%$. The useful mechanism is the combination of residual identity channels, local convolutional evidence, conductance/V1 state, and shallow predictive residual correction. The $3$-stage row is also important: more local memory is not automatically better, and undamped repeated correction can overfit or destabilize the class field. On MNIST, the same family transfers, with the single dopamine residual reaching $98.23\%$ and the second stage slightly lowering accuracy to $98.22\%$.
This result answers part of the architectural criticism. The model is no longer just a rigid MLP/CNN/RNN/attention analogy with fixed random features; it contains retina-like normalization, V1-like competition, conductance-pole cellular dynamics, modulatory residual memory, and a predictive-coding correction loop. The boundary is equally clear. The best Fashion row still consumes about $1.42$ GB in dense feature/readout memory, and it remains far below heavily tuned backprop SOTA on MNIST/FashionMNIST. The next liquid architecture must therefore make the residual stack local and streaming: solve many small region/cell normal equations, sparse-select conductance gates and feedback projections, and expose intermediate predictive targets, rather than fitting one dense global readout over all cellular features.
\subsection{Bio-plasticity continual learning without replayed gradients}
The dense-readout limitation suggests a different validation regime. Biological learning is not an offline i.i.d. fit over a stationary dataset; it is sequential plasticity under interference. The runner \texttt{apps\_industrial\_breakthrough/bio\_plasticity\_continual\_benchmark.py} therefore tests class-incremental MNIST/FashionMNIST. The learner receives five tasks with two classes per task and is evaluated after each task on all classes seen so far, then on the full ten-class test set. Backprop controls are small MLP/CNN models trained sequentially with AdamW. Two control regimes are reported: no replay, which exposes catastrophic forgetting, and equal-exemplar replay, which stores the same number of old images per class as the bio learner stores local center states.
The bio learner uses the same fixed retina/V1/conductance/local-operator front end as Table~\ref{tab:bio-conductance-vision}, but projects the temporary feature stack into a compact state
\[
z_i
=
\frac{\tanh\!\left(P[d_i,v_i,c_i,\ell_i]\right)}
{\|\tanh\!\left(P[d_i,v_i,c_i,\ell_i]\right)\|_2},
\]
where $d_i$ are DCT coefficients, $v_i$ are V1 inhibited-energy features, $c_i$ are conductance scan states, and $\ell_i$ are fixed local-convolution responses. The stored biological memory is not the raw image and not the full feature stack. When class $c$ arrives, the learner deposits a small diverse center set
\[
\mathcal{C}_{c,t}
=
\operatorname{FPS}_K\{z_i:y_i=c,\ i\in\mathcal{T}_t\},
\]
using farthest-point selection in the compact state space. After each task, only the accumulated centers solve a local normal equation
\[
W_t^\star
=
\arg\min_W
\|[{\bf 1},Z_{\mathcal{C}_t}]W-Y_{\mathcal{C}_t}\|_F^2
+
\lambda\|W_{1:}\|_F^2.
\]
Thus old classes are retained by deposited center states and a small algebraic readout, not by replaying old images through backpropagation. The stronger variant replaces the center buffer by a streaming covariance/eligibility field. For each task it updates only local sufficient statistics
\[
G_t
=
G_{t-1}+\sum_{i\in\mathcal{T}_t}[1,z_i]^\top[1,z_i],
\qquad
B_t
=
B_{t-1}+\sum_{i\in\mathcal{T}_t}[1,z_i]^\top y_i,
\]
and then solves the modulatory controller
\[
W_t^\star
=
\arg\min_W
\sum_{\tau\le t}\|[{\bf 1},Z_{\mathcal{T}_\tau}]W-Y_{\mathcal{T}_\tau}\|_F^2
+
\lambda\|W_{1:}\|_F^2.
\]
This is recursive least squares written as a local co-activity field: $G_t$ is an eligibility covariance accumulated from presynaptic cellular states, $B_t$ is the dopamine/label-modulated cross-covariance, and the solve minimizes a quadratic control energy without reverse-mode gradients or raw-image replay. We also include ablations: max/mean RBF prototype voting, diagonal Gaussian local statistics, a class-subspace attractor energy, and a multi-head attention-fusion controller. The last two are useful negative results under this budget; splitting the state into weak heads or class subspaces did not beat the single stable covariance field.
\begin{table}[h]
\centering
\scriptsize
\setlength{\tabcolsep}{3pt}
\begin{tabular}{@{}llrrrr@{}}
\toprule
Dataset/order & Profile & Final acc. & Train ms & Est. MB & Backprop \\
\midrule
Fashion canonical & MLP no replay, $3$ ep/task & $19.84\%$ & $212.8$ & $0.65$ & yes \\
Fashion canonical & CNN no replay, $3$ ep/task & $19.91\%$ & $10222.1$ & $0.91$ & yes \\
Fashion canonical & MLP equal replay, $3$ ep/task & $82.89\%$ & $345.0$ & $16.00$ & yes \\
Fashion canonical & CNN equal replay, $3$ ep/task & $82.93\%$ & $16772.5$ & $16.26$ & yes \\
Fashion canonical & Bio center linear readout, $D=768$ & $86.38\%$ & $2091.9$ & $17.29$ & no \\
Fashion canonical & Bio covariance field, $D=2048$ & $89.37\%$ & $125.3$ & $16.09$ & no \\
Fashion canonical & Bio covariance field, $D=4096$ & $89.98\%$ & $673.8$ & $64.19$ & no \\
Fashion canonical & Bio dendritic covariance, $3{\times}3072$ & $\mathbf{90.13\%}$ & $927.9$ & $108.42$ & no \\
\midrule
Fashion shuffled & MLP equal replay, $3$ ep/task & $80.88\%$ & $346.2$ & $16.00$ & yes \\
Fashion shuffled & CNN equal replay, $3$ ep/task & $80.90\%$ & $16737.9$ & $16.26$ & yes \\
Fashion shuffled & Bio center linear readout, $D=768$ & $86.39\%$ & $1457.2$ & $17.29$ & no \\
Fashion shuffled & Bio covariance field, $D=4096$ & $89.98\%$ & $681.3$ & $64.19$ & no \\
Fashion shuffled & Bio dendritic covariance, $3{\times}3072$ & $\mathbf{90.13\%}$ & $926.8$ & $108.42$ & no \\
\midrule
Fashion full canonical & Bio MPS covariance, $D=4096$ & $91.54\%$ & $983.8$ & $64.19$ & no \\
Fashion full canonical & Bio MPS covariance, $D=8192$ & $\mathbf{92.05\%}$ & $4498.9$ & $256.38$ & no \\
Fashion full shuffled & Bio MPS covariance, $D=8192$ & $\mathbf{92.06\%}$ & $4511.5$ & $256.38$ & no \\
Fashion full canonical & Bio streaming MPS covariance, $D=12288$ & $\mathbf{92.32\%}$ & $40241.0$ & $576.63$ & no \\
Fashion full shuffled & Bio streaming MPS covariance, $D=12288$ & $92.29\%$ & $40272.4$ & $576.63$ & no \\
Fashion full canonical & Bio streaming MPS covariance, $D=16384$ & $92.24\%$ & $73964.3$ & $1024.8$ & no \\
\midrule
MNIST canonical & MLP equal replay, $3$ ep/task & $93.37\%$ & $347.6$ & $16.00$ & yes \\
MNIST canonical & CNN equal replay, $3$ ep/task & $96.46\%$ & $16751.5$ & $16.26$ & yes \\
MNIST canonical & Bio center linear readout, $D=768$ & $97.72\%$ & $1482.1$ & $17.29$ & no \\
MNIST canonical & Bio covariance field, $D=4096$ & $\mathbf{98.33\%}$ & $674.8$ & $64.19$ & no \\
\midrule
MNIST shuffled & MLP equal replay, $3$ ep/task & $93.18\%$ & $349.6$ & $16.00$ & yes \\
MNIST shuffled & CNN equal replay, $3$ ep/task & $95.86\%$ & $16661.4$ & $16.26$ & yes \\
MNIST shuffled & Bio center linear readout, $D=768$ & $97.73\%$ & $1505.4$ & $17.29$ & no \\
MNIST shuffled & Bio covariance field, $D=4096$ & $\mathbf{98.31\%}$ & $675.1$ & $64.19$ & no \\
\bottomrule
\end{tabular}
\caption{Bio-plasticity class-incremental learning on real MNIST/FashionMNIST subsets. Unless marked ``full'', each run uses $10{,}000$ training and $10{,}000$ test examples, five two-class tasks, and $3$ backprop epochs per task for the replay controls. Full Fashion rows use all $60{,}000$ training images and the same $10{,}000$ test images. Center rows store $512$ compact states per class. Covariance-field rows store only sufficient statistics $G_t,B_t$ over the fixed cellular state and no raw exemplars. The dendritic covariance rows were selected by a separate $250$-examples-per-class validation split over degree, projection depth, dendrite count, neuron count, ridge, and probability-fusion temperature, then refit on the full $10{,}000$ training examples. MPS capacity rows use a larger $14{,}992$-dimensional sensory stack and \texttt{torch.mps} for projection/covariance solves. Streaming MPS rows keep the sensory operators, image-to-state projection, covariance accumulation, solve, and evaluation on MPS and avoid the earlier host-side feature matrix; the memory column reports the covariance field, while projection/cache footprints are listed in the artifacts. Train time for the bio rows is the plasticity/readout update after fixed feature extraction or state caching; feature extraction/cache and projection time are recorded separately in the artifacts.}
\label{tab:bio-plasticity-continual}
\end{table}
Table~\ref{tab:bio-plasticity-continual} is the strongest no-backprop learning result in this branch so far. It is not an offline SOTA classifier claim. It is a continual-learning claim under a specific replay/memory protocol: local cellular evidence plus algebraic plasticity resists class-incremental interference better than the bounded backprop controls, including equal-exemplar replay. On FashionMNIST canonical, the memory-matched covariance field reaches $89.37\%$ with $16.09$ MB of sufficient-statistic state, while the equal-replay CNN reaches $82.93\%$ with $16.26$ MB. The larger $D=4096$ covariance field reaches $89.98\%$. A validation-selected dendritic ensemble of three independent $3072$-neuron covariance fields with probability fusion reaches $90.13\%$ on both canonical and shuffled FashionMNIST, crossing the previous single-field ceiling. The larger MPS capacity runner \texttt{apps\_industrial\_breakthrough/bio\_plasticity\_mps\_capacity\_benchmark.py} uses $20$ DCT modes, $12$ V1 orientations, $4$ scales, $384$ conductance cells, $96$ local convolution filters, and an $8192$-neuron covariance field. On the full $60{,}000/10{,}000$ FashionMNIST protocol it reaches $92.05\%$ canonical and $92.06\%$ shuffled, with $256.38$ MB of covariance state; the MPS projection and solve take about $7.5$ s and $4.5$ s respectively after about $38.6$ s of feature extraction.
The first implementation still underused the GPU because it materialized a multi-GB host feature matrix and then used MPS mostly for projection and dense solves. The streaming correction, \texttt{apps\_industrial\_breakthrough/bio\_plasticity\_streaming\_mps\_benchmark.py}, caches the DCT, V1/Gabor, local-convolution, and conductance operators on MPS, maps image batches directly to projected cellular state on MPS, accumulates covariance fields on MPS, and caches only the projected state. This reduces the full Fashion image-to-state cache pass to about $6.6$--$7.2$ s for the $D=8192$--$12288$ rows, then exposes the true bottleneck: the covariance solve. The $D=8192$ streaming row reaches $92.10\%$, the $D=12288$ row reaches $92.29\%$ with a float16 state cache and $92.32\%$ with a float32 state cache, and shuffled order reaches $92.29\%$. Pushing to $D=16384$ lowers final accuracy to $92.24\%$ while increasing the solve to about $74$ s, so raw covariance width is now hitting a conditioning/credit-allocation wall. The next architectural step should not be another global dense field; it should use local/block covariance fields, low-rank Woodbury updates, gated dendritic subfields, or residual-modulated cell groups that preserve GPU residency without an $O(D^3)$ global solve. On MNIST, covariance-field plasticity reaches $98.33\%$ canonical and $98.31\%$ shuffled, versus $96.46\%$ and $95.86\%$ for equal-replay CNN. No-replay backprop collapses to about $18$--$20\%$ final accuracy on both datasets, confirming that the task is measuring interference rather than ordinary stationary classification.
The next MPS experiment moved from algebraic readouts over fixed cellular states to a genuinely trainable neural network without reverse-mode differentiation. The runner \texttt{apps\_industrial\_breakthrough/bio\_feedback\_alignment\_mps\_benchmark.py} trains a two-hidden-layer local-feedback MLP on real MNIST/FashionMNIST images. The input is a fixed retinal/operator sensory stack: normalized pixels; $2\times$ pooled identity channels; a low-frequency DCT block; and a local convolution bank with signed rectified responses and pooled/statistical summaries. For the strongest FashionMNIST row this gives $7832$ sensory channels. The trainable network is
\[
x\in\mathbb{R}^{7832}
\xrightarrow{\tanh(W_1x+b_1)}
h_1\in\mathbb{R}^{4096}
\xrightarrow{\tanh(W_2h_1+b_2)}
h_2\in\mathbb{R}^{2048}
\xrightarrow{W_3h_2+b_3}
\hat y\in\mathbb{R}^{10}.
\]
\noindent\textbf{Architecture diagram.}
\begin{figure}[H]
\centering
\fbox{%
\begin{minipage}{0.96\linewidth}
\centering
\setlength{\unitlength}{1mm}
\begin{picture}(150,72)
\put(2,44){\fbox{\parbox[c][16mm][c]{28mm}{\centering Retinal/operator\\sensory stack\\$x\in\mathbb{R}^{7832}$}}}
\put(6,33){\scriptsize pixels + pool + DCT + local conv}
\put(43,58){\circle*{2.0}}
\put(43,52){\circle*{2.0}}
\put(43,46){\circle*{2.0}}
\put(43,40){\circle*{2.0}}
\put(43,34){\circle*{2.0}}
\put(50,55){\circle*{2.0}}
\put(50,49){\circle*{2.0}}
\put(50,43){\circle*{2.0}}
\put(50,37){\circle*{2.0}}
\put(37,24){\fbox{\parbox[c][8mm][c]{21mm}{\centering $h_1$: 4096\\tanh cells}}}
\put(82,55){\circle*{2.0}}
\put(82,49){\circle*{2.0}}
\put(82,43){\circle*{2.0}}
\put(82,37){\circle*{2.0}}
\put(89,52){\circle*{2.0}}
\put(89,46){\circle*{2.0}}
\put(89,40){\circle*{2.0}}
\put(76,24){\fbox{\parbox[c][8mm][c]{21mm}{\centering $h_2$: 2048\\tanh cells}}}
\put(116,42){\fbox{\parbox[c][16mm][c]{18mm}{\centering logits\\$\hat y\in\mathbb{R}^{10}$}}}
\put(116,29){\scriptsize class readout}
\put(31,52){\vector(1,0){9}}
\put(54,49){\vector(1,0){25}}
\put(93,46){\vector(1,0){21}}
\put(34,56){\scriptsize $W_1$}
\put(64,53){\scriptsize $W_2$}
\put(101,50){\scriptsize $W_3$}
\put(116,64){\vector(-1,0){25}}
\put(73,64){\vector(-1,0){24}}
\put(95,66){\scriptsize fixed feedback $B_2$}
\put(47,66){\scriptsize fixed feedback $B_1$}
\put(102,61){\scriptsize innovation $e=\mathrm{softmax}(\hat y)-y$}
\put(37,12){\fbox{\parbox[c][9mm][c]{25mm}{\centering local rule\\$\Delta W_1\propto x^\top\delta_1$}}}
\put(73,12){\fbox{\parbox[c][9mm][c]{28mm}{\centering local rule\\$\Delta W_2\propto h_1^\top\delta_2$}}}
\put(111,12){\fbox{\parbox[c][9mm][c]{25mm}{\centering local rule\\$\Delta W_3\propto h_2^\top e$}}}
\put(49,24){\vector(0,-1){3}}
\put(86,24){\vector(0,-1){3}}
\put(124,42){\vector(0,-1){20}}
\end{picture}
\end{minipage}}
\caption{No-backprop local-feedback MLP used in Table~\ref{tab:bio-local-feedback-mps-training}. The forward path is an ordinary two-hidden-layer neural classifier, but the backward path is not reverse-mode differentiation. The output innovation is broadcast through fixed random feedback matrices $B_1,B_2$; each layer updates only from its presynaptic activity and its local postsynaptic/modulatory signal.}
\label{fig:bio-local-feedback-mps-architecture}
\end{figure}
Figure~\ref{fig:bio-local-feedback-mps-architecture} makes the key architectural distinction explicit. The forward pass is conventional enough to compare against backprop-trained MLP/CNN controls, but the credit path is a broadcast-modulatory path rather than a reverse-mode computational graph.
No autograd graph is built for the OSNR/bio rows. The output innovation is
\[
e=\mathrm{softmax}(\hat y)-y,
\]
and each hidden layer receives a fixed random feedback projection rather than the transpose of downstream weights:
\[
\delta_2=(eB_2)\odot(1-h_2^2),
\qquad
\delta_1=(eB_1)\odot(1-h_1^2).
\]
The local plasticity updates are the three-factor eligibility rules
\[
\Delta W_3=-\eta h_2^\top e,\qquad
\Delta W_2=-\eta h_1^\top\delta_2,\qquad
\Delta W_1=-\eta x^\top\delta_1,
\]
with optional weight decay, no reverse-mode chain rule, no stored computation graph, and epochwise plasticity decay $\eta_t=\eta_0\gamma^t$. The strongest full FashionMNIST run uses all $60{,}000$ training images, $10{,}000$ test images, batch size $512$, $\eta_0=0.012$, $\gamma=0.94$, feedback scale $1.0$, no momentum, rank-free online updates, and MPS tensors. It reaches $92.34\%$ online test accuracy, while the same trained representation with a final closed-form ridge readout reaches $92.41\%$. Under the same runner, the bounded $12$-epoch backprop controls reach $88.74\%$ for the small MLP and $91.14\%$ for the shallow CNN. The exact metrics are in \texttt{apps\_industrial\_breakthrough/bio\_feedback\_alignment\_mps\_outputs\_fashion60k\_h4096\_decay094/bio\_feedback\_alignment\_mps\_metrics.json}.
\begin{table}[h]
\centering
\small
\begin{tabular}{llcccc}
\toprule
Dataset/protocol & Profile & Accuracy & Train ms & Est. memory & Backprop \\
\midrule
Fashion full & Local-feedback MLP, online & $\mathbf{92.34\%}$ & $54395.9$ & $717.95$ MB & no \\
Fashion full & Local-feedback MLP, ridge readout & $\mathbf{92.41\%}$ & $55630.7$ & $717.95$ MB & no \\
Fashion full & $3$-member local-feedback ensemble, online & $92.44\%$ & $199815.0$ & $2153.8$ MB & no \\
Fashion full & $3$-member local-feedback ensemble, ridge & $\mathbf{92.64\%}$ & $199815.0$ & $2153.8$ MB & no \\
Fashion full & Prior cross-run OSNR/local-feedback fusion & $93.23\%$ & $853738.1$ & $2919.9$ MB & no \\
Fashion full & Heterogeneous residual OSNR/local-feedback fusion & $\mathbf{93.27\%}$ & $557494.8$ & $4464.3$ MB & no \\
Fashion full & Small MLP control, $12$ epochs & $88.74\%$ & $983.3$ & $2.60$ MB & yes \\
Fashion full & Small CNN control, $12$ epochs & $91.14\%$ & $9118.6$ & $3.63$ MB & yes \\
\midrule
MNIST full & Local-feedback MLP, online & $98.71\%$ & $32142.5$ & $529.50$ MB & no \\
MNIST full & Local-feedback ensemble, ridge, $5$ members & $98.89\%$ & $493270.3$ & $3864.4$ MB & no \\
MNIST full & Streaming OSNR covariance, $2\times8192$ & $98.90\%$ & $62464.5$ & $1976.1$ MB & no \\
MNIST full & Margin-weighted OSNR/local-feedback fusion & $\mathbf{99.14\%}$ & $644565.9$ & $4464.3$ MB & no \\
MNIST full & Small CNN control, $12$ epochs & $99.09\%$ & $10745.3$ & $3.63$ MB & yes \\
MNIST full & Local-feedback trainable conv sheet & $98.00\%$ & $42217.1$ & $98.65$ MB & no \\
\bottomrule
\end{tabular}
\caption{Real neural training on MPS without reverse-mode differentiation. The local-feedback MLP rows train hidden weights by fixed-feedback three-factor plasticity, not by backpropagation. The FashionMNIST result is the first full-data stationary-vision run in this project where an online no-backprop neural network beats the bounded shallow CNN backprop control. The best Fashion fusion row uses a wider $6144/2048$ local-feedback branch, a damped class-local residual-memory readout over trained $h_2$ states, and two streaming OSNR covariance sources; the saved-logit postprocess applies fixed label-free margin weighting to centered logits. No labels, learned fusion weights, or reverse-mode graph are used in that postprocess. The MNIST fusion row is the first full-data run here to cross the same internal shallow-CNN control: it fuses the five-member local-feedback logits with streaming OSNR covariance-state logits, with no learned fusion weights and no reverse-mode graph. This is still not a public MNIST SOTA claim.}
\label{tab:bio-local-feedback-mps-training}
\end{table}
The follow-up ensemble runner \texttt{apps\_industrial\_breakthrough/bio\_feedback\_alignment\_ensemble\_mps\_benchmark.py} tests whether the remaining error is primarily local-credit variance. Three independently seeded $7832\to4096\to2048\to10$ local-feedback learners improve the online result to $92.44\%$, and averaging their closed-form ridge logits gives the current best stationary FashionMNIST no-backprop row, $92.64\%$. Pushing to five members raises online averaging only to $92.47\%$ and lowers ridge averaging to $92.43\%$, so naive ensembling is not a route to a large jump. It reduces some variance, but the shared architecture still makes correlated errors.
The stronger MNIST experiment is \texttt{apps\_industrial\_breakthrough/bio\_osnr\_no\_backprop\_fusion\_mps\_benchmark.py}. It deliberately fuses mechanisms instead of merely widening one model. The first source is a five-member local-feedback ensemble with $11348$-D sensory states, $4096/2048$ hidden cells, $24$ epochs, $\eta_0=0.01$, $\gamma=0.96$, and no autograd; its ridge-logit average reaches $98.89\%$. The second source is a streaming OSNR/V1/conductance covariance field with a $20032$-D operator feature stack, a two-head $8192$-cell dendritic projection, and an online covariance solve; it reaches $98.90\%$. A single $16384$-cell streaming source reaches $98.87\%$. Probability averaging is not enough---the all-source probability fusions reach only $98.83\%$ and $98.80\%$---but centered-logit fusion exposes complementary evidence. Centering each source logit vector per example and averaging the four sources \{local online, local ridge, $2\times8192$ streaming, $16384$ streaming\} reaches $99.13\%$ on the full $60{,}000/10{,}000$ MNIST protocol. The deterministic postprocess runner \texttt{apps\_industrial\_breakthrough/bio\_osnr\_fusion\_postprocess.py} then evaluates label-free confidence rules on the saved logits; margin-weighted centered fusion reaches $99.14\%$. This crosses the bounded $12$-epoch shallow-CNN control at $99.09\%$ without reverse-mode differentiation. The result is important because it says the bottleneck is not only local-feedback seed variance: the OSNR covariance field makes different errors from the trainable local-feedback network.
The same transfer now holds on FashionMNIST. A direct full-data fusion run with the previous $4096/2048$ three-member local-feedback ensemble and streaming sources \texttt{single12288,dendrites2\_8192} reaches $93.10\%$ when the local ridge logits are centered and averaged with the prior-best \texttt{single12288} OSNR source; the all-source centered row reaches $92.90\%$. A more ambitious $6144/3072$ local-feedback branch is not better alone---its ridge ensemble is $92.60\%$ and its online ensemble is $92.22\%$---but it is more complementary to the OSNR stream. Fixed label-free margin-weighted fusion of \{local online, local ridge, \texttt{single12288}\} reaches $93.19\%$. A subsequent saved-logit cross-fusion, generated by \texttt{apps\_industrial\_breakthrough/bio\_osnr\_cross\_fusion\_postprocess.py} with \texttt{--keep\_duplicates} and stored in \texttt{apps\_industrial\_breakthrough/bio\_osnr\_no\_backprop\_cross\_fusion\_outputs\_fashion\_valid\_best.json}, fixed-centers and equally averages the seven saved source logits from those two runs and reaches $93.23\%$, moving the Fashion no-backprop boundary by $+0.59$ points over the previous $92.64\%$ ensemble row.
The next MPS sprint tested whether that boundary was caused by too little architecture diversity, weak residual control, or weak activation modeling. The new runner \texttt{apps\_industrial\_breakthrough/bio\_spline\_residual\_feedback\_mps\_benchmark.py} compares tanh, conductance-softsign, normalized RBF-spline, sinusoidal-pole, mixed-cell, and fixed sensory-residual variants under the same local-feedback rule. On the $10{,}000/3{,}000$ FashionMNIST sweep, softsign plus a fixed sensory residual was the best small-split ridge row ($88.47\%$), but under the stronger full-data no-normalization recipe it fell to $91.98\%$ ridge versus $92.39\%$ for the homogeneous tanh control. Thus the apparent activation/residual-skip gain was not robust. A readout-mirror feedback variant, where $h_2$ receives current classifier weights as top-down apical feedback without autograd, also underperformed on the small split ($88.60\%$ ridge). Finally, an EGGROLL-inspired antithetic low-rank refinement over the trained $W_2$ matrix reduced reward-batch cross-entropy but did not improve held-out accuracy, indicating that naive weight-space evolution is optimizing the wrong local objective.
The useful improvement came from a damped residual-memory controller over trained $h_2$ states. With the original $4096/2048$ local-feedback network, $64$ high-residual centers per class, residual ridge $1.0$, $\gamma=0.1$, and scale $0.1$, the single-run FashionMNIST row improves from $92.41\%$ ridge to $92.52\%$ residual memory; $128$ centers worsens to $92.48\%$, and a wider/lower-amplitude setting gives $92.51\%$. We then extended \texttt{bio\_osnr\_no\_backprop\_fusion\_mps\_benchmark.py} so residual-memory logits become first-class fusion sources. The $4096/2048$ residual fusion with three local members and two OSNR streams reaches $93.21\%$ after deterministic label-free postprocessing. A wider $6144/2048$ local-feedback run with the same residual controller and two OSNR streams reaches $93.22\%$ in-run and $93.27\%$ after fixed margin-weighted centered-logit postprocessing, recorded in \texttt{apps\_industrial\_breakthrough/bio\_osnr\_no\_backprop\_fusion\_mps\_outputs\_fashion60k\_lf3\_h6144\_residual64\_stream2/}. Cross-run all-source fusion does not improve further ($93.21\%$); a label-selected diagnostic subset reaches $93.35\%$ but is explicitly not a benchmark claim. Thus the latest positive result is small but mechanistic: local residual memory adds a complementary error mode, while unstructured same-architecture columns, naive activation swaps, readout mirroring, and raw EGGROLL weight perturbations do not break the ceiling. The next version should make complementarity endogenous, by adding local predictive targets and topographic residual pathways inside the cellular network rather than fusing two finished systems after the fact.
\noindent\textbf{Fusion architecture diagram.}
\begin{figure}[H]
\centering
\fbox{\begin{minipage}{0.97\linewidth}
\centering
\setlength{\unitlength}{1mm}
\begin{picture}(166,82)
\put(3,54){\fbox{\parbox[c][12mm][c]{21mm}{\centering image\\$28\times28$}}}
\put(27,60){\vector(1,0){8}}
\put(37,52){\fbox{\parbox[c][18mm][c]{29mm}{\centering local sensory\\pixels/pool\\DCT/conv\\$11348$}}}
\put(69,60){\vector(1,0){7}}
\put(78,52){\fbox{\parbox[c][18mm][c]{30mm}{\centering $5$ local NNs\\$4096/2048$\\fixed $B_1,B_2$\\plasticity}}}
\put(110,60){\vector(1,0){7}}
\put(119,53){\fbox{\parbox[c][13mm][c]{27mm}{\centering online/ridge\\logits}}}
\put(3,12){\fbox{\parbox[c][12mm][c]{21mm}{\centering same image\\no replay}}}
\put(27,18){\vector(1,0){8}}
\put(37,8){\fbox{\parbox[c][18mm][c]{29mm}{\centering OSNR/V1\\DCT/Gabor\\conductance\\local conv}}}
\put(69,18){\vector(1,0){7}}
\put(78,8){\fbox{\parbox[c][18mm][c]{30mm}{\centering covariance\\$2\times8192$\\$16384$ cells\\$G_t,B_t$ solve}}}
\put(110,18){\vector(1,0){7}}
\put(119,11){\fbox{\parbox[c][13mm][c]{27mm}{\centering OSNR\\logits}}}
\put(133,53){\line(0,-1){11}}
\put(133,24){\line(0,1){10}}
\put(133,42){\vector(0,-1){5}}
\put(133,34){\vector(0,1){4}}
\put(119,35){\fbox{\parbox[c][10mm][c]{27mm}{\centering center logits\\average}}}
\put(146,40){\vector(1,0){5}}
\put(153,34){\fbox{\parbox[c][12mm][c]{12mm}{\centering $99.13\%$\\MNIST}}}
\put(76,73){\scriptsize branch A: trainable weights, local feedback only}
\put(76,2){\scriptsize branch B: operator states, covariance plasticity only}
\put(75,38){\scriptsize fixed fusion, no learned classifier}
\end{picture}
\end{minipage}}
\caption{No-backprop OSNR/local-feedback fusion architecture for the full MNIST $99.13\%$ result. The top branch trains neural hidden weights by fixed-feedback local plasticity. The bottom branch builds operator-state evidence and solves local covariance fields. Fusion is fixed centered-logit averaging, not an extra trained classifier, so the positive result measures complementarity between two no-backprop evidence streams.}
\label{fig:bio-osnr-fusion-architecture}
\end{figure}
We also implemented a more explicitly spatial no-backprop CNN in \texttt{apps\_industrial\_breakthrough/bio\_local\_feedback\_conv\_mps\_benchmark.py}. Its visual sheet is \texttt{image -> tanh(conv5x5) -> 2x average pool -> tanh(hidden) -> logits}. The convolutional weights are trainable without \texttt{loss.backward()}: each update forms local image patches with \texttt{torch.nn.functional.unfold}, multiplies them by a neuromodulatory membrane delta projected from the output innovation, and applies the resulting local eligibility tensor to the $5\times5$ filters. On a $64$-channel, $2048$-hidden MNIST full run, this reaches $98.00\%$. That is a useful proof that trainable spatial filters can be updated by local tensor rules on MPS, but it is not yet the winning architecture. We then added residual readout variants and synaptic-homeostasis/validation-restoration controls to the local-feedback MLP. The residual readout variants helped the $2000$-example smoke test but did not improve the full FashionMNIST result: \texttt{x\_h1\_h2} finished at $92.22\%$ and \texttt{h1\_h2} at $92.16\%$, below the $92.34\%$ \texttt{h2}-only online row. Homeostatic normalization with a $5\%$ validation split finished at $92.30\%$ online. The newer residual-memory fusion results above supersede the earlier conclusion that residual control only matches ridge; the corrected statement is that a carefully damped residual controller helps, but only by about $0.1$ point as a single-network readout and about $0.04$ point at the fusion frontier.
The first heterogeneous operator-cell moonshot, \texttt{apps\_industrial\_breakthrough/bio\_heterogeneous\_operator\_dopamine\_mps\_benchmark.py}, explicitly mixes leaky tanh cells, conductance-like softsign cells, oscillatory pole cells, and sparse event cells. It also adds regional predictive heads that can provide dopamine-like local class-prediction errors to each hidden layer. The small-split tuning showed that injecting regional heads into the final logits hurt, and that regional modulators did not yet improve over global feedback. The best full FashionMNIST heterogeneous run therefore used heterogeneous cells with global feedback only. It reached $92.23\%$ online and $92.37\%$ with a ridge readout, below the homogeneous local-feedback ensemble. This is a useful negative result: heterogeneity is likely necessary at scale, but the current mixture of cell operators and regional signals is not sufficient. We also tested a class-local high-residual RBF controller over trained hidden states as a dopamine residual memory. The undamped residual controller overcorrected badly on the small split, while a damped variant matched but did not improve the ridge readout. This repeats the earlier warning from dense memory kernels: residual controllers need local structure and validation gates, not just high-error centers in one global state space.
The current interpretation is sharper than before: no-backprop training can beat bounded backprop controls on full FashionMNIST when the sensory operator stack is rich and the local-feedback dynamics are stabilized, and a fixed fusion of local-feedback and OSNR covariance-state evidence now beats the bounded shallow CNN control on full MNIST. It still does not crush public SOTA. The next serious step is a deeper topographic local-feedback stack with residual identity paths, normalization/homeostatic targets, validation-selected plasticity schedules, and local predictive losses for intermediate layers, rather than another global covariance solve, a single random feedback projection, a wider final readout, a larger same-architecture ensemble, or unstructured heterogeneous cell mixing.
The first full-color vision stress is \texttt{apps\_industrial\_breakthrough/bio\_physical\_resnet\_cifar\_mps\_benchmark.py}. Unlike the older cross-architecture runners, it keeps CIFAR-10 as RGB $32\times32$ images. The physical model is a residual topographic reservoir
\[
x \rightarrow C_1(5\times5) \rightarrow \operatorname{pool}
\rightarrow C_2(3\times3) \rightarrow \operatorname{pool}
\rightarrow C_3(3\times3) \rightarrow s(x) \rightarrow h \rightarrow \ell,
\]
where each convolutional sheet uses heterogeneous cell channels: tanh membrane cells, conductance-softsign cells, Gaussian spline-like cells, and rational-pole cells. When local training is enabled, each sheet receives a fixed direct neuromodulatory projection of the output innovation, and the convolutional update is the local product of unfolded presynaptic patches with the broadcast membrane delta. No reverse-mode graph is built for the physical rows. The final readout can either use online logits or a closed-form ridge solve over the physical state summaries.
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{lcccc}
\toprule
CIFAR-10 protocol & Model & Accuracy & Recorded time & Backprop \\
\midrule
$10k/3k$ & random physical reservoir, $32/64/96$ sheets & $49.80\%$ & $1.51$ s & no \\
$10k/3k$ & locally trained physical reservoir, $32/64/96$, $8$ epochs & $50.00\%$ ridge / $23.97\%$ online & $4.63$ s & no \\
$10k/3k$ & random physical reservoir, $64/128/192$ sheets & $52.87\%$ & $2.26$ s & no \\
$10k/3k$ & projected spatial-state variant, $64/128/192$ sheets & $47.27\%$ & $2.79$ s & no \\
$50k/10k$ & locally trained physical reservoir, $32/64/96$, $8$ epochs & $52.91\%$ ridge / $30.86\%$ online & $19.27$ s & no \\
$50k/10k$ & random physical reservoir, $64/128/192$ sheets & $57.33\%$ & $2.51$ s & no \\
$50k/10k$ & TinyResNet, width $64$, $8$ epochs & $82.29\%$ & $108.33$ s & yes \\
\bottomrule
\end{tabular}
\caption{Full RGB CIFAR-10 physical-ResNet stress on MPS. The physical rows use fixed heterogeneous ODE-like cell nonlinearities and closed-form readouts or local eligibility updates, not \texttt{loss.backward()}. The result is not a SOTA win: the best no-backprop physical reservoir reaches $57.33\%$, far below the small backprop TinyResNet at $82.29\%$. The useful signal is that a fixed physical reservoir plus ridge readout already extracts meaningful CIFAR evidence quickly, while the current direct-feedback local convolutional plasticity does not improve the reservoir and spatial random projections actually hurt.}
\label{tab:cifar-physical-resnet}
\end{table}
This experiment changes the biological-learning diagnosis. The failure is not that physical states are useless; the $57.33\%$ full-CIFAR row is far above chance and comes from a fixed heterogeneous physical reservoir. The failure is credit assignment inside the reservoir. The current local convolutional dopamine rule optimizes online logits weakly and does not make the final physical state more linearly separable than the untrained reservoir. The next no-backprop architecture therefore needs local predictive targets, contrastive/target-propagation-like regional objectives, or layerwise self-supervised physical fields before the supervised dopamine signal, rather than only a fixed output-error broadcast to every sheet.
The first response to that diagnosis is \texttt{apps\_industrial\_breakthrough/bio\_scattering\_patch\_cifar\_mps\_benchmark.py}. This runner keeps the same full RGB CIFAR-10 protocol, but replaces global-error-driven convolutional plasticity by a sensory growth model. The state contains color grid statistics, RGB DCT coefficients, color-opponent V1-like complex Gabor energy, class-balanced Hebbian patch filters, signed feature hashing into a compact cortical field, and the stronger random physical branch from Table~\ref{tab:cifar-physical-resnet}. The learned non-readout objects are local image patches sampled in a class-balanced way, or optionally from low-margin hard examples. The classification readout is still closed-form ridge. No reverse-mode graph is built for the physical rows.
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{lcccc}
\toprule
CIFAR-10 protocol & No-backprop sensory state & Accuracy & Recorded time & Memory \\
\midrule
$10k/3k$ & scattering/patch only, $32$ filters/class, hash $4096$ & $48.37\%$ & $5.34$ s & $470.5$ MB \\
$10k/3k$ & hard-example synaptogenesis variant & $48.30\%$ & $6.54$ s & $470.5$ MB \\
$10k/3k$ & fused physical+scattering, $24$ filters/class, hash $4096$, ridge $100$ & $56.23\%$ & $10.19$ s & $868.2$ MB \\
$10k/3k$ & fused physical+scattering, $8$ filters/class, hash $2048$, ridge $100$ & $57.90\%$ & $4.45$ s & $573.0$ MB \\
$10k/3k$ & fused physical+scattering, $12$ filters/class, hash $2048$, ridge $300$ & $58.10\%$ & $6.59$ s & $573.0$ MB \\
$50k/10k$ & fused physical+scattering, $8$ filters/class, hash $2048$, ridge $100$ & $65.21\%$ & $12.45$ s & $2317.3$ MB \\
$50k/10k$ & fused physical+scattering, $12$ filters/class, hash $2048$, ridge $100$ & $\mathbf{65.60\%}$ & $13.97$ s & $2317.3$ MB \\
$50k/10k$ & fused physical+scattering, $12$ filters/class, hash $2048$, ridge $300$ & $65.23\%$ & $13.97$ s & $2317.3$ MB \\
\bottomrule
\end{tabular}
\caption{First full-color CIFAR-10 architecture lift after the negative physical-ResNet stress. The fused physical+scattering reservoir improves the best full no-backprop CIFAR-10 result from $57.33\%$ to $65.60\%$, an $8.27$ point absolute gain, while remaining below the bounded TinyResNet backprop control at $82.29\%$. The negative rows are equally important: patch/Gabor scattering without the physical reservoir overfits badly, hard-example patch growth does not help, and a larger hash field or too many patch filters can reduce test accuracy.}
\label{tab:cifar-scattering-patch}
\end{table}
The interpretation is more constructive than the previous negative result. The strong row is not a trained deep CNN in disguise; it is a fixed physical branch plus local sensory fields and a ridge readout. It therefore validates that OSNR-style operator states, color-opponent scattering, and local patch growth can add substantial linearly decodable evidence without backpropagation. It does not validate the full replacement thesis yet. Accuracy remains $16.69$ points below the small TinyResNet control, and the best row still relies on a global algebraic readout rather than a fully local multilayer credit mechanism. A follow-up nonlinear mixed-cell readout expansion is negative on the $10k/3k$ protocol: concatenating a $4096$-cell expansion drops the tuned fused state to $55.13\%$, and expansion-only drops to $53.63\%$. The next step should not be a larger patch bank or generic random nonlinear readout. It should use the $65.60\%$ fused state as the sensory substrate, then add local predictive targets between regions so the hidden physical branch itself is shaped by non-terminal, reward-gated objectives.
We then pushed the architecture in two directions motivated by recent no-backprop and biologically plausible learning work: local target fields and population codes. The local target runner, \texttt{apps\_industrial\_breakthrough/bio\_forward\_target\_cifar\_mps\_benchmark.py}, learns layerwise class-prototype fields by local covariance/ridge solves and mixed ODE-like cells. It is a negative result: the target layers learn their own prototypes on the training set but do not improve test accuracy. The stronger direction is \texttt{apps\_industrial\_breakthrough/bio\_population\_columns\_cifar\_mps\_benchmark.py}, which trains independent physical/scattering columns and fuses their logits by a simple population mean or a small validation ridge head.
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{lcccc}
\toprule
CIFAR-10 protocol & Model & Accuracy & Recorded time & Backprop \\
\midrule
$10k/3k$ & forward-only target stack, two $2048$-cell target layers & $57.13\%$ & $7.43$ s & no \\
$10k/3k$ & Fisher-selective patch synaptogenesis, single column & $58.27\%$ & $4.75$ s & no \\
$50k/10k$ & Fisher-selective patch synaptogenesis, single column & $65.35\%$ & $13.77$ s & no \\
$10k/3k$ & $4$ independent physical/scattering columns, mean logits & $62.40\%$ & $19.01$ s & no \\
$10k/3k$ & $4$ columns, $2$ deterministic train/test views, no validation holdout & $65.73\%$ & $29.70$ s & no \\
$10k/3k$ & $4$ columns, $4$ deterministic train/test views, no validation holdout & $66.83\%$ & $152.73$ s & no \\
$10k/3k$ & $8$ heterogeneous hard-margin columns, $4$ test views & $67.77\%$ & $117.62$ s & no \\
$10k/3k$ & heterogeneous columns plus ES-CNN logits, confidence fusion audit & $68.83\%$ & $504.52$ s & no \\
$10k/3k$ & local-feedback MLP over retinal/DCT/conv features & $55.80\%$ & $11.03$ s & no \\
$10k/3k$ & auxiliary local-error RGB CNN, local sheet heads & $32.63\%$ & $95.94$ s & no \\
$10k/3k$ & normalized residual CNN, binary direct feedback, ridge over best state & $55.73\%$ & $207.70$ s & no \\
$10k/3k$ & normalized residual CNN, readout-aligned head feedback, online best & $46.60\%$ & $275.08$ s & no \\
$10k/3k$ & normalized residual CNN, readout-aligned head feedback, ridge over best state & $54.77\%$ & $275.08$ s & no \\
$10k/3k$ & normalized residual CNN, head/deep-sheet EGGROLL ES, online & $44.40\%$ & $333.76$ s & no \\
$10k/3k$ & normalized residual CNN, head/deep-sheet EGGROLL ES, ridge over evolved state & $59.13\%$ & $333.76$ s & no \\
$10k/3k$ & normalized residual CNN, all-conv EGGROLL ES, online & $46.20\%$ & $372.88$ s & no \\
$10k/3k$ & normalized residual CNN, all-conv EGGROLL ES, ridge over evolved state & $60.73\%$ & $372.88$ s & no \\
$50k/10k$ & $4$ independent physical/scattering columns, mean logits & $67.94\%$ & $59.17$ s & no \\
$50k/10k$ & $8$ independent physical/scattering columns, mean logits & $\mathbf{68.63\%}$ & $241.91$ s & no \\
$50k/10k$ & $8$ independent physical/scattering columns, validation ridge fusion & $68.24\%$ & $242.18$ s & no \\
$50k/10k$ & $4$ columns, $2$ deterministic train/test views, no validation holdout & $69.42\%$ & $166.66$ s & no \\
$50k/10k$ & $8$ columns, $2$ deterministic train/test views, no validation holdout & $69.79\%$ & $413.39$ s & no \\
$50k/10k$ & $8$ heterogeneous hard-margin columns, mean logits & $71.33\%$ & $449.14$ s & no \\
$50k/10k$ & $8$ heterogeneous hard-margin columns, core ridge fusion & $\mathbf{72.06\%}$ & $449.29$ s & no \\
$50k/10k$ & $8+4$ heterogeneous hard-margin runs, fixed logit fusion audit & $72.99\%$ & $646.87$ s & no \\
$50k/10k$ & $8+4$ heterogeneous runs, hard-state spline controller & $72.37\%$ & $551.61$ s & no \\
$50k/10k$ & $8+4$ heterogeneous runs, spline-controller sweep audit & $72.57\%$ & $551.93$ s & no \\
$50k/10k$ & $8+4$ runs plus SSL heads, validation-selected spline controller & $72.66\%$ & $565.22$ s & no \\
$50k/10k$ & $8+4+4$ heterogeneous/augmented runs, hard-state spline audit & $\mathbf{73.69\%}$ & $1126.21$ s & no \\
$50k/10k$ & TinyResNet, width $64$, $8$ epochs & $82.29\%$ & $108.33$ s & yes \\
\bottomrule
\end{tabular}
\caption{Post-failure CIFAR-10 architecture search. The target-field stack, Fisher-selective patch growth, trainable direct-feedback MLP, and auxiliary local-error CNN are useful negative controls. The first large positive jump after the $65.60\%$ single-column result is a population code: independent no-backprop physical/scattering columns reach $69.79\%$ on full CIFAR-10 when full-train fitting and two deterministic train/test views are used. The newer heterogeneous hard-margin population route improves the clean full-data no-backprop result to $72.06\%$; a rerun with saved core logits plus a prechosen hard-state spline controller reaches $72.37\%$. A validation-selected post-training controller over the population columns plus self-supervised residual heads reaches $72.66\%$. A later all-in population/augmentation audit reaches $73.69\%$ with physical-only inference and no reverse-mode training. The $72.99\%$ cross-run logit fusion remains a diagnostic audit until source sets are pre-registered or selected on validation. This is still far from SOTA and below the bounded TinyResNet control, but it is the strongest evidence so far that structured diversity of locally learned physical columns matters more than simply widening one column or attaching a naive local-error CNN.}
\label{tab:cifar-population-columns}
\end{table}
The population result changes the next experiment. Generic target-field readouts, hard-example patch growth, Fisher-selected patches, random nonlinear readout depth, and a direct-feedback MLP are insufficient. The new auxiliary local-error RGB CNN in \texttt{apps\_industrial\_breakthrough/bio\_cifar\_auxiliary\_local\_feedback\_mps\_benchmark.py} is also negative: a $10k/3k$ run with four trainable convolutional sheets, local class heads, and local-head feedback ends at $32.63\%$ after $30$ epochs, with an unstable transient peak near $41.77\%$. We then tested a literature-inspired normalized residual CNN in \texttt{apps\_industrial\_breakthrough/bio\_cifar\_cbdfa\_reservoir\_mps\_benchmark.py}, adding batch-style homeostatic normalization, residual same-shape blocks, crop/flip/cutout augmentation, leaky/spline cell variants, direct binary feedback, rec-LRA-style reachable targets, and readout-aligned top-down head feedback. The best pure local-feedback CNN result is still only $46.60\%$ online and $54.77\%$ with a closed-form ridge readout over the best checkpointed state; the strongest ridge variant with random binary feedback reaches $55.73\%$.
We therefore added the missing evolutionary test to the same runner. After local no-backprop training, the script now starts from the best checkpoint and applies forward-only antithetic low-rank EGGROLL-style perturbations to either the readout plus deepest convolutional sheets or all convolutional sheets. Each ES pair is scored only by reward-subset cross-entropy; no reverse-mode graph, layerwise backpropagated gradient, or test-set selection is used. The smoke test verifies that sign-scored antithetic ES moves the actual CNN state: on $2k/1k$, online accuracy rises from $36.10\%$ to $37.90\%$. On the harder $10k/3k$ protocol, head/deep-sheet ES improves the online checkpoint from $41.93\%$ to $44.40\%$ and lifts the ridge readout over the evolved state to $59.13\%$. Perturbing all convolutional sheets is stronger: online accuracy reaches $46.20\%$ and the ridge readout reaches $60.73\%$. This is the first positive CIFAR evidence in this branch that evolutionary, no-backprop weight-space refinement can improve a trainable convolutional physical network, but it is not a SOTA result and it still trails the fixed physical/scattering population code at $69.79\%$.
The failure is now more specific. Simply adding normalized direct feedback, reachable local targets, spline-like activations, readout-aligned feedback, or low-rank evolutionary perturbations does not reproduce the strong feature geometry of the fixed physical/scattering reservoir. Independent columns help because they preserve different local patch samples, random physical poles, and feature hash collisions. ES helps when it can perturb the convolutional sheets, but the present reward signal is still a shallow terminal classifier objective over a weak state.
We then changed the population route itself. The updated population runner can save logits, vary member architecture, and let later columns sample class-balanced patches from low-margin training examples discovered by the first columns. The heterogeneous recipe cycles filters per class, patch sizes, DCT ranks, hash dimensions, physical pole-bank widths, hidden physical dimensions, and ridge values; after three columns, subsequent patch banks are grown from the lowest-margin $35\%$ of core examples. This is a different mechanism from simply adding more identical columns. On the $10k/3k$ split, eight heterogeneous hard-margin columns reach $67.77\%$, and a label-free confidence fusion audit with the ES-CNN logits reaches $68.83\%$. On full CIFAR-10, the same eight-column heterogeneous/hard run reaches $71.33\%$ by mean logits and $72.06\%$ by a core ridge fusion over column logits. A second independent four-column offset run reaches $71.40\%$ by core fusion; fixed normalized averaging of the two runs' core and pairwise-fusion logits reaches $72.99\%$. The latter is recorded as a post-hoc audit, not yet as a clean benchmark claim, because the source set should be pre-registered or selected on validation before being treated as a final headline.
We then cross-pollinated this route with post-training model fusion, optimal-control language, and spline theory in \texttt{apps\_industrial\_breakthrough/bio\_cifar\_control\_spline\_fusion\_audit.py}. The saved member logits are treated as a columnar state: normalized class voltages, margins, entropies, votes, and disagreement terms. A confidence-bin reliability gate estimates source weights from core examples. A closed-form ridge controller then maps this state to class drives, analogous to an algebraic LQR-style terminal controller over a fixed dynamical state. Finally, a hard-state RBF spline residual appends kernels centered on core examples that are wrong or low margin, so that the controller can correct local residual geometry without reverse-mode gradients. On the rerun full-data sources, the eight-member population with saved core logits reaches $71.35\%$ by mean logits and $72.11\%$ by core ridge; the four-member offset run reaches $70.62\%$ by mean logits and $71.48\%$ by core ridge. Combining all twelve members, normalized mean logits reach $71.75\%$, reliability gating reaches $71.77\%$, the closed-form control feature ridge reaches $71.98\%$, and the prechosen $512$-center hard-state spline controller reaches $72.37\%$. A diagnostic sweep with $1024$ centers and a tighter scale reaches $72.57\%$, while $2048$ centers does not improve it ($72.54\%$). The diagnostic sweep is not a clean benchmark claim because the hyperparameter was chosen after seeing the test result.
The next cross-pollinated audit, \texttt{apps\_industrial\_breakthrough/bio\_cifar\_ssl\_jepa\_attention\_fusion\_audit.py}, explicitly tests whether pretraining and post-training ideas can add residual information without backpropagation. It constructs four closed-form self-supervised heads on full CIFAR: a spline/pole random-feature head, a JEPA-style deterministic-view prediction head, a denoising/diffusion-style latent recovery head, and a low-margin landmark-attention head. These heads are weak as standalone classifiers, reaching only $46.17\%$, $45.19\%$, $45.56\%$, and $46.52\%$ respectively for the $512$-latent/$384$-attention run. Naively appending them as equal experts hurts: mean fusion falls to $69.22\%$, reliability fusion to $69.82\%$, and the hybrid spline controller to $72.11\%$. The useful result appears only when the post-training controller treats them as candidate residual sources and chooses source set plus spline hyperparameters on held-out core examples. With a $5000$-example core-validation split, the selected source is all twelve population columns plus all four self-supervised heads, with a $1024$-center RBF spline controller, scale $0.5$, and ridge $100$. Refit on all $50k$ core examples, this reaches $72.66\%$ on the CIFAR-10 test set. A larger $768$-latent/$512$-attention profile improves the standalone attention head to $47.81\%$ but falls to $72.52\%$ after validation-selected fusion; a $10k$ validation split selects only the attention head as residual source and reaches $72.63\%$. Thus the clean gain is real but small, and capacity scaling over these shallow SSL heads overfits rather than compounding.
The next all-in push asked whether the gap to a frozen AlexNet sensory prior is mainly a missing post-training controller or a missing representation. We first added eight more heterogeneous hard-margin physical columns with a new seed offset. They reach $71.33\%$ by mean logits and $72.34\%$ by core ridge fusion, so identical population scaling is already saturated. A second four-column block with four deterministic train and test views reaches only $71.49\%$ by core ridge, showing that simple view augmentation is also insufficient. However, fusing the original twelve columns with this augmented four-column block and a narrower $4096$-center hard-state spline controller gives the strongest clean physical-only CIFAR row so far: $73.69\%$. Adding all twenty columns is worse: the core ridge remains only $72.27\%$ and the large RBF controller collapses to $70.53\%$, which is a conditioning/overfit failure rather than a capacity gain.
We then made the AlexNet-geometry bridge explicit in \texttt{apps\_industrial\_breakthrough/bio\_cifar\_alexnet\_logit\_bridge\_audit.py}. The audit uses the frozen AlexNet CIFAR logits only as train targets, then evaluates a student whose inference path contains only physical-column logits. This distinction matters: teacher-guided rows are physical-only at inference, but they are not fully independent from scratch because their targets came from an external ImageNet-pretrained model. Over the best sixteen-column physical state, the label-trained hard-state spline reaches $73.69\%$; replacing the label target by AlexNet logits is worse at $73.42\%$; mixing AlexNet logits with the label target improves only to $73.98\%$. Thus the current gap to the $86.22\%$ frozen-AlexNet readout is not a missing final controller. The physical columns do not yet contain enough of the AlexNet-class sensory geometry, and soft teacher logits can only add a fractional correction.
The all-in pretraining stress uses a different protocol and must not be mixed with the from-scratch physical-column claim. In \texttt{apps\_industrial\_breakthrough/bio\_cifar\_frozen\_alexnet\_prior\_audit.py}, the only locally cached external model was an ImageNet-pretrained AlexNet checkpoint. We froze it, extracted CIFAR features on MPS, and fitted only closed-form ridge readouts or the same spline/control fusion heads. There is no CIFAR backpropagation, but the sensory prior was trained externally and is therefore marked as an external-pretrained-prior result in Table~\ref{tab:cifar-external-alexnet}. The result crosses the requested $80\%$ line easily: the frozen AlexNet multi-layer ridge readout reaches $84.78\%$ with one deterministic view, $85.79\%$ with two views, $86.02\%$ with four views, and $86.22\%$ with eight views. Adding the weaker physical/scattering population logits to the AlexNet logits hurts the best readout, although the fused spline controller still reaches $83.23\%$ at eight views and the fused control ridge reaches $80.57\%$. The interpretation is precise: current post-training spline control can exploit a strong pretrained sensory cortex, but it does not yet make the weaker from-scratch physical columns competitive with that external prior.
\begin{table}[ht]
\centering
\scriptsize
\begin{tabular}{lccccc}
\toprule
CIFAR-10 protocol & External-pretrained sensory prior & Post-training readout & Views & Accuracy & CIFAR backprop \\
\midrule
$50k/10k$ & frozen ImageNet AlexNet & ridge over multi-layer features & $1$ & $84.78\%$ & no \\
$50k/10k$ & frozen ImageNet AlexNet & ridge over multi-layer features & $2$ & $85.79\%$ & no \\
$50k/10k$ & frozen ImageNet AlexNet & ridge over multi-layer features & $4$ & $86.02\%$ & no \\
$50k/10k$ & frozen ImageNet AlexNet & ridge over multi-layer features & $8$ & $\mathbf{86.22\%}$ & no \\
$50k/10k$ & frozen ImageNet AlexNet plus $12$ physical columns & control ridge over logits & $8$ & $80.57\%$ & no \\
$50k/10k$ & frozen ImageNet AlexNet plus $12$ physical columns & hard-state spline controller & $8$ & $83.23\%$ & no \\
$50k/10k$ & AlexNet logits as training target only & $16$-column physical spline student & -- & $73.98\%$ & no \\
\bottomrule
\end{tabular}
\caption{External-pretrained-prior CIFAR-10 stress. The AlexNet weights are a locally cached ImageNet-pretrained prior, frozen during all CIFAR experiments. The readouts/controllers are closed-form and do not use CIFAR backpropagation. These rows show that the post-training OSNR/spline controller can cross $80\%$ when given a strong pretrained sensory cortex. The final row removes AlexNet from the inference path but still uses its logits as a training target, so it is teacher-guided physical-only inference rather than a fully independent from-scratch no-backprop result.}
\label{tab:cifar-external-alexnet}
\end{table}
For reproducibility, all CIFAR rows above use torchvision CIFAR-10 stored under \texttt{artifacts/torchvision\_data}, source-order seed \texttt{20260601}, the full $50{,}000/10{,}000$ train/test split, PyTorch \texttt{2.12.0}, torchvision \texttt{0.27.0}, and no CIFAR reverse-mode training in the reported readouts/controllers. The physical population logits used by the $80\%+$ AlexNet fusion are generated by the following two MPS runs; the \texttt{runpy} wrapper is intentional because direct script execution can hide the MPS backend in this environment:
\begin{verbatim}
venv/bin/python -c "import runpy, sys; sys.argv=[
'bio_population_columns_cifar_mps_benchmark.py',
'--train','50000','--test','10000','--val','0','--device','mps',
'--members','8','--train_views','2','--test_views','2',
'--member_recipe','hetero','--hard_after','3','--hard_fraction','0.35',
'--member_seed_offset','0','--save_logits','--save_core_logits',
'--output_dir',
'apps_industrial_breakthrough/bio_population_columns_cifar_mps_outputs_cifar10_full_m8_hetero_hard_aug2_corelogits'
]; runpy.run_path('apps_industrial_breakthrough/bio_population_columns_cifar_mps_benchmark.py', run_name='__main__')"
venv/bin/python -c "import runpy, sys; sys.argv=[
'bio_population_columns_cifar_mps_benchmark.py',
'--train','50000','--test','10000','--val','0','--device','mps',
'--members','4','--train_views','2','--test_views','2',
'--member_recipe','hetero','--hard_after','3','--hard_fraction','0.35',
'--member_seed_offset','100000','--save_logits','--save_core_logits',
'--output_dir',
'apps_industrial_breakthrough/bio_population_columns_cifar_mps_outputs_cifar10_full_m4_hetero_hard_aug2_offset100k_corelogits'
]; runpy.run_path('apps_industrial_breakthrough/bio_population_columns_cifar_mps_benchmark.py', run_name='__main__')"
\end{verbatim}
Each population member builds a local physical/scattering state from RGB color statistics, RGB DCT coefficients, color-opponent Gabor energy, class-balanced Hebbian image patches, a signed hash field, and a random physical branch. The heterogeneous recipe cycles filters/class, patch sizes, DCT ranks, hash dimensions, physical widths, hidden dimensions, and ridge values; after member three, hard-margin synaptogenesis samples patches from the lowest-margin $35\%$ of the core examples. The saved arrays \texttt{member\_core\_logits}, \texttt{member\_test\_logits}, \texttt{y\_core}, and \texttt{y\_test} are the only population inputs used by the downstream fusion audits.
The $86.22\%$ frozen-AlexNet row and the $80.57\%$/$83.23\%$ AlexNet-plus-physical fusion rows are then reproduced by:
\begin{verbatim}
venv/bin/python apps_industrial_breakthrough/bio_cifar_frozen_alexnet_prior_audit.py \
--train 50000 --test 10000 --device mps --batch 256 \
--views 8 --feature multi --ridge 300 \
--fusion_ridge 30 --rbf_centers 1024 --rbf_scale 0.5 --rbf_ridge 100 \
--logits_npz apps_industrial_breakthrough/bio_population_columns_cifar_mps_outputs_cifar10_full_m8_hetero_hard_aug2_corelogits/bio_population_columns_cifar_mps_logits.npz \
--logits_npz apps_industrial_breakthrough/bio_population_columns_cifar_mps_outputs_cifar10_full_m4_hetero_hard_aug2_offset100k_corelogits/bio_population_columns_cifar_mps_logits.npz \
--output_dir apps_industrial_breakthrough/bio_cifar_frozen_alexnet_prior_outputs_full_m12_multi_v8
\end{verbatim}
The only external prior in this command is the locally cached ImageNet AlexNet checkpoint \texttt{\textasciitilde/.cache/torch/hub/checkpoints/alexnet-owt-7be5be79.pth}. The script constructs \texttt{torchvision.models.alexnet(weights=AlexNet\_Weights.IMAGENET1K\_V1)}, freezes every parameter, resizes CIFAR images to $224\times224$, applies ImageNet normalization, averages eight deterministic views, and concatenates \texttt{fc6}, \texttt{fc7}, and ImageNet logits into a $9192$-dimensional feature vector. The CIFAR readout is a closed-form ridge solve with ridge $300$ and targets $2\,\mathrm{onehot}(y)-1$. For the fusion rows, the AlexNet CIFAR logits are appended as one additional source to the twelve physical-column logit sources; normalized source voltages, margins, entropies, votes, and disagreement features feed either a closed-form control ridge or a $1024$-center hard-state RBF spline controller with scale $0.5$ and ridge $100$. No CIFAR loss is backpropagated through AlexNet or through the fusion controller.
The clean $73.69\%$ physical-only row and the $73.98\%$ teacher-guided bridge row require one more four-column augmented population block and then two algebraic audits:
\begin{verbatim}
venv/bin/python -c "import runpy, sys; sys.argv=[
'bio_population_columns_cifar_mps_benchmark.py',
'--train','50000','--test','10000','--val','0','--device','mps',
'--members','4','--train_views','4','--test_views','4',
'--member_recipe','hetero','--hard_after','3','--hard_fraction','0.35',
'--member_seed_offset','300000','--save_logits','--save_core_logits',
'--output_dir',
'apps_industrial_breakthrough/bio_population_columns_cifar_mps_outputs_cifar10_full_m4_hetero_hard_aug4_offset300k_corelogits'
]; runpy.run_path('apps_industrial_breakthrough/bio_population_columns_cifar_mps_benchmark.py', run_name='__main__')"
venv/bin/python apps_industrial_breakthrough/bio_cifar_control_spline_fusion_audit.py \
--device auto --ridge 30 --rbf_centers 4096 --rbf_scale 0.125 --rbf_ridge 10 \
--logits_npz apps_industrial_breakthrough/bio_population_columns_cifar_mps_outputs_cifar10_full_m8_hetero_hard_aug2_corelogits/bio_population_columns_cifar_mps_logits.npz \
--logits_npz apps_industrial_breakthrough/bio_population_columns_cifar_mps_outputs_cifar10_full_m4_hetero_hard_aug2_offset100k_corelogits/bio_population_columns_cifar_mps_logits.npz \
--logits_npz apps_industrial_breakthrough/bio_population_columns_cifar_mps_outputs_cifar10_full_m4_hetero_hard_aug4_offset300k_corelogits/bio_population_columns_cifar_mps_logits.npz \
--output_dir apps_industrial_breakthrough/bio_cifar_control_spline_fusion_outputs_full_m16_corelogits_aug4_rbf4096_s0125_r10
venv/bin/python apps_industrial_breakthrough/bio_cifar_alexnet_logit_bridge_audit.py \
--teacher_npz apps_industrial_breakthrough/bio_cifar_frozen_alexnet_prior_outputs_full_m12_multi_v8/bio_cifar_frozen_alexnet_prior_audit_logits.npz \
--rbf_centers 4096 --rbf_scale 0.125 --ridge 10 --teacher_label_mix 1.0 \
--logits_npz apps_industrial_breakthrough/bio_population_columns_cifar_mps_outputs_cifar10_full_m8_hetero_hard_aug2_corelogits/bio_population_columns_cifar_mps_logits.npz \
--logits_npz apps_industrial_breakthrough/bio_population_columns_cifar_mps_outputs_cifar10_full_m4_hetero_hard_aug2_offset100k_corelogits/bio_population_columns_cifar_mps_logits.npz \
--logits_npz apps_industrial_breakthrough/bio_population_columns_cifar_mps_outputs_cifar10_full_m4_hetero_hard_aug4_offset300k_corelogits/bio_population_columns_cifar_mps_logits.npz \
--output_dir apps_industrial_breakthrough/bio_cifar_alexnet_logit_bridge_outputs_full_m16_rbf4096_s0125_r10_mix10
\end{verbatim}
The bridge audit reports both label-only and teacher-guided rows from the same physical state. The label-only hard-state spline target gives $73.69\%$. AlexNet-logit targets alone give $73.42\%$. The mixed target \texttt{AlexNet logits + 1.0*(2*onehot-1)} gives $73.98\%$ while using only physical-column logits at inference. This is why the bridge row is reported separately: it removes AlexNet from the inference path, but it does not remove the external teacher from training.
The immediate biological/predictive follow-up was to ask whether the missing AlexNet-like sensory geometry can be built internally by local physical pretraining rather than by another final controller. Two new MPS runners test this directly. The stricter runner, \texttt{apps\_industrial\_breakthrough/bio\_cifar\_physical\_predictive\_pretrain\_mps.py}, splits CIFAR into a $4\times4$ cortical sheet. Each local column receives RGB patch descriptors and an unsupervised Hebbian patch bank, then passes them through four fixed heterogeneous cell branches: a tanh leak cell, conductance-softsign cell, damped oscillatory pole cell, and signed Gaussian event cell. The column states are laterally diffused and inhibited on the sheet. Learning before labels is a closed-form local predictive map: for each region, north/south/east/west/global neighboring states and coordinates predict a deterministic target sensory view by a ridge normal equation. Labels enter only in the final ridge or spline readout. The strongest $10k/3k$ run used $96$ unsupervised filters per kernel and $256$ hidden units per column:
\begin{verbatim}
venv/bin/python -c "import runpy, sys; sys.argv=[
'bio_cifar_physical_predictive_pretrain_mps.py',
'--train','10000','--test','3000','--device','mps','--batch','384',
'--grid','4','--branches','4','--branch_dim','64','--patch_filters','96',
'--prediction_ridge','300','--class_ridge','300','--rbf_centers','512',
'--output_dir',
'apps_industrial_breakthrough/bio_cifar_physical_predictive_pretrain_mps_outputs_10k_g4_b64_pf96'
]; runpy.run_path('apps_industrial_breakthrough/bio_cifar_physical_predictive_pretrain_mps.py', run_name='__main__')"
\end{verbatim}
It reaches only $55.17\%$ from the source state, $55.37\%$ from the predicted target state, and $54.03\%$ after predictive-logit spline control. The predictive residual state collapses to $45.43\%$, so the residual channel is not a useful class geometry. That early $10k/3k$ result was not the end of the path, however. After the pair-graph audit below showed that late logit-level residuals were saturating, we returned to this runner and scaled the representation itself on the full $50k/10k$ CIFAR protocol with guarded readout profiles, using the new \texttt{--profiles} option to omit the oversized residual feature solve. With $64$ unsupervised patch filters/kernel and $128$ hidden units per cortical region, the full run reaches $61.90\%$ from the source state and $62.46\%$ after the predictive-logit spline controller. Scaling to $96$ filters/kernel and $256$ hidden units per region raises the controller to $65.29\%$. Scaling once more to $128$ filters/kernel and $384$ hidden units per region gives the strongest from-scratch predictive-column source so far: source-state ridge $65.98\%$, target-state ridge $65.73\%$, predicted-target-state ridge $66.01\%$, and predictive-logit spline controller $66.20\%$. The key observation is that, at full scale, the predicted target state slightly exceeds the raw source state; the local view-prediction operator is no longer only smoothing away class geometry.
The second runner, \texttt{apps\_industrial\_breakthrough/bio\_cifar\_scattering\_predictive\_geometry\_mps.py}, starts from the stronger existing OSNR/V1/scattering substrate: color/DCT statistics, color-opponent Gabor energy, class-balanced Hebbian patch filters, signed hashing, and the fixed physical branch. It then fits a closed-form mixed-cell view-prediction map in a latent state and adds a hard-state feature-space RBF spline readout. The two registered $10k/3k$ profiles were:
\begin{verbatim}
venv/bin/python -c "import runpy, sys; sys.argv=[
'bio_cifar_scattering_predictive_geometry_mps.py',
'--train','10000','--test','3000','--device','mps','--batch','256',
'--filters_per_class','12','--hash_dim','4096','--physical_c1','64',
'--physical_c2','128','--physical_c3','192','--physical_hidden','2048',
'--latent_dim','1024','--target_view','1','--prediction_ridge','100',
'--class_ridge','300','--feature_rbf_centers','512',
'--feature_rbf_scale','0.5','--feature_rbf_ridge','300',
'--rbf_centers','512','--output_dir',
'apps_industrial_breakthrough/bio_cifar_scattering_predictive_geometry_mps_outputs_10k_f12_l1024_v1_frbf512'
]; runpy.run_path('apps_industrial_breakthrough/bio_cifar_scattering_predictive_geometry_mps.py', run_name='__main__')"
venv/bin/python -c "import runpy, sys; sys.argv=[
'bio_cifar_scattering_predictive_geometry_mps.py',
'--train','10000','--test','3000','--device','mps','--batch','256',
'--filters_per_class','24','--hash_dim','4096','--physical_c1','64',
'--physical_c2','128','--physical_c3','192','--physical_hidden','2048',
'--latent_dim','512','--target_view','1','--prediction_ridge','100',
'--class_ridge','300','--feature_rbf_centers','512',
'--feature_rbf_scale','0.5','--feature_rbf_ridge','300',
'--rbf_centers','512','--output_dir',
'apps_industrial_breakthrough/bio_cifar_scattering_predictive_geometry_mps_outputs_10k_f24_l512_v1_frbf512'
]; runpy.run_path('apps_industrial_breakthrough/bio_cifar_scattering_predictive_geometry_mps.py', run_name='__main__')"
\end{verbatim}
The $12$-filter profile reaches $59.83\%$ with the source scattering state, $60.40\%$ after feature-space hard-state RBF splines, and $60.50\%$ after predictive-logit spline control. The predictive latent itself reaches only $51.33\%$, and concatenating source plus predictive geometry falls to $55.77\%$. The $24$-filter profile is similar: source ridge $59.03\%$, feature RBF $59.53\%$, source-plus-predictive geometry $56.67\%$, and predictive-logit spline control $60.73\%$. The conclusion is not that predictive pretraining is useless in general. It is that one-shot deterministic-view prediction in a global latent space does not build the missing sensory hierarchy. It smooths or compresses away class geometry while the RBF spline recovers only a small local correction. The next internal-pretraining attempt must therefore use interacting columns with local target selection, contrastive negative states, or distance-forward residual objectives that preserve discriminative patch identity, rather than a single view-prediction ridge map attached after a fixed scattering encoder.
The clean conclusion at this point was that population diversity plus residual-style hard synaptogenesis moves the full no-backprop CIFAR frontier from $69.79\%$ to $73.69\%$ when fixed post-training spline control is allowed, while an external-teacher bridge reaches $73.98\%$ with physical-only inference. The later full-scale predictive-column runs sharpen this story further. Appending the three logits from the $4\times4$, $384$-hidden/region predictive physical column to the $14$-source physical population and fitting the same closed-form hard-state spline controller reaches $\mathbf{75.51\%}$ on the full CIFAR-10 test set, with no external pretrained model and no reverse-mode CIFAR training. This beats the prospective pair-graph frontier below at $74.75\%$. The controller sweep is also informative: RBF scale $0.125$ gives $75.36\%$, scale $0.20$ gives $75.38\%$, ridge $5$ gives $74.54\%$, ridge $20$ gives $75.28\%$, $2048$ centers gives $75.03\%$, and $8192$ centers overfits to $74.74\%$; the best remains $4096$ centers, scale $0.15$, ridge $10$. Adding the old top-six pair specialists on top of the predictive source drops to $74.99\%$, so the new gain is not a late residual-pair effect. It is a complementary early physical representation effect. This is still not SOTA and still below the small TinyResNet control at $82.29\%$, but it is the first full-CIFAR result in this branch where internal no-backprop predictive column formation gives a clear jump beyond the hand-built residual-controller frontier. The next serious architecture should therefore deepen this early route: multiple predictive sensory views, local contrastive negatives, recurrent column settling, and validation-selected predictive targets should be learned inside the column state before logit compression, rather than only voting at the end, training isolated sheet-local heads, or broadcasting one-step readout feedback. This also gives a concrete bridge to the broader biological thesis: pretraining can build sensory columns, post-training controllers can act as fast neuromodulatory adaptation, and spline/control residuals can target the hard state manifold without storing a reverse-mode computation graph.
The next MPS run tested whether the useful augmented block was a one-off or a reproducible high-information physical column. A second full $50k/10k$ four-member population with train views $4$, test views $8$, and seed offset $700000$ again produced a strong member-2 column ($71.61\%$ test), close to the previous offset-$500000$ member-2 column ($71.75\%$). However, adding both view-rich member-2 columns to the controller dropped the spline row to $73.84\%$, and replacing the old member by the new one reached only $73.91\%$. The conclusion is that this member type is reproducible but highly correlated across seeds; source diversity, not raw ensemble count, controls the residual spline gain.
We then pushed the same mechanism harder by increasing the sensory orbit inside the standout member rather than adding more columns. This required two implementation changes. First, \texttt{bio\_population\_columns\_cifar\_mps\_benchmark.py} now accepts \texttt{--member\_ids}, so a targeted run can instantiate only the heavy heterogeneous member-2 architecture. Second, \texttt{bio\_scattering\_patch\_cifar\_mps\_benchmark.py} now detects oversized MPS ridge systems and accumulates the normal equations from CPU-held feature chunks streamed through MPS. The original single \texttt{train.T @ train} path hits an MPSGraph tensor-dimension limit for the $400{,}000$-row train-view-$8$ system; the streamed path preserves the same closed-form ridge objective without backpropagation.
The targeted full-CIFAR member is reproduced by:
\begin{verbatim}
venv/bin/python -c "import runpy, sys; sys.argv=[
'bio_population_columns_cifar_mps_benchmark.py',
'--train','50000','--test','10000','--val','0','--device','mps',
'--members','3','--member_ids','2',
'--train_views','8','--test_views','8',
'--member_recipe','hetero','--hard_after','3','--hard_fraction','0.35',
'--member_seed_offset','900000','--save_logits','--save_core_logits',
'--output_dir',
'apps_industrial_breakthrough/bio_population_columns_cifar_mps_outputs_cifar10_full_member2_aug8train8_offset900k_corelogits'
]; runpy.run_path('apps_industrial_breakthrough/bio_population_columns_cifar_mps_benchmark.py', run_name='__main__')"
\end{verbatim}
The selected member uses the heterogeneous recipe's third architecture: $16$ filters/class, patch sizes $3$ and $7$, DCT keep $12$, hash dimension $4096$, physical widths $(96,128,192)$, hidden physical dimension $2048$, and member ridge $180$. Its state dimension is $6976$. With eight deterministic train views and eight deterministic test views it reaches $72.03\%$ by mean logits and $72.05\%$ by core ridge fusion, with no reverse-mode graph.
The best current physical-only post-training controller combines four sources: the original eight heterogeneous hard-margin columns, the old four-column train-view-$4$ offset-$300000$ block, the old offset-$500000$ view-rich member-2 source, and the new train-view-$8$ member-2 source:
\begin{verbatim}
venv/bin/python -c "import runpy, sys; sys.argv=[
'bio_cifar_control_spline_fusion_audit.py',
'--device','mps','--ridge','30',
'--rbf_centers','4096','--rbf_scale','0.15','--rbf_ridge','10',
'--logits_npz',
'apps_industrial_breakthrough/bio_population_columns_cifar_mps_outputs_cifar10_full_m8_hetero_hard_aug2_corelogits/bio_population_columns_cifar_mps_logits.npz',
'--logits_npz',
'apps_industrial_breakthrough/bio_population_columns_cifar_mps_outputs_cifar10_full_m4_hetero_hard_aug4_offset300k_corelogits/bio_population_columns_cifar_mps_logits.npz',
'--logits_npz',
'apps_industrial_breakthrough/bio_population_columns_cifar_mps_outputs_cifar10_full_m4_hetero_hard_aug8_offset500k_member2_corelogits/bio_population_columns_cifar_mps_logits.npz',
'--logits_npz',
'apps_industrial_breakthrough/bio_population_columns_cifar_mps_outputs_cifar10_full_member2_aug8train8_offset900k_corelogits/bio_population_columns_cifar_mps_logits.npz',
'--output_dir',
'apps_industrial_breakthrough/bio_cifar_control_spline_fusion_outputs_full_m14_m8plus_aug4plus_aug8member2plus_train8_rbf4096_s015_r10'
]; runpy.run_path('apps_industrial_breakthrough/bio_cifar_control_spline_fusion_audit.py', run_name='__main__')"
\end{verbatim}
The resulting controller state has $14$ physical-column logit sources. Normalized mean logits reach $72.28\%$, reliability-gated members $72.38\%$, and the closed-form control feature ridge $72.45\%$ over $219$ state features. The hard-state RBF spline residual, with $4096$ centers selected from wrong or low-margin core states, reaches $\mathbf{74.51\%}$ on the full CIFAR-10 test set. The fitted RBF variance is $\sigma^2=49.5651$, the final feature dimension is $4315$, the controller fit time is $1.306$ s, and the controller memory estimate is $1059.1$ MB. A narrow sweep confirms that this is a locality-controlled effect rather than a capacity-only effect: scale $0.20$ gives $74.45\%$, scale $0.125$ gives $74.24\%$, RBF ridge $30$ at scale $0.15$ gives $74.35\%$, and $8192$ centers at scale $0.15$ gives $74.50\%$.
\begin{table}[ht]
\centering
\scriptsize
\begin{tabular}{lccc}
\toprule
Full CIFAR-10 physical-only row & Sources / mechanism & Controller & Accuracy \\
\midrule
Previous population spline frontier & $16$ columns, aug2+aug4 blocks & $4096$ RBF, scale $0.125$ & $73.69\%$ \\
Previous best replacement source & $8$ base + aug4 + offset-$500000$ member 2 & $4096$ RBF, scale $0.25$ & $74.09\%$ \\
New train-view-$8$ member alone & one targeted member-2 source & mean/core ridge & $72.03/72.05\%$ \\
Train-view-$8$ physical fusion & $8$ base + aug4 + two complementary member-2 sources & $4096$ RBF, scale $0.15$ & $74.51\%$ \\
Residual-pair margin controller & same $14$ sources + six directed pair-margin columns & base-only $4096$ RBF + pair linear state & $74.70\%$ \\
Prospective pair-graph controller & same sources + one-step ungated pair settling & base-only $4096$ RBF + prospective pair state & $\mathbf{74.75\%}$ \\
Full predictive physical column + $14$-source population & $4{\times}4$ columns, $128$ patch filters/kernel, $384$ hidden/region, source/target/predicted logits & $4096$ RBF, scale $0.15$, ridge $10$ & $\mathbf{75.51\%}$ \\
\bottomrule
\end{tabular}
\caption{Current full-CIFAR no-backprop physical-column frontier. All rows use only OSNR/V1/scattering/physical-column logits at inference and closed-form readouts/controllers; no CIFAR reverse-mode training is used. The train-view-$8$ gain comes from complementary high-information physical columns and retuned spline locality. The residual-pair gain comes from six confusion-local binary residual columns used only as controller coordinates while the nonlinear RBF geometry remains anchored to the original $14$-source state. The prospective pair-graph gain then adds one weak ungated settling step over the directed confusion margins. The new best row appends internally pretrained predictive physical-column logits, not an external pretrained teacher. This remains below the TinyResNet backprop control and is therefore a stronger research waypoint, not a SOTA claim.}
\label{tab:cifar-physical-frontier-trainview8}
\end{table}
The next step converted the manual source search into a clean validation-gated protocol. The new runner \texttt{apps\_industrial\_breakthrough/bio\_cifar\_adaptive\_column\_search\_mps.py} loads candidate physical-column logit banks generated on a shared $45k/5k/10k$ core/validation/test split, treats every saved member as an individual candidate source, and greedily accepts only the source that improves held-out validation accuracy under a closed-form controller or hard-state spline residual. Final controller hyperparameters are also selected on validation and then evaluated once on the test set. We also patched the population runner with \texttt{--shuffle\_split} and \texttt{--split\_seed}, because the first fixed-last-$5k$ validation split was too brittle.
The first clean bank used the fixed last-$5k$ validation split: an eight-member train-view-$2$/test-view-$4$ base bank, a train-view-$8$ member-2 source at offset $900000$, a train-view-$4$ member-2 source at offset $500000$, and a member-6 source. The selector chose sources $[8,9,5]$ and reached $74.24\%$ validation with a $1024$-center search spline, but only $73.49\%$ on the held-out test split. Full refitting the selected three-source topology on all $50k$ examples reached $73.56\%$; a stricter gate that kept only sources $[8,9]$ reached $73.61\%$ after full refit. This ruled out the fixed-last-$5k$ protocol as a reliable selector.
With deterministic shuffled validation (\texttt{--shuffle\_split --split\_seed 20260602}), validation/test agreement improved. The base eight-member bank plus the two member-2 high-view sources selected $[8,9,2,1]$, reached $73.81\%$ on the $45k$-core test protocol, and reached $73.84\%$ after full $50k$ refit. Adding the shuffled aug4 offset-$300000$ family gave a closer manual-family selector: it selected $[12,13,2,10,5,4,8]$ and reached $73.78\%$ on the $45k$ protocol. Refit on all $50k$ with the validation-selected $4096$-center, scale-$0.20$, ridge-$10$ spline reached $74.23\%$; the diagnostic scale-$0.15$ row reached $74.19\%$. Thus adaptive source selection is now clean and reproducible, but it still does not beat the manually discovered $74.51\%$ source topology.
\begin{center}
\scriptsize
\textbf{Adaptive validation-gated CIFAR selection.}\\[0.35em]
\begin{tabular}{lccc}
\toprule
Adaptive protocol & Validation-selected sources & Full-refit test & Interpretation \\
\midrule
Fixed last-$5k$ split & $[8,9,5]$ & $73.56\%$ & validation overfit \\
Fixed split, stricter gate & $[8,9]$ & $73.61\%$ & duplicate member-2 pair is robust but limited \\
Shuffled split, base + two member-2 sources & $[8,9,2,1]$ & $73.84\%$ & better validation/test alignment \\
Shuffled manual-family bank & $[12,13,2,10,5,4,8]$ & $\mathbf{74.23\%}$ & closest clean adaptive topology \\
Manual source topology from Table~\ref{tab:cifar-physical-frontier-trainview8} & preselected family & $\mathbf{74.51\%}$ & current frontier, not validation-selected \\
\bottomrule
\end{tabular}
\end{center}
The clean adaptive protocol improves reproducibility and removes direct test-set source selection, but it does not yet beat the manual $74.51\%$ topology. The next architectural bottleneck is therefore not only selecting among already generated columns; new columns must be grown from validation residuals, disagreement fields, and low-margin source complementarity.
We then executed that next step and separated source-generation failure from architecture-level credit assignment. The population runner now accepts external hard-source logits through \texttt{--hard\_source\_npz}. Given one or more prior physical-column banks on the same core split, it builds a normalized ensemble, scores examples by wrong prediction, low margin, residual norm, residual-plus-error, or a specified class-confusion pair, and uses the selected examples for hard synaptogenesis. We also patched Fisher-selective patch growth so \texttt{candidate\_filters\_per\_class > filters\_per\_class} can score candidate filters grown from the residual hard set rather than silently falling back to generic balanced examples.
The broad residual-growth test used the shuffled $45k/5k/10k$ split and the current manual-family validation banks as the hard-source field. With \texttt{error\_residual}, fraction $0.35$, four high-capacity residual-grown members, train views $4$, test views $8$, and member ids $30$--$33$, the external hard field selected $15750/45000$ core examples. The members reached $71.03\%$, $68.15\%$, $68.94\%$, and $67.94\%$ on the test set; mean logits reached $70.63\%$ and core ridge $71.28\%$. When these four residual-grown columns were appended to the shuffled manual-family validation bank, the clean selector still chose only old-family sources $[12,13,2]$ and reached $73.99\%$ on the $45k$ protocol. Thus global residual hard-example flooding creates weaker variants of the same feature family rather than orthogonal corrections.
We therefore narrowed the biology-inspired synaptogenesis to the dominant confusion manifold. The top shuffled-core confusions of the current manual family are $5\!\to\!3$ and $3\!\to\!5$. A random hard-patch class-pair run for $5\!\to\!3$ selected only $1556/45000$ examples and produced one useful member at $71.47\%$ plus one weak member at $68.41\%$; validation still rejected both as residual sources. The hard-aware Fisher version was more interesting: the split source reached $71.76\%$ validation and $71.45\%$ test, and the adaptive selector accepted it after the two high-view member-2 sources, raising validation to $74.30\%$. However, its held-out test accuracy dropped to $73.65\%$, and the full $50k$ refit of the same Fisher $5\!\to\!3$ source reached only $71.27\%$. Adding that full Fisher source to the selected-seven full-refit family reached $74.15\%$ at RBF scale $0.15$ and $74.12\%$ at scale $0.20$, below both the clean adaptive $74.23\%$ and the manual $74.51\%$ frontier.
\begin{center}
\scriptsize
\textbf{Residual-grown CIFAR physical columns.}\\[0.35em]
\begin{tabular}{lccc}
\toprule
Residual route & Source-generation result & Fusion / selection result & Interpretation \\
\midrule
Broad error-residual hard field & best member $71.03\%$, core fusion $71.28\%$ & rejected by validation selector & hard cloud too generic \\
$5\!\to\!3$ class-pair random patches & best member $71.47\%$, core fusion $71.30\%$ & rejected by validation selector & sharper but not orthogonal \\
$5\!\to\!3$ hard-aware Fisher, $45k$ split & $71.76\%$ validation / $71.45\%$ test & selected on validation, $73.65\%$ test & validation overfit \\
$5\!\to\!3$ hard-aware Fisher, full refit & $71.27\%$ single source & selected-seven plus source $74.15\%$ & below frontier \\
\bottomrule
\end{tabular}
\end{center}
This is a useful negative result. Residual-aware synaptogenesis is necessary as a mechanism, but patch-level residual growth alone does not create the missing representation. The residual evidence overfits unless the new source changes the underlying layer geometry.
The next architecture-level audit is \texttt{apps\_industrial\_breakthrough/bio\_cifar\_forward\_projection\_fusion\_audit.py}. It keeps the physical-column population as the sensory substrate but adds actual no-backprop hidden layers above it. Each hidden layer receives a class prototype field plus the current readout innovation projected into that same prototype space, solves a local ridge/covariance system from its input state to that target field, applies heterogeneous OSNR cell nonlinearities, and exposes the new hidden state to the next layer. This is closer to feedback alignment, direct feedback alignment, equilibrium propagation, and prospective-configuration thinking than the previous final-controller-only audits \cite{lillicrap2016randomfeedback,nokland2016directfeedbackalignment,scellier2017equilibriumpropagation,song2024prospectiveconfiguration}: hidden states are trained locally, but no reverse-mode graph or weight transport is used.
On the exact $14$-source family that gives the $74.51\%$ frontier, the base control feature ridge over $219$ state features reaches $72.44\%$. A two-layer forward-projection stack with $512+512$ target dimensions improves the linear ridge to $73.22\%$, and a $1024+1024$ stack improves it to $73.35\%$. This is a real layerwise no-backprop credit-assignment gain. However, after adding the same hard-state RBF controller, the $512+512$ stack reaches only $74.30\%$ and the $1024+1024$ stack reaches $74.22\%$, both below the old $74.51\%$ hard-state spline over the raw control state. On the selected-seven full-refit family, the $512+512$ stack similarly improves the linear ridge from $73.23\%$ to $73.56\%$ but reaches only $74.02\%$ with the RBF controller. The conclusion is precise: local forward-projection layers improve linear credit assignment, but the present prototype/residual target fields do not yet create a better nonlinear hard-state geometry than the original spline controller. The next no-backprop architecture must move the target field inside the sensory columns themselves---for example local predictive targets, lateral recurrent settling, learned feedback/projection pathways, or prospective equilibrium states---rather than only projecting class residuals after the logit-level population has already compressed the image.
\paragraph{Early-pipeline DFC/OSNR sensory-column audit.}
The next follow-up moved the feedback/control signal before the logit-level population compression. The runner \texttt{apps\_industrial\_breakthrough/bio\_cifar\_dfc\_osnr\_sensory\_mps.py} starts from the no-backprop CIFAR sensory field used by the scattering/patch reservoir: color grid statistics, RGB DCT coefficients, color-opponent Gabor energies, class-balanced or Fisher-selected local patch filters, signed feature hashing, and an optional fixed physical branch. It then adds explicit Deep-Feedback-Control-inspired compartments \cite{meulemans2021deepfeedbackcontrol,guerguiev2017segregateddendrites,sacramento2018dendriticmicrocircuits}. For layer $\ell$,
\[
v_{\ell}^{\mathrm{ff}} = r_{\ell-1} W_\ell + b_\ell,\qquad
v_{\ell}^{\mathrm{ctrl}} = v_{\ell}^{\mathrm{ff}} + u Q_\ell,\qquad
r_\ell = \phi(v_{\ell}^{\mathrm{ctrl}}),
\]
where $\phi(z)=z/\sqrt{1+z^2}$ in the main run, followed by per-sample RMS normalization. The controller uses the class innovation $e=t-\hat y$ with $t=2\,\mathrm{onehot}(y)-1$, initializes $u_0=\lambda e$, and settles for a few iterations by recomputing the controlled network and refreshing $u$. Each layer then receives only its local presynaptic state and local basal--apical voltage gap:
\[
\Delta W_\ell \propto r_{\ell-1}^{\top}(v_{\ell}^{\mathrm{ctrl}}-v_{\ell}^{\mathrm{ff}}),
\qquad
\Delta b_\ell \propto \langle v_{\ell}^{\mathrm{ctrl}}-v_{\ell}^{\mathrm{ff}}\rangle.
\]
The output head uses the output innovation directly. No \texttt{loss.backward()} call, reverse-mode graph, or layerwise adjoint is constructed. To avoid a blind random controller, the implementation also tests two DFC geometry safeguards: a closed-form ridge initialization of the hidden readout, and a fixed closed-form sensory ridge skip so the controller starts from a meaningful output geometry. The scaled MPS run used the known strong $10k/3k$ physical-sensory front end: $12$ filters per class selected from $48$ Fisher candidates, hash dimension $2048$, DCT rank $8$, $8\times3$ Gabor bank, physical branch $(64,128,192,2048)$, giving a $4864$-dimensional sensory state. Above it, the DFC stack used $1024+512$ hidden cells, readout-mirror feedback, fixed sensory ridge skip, three settling steps, $\lambda=0.22$, controller decay $0.20$, update clip $0.04$, ridge $100$, and MPS execution through the project \texttt{runpy} wrapper.
\begin{center}
\scriptsize
\textbf{Early-pipeline DFC/OSNR sensory-column audit.}\\[0.35em]
\begin{tabular}{lccc}
\toprule
Protocol & Baseline / reference & Best output & Interpretation \\
\midrule
$10k/3k$ physical sensory ridge & $56.37\%$ & -- & strong fixed front end \\
DFC $1024+512$, ridge skip, mirror feedback & $56.37\%$ & $48.20\%$ feedforward & controlled training memorizes but does not transfer \\
Same DFC hidden state, ridge readout & $56.37\%$ & $32.93\%$ & hidden state is not reusable class geometry \\
Sensory plus DFC hidden ridge & $56.37\%$ & $55.27\%$ & hidden columns slightly hurt the sensory state \\
Controlled-label energy inference, smoke split & $35.00\%$ & $17.80\%$ & PC-style label search is not calibrated \\
Single-column mirror-feedback smoke split & $35.00\%$ & $31.20\%$ fusion & exact last-layer feedback alone is insufficient \\
Closed-form target-solve smoke split & $35.00\%$ & $32.80\%$ fusion & algebraic hidden target projection still loses sensory information \\
\bottomrule
\end{tabular}
\end{center}
This is an important negative result. Moving the innovation earlier is necessary, but the simple DFC transplant is not sufficient. In the scaled run the controlled training phase reaches essentially perfect training control, yet autonomous feedforward test accuracy falls below the fixed sensory ridge. The failure mode is therefore not merely ``feedback was too late.'' The present controller can force hidden voltages during the teaching phase but does not create a stable sensory representation that works when the target is absent. The next cellular architecture must learn feedback pathways and local predictive targets inside the sensory columns themselves, or pretrain columns to reproduce high-information feature geometry before class control is applied. Fixed OSNR sensory features plus a late-added DFC controller are not enough to beat the current physical-column frontier.
\paragraph{Predictive coding matches backpropagation in the correct regime.} The repeated theme above---local error broadcast forces hidden activity during teaching but fails to co-adapt layers into a transferable representation---motivated an isolated, controlled study of whether a principled local credit-assignment rule can actually \emph{equal} backpropagation rather than merely approach it. We implemented a genuine two-phase predictive-coding network (PCN) in \texttt{bio\_growth/closed\_form\_neat\_predcoding.py} and its successors \texttt{\_predcoding2.py}--\texttt{\_predcoding4.py}: an inference phase settles value nodes to minimize the free energy $F=\tfrac12\sum_\ell\|\varepsilon_\ell\|^2$ with the output clamped, followed by a purely \emph{local} Hebbian weight update $\Delta W_\ell \propto \varepsilon_\ell\,\phi(x_{\ell-1})^{\top}$ that uses only the local error node $\varepsilon_\ell$ and the presynaptic activity---no global backward pass and no layerwise adjoint \cite{rao1999predictivecoding,whittington2017predictivebackprop}. On a teacher--student task with dimensions $[50,128,128,10]$, the naive hard-clamp PCN \emph{lost}: $0.595$ test accuracy versus $0.725$ for backpropagation. Per our standing rule (either succeed or understand exactly why), we diagnosed the gap rather than abandoning the rule. It is \emph{not} a loss confound: backpropagation with the PCN's own mean-squared free energy reaches $0.713$, essentially matching backpropagation with cross-entropy ($0.725$), while the PCN still sat at $0.55$. It is \emph{not} non-convergence: lengthening the inference phase from $25$ to $50$ to $100$ settling steps did not help, and the hidden residual was already small and stable.
The decisive diagnostic was a gradient-alignment unit test: holding weights fixed, we measured the per-layer cosine between the PCN free-energy gradient and the true autograd backpropagation-MSE gradient. Hard clamping gives $\cos(W)=[0.971,0.970,0.916]$---the alignment degrades precisely at the output-adjacent layer; a small target nudge ($\beta=0.1$) gives a uniform $[0.985,0.986,0.987]$; and the zero-divergence inference-learning (Z-IL) schedule of Song et al. \cite{song2020zil} gives $[1.000,1.000,1.000]$, i.e. predictive coding computes \emph{exactly} the backpropagation gradient, locally. The exact cause is therefore the well-known boundary condition of the Whittington--Bogacz equivalence: PC$\,\approx\,$backprop holds near small output error or under the correct inference schedule, and hard-clamping a one-hot target on an untrained network is the worst case, deviating the top-layer gradient (cosine $0.92$). The end-to-end run closes the loop under an identical optimizer and training loop for all four methods: backpropagation-MSE $0.678$, PC-Z-IL $0.676$, PC-nudged ($\beta=0.1$) $0.674$, and the artefactual PC-hard $0.563$. Local error/value-node dynamics with no global backward pass thus \emph{match} backpropagation to within $0.003$--$0.004$ once run in the theoretically correct regime, confirmed both at the gradient level (cosine $1.000$) and end to end (accuracy parity). Metrics are saved in \texttt{bio\_growth/closed\_form\_neat\_outputs/metrics\_predcoding\{,2,3,4\}.json}. One honesty caveat must be stated plainly: classical predictive-coding feedback uses $W^{\top}$ (symmetric weights, i.e. weight transport), so this experiment establishes the absence of a \emph{global backward pass}, not the absence of weight transport---the latter is the separate direct-feedback-alignment/random-feedback result already reported above \cite{lillicrap2016randomfeedback,nokland2016directfeedbackalignment}. The significance for this branch is that predictive coding supplies, via its inference phase, the layer co-adaptation that the greedy forward-projection and local-target audits lacked: the global target reaches every layer through purely local errors, which is why those greedy methods matched only within a point or two while PC reaches full parity. The natural next step is to extend the PCN inference--plasticity loop to convolutional columns---closing the co-adaptation gap the early-pipeline DFC and forward-projection runs left open---and then to use it as the learning rule inside per-area grown topologies alongside the closed-form Gram memory. Our independent convolutional experiments reproduce the depth pathology that this literature now formalizes: the downward prediction-error wave attenuates with depth, with the per-layer cosine against autograd backpropagation falling from $1.0$ at the output to $\approx 0$ at conv1 even at $300$ settling steps---i.e. exponential signal decay \cite{goemaere2025epc} and the Exploding-and-Vanishing Prediction Errors (EVPE) and PE-imbalance failure mode \cite{ha2026metapcn}. Depth-dependent precision weighting partially restores early-layer alignment (conv1 cosine $0.03 \to 0.46$ as the gain rises), consistent with restoring Friston's precision matrix; naive per-sample error normalization fails, whereas the principled meta-PE plus weight-variance route \cite{ha2026metapcn} is the stable form. The closed-form equilibrium route \cite{baskakovs2026hgf}---which replaces iterative relaxation with a direct equilibrium solve over states, weights, and precisions---coincides with this project's closed-form-solve thesis, exactly as in our OSNR inner solve and Gram memory. Feedforward/amortized initialization \cite{millidge2022pcbeyondbackprop} was already used here and fixes the forward state estimate but not the credit-assignment wave itself.
\paragraph{Predictive coding on convolutions: relaxation loses, with an exact diagnosis.} We made the convolutional study quantitative on CIFAR-10 with a deliberately simple architecture---plain strided convolutions, no batch-norm or max-pool, so the free energy is well defined---and an identical four-convolution-plus-linear stack for every method (\texttt{bio\_growth/closed\_form\_neat\_predcoding\_conv.py}, \texttt{\_conv\_align.py}, all runs on Apple MPS). Nudged relaxation predictive coding \emph{lost}: backpropagation-CE $0.541$ and backpropagation-MSE $0.545$ versus PC-nudged ($T=15$) $0.399$. The same gradient-alignment unit test that nailed the MLP localizes the cause exactly: holding weights fixed, the per-parameter cosine between the PC free-energy gradient and the autograd backpropagation-MSE gradient across layers $[\mathrm{conv}1,\dots,\mathrm{conv}4,\mathrm{linear}]$ is $[0.00,-0.02,0.78,1.00,1.00]$ at $T=15$ and $[0.08,0.77,0.96,1.00,1.00]$ at $T=300$. The downward error wave aligns perfectly at the output but attenuates toward the input, so conv1 stays near zero even at $300$ settling steps---the documented exponential signal-decay / EVPE pathology already cited above \cite{goemaere2025epc,ha2026metapcn}, now reproduced on convolutions rather than merely cited.
\paragraph{Depth-precision helps but convergence is the wall.} A depth-dependent precision gain---scaling each hidden layer's inference step by $\mathrm{gain}^{(\text{depth from output})}$, a Friston precision-weighting of the error wave---monotonically restores early-layer alignment: conv1 cosine rises $0.026 \to 0.163 \to 0.456$ as the gain goes $1\to2\to3$ at $T=150$ (\texttt{bio\_growth/closed\_form\_neat\_predcoding\_conv\_prec.py}). Naive per-sample RMS error normalization instead \emph{breaks} the rule (output cosine $\to -0.89$), since dividing out per-sample magnitude destroys the descent direction; the principled meta-PE variant \cite{ha2026metapcn} is the stable form. Critically, at low settling counts ($T=40$--$60$) conv1 never recovers at any gain, and only $T\ge400$ with gain $3$ aligns all layers ($[0.79,0.94,0.98,1.0,1.0]$): precision accelerates but the true bottleneck is convergence, which is precisely what motivates the closed-form one-sweep route below.
\paragraph{The predictive-coding $\times$ neuroevolution bridge.} Because predictive coding is purely local message passing, it trains an \emph{arbitrary evolved topology} with no global backward pass---exactly the irregular wiring neuroevolution produces---whereas backpropagation needs a clean adjoint over the unrolled graph. On a residual-style skip-DAG with multi-parent fan-in (parents $\{1\!:\![0],\,2\!:\![1,0],\,3\!:\![2,1],\,4\!:\![3,2],\,5\!:\![4,2]\}$) in a teacher--student task, backpropagation-CE reaches $0.5806$, the fair backpropagation-MSE control (predictive coding minimizes the MSE free energy) reaches $0.4764$, and PC-nudged with only local updates and no global backward pass reaches $0.5874$ (\texttt{bio\_growth/closed\_form\_neat\_predcoding\_graph.py}). Predictive coding thus matches backpropagation-CE within single-seed noise and \emph{beats} the MSE control, consistent with the prospective-configuration advantage---relaxation settles into a better activity configuration before plasticity \cite{song2024prospectiveconfiguration}. This validates the design slogan ``neuroevolution evolves the topology, predictive coding learns the weights.'' The corollary architecture is shallow per-area predictive coding (faithful where relaxation converges) composed hierarchically, with closed-form solves for deep credit assignment---the cortical picture: skip/residual links become evolvable prediction edges, and attention becomes precision-weighting of error channels (Feldman--Friston).
\paragraph{Two honest negatives.} First, evolving the topology with predictive coding as the inner learner (no backpropagation anywhere) showed no gain on this configuration---but because an MSE/capacity ceiling was binding, not because the bridge mechanism failed: evolution settled on a near-linear network where backpropagation-MSE ($0.48$) trails backpropagation-CE ($0.63$) and predictive coding again matched the MSE control, so a clean PC-NEAT win needs a task where depth or topology is genuinely required (\texttt{bio\_growth/closed\_form\_neat\_pc\_neat.py}). Second, predictive coding with a categorical (cross-entropy) readout is implementable and computes the \emph{exact} cross-entropy output gradient locally, but applying the full cross-entropy force with no nudge sits in the large-error regime ($0.653$ versus a strengthened backpropagation-CE $0.716$ that also carried an initialization/optimizer confound), so the clean apples-to-apples mechanism result remains the MLP parity established above (\texttt{bio\_growth/closed\_form\_neat\_predcoding\_ce.py}).
\paragraph{Closed-form precision predictive coding: the deep fix in one sweep.} The deep credit-assignment problem is removed not by longer relaxation but by computing the predictive-coding equilibrium directly. A single local downward sweep of the error nodes---$\delta_L=\mathrm{softmax}(\mathrm{out})-y$ for the categorical readout, $\delta_\ell=\phi'(z_\ell)\odot(W_{\ell+1}^{\top}\delta_{\ell+1})$, with the local Hebbian update $\Delta W_\ell=\delta_\ell\,\mathrm{acts}_{\ell-1}^{\top}$---needs no relaxation, no global autodiff, and is exact at all depths by construction, with no wave to attenuate (\texttt{bio\_growth/closed\_form\_neat\_predcoding\_hgf.py}). On a matched baseline (same initialization and same Adam, differing only in autograd-backward versus local sweep), backpropagation-CE $0.658$ equals closed-form predictive coding ($\Pi=I$) $0.662$: the local one-sweep \emph{is} backpropagation, and the higher $0.716$ run from the earlier categorical test simply had a better, adoptable initialization. Per-unit precision whitening ($\Pi=1/\mathrm{var}(\delta)$) \emph{hurt} ($0.632$)---an honest negative: the lever for beating backpropagation is prospective configuration, not error whitening. This closed-form / HGF route \cite{baskakovs2026hgf} is this project's closed-form-solve thesis applied to the cortical learning rule, and supplies the deep-capable per-area learner for the planned federated cortex. Tying the arc together: predictive coding equals backpropagation at both the gradient and accuracy level; it trains arbitrary evolved topologies with purely local updates (the neuroevolution bridge); deep credit assignment is recovered cheaply by the closed-form one-sweep route; and the residual MLP gaps were matters of regime, tuning, and confound rather than mechanism.
\paragraph{AlexNet feature-geometry distillation audit.}
The next experiment tested the latter hypothesis directly: if the physical columns are missing the representation geometry of a strong sensory cortex, can a no-backprop closed-form map teach those columns to imitate the frozen AlexNet feature geometry and then remove AlexNet at inference time? The runner is \texttt{apps\_industrial\_breakthrough/bio\_cifar\_alexnet\_feature\_geometry\_distill\_mps.py}. Its source state is the same early CIFAR sensory stack used above: color/grid statistics, RGB DCT modes, color-opponent Gabor energies, Fisher-selected class-balanced patch filters, signed hashing, optional local normalization, and the fixed physical branch. The main state has $12$ filters/class selected from $48$ candidates, patch sizes $5$ and $7$, pooling grids $4$ and $2$, DCT rank $8$, an $8\times3$ Gabor bank, hash dimension $2048$, physical widths $(64,128,192)$, physical hidden dimension $2048$, and final physical state dimension $4864$.
The teacher is the locally cached ImageNet-pretrained torchvision AlexNet \texttt{AlexNet\_Weights.IMAGENET1K\_V1}, frozen throughout. CIFAR images are resized to $224\times224$, ImageNet-normalized, and passed through deterministic views. The main teacher target concatenates \texttt{fc6}, \texttt{fc7}, and ImageNet logits, giving a $9192$-dimensional multi-layer feature vector; this is multiplied by a fixed signed random projection to a $2048$-dimensional teacher sketch, streamed on the full $50k/10k$ run so the dense AlexNet feature matrix is not kept in memory. The physical-to-teacher map is a ridge normal equation
\[
A_\star=\arg\min_A \|S_{\mathrm{phys}} A-T_{\mathrm{Alex,sketch}}\|_F^2+\lambda\|A\|_F^2,
\]
followed by a second closed-form ridge readout from the predicted teacher sketch to CIFAR labels. The implementation also tests a fixed nonlinear OSNR source lift, AlexNet-sketch prototype logits, physical-plus-distilled feature fusion, and post-hoc logit blending. No CIFAR \texttt{loss.backward()} call, reverse-mode graph, or layerwise adjoint is used; teacher-guided rows use AlexNet only as a training target and the inference path is CIFAR image $\to$ OSNR physical state $\to$ closed-form distilled map/readout.
\begin{center}
\scriptsize
\textbf{AlexNet feature-geometry distillation into physical CIFAR states.}\\[0.35em]
\begin{tabular}{lcccc}
\toprule
Protocol & Physical baseline & Teacher/reference & Best physical-only row & Interpretation \\
\midrule
$10k/3k$, two AlexNet views, sketch $2048$ & $58.37\%$ & $70.43\%$ frozen-teacher upper & $60.03\%$ blend & small signal \\
$10k/3k$, four AlexNet views, sketch $2048$ & $58.37\%$ & $71.23\%$ frozen-teacher upper & $59.80\%$ blend & stronger teacher, no transfer gain \\
$10k/3k$, nonlinear source lift $2048$ & $58.37\%$ & $70.43\%$ frozen-teacher upper & $58.07\%$ blend & overfits teacher sketch \\
$10k/3k$, AlexNet-sketch prototypes & $58.37\%$ & $51.13\%$ prototype upper & $60.03\%$ feature blend & prototype target too weak \\
$50k/10k$, two views, streamed sketch & $64.93\%$ & $77.55\%$ frozen-teacher upper & $65.01\%$ feature fusion / $64.93\%$ best blend & negligible full-data gain \\
$50k/10k$, append all $5$ distillation sources to $14$-source frontier & $74.51\%$ frontier & -- & $73.93\%$ RBF controller & hurts controller \\
$50k/10k$, append best blended distillation source only & $74.51\%$ frontier & -- & $74.18\%$ RBF controller & not complementary \\
$50k/10k$, append physical+distilled feature source only & $74.51\%$ frontier & -- & $74.17\%$ RBF controller & not complementary \\
\bottomrule
\end{tabular}
\end{center}
The full run is therefore a controlled rejection of global teacher-sketch regression as the next breakthrough route. The full-CIFAR physical state reaches only $64.93\%$ by its direct ridge readout, while the frozen AlexNet sketch is a $77.55\%$ teacher. The distilled feature map aligns enough to give a weak $65.01\%$ physical-only row, but it does not create a source that improves the then-current $74.51\%$ physical-only controller; adding all five fused distillation source logits drops the hard-state RBF row to $73.93\%$, and adding only the best distilled source still drops it to $74.18\%$. The important diagnosis is that coarse feature-geometry imitation is not the same as acquiring class-separable sensory geometry. The next serious architecture should put the target inside the columns before the global sketch/readout: local patch-level contrastive targets, class/disagreement-specific residual columns, learned feedback paths, recurrent settling, or prospective equilibrium targets that are selected and validated before logit compression.
\paragraph{Pairwise residual sensory-column audit.}
The next MPS experiment implemented the most direct follow-up to that diagnosis: instead of regressing a global teacher sketch after the image has already been compressed, grow new physical columns on the current frontier's dominant directed confusions. The runner is \texttt{apps\_industrial\_breakthrough/bio\_cifar\_pairwise\_residual\_columns\_mps.py}. It loads the exact $14$-source frontier logits, computes the core confusion matrix, and selects directed residual pairs. The main top-six run selects
\[
(5{\to}3),\ (3{\to}5),\ (0{\to}8),\ (2{\to}6),\ (4{\to}7),\ (9{\to}1),
\]
corresponding to cat/dog, airplane/ship, bird/frog, deer/horse, and truck/automobile confusions in CIFAR-10 class order. For each pair $a{\to}b$, the column samples local RGB patches from hard $a$ examples confused as $b$ and from competing $b$ examples, with patch sizes $5$ and $7$, $48$ candidate filters per side, and $16$ selected filters per pair by a pairwise Fisher score. The selected patches feed the same fixed OSNR/V1/scattering/physical state as the earlier sensory audits: pooling grids $4$ and $2$, DCT rank $8$, signed hash dimension $2048$, physical widths $(64,128,192)$, physical hidden dimension $2048$, one deterministic train/test view, and final state dimension $4864$. Labels enter only through closed-form readouts: a multiclass ridge with ridge $120$, and a pair-local binary margin ridge with ridge $30$, $10$, or $5$. No \texttt{loss.backward()} call, reverse-mode graph, or layerwise adjoint is used.
Two implementation details were required to make the audit meaningful. First, \texttt{apps\_industrial\_breakthrough/bio\_cifar\_gate\_pairwise\_sources.py} converts the full pairwise bank into binary-only, multiclass-only, gated, or ungated source NPZ files with the same \texttt{member\_core\_logits}/\texttt{member\_test\_logits} interface as the population columns. Second, \texttt{apps\_industrial\_breakthrough/bio\_cifar\_control\_spline\_fusion\_audit.py} now has \texttt{--rbf\_feature\_members}. With \texttt{--rbf\_feature\_members 14}, the RBF spline centers and distances are computed only from the original $14$ frontier sources, while the pairwise residual columns are appended only to the final linear/control feature block. This separation is crucial: when the pair columns are allowed to define the RBF geometry, the RBF row drops below the frontier; when they act as residual controller coordinates over the old nonlinear manifold, they improve it.
\begin{center}
\scriptsize
\textbf{Pairwise residual CIFAR sensory-column audit.}\\[0.35em]
\begin{tabular}{lccc}
\toprule
Protocol & Control ridge & RBF controller & Interpretation \\
\midrule
$14$-source frontier parity, \texttt{rbf\_feature\_members=14} & $72.45\%$ & $74.52\%$ & switch reproduces old geometry \\
Top-six binary residuals, pair ridge $30$, all sources define RBF & $72.81\%$ & $73.51\%$ & pair coordinates poison RBF locality \\
Top-six binary residuals, top-$3$ gate, base-only RBF & $73.00\%$ & $74.57\%$ & useful linear signal, modest RBF gain \\
Top-six binary residuals, ungated, pair ridge $30$, base-only RBF & $72.81\%$ & $74.66\%$ & better as residual coordinates than gated experts \\
Top-six binary residuals, ungated, pair ridge $10$, base-only RBF & $72.80\%$ & $\mathbf{74.70\%}$ & new physical-only no-backprop frontier \\
Top-six binary residuals, ungated, pair ridge $5$, base-only RBF & $72.73\%$ & $74.66\%$ & over-sharp pair margins do not help \\
Top-six wider filters, $32/96$ selected/candidate, base-only RBF & $72.98\%$ & $74.29\%$ & more filter capacity is less complementary \\
Top-ten binary residuals, top-$3$ gate, base-only RBF & $72.66\%$ & $74.43\%$ & naive pair expansion adds noisy residuals \\
Top-six multiclass+binary residuals, base-only RBF & $72.17\%$ & $74.28\%$ & multiclass pair readouts dilute the margin signal \\
\bottomrule
\end{tabular}
\end{center}
The new best row fixes $181$ mistakes made by the old $74.51\%$ RBF frontier and breaks $162$, for a net gain of $19$ CIFAR-10 test examples and $74.70\%$ total accuracy. The fixes are concentrated in the targeted confusion families: among old $5{\to}3$ mistakes, $35$ are corrected to class $5$; among old $3{\to}5$ mistakes, $23$ are corrected to class $3$. The conclusion is narrow but important. Residual targets must enter before or alongside the controller, but not every residual signal should redefine the nonlinear state manifold. The useful architecture is a two-substrate controller: a stable base physical manifold supplies the hard-state spline neighborhoods, while small pair-specific biological residual columns supply local margin coordinates. The next serious step is to turn these pairwise columns from post-hoc residual readouts into interacting recurrent sensory columns with local contrastive/prospective targets and validation-selected pair recruitment, rather than increasing pair count or filter count blindly.
We then tested exactly that next interaction mechanism in \texttt{apps\_industrial\_breakthrough/bio\_cifar\_pairwise\_prospective\_fusion\_audit.py}. The audit keeps the same $14$-source physical manifold and the same six ridge-$10$ binary pair columns, but adds a directed class-confusion graph over the pair margins. The design is deliberately tied to the biological credit-assignment literature: predictive coding makes residual/error units explicit \cite{rao1999predictivecoding}; feedback alignment shows that exact weight transport is not mandatory \cite{lillicrap2016randomfeedback}; segregated dendrites and dendritic cortical microcircuits turn local apical/basal voltage gaps into credit signals \cite{guerguiev2017segregateddendrites,sacramento2018dendriticmicrocircuits}; e-prop separates local eligibility traces from delayed learning signals in recurrent spiking networks \cite{bellec2020eprop}; and prospective configuration reverses the order of learning by first inferring the neural state that should exist after learning, then consolidating the weights \cite{song2024prospectiveconfiguration}. For each pair $a{\to}b$, the saved pair column supplies a local margin $m_{ab}$. Starting from the base physical population mean logits $z$, the prospective update nudges the class-voltage difference toward that local margin,
\[
e_{ab}=m_{ab}-(z_a-z_b),\qquad
z_a\leftarrow z_a+\eta e_{ab},\qquad
z_b\leftarrow z_b-\eta e_{ab}.
\]
The script exposes the original pair margins, base margins, edge residuals, gate indicators, settled logits, settled pair margins, and residual errors to the same closed-form ridge controller. The RBF spline centers are still selected only from the original $14$-source physical control state. This is therefore a prospective local graph field over pairwise biological residual columns, not a backpropagated hidden layer.
\begin{center}
\scriptsize
\textbf{Prospective pair-graph CIFAR audit.}\\[0.35em]
\begin{tabular}{lccc}
\toprule
Prospective protocol & Control ridge & RBF controller & Interpretation \\
\midrule
Soft-gated settling, $\eta=\{0.15,0.30,0.50\}$, $3$ steps & $72.91\%$ & $74.52\%$ & linear signal, over-constrained RBF readout \\
Soft-gated settling, $\eta=\{0.05,0.10,0.20\}$, $1$ step & $72.91\%$ & $74.55\%$ & still below pair-margin frontier \\
Ungated settling, $\eta=\{0.02,0.05,0.10\}$, $1$ step & $72.81\%$ & $74.71\%$ & perturbation too weak \\
Ungated settling, $\eta=\{0.05,0.10,0.20\}$, $1$ step & $72.79\%$ & $\mathbf{74.75\%}$ & new physical-only no-backprop frontier \\
Ungated settling, $\eta=\{0.10,0.20,0.30\}$, $1$ step & $72.82\%$ & $74.72\%$ & stronger field does not compound \\
Ungated settling, $\eta=\{0.05,0.10,0.20\}$, $2$ steps & $72.81\%$ & $74.71\%$ & over-relaxation loses the gain \\
Best setting, RBF scale $0.14/0.16$ & -- & $74.38/74.59\%$ & old scale $0.15$ remains optimal \\
Best setting, RBF ridge $5/20$ & -- & $74.43/74.51\%$ & ridge $10$ remains optimal \\
Restored legacy rerun, same top-six protocol & $72.79\%$ & $\mathbf{74.75\%}$ & confirms reproducibility after code patch \\
Top-seven/top-eight/top-ten binary banks, legacy features & $72.83/72.73/72.61\%$ & $74.64/74.63/74.58\%$ & extra residual pairs add noise \\
Rich dendritic gates: ungated, soft, top-$2$, top-$3$, soft-top-$3$ & $72.81\%$ & $74.45\%$ & late multigate compartments overfit the spline solve \\
Focused scale/ridge sweep: $s=0.145/0.155$, ridge $8/12$ & -- & $74.46/74.56$, $74.47/74.62\%$ & no retuning beats $s=0.15$, ridge $10$ \\
\bottomrule
\end{tabular}
\end{center}
The best prospective graph row fixes $37$ mistakes made by the $74.70\%$ pair-margin controller and breaks $32$, for a net gain of five additional test examples. Relative to the old $74.51\%$ physical frontier, it fixes $205$ mistakes and breaks $181$, for a net gain of $24$ examples. After the literature audit we pushed this route harder. First, we patched \texttt{bio\_cifar\_pairwise\_prospective\_fusion\_audit.py} with an opt-in rich dendritic feature mode exposing simultaneous ungated, soft, top-$2$, top-$3$, and soft-top-$3$ apical gates; this increased the prospective feature dimension from $429$ to $1207$ but dropped the RBF controller to $74.45\%$. Second, we added \texttt{--max\_members} to \texttt{bio\_cifar\_gate\_pairwise\_sources.py} and tested top-seven, top-eight, and top-ten ungated binary residual banks from the already generated top-ten pair columns; all underperformed the top-six bank. Third, a full-CIFAR distance-forward contrastive attractor audit, \texttt{apps\_industrial\_breakthrough/bio\_cifar\_distance\_forward\_contrastive\_mps.py}, used two deterministic views, $16$ filters/class, a $16256$-dimensional physical state, a $2048$-dimensional mixed-cell latent, $192$ class centers/class, top-$12$ center scores, rank-$20$ class subspaces, and a $2048$-center RBF controller on MPS. Its base physical ridge reached $68.18\%$, the compact distance-forward ridge reached $46.09\%$, and the best distance-forward spline controller reached only $68.46\%$. Thus class-attractor goodness is not yet a replacement sensory geometry; at this stage it is weaker than the population-column manifold.
The main lesson is architectural rather than numerical: a weak ungated local graph field is more useful than confidence-gated, multi-step, larger-pair, or late multigate relaxations. That matches the biological hypothesis better than a hard gate: pair columns should act like local voltage/neuromodulatory perturbations that the global controller can choose to use, not like externally forced class switches. The final-logit route is now saturated. The next experiment should move this prospective graph one level earlier: the pair residual columns should exchange graph messages while their patch/filter states are being formed, with validation-selected pair recruitment, local contrastive targets in the spirit of Forward-Forward goodness \cite{hinton2022forwardforward}, and branch-local predictive residuals rather than exposing only settled logits to the final controller. The broader biological framing follows the review of no-backprop credit-assignment mechanisms in \cite{lillicrap2020backpropbrain}: the useful ingredients are local eligibility, structured feedback or residual channels, and compartmental state differences, not a scalar global reward signal alone.
We then tested one early-source version of that idea. In \texttt{bio\_cifar\_pairwise\_residual\_columns\_mps.py}, the pair-column generator already exposes two pre-logit biological signals: an apical projection of the source-population state, and filter-score weights derived from the source-population conflict/innovation field. Combining \texttt{apical\_mode=all}, a $512$-dimensional fixed apical projection, and \texttt{filter\_weight\_mode=all} with strength $1.0$ slightly improves the top-six pairwise residual source mean to $70.74\%$ (the earlier apical-all source was $70.61\%$). However, the improvement does not transfer to the stable base manifold controller. The follow-up prospective audit gives $73.98\%$ combined pair control, $73.58\%$ prospective graph control, $74.37\%$ base-only RBF plus pair control, and $74.04\%$ base-only RBF plus prospective graph, all below the previous $74.70\%$ pair-margin controller. Thus stronger local binary pair margins are not automatically better global residual coordinates; pair information must be recruited by validation-stable marginal innovation, not by maximizing pair-column standalone strength.
The next set of experiments moved the predictive target into the early physical columns. The runner \texttt{apps\_industrial\_breakthrough/bio\_cifar\_physical\_predictive\_pretrain\_mps.py} partitions each CIFAR image into a $4\times4$ cortical sheet. Each column receives pooled RGB patch statistics, local means, standard deviations, centered energy, edge means/maxima, and an unsupervised Hebbian patch bank. The best full-core predictive source uses $128$ filters for each $3\times3$ and $5\times5$ kernel family, four heterogeneous ODE-inspired branches of width $96$ per column, two lateral diffusion/inhibition steps, source view $0$, and target view $1$. The four branch nonlinearities are a leak/tanh cell, a conductance softsign cell, a damped oscillatory pole cell, and a signed Gaussian event cell. For each region $r$, the local predictive map is fitted by a closed-form ridge solve from north/south/east/west/global source-column context plus coordinates to the target-view state,
\[
\hat h^{(1)}_r
=
\argmin_{A_r}
\left\|
C_r(h^{(0)}) A_r - h^{(1)}_r
\right\|_2^2
+\lambda \|A_r\|_F^2,
\]
with no reverse-mode graph. Labels enter only after this self-supervised predictive step, through closed-form class readouts and a hard-state RBF spline controller over the saved source/prediction logits.
\begin{center}
\scriptsize
\begin{tabular}{@{}p{0.30\linewidth}p{0.18\linewidth}p{0.18\linewidth}p{0.25\linewidth}@{}}
\toprule
Experiment & Standalone predictive controller & Fused full-CIFAR test & Interpretation \\
\midrule
$4\times4$, branch $96$, source+target+prediction profiles & $66.20\%$ & $75.51\%$ & first early-predictive lift over $74.75\%$ prospective graph \\
Source+prediction pair only, same branch $96$ source & -- & $\mathbf{75.84\%}$ & target-state logit was noisy; the useful signal is the phase-separated source/prediction pair \\
Multi-view targets $1,2,3$ plus one settling step & $67.50\%$ & $75.19\%$ all profiles, $75.84\%$ source+prediction subset & more views improve the internal controller but add redundant/noisy final sources \\
Contrastive rolled-target score states & $64.39\%$ & $74.96\%$ & simple negative-roll contrast is too weak \\
Regional predictive lift, $32$ fixed mixed-cell features per stream/region & $67.66\%$ internal controller, $63.60\%$ lifted-state ridge & $74.71\%$ all profiles & compressed residual geometry helps the internal logit controller but is noisy for final hard-state neighborhoods \\
Wider branch $128$, $160$ filters/kernel & $67.63\%$ & $74.92\%$ & raw capacity improves standalone accuracy but hurts complementarity \\
One $12$-view member-2 population source & $72.27\%$ & $75.78\%$ when appended & deterministic view scaling is saturated and expensive \\
Residual-hard Fisher patch growth from current frontier errors & $70.78\%$ & $75.72\%$ when appended & hard-example patch growth alone does not create the missing geometry \\
Clean validation factory round 5, coarse-wide plus fine prediction sources & best member $67.28\%$ & $75.85\%$ & validation-selected robust stack; simple extra RGB-prediction seeds are rejected \\
Retinal predictive targets in fine $8\times8$ columns & target $4/5/6/7$ predicted ridges $66.73/66.13/64.81/66.24\%$ & $\mathbf{76.12\%}$ fixed validation-selected stack & opponent, edge, and local-contrast target codes add early biological sensory geometry \\
PGPE source-topology search over retinal stack & validation $76.98\%$ & $\mathbf{76.21\%}$ & non-greedy evolved source weights beat equal mean fusion without backprop \\
Sparse PGPE refinement, top-$10$ sources & validation $76.92\%$ & $\mathbf{76.36\%}$ measured held-out test & sparser topology generalizes better but is not the validation-best selection \\
Wide/consensus PGPE audits & validation $77.04$--$77.06\%$ & $75.94$--$76.09\%$ & pure validation-label source search can overfit the $5$k validation split \\
Shared multi-retinal generator, kernels $3/5/7$ & best member $65.55\%$, controller $66.48\%$ & not promoted to fusion bank & richer targets plus stronger diffusion/inhibition slow the run and reduce source separability \\
Cell-family proxy screen, $20$k train edge target & mixed $59.58\%$ vs conductance $57.90\%$, HH $56.86\%$, spline $35.88\%$ & proxy only & current mixed pole bank remains the best transfer family; naive compact spline windows are misaligned \\
Precision/residual lift proxy, same edge target & precision lift $49.54\%$/$49.20\%$, regional lift $48.82\%$ & proxy only & compact residual lifts destroy separability; use target-bank/reliability selection instead \\
Target-bank proxy and promoted target $7$ & proxy target $7$ $61.46\%$; full promoted target $7$ $66.67\%$ & PGPE $76.10\%$, fixed add $75.84\%$ & proxy finds a strong standalone source, but complementarity must be selected separately \\
Neuromodulated reliability gate, existing bank & no positive dopamine additions & preserves weighted PGPE top-$10$: $76.92\%$ validation, $\mathbf{76.36\%}$ test & fold-stable marginal innovation rejects non-complementary sources \\
Dopamine-generated retinal/V1 sources & retained $4{\to}5$ and $7{\to}5$ source paths & $\mathbf{77.05\%}$ test at $77.92\%$ validation with three retained additions & source generation optimized for marginal innovation, not standalone accuracy \\
\bottomrule
\end{tabular}
\end{center}
The important clean-validation follow-up patched the predictive runner with a true \texttt{45k/5k/10k} core/validation/test split and saved \texttt{member\_val\_logits}. On the first validation-aware population bank, \texttt{apps\_industrial\_breakthrough/bio\_cifar\_adaptive\_column\_search\_mps.py} selects five sources when the clean predictive pair is available: the two strong member-2 multi-view sources, member $2$ of the base eight-member population, predictive source-state member $16$, and member $2$ of the offset population. This reaches $74.63\%$ held-out test accuracy. The matched no-predictive control selects seven purely population/residual sources and reaches only $73.78\%$. Thus the source/prediction column is not just a test-set tuning artifact: under clean validation it contributes a $+0.85$ point held-out gain and reduces the number of selected sources. However, this rigorous clean result was still below the exploratory full-core $75.84\%$ frontier, so the next experiments moved upstream again instead of tuning only the final controller.
The second clean-validation factory tested view direction, patch scale, cell-type diversity, topology, and targeted class-pair synaptogenesis. Reversing the original blur-like prediction direction (source view $1$ to target view $0$) reaches only $65.80\%$ internal predictive control. A high-pass source view $3$ to identity target reaches $66.37\%$. Adding a $7\times7$ patch scale lowers the local prediction residual from about $0.579$ to $0.538$, but classification falls to $65.74\%$, proving that low reconstruction residual is not the right source-selection objective. The most useful early change before the retinal pass is heterogeneity of cell branches: six branches of width $64$ reach $66.64\%$, six branches of width $80$ reach $66.71\%$, and wider $4\times4$ branch-diverse sources plus fine $8\times8$ predicted targets yield the clean validation-selected round-five stack. This stack selects indices $[12,63,13,60,54,5,64,0,2]$ from the shared clean candidate bank and reaches $75.85\%$ test by mean fusion. Adding two more simple RGB-prediction seeds with the same fixed source-order split does not change the selected stack: round six selects the same indices and again reaches $75.85\%$. Thus seed diversity alone is saturated.
The next successful upstream change is biological target shaping. We extended \texttt{bio\_cifar\_physical\_predictive\_pretrain\_mps.py}'s deterministic sensory views from raw/blur/color/high-pass variants to four retinal-style signed targets: opponent center-surround (view $4$), Sobel edge channels (view $5$), local contrast normalization (view $6$), and DoG/opponent channels (view $7$). All runs use the same clean split, source-order seed \texttt{20260601}, MPS one-line \texttt{runpy} launch, $8\times8$ columns, eight branches of width $32$, $128$ unsupervised patch filters per kernel, source view $0$, guarded \texttt{source\_state,predicted\_target\_state} profiles, prediction ridge $300$, and class ridge $300$. The independent single-target predicted-state readouts are $66.73\%$ for opponent center-surround, $66.13\%$ for edge, $64.81\%$ for local contrast, and $66.24\%$ for DoG/opponent. Greedy selection over all sources over-selects the edge member and reaches only $75.24\%$ test, so the robust protocol fixes the validated round-five stack and evaluates retinal additions under the same held-out validation criterion. Adding predicted retinal members $[67,69,71]$---opponent, edge, and local contrast---raises validation from $76.58\%$ to $76.80\%$ and held-out test from $75.85\%$ to $\mathbf{76.12\%}$. The fixed stack is reproduced by \texttt{bio\_cifar\_adaptive\_column\_search\_mps.py --fixed\_selected 12,63,13,60,54,5,64,0,2,67,69,71} with the round-five bank plus the target-$4/5/6/7$ retinal logits; the saved artifact is \texttt{apps\_industrial\_breakthrough/bio\_cifar\_adaptive\_column\_search\_outputs\_cleanval\_round8\_fixed\_retinal\_stack/}.
The negative follow-ups are equally important. A richer multi-scale patch bank with kernels $1,3,5,7$ lowers the fine-grid prediction residual to $0.3596$ but drops source/predicted readouts to $60.65/64.74\%$, again showing that reconstruction fidelity can chase nuisance detail. Extra diffusion/inhibition smoothing gives source/predicted readouts $61.98/65.60\%$. A shared-source multi-target retinal run over views $4,5,6$ has a respectable internal controller ($66.50\%$) but weaker individual members ($65.95/65.14/64.79\%$). A second edge seed ($65.43\%$ predicted), a second opponent seed ($66.42\%$ predicted), and a coarse-wide $4\times4$ edge target ($65.56\%$ predicted) do not improve the validation-selected fixed stack. Linear closed-form control over member logits, margins, entropies, and votes is also worse than mean fusion across ridge values. The conclusion is now sharper than before: the path forward is not lower residual, more late controller capacity, or more same-family seeds. The gain comes from selecting biologically meaningful local predictive targets before logit compression. The next credible model should learn or validate retinal/V1 target banks inside recurrent columns, then use multiple validation folds or online neuromodulatory reliability to decide which local target populations are retained.
We therefore ran the first explicit PGPE-style architecture search on this stack with \texttt{apps\_industrial\_breakthrough/bio\_cifar\_pgpe\_logit\_arch\_search.py}. The search space is deliberately small but principled: each saved physical-column logit member is a candidate cortical source, the genome is a sparse nonnegative source-weight vector, and each candidate is scored only by a forward validation pass over normalized logits. Antithetic parameter-based exploration updates the source-weight logits; no reverse-mode graph, layerwise gradient, or CIFAR test labels are used. Starting from the fixed retinal stack and searching $180$ steps with population $48$, top-$12$ source support, $\sigma=0.28$, learning rate $0.06$, and temperature $0.35$ raises validation to $76.98\%$ and held-out test to $\mathbf{76.21\%}$. The selected topology is $[13,69,60,64,54,5,67,63,2,71,12,0]$ with weights approximately $[0.263,0.114,0.107,0.100,0.098,0.073,0.071,0.047,0.039,0.037,0.029,0.022]$. The leading source is the offset member-2 population anchor; the next sources are edge/opponent/local-contrast retinal predictive members and coarse-wide source states. This is a small numerical gain, but a meaningful research turn: architecture/topology search over closed-form physical sources can exploit the speed of OSNR evaluation in a way that ordinary NEAT/PGPE over backprop-trained networks usually cannot. The next loop should broaden the genome from source weights to source-generating architecture: retinal target type, grid, branch width, branch nonlinearities, diffusion, and local predictive objective should become mutable genes, while the inner readouts remain algebraic.
The follow-up topology loop clarifies both the promise and the failure mode. A tighter top-$10$ PGPE refinement seeded from the same retinal stack uses $240$ steps, population $64$, $\sigma=0.22$, learning rate $0.04$, and temperature $0.28$. It lowers validation to $76.92\%$ but raises the measured held-out test accuracy to $\mathbf{76.36\%}$ with selected sources $[13,64,54,12,71,5,60,2,69,63]$. This cannot be treated as a validation-selected frontier, but it is a useful generalization clue: sparse source support can remove noisy retinal members. The opposite top-$16$ branch reaches $77.04\%$ validation but falls to $75.94\%$ test, and the four-run consensus audit \texttt{bio\_cifar\_pgpe\_consensus\_eval.py} reaches $77.06\%$ validation but only $76.09\%$ test. A robust top-$10$ run with a validation-half stability penalty and an $8192$-sample core reward anchor reaches $76.94\%$ validation and $76.12\%$ test. The diagnosis is therefore precise: source-weight evolution is real and cheap, but a single $5$k validation split is too small to drive an unconstrained source-topology search. Future topology search must either use multiple clean validation folds, an online neuromodulatory reliability field, or a source-generating proxy objective before any test-set audit.
We also moved one step earlier in the pipeline and tested whether a single richer physical generator could replace separate retinal source runs. The artifact \texttt{bio\_cifar\_physical\_predictive\_pretrain\_mps\_outputs\_cleanval\_round9\_multiretina\_g8\_b8x32\_pf160\_k357\_diff3/} uses the same clean split and source-order seed, but adds patch kernels $3,5,7$, $160$ unsupervised filters per kernel, diffusion steps $3$, diffusion $\alpha=0.12$, inhibition $0.22$, and target views $4,5,6,7$ in one shared-source run. This is a negative architecture result: source-state accuracy is only $61.96\%$, predicted target-state accuracies are $65.18\%$, $64.48\%$, $64.16\%$, and $65.55\%$, and the predictive-logit spline controller reaches $66.48\%$. The run is also slower because each full-grid target requires many local algebraic solves. Thus ``more retinal biology'' by itself is not the answer. The next source-generating loop should use a cheap proxy stage and mutate one biologically meaningful factor at a time around the previous winning operator family: retinal target, branch nonlinearity/pole family, diffusion schedule, skip/residual precision, and local prediction objective.
The first such proxy stage added a guarded \texttt{--cell\_family} switch to the physical predictive runner while keeping the default mixed transfer bank unchanged. On a clean $18$k/$2$k/$5$k core/validation/test proxy with $8\times8$ columns, eight branches of width $24$, edge target view $5$, and the same source-order seed, the original mixed family remains best: predicted edge-state accuracy is $59.58\%$. Conductance-style reversal gates reach $57.90\%$, a Hodgkin--Huxley-inspired algebraic gate reaches $56.86\%$, and compact cubic spline windows collapse to $35.88\%$. This is an important negative result for first-principles design. Biological names alone do not help; the transfer family must be matched to the descriptor distribution and preserve class-separable geometry. The next knob should therefore be precision-balanced residual/skip routing around the existing mixed cells, not a full-scale promotion of these naive alternative transfer laws.
That precision-routing branch also failed in the first proxy. We added \texttt{precision\_predictive\_lift\_state}, a compact predictive-coding lift that scales local residual streams by inverse residual energy before fixed mixed-cell projection. With precision floor $0.05$ it reaches only $49.54\%$; damping the precision floor to $0.25$ reaches $49.20\%$. The matched unweighted \texttt{regional\_predictive\_lift\_state} reaches $48.82\%$. Since the plain predicted edge target remains $59.58\%$, the failure is not just over-amplified precision; compact residual-lift projections are losing class geometry. The next credible branch is therefore not more residual lifting. It is target-bank/reliability search: choose which retinal/V1 predictive targets to create and retain using validation folds, source recurrence, or an online neuromodulatory reliability signal.
The first target-bank proxy is more encouraging but also exposes the next bottleneck. With the same $18$k/$2$k/$5$k proxy, mixed cells, and plain predicted-target readouts, views $4,5,6,7$ score $59.98\%$, $59.58\%$, $58.98\%$, and $61.46\%$ respectively; the predictive-logit controller reaches $61.56\%$. We promoted the proxy winner, target view $7$, to the full clean $45$k/$5$k/$10$k run with branch width $32$ and seed \texttt{20260661}. The promoted target reaches $66.67\%$ as a standalone predicted-state readout, improving over the previous target-$7$ seed. However, appending it to the existing source bank does not improve the final topology: PGPE with the new member reaches $76.98\%$ validation and $76.10\%$ test, while fixed mean fusion of the previous retinal stack plus the new member falls to $75.84\%$. The proxy therefore works for finding stronger standalone source generators, but standalone strength is not equivalent to final-stack complementarity. The next target-bank search must score both source quality and marginal innovation against the current retained population.
We therefore implemented the first explicit neuromodulated reliability gate in \texttt{apps\_industrial\_breakthrough/bio\_cifar\_neuromodulated\_reliability\_search.py}. The retained source population defines the current cortical state. Each candidate is evaluated by a local dopamine scalar composed of fold-stable marginal gain, rescue/harm innovation on current errors, source disagreement, and a redundancy penalty against the retained logit field. A later patch also supports weighted retained populations, tunable dopamine coefficients, and small candidate gate strengths $\eta$, so the audit can score a PGPE-weighted population rather than only equal mean fusion. On the current full source bank including the promoted target-$7$ member, the gate correctly rejects all additions to the validation-selected $12$-source retinal stack: the top candidate has negative dopamine, negative fold gain, and the final stack remains $76.80\%$ validation and $76.12\%$ test. On the PGPE top-$10$ weighted population, it reproduces the current best measured test point, $76.92\%$ validation and $\mathbf{76.36\%}$ test, and again rejects all additions; the top candidate has negative dopamine ($-0.00919$) and negative mean fold gain ($-0.02$ points). The first implementation over-penalized redundancy, however: a new source with positive gain on all five validation folds was scored negative because the redundancy coefficient was too large. The corrected default keeps fold gain, minimum fold gain, and rescue/harm as the primary neuromodulatory signal and reduces the redundancy coefficient from $0.015$ to $0.003$.
We then optimized source generation directly for this dopamine signal. Two full clean-split MPS runs generated biologically shaped retinal candidates using \texttt{source\_state,predicted\_target\_state} profiles, $8\times8$ columns, eight mixed branches of width $32$, $128$ unsupervised patch filters per kernel, patch kernels $3,5$, prediction ridge $300$, and class ridge $300$. A high-pass source to DoG/opponent target ($3{\to}7$, seed \texttt{20260671}) reached $62.03\%$ source-state accuracy and $66.07\%$ predicted-target accuracy but was not retained. An opponent center-surround source to Sobel-edge target ($4{\to}5$, seed \texttt{20260672}) reached $61.05\%$ source-state accuracy and $65.85\%$ predicted-target accuracy; after the corrected dopamine score, its predicted member $79$ has fold gains $[0.007,0.002,0.009,0.003,0.001]$, $\eta=0.10$, and raises the PGPE-weighted base from $76.92\%$ to $77.36\%$ validation. Adding its paired source-state member $78$ with $\eta=0.01$ gives $77.44\%$ validation and $76.80\%$ held-out test. A seed-diverse repeat of the same $4{\to}5$ map (seed \texttt{20260673}) gives $60.54\%$ source-state and $66.11\%$ predicted-target accuracy. With all six generated candidates visible and a fixed two-addition retention budget, the dopamine gate selects $[79,80]$ and reaches the new validation-selected full-CIFAR physical/no-backprop row: $77.68\%$ validation and $\mathbf{76.98\%}$ held-out test. Allowing a third same-family addition raises validation to $77.92\%$ but lowers test to $76.79\%$; thus the new lesson is not ``add every positive dopamine source''. It is that source generation must be driven by marginal cortical innovation, while source retention needs biological consolidation constraints before another member of the same sensory family is kept.
The next retention patch made that constraint explicit. \texttt{bio\_cifar\_neuromodulated\_reliability\_search.py} now exposes \texttt{--max\_additions\_per\_source\_path} and \texttt{--min\_candidate\_fold\_gain}. With \texttt{max\_additions\_per\_source\_path=1}, the all-visible generated bank stops after $[79,80]$ because the remaining $3{\to}7$ candidate has negative minimum fold gain. We then tested a new reversed retinal family, edge source to opponent target ($5{\to}4$, seed \texttt{20260674}). This source is weak as a standalone physical generator ($55.91\%$ source state, $61.35\%$ predicted target). Without the nonnegative-fold guard, validation accepts source-state member $82$ and rises to $77.82\%$, but held-out test drops to $76.88\%$; the accepted row has one negative validation fold. With both guards enabled---one retained member per generated source path and \texttt{min\_candidate\_fold\_gain=0}---the selector rejects that family and again returns $[79,80]$ with $77.68\%$ validation and $\mathbf{76.98\%}$ test.
The next source-family pass kept the same guards and moved to adjacent retinal/V1 target directions. Local-contrast source to edge target ($6{\to}5$, seed \texttt{20260675}) is a negative result: the standalone source and predicted-target readouts are only $53.74\%$ and $59.52\%$. DoG/opponent source to edge target ($7{\to}5$, seed \texttt{20260676}) is the first positive post-guard family: source-state accuracy is $60.13\%$, predicted-target accuracy is $63.93\%$, and source-state member $86$ is retained with $\eta=0.02$, fold gains $[0.002,0.002,0.003,0.001,0.004]$, and dopamine $+0.00307$. This raises the guarded frontier to $77.92\%$ validation and $\mathbf{77.05\%}$ held-out test with selected additions $[79,80,86]$. A nearby DoG/opponent source to local-contrast target ($7{\to}6$, seed \texttt{20260677}) reaches $59.63\%$ source-state and $63.37\%$ predicted-target accuracy but is rejected after the frontier: its best member has only $+0.06$ point mean validation gain, a negative fold, and negative dopamine. We then tested stronger edge predictors as controls. High-pass source to edge target ($3{\to}5$, seed \texttt{20260678}) reaches $61.66\%$ source-state and $65.34\%$ predicted-edge accuracy; raw source to edge target ($0{\to}5$, seed \texttt{20260679}) reaches $62.26\%$ and $65.71\%$. Both are rejected after $[79,80,86]$: the best raw/high-pass member has negative mean gain and negative dopamine. A seed repeat of the accepted $7{\to}5$ family (seed \texttt{20260680}) has nearly matched standalone readouts ($59.95\%$ source, $63.89\%$ predicted target) but is also rejected after the frontier with negative mean gain and a negative fold. The current rule is therefore sharper: promote a new physical source only if it is marginally useful, not already represented by a retained local source path, and nonnegative on every validation fold; among the tested directions, the first edge prediction from opponent/DoG sources is useful, while reversed edge-to-opponent, local-contrast targets, raw/high-pass edge predictors, and seed repeats are not yet useful.
The next full-scale MPS batch tested whether the failure was caused by an overly narrow cell law or an overly blunt global retention dose. First, we promoted the explicit cell-family ablation to the full clean split around the successful $4{\to}5$ objective. The Hodgkin--Huxley-style algebraic gate (seed \texttt{20260681}) reaches $58.99\%$ source-state accuracy, $64.59\%$ predicted-target accuracy, and $61.35\%$ internal spline-control accuracy; after the frontier its best member has negative dopamine ($-0.00263$). Compact spline-window cells (seed \texttt{20260682}) collapse to $43.20\%$ source-state, $42.69\%$ predicted-target, and $43.77\%$ control accuracy. Conductance/reversal-potential cells (seed \texttt{20260683}) are the only plausible alternative: $61.19\%$ source-state, $65.93\%$ predicted-target, and $63.64\%$ control accuracy. They still fail the robust retention criterion after the frontier: the best conductance member has $+0.08$ point mean validation gain but a $-0.20$ point minimum fold gain and dopamine $-0.00131$. Thus naive biological transfer-law names are not enough; the mixed pole bank remains the best matched transfer family in this CIFAR column descriptor distribution.
We then moved the change from cell law to capacity allocation. The previous strongest from-scratch predictive column used a coarser $4\times4$ cortical sheet with wider $4\times96$ branches, so we applied that architecture to the dopamine-positive retinal objectives and added \texttt{--export\_controller\_member} to \texttt{bio\_cifar\_physical\_predictive\_pretrain\_mps.py}. This patch exports the internal predictive-logit spline controller as a normal source-bank member with core, validation, and test logits; otherwise the reliability bank only saw the individual \texttt{source\_state} and \texttt{predicted\_target\_state} readouts. On $4{\to}5$ (seed \texttt{20260684}), the coarser/wider source and predicted-state members reach only $64.48\%$ and $64.13\%$, but the internal controller reaches $66.44\%$ test and $66.18\%$ validation. Nevertheless, after $[79,80,86]$ the best exported-bank member is the source-state member, not the controller: it raises validation to $78.04\%$ with $+0.12$ point mean fold gain, but has one $-0.20$ point fold. A finer low-$\eta$ grid reduces the damage but still gives a negative minimum fold ($-0.10$ point) and no retained addition. The same $4\times4$ architecture on $7{\to}5$ (seed \texttt{20260685}) reaches $63.02\%$ source-state, $63.00\%$ predicted-state, and $64.59\%$ control accuracy; its best low-$\eta$ after-frontier probe gives only $+0.02$ point mean gain and a $-0.10$ point fold. These runs show that coarser/wider physical columns can improve the internal controller, but their errors are not yet fold-stable complements to the retained population.
Finally, we added \texttt{apps\_industrial\_breakthrough/bio\_cifar\_neuromodulated\_context\_gate.py} to test a more biological local consolidation rule. Instead of one global $\eta$ per source, the runner learns a frozen context table on the core split over base prediction, candidate prediction, and confidence-margin bins. Each context chooses its retention dose by closed-form grid search; the table is then evaluated on validation and test with no reverse-mode graph. On the after-frontier candidate set from the HH, spline, conductance, and coarser/wider runs, the permissive $4$-bin gate improves validation only to $77.98\%$ and lowers held-out test to $77.04\%$; the best candidate is the conductance predicted-target member. Stricter $3$-bin/support-$200$ and $2$-bin/support-$500$ gates become conservative and leave the $77.92\%/77.05\%$ frontier unchanged. Moving the same local gate earlier, from the PGPE top-$10$ base over all generated candidates, gives $76.98\%$ validation and $76.33\%$ test from a $76.92\%/76.36\%$ base. The conclusion is sharp: neither global dopamine nor simple local context gating is the active bottleneck now. The next lift must create genuinely new upstream source geometry, likely by changing the physical column objective or multi-stage recurrent target formation before logit compression, rather than by repeatedly reweighting the existing edge-family sources.
The first such upstream-geometry test is a multi-target recurrent formation run. The artifact \texttt{bio\_cifar\_physical\_predictive\_pretrain\_mps\_outputs\_cleanval\_g4\_b4x96\_pf128\_source4\_targets567\_settle1\_seed20260686\_exportctrl/} keeps the coarser $4\times4$, $4\times96$ mixed-cell architecture, uses source view $4$, predicts target views $5,6,7$, and then applies one settled-source update with $\beta=0.25$ before the readouts. This is the first positive upstream signal after the selector failures: source-state accuracy rises to $65.51\%$, the settled source reaches $65.56\%$, and the internal predictive-logit spline controller reaches $67.55\%$ test with $67.24\%$ linear control. Post-frontier retention still rejects it after $[79,80,86]$: the strongest standalone controller member has one $-0.20$ point validation fold. Placing the settled source earlier is more useful. Starting from the PGPE top-$10$ base, the guarded dopamine gate retains settled-source member $100$ with $\eta=0.10$, all five folds positive, and raises validation to $77.32\%$ and test to $76.81\%$. Continuing the generated-source search from that base selects $[78,90,86,80]$ and gives $77.64\%$ validation and $77.07\%$ held-out test. This is a new measured held-out high for the branch, but it is not a validation-selected frontier because validation remains below $77.92\%$. The useful scientific signal is ordering: recurrent multi-target formation can create a generalizing source that changes which later edge-family members are useful. The next serious run should deepen this source-family, not the selector: sweep settled target sets, recurrent $\beta$, and second-stage targets while keeping the validation rule fixed.
The follow-up sweep confirms that this is a source-geometry problem, not a pure retention problem. All runs used the same clean $45$k/$5$k/$10$k split, \texttt{source\_order\_seed=20260601}, Apple MPS, no reverse-mode graph, $4\times4$ columns, four mixed branches of width $96$, $128$ unsupervised patch filters per kernel, patch kernels $3,5$, prediction ridge $300$, class ridge $300$, control ridge $30$, $512$ RBF centers, and exported controller logits. Removing target view $6$ (\texttt{source4\_targets57\_settle1\_seed20260687}) improves some raw readouts but lowers the controller to $67.13\%$ and is not retained after the frontier; from the earlier PGPE base it gives only $77.18\%$ validation and $76.84\%$ test. Increasing recurrence to two settled steps (\texttt{source4\_targets567\_settle2\_seed20260689}) lowers the source/predicted readouts and controller to $67.06\%$; the selector can extract a tiny validation-only after-frontier gain ($77.96\%$ validation) but held-out test falls to $77.02\%$. Stronger settling with $\beta=0.40$ is rejected ($77.92\%/77.03\%$), and weaker settling with $\beta=0.15$ reaches $78.00\%$ validation but drops test to $76.98\%$, showing that residual RMS improvements are not sufficient when the induced class geometry is wrong. The reciprocal directed graph, source view $5$ predicting $4,6,7$ (seed \texttt{20260691}), is a hard negative: source/predicted readouts are about $60\%$ and the controller reaches only $61.79\%$. Thus the useful column is directed: view $4$ is a good source for the $5,6,7$ target bank, but the reverse source is not.
The successful push is seed-diverse cortical population formation around the same directed source graph. A repeat of the source-$4$, targets-$5,6,7$, one-step $\beta=0.25$ architecture with seed \texttt{20260692} produces source/predicted/settled readouts $65.79\%$, $65.41\%$, $64.94\%$, $65.32\%$, and $66.01\%$, plus a $67.47\%$ predictive-logit controller. After the existing validation-selected frontier $[79,80,86]$, a one-member gate retains member $97$ (predicted target state $v5$) with $\eta=0.06$ and raises validation/test to $78.10\%/77.06\%$. Allowing the same dopamine rule to add a second member from this source retains member $100$ (settled source state) with $\eta=0.018$ and reaches the new full-CIFAR physical/no-backprop frontier: \textbf{$78.18\%$ validation and $77.11\%$ held-out test}. The final retained source indices are $[13,64,54,12,71,5,60,2,69,63,79,80,86,97,100]$ with weights approximately $[0.215,0.084,0.072,0.071,0.059,0.052,0.049,0.046,0.046,0.040,0.081,0.090,0.018,0.059,0.018]$. This is not a SOTA CIFAR result, but it is a clean no-backprop improvement over the previous validation-selected $77.92\%/77.05\%$ frontier and over the previous measured $77.07\%$ held-out high. Importantly, continuing to add old generated-source candidates after this base raises validation to $78.28\%$ but lowers test to $77.04$--$77.08\%$, and further source-$4$ seed repeats are quality-gated out: seed \texttt{20260693} has a $67.03\%$ controller and contributes nothing after the new base, while seed \texttt{20260694} drops to a $66.55\%$ controller. The working rule is therefore specific: generate a population of directed multi-target physical columns, retain only class-aligned seed members that improve every validation fold at small $\eta$, and reject validation-only additions even when their local residuals improve.
We then hardened the consolidation rule itself. The new runner \texttt{apps\_industrial\_breakthrough/bio\_cifar\_neuromodulated\_resample\_consolidation.py} keeps the same no-backprop source logits, but accepts a candidate only if its small-$\eta$ gain survives many random validation resamples. A candidate must have enough full-validation support, a sufficiently high resample win rate, and a non-catastrophic lower-tail gain before it is consolidated. Re-auditing the seed-\texttt{20260692} source under this rule still selects members $97$ and $100$, now with weights ending in $0.0588$ and $0.0200$, and raises the reproducible frontier to $78.22\%$ validation and $77.13\%$ held-out test. The first selected member has full validation gain $+0.18$ points, resample lower-tail gain $+0.036$ points, $90.1\%$ resample win rate, and support $+9$ examples. This is a better rule than the deterministic five-fold gate because it rejects candidates whose apparent gain is carried by a few validation examples.
Finally, we implemented an explicit two-stage predictive hierarchy in \texttt{bio\_cifar\_physical\_predictive\_pretrain\_mps.py}. The new \texttt{--second\_stage\_target\_views} option first fits the usual source-$4$ maps to targets $5,6,7$, settles the source state, and then fits a second set of local ridge maps from the settled source to the same targets before readout. A two-stage-only seed \texttt{20260695} improves physical residuals for all targets ($v5:0.773{\to}0.756$, $v6:0.833{\to}0.813$, $v7:0.697{\to}0.688$) but is rejected after the consolidated base; its best marginal row has only $+0.04$ validation points, support $+2$, and a negative resample lower tail. Exporting both first-stage and second-stage profiles from the same seed lifts the internal controller to $67.85\%$, but it is still redundant after the consolidated source-\texttt{20260692} base. A second hybrid seed, \texttt{20260696}, has weaker standalone readouts and controller ($67.35\%$), yet its first-stage predicted target-$6$ member is marginally complementary. The resampled gate retains this member with $\eta=0.015$, full validation gain $+0.10$ points, lower-tail gain $0.00$, $85.9\%$ resample win rate, and support $+5$ examples, producing the new full-CIFAR physical/no-backprop frontier: \textbf{$78.32\%$ validation and $77.14\%$ held-out test}. The final retained source indices are $[13,64,54,12,71,5,60,2,69,63,79,80,86,97,100,113]$. A targeted v6-only hierarchy (seed \texttt{20260697}) gives the best v6 residual in the batch ($0.828{\to}0.799$) but weak class readouts and is rejected. The scientific conclusion is sharper than the numerical gain: physically better target reconstruction is not enough; useful no-backprop source formation requires class-aligned multi-target context plus resampled neuromodulatory consolidation.
The next push made that conclusion explicit. We added two supervised-but-still-local training signals to \texttt{bio\_cifar\_physical\_predictive\_pretrain\_mps.py}. The option \texttt{--class\_align\_strength} injects a training-only class-centroid dopamine signal into the target states used by the local predictive ridge maps; validation/test features still use only the image-derived source state and the learned maps. A full source-$4$, targets-$5,6,7$ run with strength $0.20$ gives source/predicted/settled readouts $64.97\%$, $65.08\%$, $64.73\%$, $65.36\%$, $65.18\%$ and a $67.02\%$ controller. The resampled gate rejects all exported members after the $78.32\%/77.14\%$ base; the best candidate has full validation gain $-0.06$ points, lower-tail gain $-0.109$ points, $2.6\%$ win rate, and support $-3$. Thus naively adding class centroids to reconstruction targets does not create useful marginal geometry.
The stronger variant is a local dopamine classifier field. For each cortical region, the script now fits a closed-form map from the region's neighboring source context to the global class signal, producing \texttt{local\_class\_context\_state} and \texttt{settled\_local\_class\_context\_state} profiles. This is closer to a biological three-factor rule: source activity supplies the eligibility field, labels provide a broadcast neuromodulator on the core split, and inference uses only the learned local maps. With ridge $300$, seed \texttt{20260703} reaches $66.37\%$ local-class readout, $66.32\%$ settled local-class readout, and a $67.86\%$ controller. Seed \texttt{20260704} improves to $67.00\%$, $66.84\%$, and a $67.96\%$ controller. A ridge sweep shows the regularization boundary: ridge $30$ overfits tiny-split regional train accuracy to about $99.8\%$ and hurts readout; ridge $1000$ improves the tiny-split smoke but at full scale slips to $66.93\%$, $66.66\%$, and $67.88\%$. These local dopamine fields are therefore a real standalone architectural improvement over raw source/predictive readouts, but they still do not beat the current source bank after consolidation. The resampled additive gate over seeds \texttt{20260703}/\texttt{20260704} rejects all members; the best row is candidate $131$ with $\eta=0.008$, zero full-validation gain, lower-tail gain $-0.073$ points, $38.3\%$ win rate, and support $0$. A fixed closed-form spline controller over the retained base plus all local-class members drops to $75.10\%$ test, and a context-gated dopamine table over the strongest local-class candidates leaves validation flat at $78.32\%$ while test slips to $77.13\%$. The current frontier therefore remains $78.32\%/77.14\%$, and the next necessary change is not another late gate; it is to make the local dopamine classifier field participate earlier in source formation, for example by feeding its regional error/context back into patch selection, target selection, or multi-stage recurrent state formation before logits are compressed.
We implemented that early-feedback test in the same runner. The new \texttt{--early\_class\_context\_gain}, \texttt{--early\_class\_context\_temperature}, and \texttt{--early\_class\_feedback\_steps} options first fit the local class-context field on the raw source columns, convert the resulting regional logits into centered class-probability mixtures over training-set class-centroid displacements, inject that dopamine-like displacement into the source state, renormalize, and diffuse/inhibit the state before any predictive maps or class readouts are fitted. Labels enter only through the core-set local class maps and centroid table; validation/test source shaping uses the learned regional logits. On a $3000/500/1200$ smoke split, moving the field earlier raises the source-state readout from $45.58\%$ to $51.08\%$ at gain $0.50$, temperature $0.45$, confirming that the feedback changes the representation rather than merely adding a late logit source. On the full clean split with the previously strong seed \texttt{20260704}, raw source is $65.79\%$, early-shaped source is $66.37\%$, early class-context is $67.00\%$, shaped local context is $67.42\%$, and predicted targets $v5/v6/v7$ reach $66.56\%/66.06\%/65.96\%$. The shaped local field's mean regional train accuracy rises from $63.88\%$ to $72.08\%$ before settling and $72.30\%$ after settling. Thus the upstream geometry hypothesis is validated. However, a global resampled additive gate still rejects all early-feedback members after the $78.32\%/77.14\%$ base; the best raw additive candidate has full validation gain $-0.04$ points and lower-tail gain $-0.073$ points. We therefore added \texttt{bio\_cifar\_predictive\_recontroller.py} for closed-form subset controllers and upgraded \texttt{bio\_cifar\_neuromodulated\_context\_gate.py} so it can inherit prior source banks and export chainable member logits. Subset controllers improve some held-out test rows but overfit validation. The only robust consolidation lift is an ultra-conservative context replacement gate with $2$ confidence bins, minimum bin support $300$, and support shrink $500$: it selects early-feedback candidate $129$, activates only two contexts with mean eta $0.0129$, and improves the current base from $78.32\%/77.14\%$ to \textbf{$78.36\%$ validation and $77.15\%$ held-out test}, with nonnegative fold minimum. This is numerically tiny, not a SOTA claim, but it is the first evidence that early physical dopamine plus context-local retention can improve both validation and held-out test beyond the resampled frontier. Chaining the exported context member through the generic source-bank normalizer changes its calibration, so the next implementation task is a calibration-aware context-source loader or a replacement-base consolidation protocol, not more blind global addition.
The calibration-aware follow-up resolves that artifact. Both \texttt{bio\_cifar\_neuromodulated\_context\_gate.py} and \texttt{bio\_cifar\_neuromodulated\_resample\_consolidation.py} now accept \texttt{--calibrated\_logits\_npz}; such members are appended to the source bank but bypass the per-member core mean/std/RMS normalization, because they are already fused logits in the ensemble's calibrated decision space. Reloading the round-60 context member this way exactly preserves its replacement-base accuracy, $78.36\%/77.15\%$. Narrow context gating over the early-dopamine members is validation-flat and lowers test to $77.14\%$; the strict resampled gate rejects the same candidates, with the best row having full validation gain $-0.08$ points, lower-tail gain $-0.109$ points, $0\%$ resample win rate, and support $-4$ examples. A broad all-nonbase context gate over the whole available bank is also exactly flat at $78.36\%/77.15\%$. We then tested a recurrent cellular version of early dopamine: \texttt{--early\_class\_feedback\_refit\_rounds} refits the local class-context field on the shaped source and applies another closed-form class-centroid displacement, while \texttt{--early\_class\_feedback\_gate} can modulate the displacement by local entropy, confidence, or margin. On the smoke split, ungated refit improves source/local readouts to $51.92\%/51.75\%$; entropy gating improves settled source and predicted-target readouts ($52.00\%$ settled source, $v5/v6/v7=51.17\%/50.83\%/50.17\%$); confidence gating collapses the source to $45.17\%$ and the controller to $36.42\%$. Full clean seed \texttt{20260704} is more decisive. Ungated refit slightly improves source and settled-source readouts ($66.45\%$, $66.65\%$) but lowers local context and controller ($67.25\%$, $67.17\%$) relative to the one-pass early-dopamine run ($67.42\%$, $67.59\%$). Entropy gating improves the physical residuals ($v5/v6/v7=0.770/0.831/0.693$) but hurts every class readout (source $66.11\%$, local context $66.90\%$, controller $66.74\%$). Context and resampled consolidation reject the recurrent-refit members after the calibrated $78.36\%/77.15\%$ base. The useful conclusion is not merely negative: residual reconstruction, uncertainty-gated dopamine, and discriminative class geometry are empirically different objectives. The next source architecture should therefore optimize local discriminative predictive targets directly---for example class-conditional residual fields, local contrastive target formation, or target selection by validation-stable class innovation---rather than adding more blind residual reconstruction or late logit gates.
We then made that target explicit with \texttt{--discriminative\_target\_mode} and \texttt{--discriminative\_target\_strength}. For each target view, the runner computes the class-conditional target-state center $c_y(r)$ for every cortical region and hidden channel, the global center $\bar c(r)$, and the residual $x_{\mathrm{target}}(r)-c_y(r)$. The tested modes are \texttt{class\_delta}, which uses $c_y-\bar c$; \texttt{class\_center}, which uses $c_y$; and \texttt{class\_residual}, which uses $(x_{\mathrm{target}}-c_y)+(c_y-\bar c)$. The surrogate field is RMS-balanced to the raw target field, then either blended with the raw target for strengths in $[0,1]$ or added for strengths above $1$. This is a training-only local target transform: validation and test states are image-derived, and the learned source-to-target maps receive no validation/test labels. On a matched $2500/500/1200$ MPS smoke split with source view $4$, targets $5,6,7$, early dopamine gain $0.50$, temperature $0.45$, and one settled step, the no-discriminative-target controller is $42.92\%$. \texttt{class\_delta} at strength $1.0$ raises individual predicted-target readouts to $51.42\%/52.17\%/53.25\%$ but leaves the controller at $45.00\%$. \texttt{class\_center} at strength $1.0$ is better: predicted-target readouts become $52.50\%/53.08\%/53.25\%$ and the controller reaches $47.67\%$. Adding regional and precision predictive lifts with \texttt{--regional\_lift\_dim=16} does not improve the individual lift readouts beyond the best predicted-target row, but it improves fusion diversity and raises the smoke controller to $50.92\%$.
The full clean result shows both the promise and the current limitation. The artifact \texttt{bio\_cifar\_physical\_predictive\_pretrain\_mps\_outputs\_cleanval\_disctarget\_center100\_lifts16\_earlydop\_g050\_t045\_seed20260704/} uses the same $45$k/$5$k/$10$k split, Apple MPS, source view $4$, targets $5,6,7$, $4\times4$ columns, four mixed branches of width $96$, $128$ patch filters, prediction ridge $300$, class/context ridge $300$, \texttt{class\_center} strength $1.0$, lift dimension $16$, and exported controller logits. The raw source, shaped source, early class-context, and local class-context rows reproduce the previous clean one-pass geometry ($65.79\%$, $66.37\%$, $67.00\%$, $67.42\%$). Pure predicted-target readouts do not improve ($66.01\%/66.26\%/65.74\%$), confirming that class-center synthesis alone is not the missing mechanism. The useful channel is lifted local predictive geometry: regional predictive lifts reach $68.77\%/68.69\%/68.14\%$, precision lifts reach $68.33\%/68.45\%/67.88\%$, and the internal predictive-logit spline controller reaches $69.31\%$ ($69.12\%$ linear control). This is a real upstream improvement over the previous $67$--$68\%$ predictive-column family. However, it is redundant with the calibrated frontier. Context gating over these new candidates after the calibrated round-60 base keeps validation flat at $78.36\%$ and lowers test to $77.14\%$; strict resampled consolidation rejects all members, with the top candidate having full validation gain $-0.08$ points, lower-tail gain $-0.109$ points, $0\%$ win rate, and support $-4$ examples. An orthogonal source view $0$ to targets $1,2,3$ run (seed \texttt{20260711}) is weaker internally: regional/precision lifts top out at $67.40\%$ and the controller reaches $68.01\%$. It is also rejected after the calibrated base, and the combined source-$4$ plus source-$0$ candidate pool remains exactly flat at $78.36\%/77.15\%$.
The immediate follow-up tested whether the lifted predictive geometry could be moved earlier by feeding it back into the source state before readout. The new \texttt{--predictive\_feedback\_*} options fit local class-context maps from source, target, prediction, residual, and multiplicative agreement streams. The compact \texttt{lift} context projects those streams to a small random branch basis per region; the resulting regional class logits are averaged over selected target views and passed through the existing class-centroid dopamine displacement. This is still a closed-form no-backprop update, but it is not the missing mechanism. On the smoke split, all-view feedback with gain $0.25$ gives feedback source/local rows $51.83\%/52.25\%$ and a $50.58\%$ controller; gain $0.50$ gives $52.08\%/52.58\%$ and a $50.42\%$ controller; view-$5$-only routing gives $51.67\%/52.25\%$ and a $49.42\%$ controller. The full clean artifact \texttt{bio\_cifar\_physical\_predictive\_pretrain\_mps\_outputs\_cleanval\_predfeedback\_g050\_center100\_lifts16\_earlydop\_g050\_t045\_seed20260704/} confirms the negative: the original regional and precision lift rows reproduce exactly, but the feedback source reaches only $66.72\%$, the feedback local-context row reaches $67.48\%$, and the controller drops to $69.09\%$. The conclusion is therefore precise: discriminative lifted predictive targets are the strongest new upstream single-family result, but post-predictive class-centroid source displacement is too blunt and still too late. The next architecture must move the lifted predictive geometry earlier into source formation, for example by letting regional lift errors select patches, target views, source-view routing, or recurrent class-conditional target fields before the first physical column state and first class readout are formed.
We then moved the signal all the way into patch formation. The new \texttt{--predictive\_patch\_growth\_*} options implement a pilot synaptogenesis pass: the runner first builds the ordinary unsupervised patch bank, collects pilot source/target columns on the training core, solves the same closed-form local predictive maps, scores each image region by raw residual, discriminative mapped residual, or entropy-weighted innovation, samples new fixed patch filters from the high-score image cells, rebuilds the final physical columns with the augmented bank, and only then fits the final readouts. This remains forward-only; the pilot uses training-core residual fields to choose fixed filters, and validation/test images only pass through the resulting filter bank. On the source-$4$ target-$5,6,7$ smoke split, mapped-residual growth with $16$ filters per kernel/view improves several target rows but overfits the small RBF controller ($49.67\%$, linear control $53.33\%$), while entropy-weighted innovation is weaker ($50.08\%$, control $51.83\%$). The best smoke is mapped-residual growth with $32$ filters per kernel/view, top-$25\%$ residual sampling, and power $1.5$: predicted target rows reach $54.50\%/53.92\%/55.08\%$, the spline controller reaches $51.08\%$, and the linear control row reaches $53.50\%$. A stricter top-$15\%$, power-$2.0$ sampler regresses to a $50.50\%$ spline controller, so lower residual alone is again not the objective.
The full clean run is a useful split decision rather than a victory. The artifact \texttt{bio\_cifar\_physical\_predictive\_pretrain\_mps\_outputs\_cleanval\_predpatch\_mapped\_f32\_top025\_center100\_lifts16\_earlydop\_g050\_t045\_seed20260704/} uses the exact previous $45$k/$5$k/$10$k split and source/target configuration, but appends the residual-grown filters to the bank before final column formation. Early source geometry improves: raw source rises from $65.79\%$ to $66.21\%$, shaped source from $66.37\%$ to $66.48\%$, and settled source from $66.30\%$ to $66.93\%$. However, the strongest lifted predictive geometry weakens: regional lifts are $68.40\%/68.17\%/67.98\%$ instead of $68.77\%/68.69\%/68.14\%$, precision lifts are $68.19\%/67.81\%/67.80\%$ instead of $68.33\%/68.45\%/67.88\%$, and the internal controller drops to $69.07\%$ (linear control $68.69\%$) instead of $69.31\%$. A validation-aware logit fusion audit shows that the grown-bank columns are nevertheless orthogonal within the physical predictive family: the original discriminative-target logits alone give $68.58\%/69.42\%$ validation/test under the same RBF fusion audit, while original plus grown-patch logits give $71.88\%/71.95\%$. But this does not survive the global calibrated source bank. Feeding the four fused members into the round-60/round-43 calibrated context gate, either normalized or marked as already calibrated, remains exactly flat at $78.36\%/77.15\%$, with no active contexts for the strongest fused control members. The conclusion is architectural: predictive residual patch growth creates useful new source evidence, but appending those filters into the same bank perturbs the best lifted target geometry and remains redundant after the larger calibrated ensemble. The next serious version should keep baseline and residual-grown patches as parallel cortical populations with source-local routing before predictive maps, rather than replacing the baseline population by concatenating filters.
That follow-up is now a negative result. We extended \texttt{bio\_cifar\_physical\_predictive\_pretrain\_mps.py} with a \texttt{parallel} patch-growth mode, a separate \texttt{grown} or \texttt{augmented} physical population, and a no-backprop confidence route fitted from local class-context solves. The grown-only route mostly rejects the grown population: on the smoke split its mean gate is $0.341$ and only $1.7\%$ of image-regions prefer the grown branch. Routed predictive residuals worsen and the controller reaches only $47.75\%$. The augmented base+grown population is no better: mean gate $0.345$, grown-preferred fraction $1.4\%$, routed predictive rows about $49$--$51\%$, and controller $48.92\%$. A relaxed resample consolidation of the full physical source bank after the calibrated $78.36\%/77.15\%$ base also selects no physical additions; the best candidate has $\eta=0.002$, full validation gain $-0.040$ points, and zero resample win rate. Thus the $71.95\%$ family-fusion signal is late-logit diversity, not evidence for state interpolation by a class-confidence gate.
We then pushed the same smoke protocol across source-generation knobs. Seed diversity and width-only scaling do not help: seed \texttt{20260711} and branch width $128$ give best predicted-target rows $52.83\%$ and $52.75\%$, with controllers $47.50\%$ and $49.92\%$. Lower early dopamine gain improves target residuals but destroys class geometry, while stronger or weaker class-center targets also regress. A \texttt{class\_residual} target law gives very low physical residuals ($0.907/0.944/0.802$ for target views $5/6/7$) and a $52.67\%$ local-context row, but predicted-target class rows collapse to $47$--$48\%$; adding it as an auxiliary neuromodulator in the main class-center run still yields only a $46.33\%$ controller. The biological lesson is concrete: physically easy target prediction is not the same as class-aligned representation formation.
Finally, we repeated the cell-law ablation in this residual-growth setting. Conductance/reversal-potential cells are the only plausible single-family alternative, with $52.42\%$ source accuracy and $54.00\%$ best predicted-target accuracy, but their controller remains $48.58\%$. Hodgkin--Huxley-style algebraic gates overfit the local class context and trail at $48.08\%$ controller. Compact spline-window cells produce smoother target residuals but collapse discriminative geometry to about $32\%$ and a $20.50\%$ controller. The mixed pole bank remains the best matched transfer law for the current CIFAR descriptors. The next architecture should therefore not promote naive biological naming, width, confidence state routing, or residual reconstruction. It should generate stronger mixed-cell physical sources with pre-registered diversity and retain them by resampled logit-level consolidation, or replace the confidence gate by a marginal predictive-innovation route that is selected before class-logit compression.
We next made the synaptogenesis reward explicitly discriminative rather than reconstructive. The same runner now supports \texttt{--predictive\_patch\_growth\_score} values \texttt{class\_error}, \texttt{class\_margin}, \texttt{class\_error\_mapped\_residual}, and \texttt{class\_margin\_mapped\_residual}. In the pilot pass, a local class-context map is fitted from the source columns to the training labels by the same regional ridge solves used for early dopamine. For each image and cortical region, the \texttt{class\_error} score is $1-p_y$, where $p_y$ is the local probability assigned to the correct class; the margin score uses the best competing class against $p_y$. These scores sample new fixed filters from regions where the source representation is locally class-hard, before the final source columns, target maps, readouts, and controllers are fitted. This is a closer three-factor biological signal than raw residual reconstruction: presynaptic image patches define candidate synapses, the local class-context solve defines a postsynaptic error field, and the sampled patch bank changes the future source representation without reverse-mode differentiation.
The smoke results identify the correct objective. Repeating the source-$4$ target-$5,6,7$ split with $24$ filters per kernel/view, top-$20\%$ sampling, power $1.5$, early dopamine gain $0.50$, temperature $0.45$, \texttt{class\_center} target strength $1.0$, and lift dimension $16$, pure \texttt{class\_error} growth reaches a $54.75\%$ regional-lift row and a $52.33\%$ spline controller on seed \texttt{20260720}. The same setting on seed \texttt{20260710} reaches a $54.25\%$ precision-lift row and the same $52.33\%$ controller. Combining class hardness with mapped residual is worse: \texttt{class\_error\_mapped\_residual} falls to a $46.42\%$ controller, and \texttt{class\_margin\_mapped\_residual} reaches only $48.08\%$. A wider \texttt{class\_error} bank with $32$ filters and top-$25\%$ sampling also regresses to a $51.67\%$ controller. Thus the local reward field itself is useful, but multiplying it by target residual geometry reintroduces the wrong objective.
The full clean result is the new strongest upstream single-family CIFAR result in this branch. The artifact \texttt{bio\_cifar\_physical\_predictive\_pretrain\_mps\_outputs\_cleanval\_classhard\_classerror\_f24\_top020\_center100\_lifts16\_earlydop\_g050\_t045\_seed20260704\_exportctrl/} uses the canonical $45$k/$5$k/$10$k split, Apple MPS, source view $4$, targets $5,6,7$, $4\times4$ columns, four branches of width $96$, $128$ base patch filters, $24$ class-error-grown filters per kernel/view, prediction and class ridges $300$, \texttt{class\_center} strength $1.0$, regional lift dimension $16$, and an exported predictive-logit spline controller. It raises the raw/source/local rows to $66.67\%$, $67.24\%$, and $67.91\%$, versus $65.79\%$, $66.37\%$, and $67.42\%$ for the earlier discriminative-lift baseline. The best regional lift reaches $69.19\%$, precision lifts reach $68.81\%/68.76\%/68.40\%$, and the exported internal controller reaches $69.99\%$ with a $69.82\%$ linear control row. This beats the previous $69.31\%$ discriminative-lift controller and the $69.07\%$ residual-growth full run, while preserving the no-backprop protocol.
A stricter promotion of the predictive-feedback variant closes that branch. The full clean artifact \texttt{bio\_cifar\_physical\_predictive\_pretrain\_mps\_outputs\_cleanval\_classhard\_predfeedback\_g025\_f24\_top020\_seed20260704\_exportctrl/} uses the same $45$k/$5$k/$10$k split, source view $4$, targets $5,6,7$, class-error growth with $24$ filters and top-$20\%$ sampling, plus a conservative predictive-feedback gain $0.25$. It preserves the upstream class-hard rows but does not improve the controller: source/local rows are $67.24\%/67.91\%$, best regional lift is $69.19\%$, and the predictive-logit spline controller reaches $69.85\%$ with $69.33\%$ linear control, below the $69.99\%$ class-hard baseline. Strict resampled consolidation after the calibrated base rejects all exported predictive-feedback members; the best candidate gives only a $+0.020$ point full-validation bump, a $-0.073$ point lower-tail gain, $45.3\%$ resample win rate, and is not retained. Thus predictive feedback is currently redundant once the class-hard local reward field and calibrated source bank are present.
The calibrated frontier audit remains negative. Appending these $16$ exported class-error members after the round-43 inherited source bank and using the round-60 context member as a calibrated base gives \texttt{base\_selected=136} and candidates $120,\ldots,135$. The context gate finds only a tiny validation bump, $78.36\%\to78.40\%$, while held-out test falls from $77.15\%$ to $77.14\%$. Strict resampled consolidation rejects every addition; the best candidate has $\eta=0.020$, full validation gain $-0.020$ points, lower-tail gain $-0.145$ points, $32.8\%$ resample win rate, and support $-1$, so the final calibrated frontier remains $78.36\%/77.15\%$. The scientific conclusion is sharper than the score: local class-hardness is the first patch-growth objective that improves all upstream source and controller geometry on full CIFAR, but the high-70s global source bank is now saturated by similar errors. The next step should not be another residual selector; it should create source diversity around the class-hardness signal itself, for example distinct local reward heads, class-pair-specific hard-region filters, or validation-stable source families whose errors differ from the round-60 calibrated base.
We also reran the self-supervised JEPA/diffusion/attention audit with the same clean split and the current source bank. The patched \texttt{bio\_cifar\_ssl\_jepa\_attention\_fusion\_audit.py} now shuffles and validates the candidate source split consistently before checking source labels. On the $17$-member current-frontier bank with \texttt{latent\_dim=1024} and $768$ attention centers, the standalone closed-form heads remain weak: spline/pole random features $47.06\%$, JEPA view prediction $47.52\%$, diffusion denoising $47.58\%$, and low-margin attention $49.34\%$. The source-bank baselines are much stronger, with mean normalized logits $69.61\%$, control ridge $72.20\%$, and hard-state RBF $73.99\%$. Appending the SSL heads is flat or worse: hybrid mean $69.86\%$, reliability-gated hybrid $70.15\%$, hybrid control ridge $72.31\%$, and hybrid hard-state RBF $73.42\%$. Validation selects the baseline hard-state RBF, not an SSL-augmented source set. This is a direct negative control for late JEPA/diffusion/attention attachments. If these ideas help OSNR, they must become source-forming local objectives inside the physical columns, not shallow heads appended after class logits.
The first diversity attempt was deliberately class- and competitor-local. The new \texttt{--predictive\_patch\_growth\_class\_focus}, \texttt{--predictive\_patch\_growth\_competitor\_focus}, and \texttt{--predictive\_patch\_growth\_focus\_leak} options restrict the pilot class-hardness field to selected true labels or best-competing labels while retaining a small off-focus sampling leak. Narrow cat/dog focus is not a win. In append mode, a true-label focus on classes $3,5$ or a competitor focus on $3,5$ collapses the fair smoke split to about $49$--$50\%$ best rows before or at the controller, far below the global class-error smoke. Keeping the focused filters as a separate \texttt{parallel} grown population avoids replacing the base bank, but it still fails: with the fair $4\times96$ branch width, shuffled split, prediction ridge $300$, and settle step, the focused grown-only row reaches $45.75\%$, concatenating it with the base source reaches $52.83\%$, the best predicted-target row is $54.92\%$, and the controller falls to $50.42\%$. A broader animal-family focus on classes $2,\ldots,7$ with leak $0.20$ is also negative: best row $54.50\%$, controller $49.58\%$. The interpretation is that class-hardness should remain a dense regional reward field; hard class subsets are too sparse and perturb the random pole projection or add weak auxiliary populations.
We also tested two more biologically plausible follow-ups on top of the class-hard growth. First, predictive-geometry dopamine feedback was added after the class-hard maps. A conservative all-view feedback gain $0.25$ preserves the old best row ($54.75\%$) and nudges the smoke controller from $52.33\%$ to $52.42\%$, but this is only marginal. Entropy-gated feedback improves the feedback-local member to $54.17\%$ but lowers the controller to $52.00\%$, so the signal is member-level diversity rather than a stable source-family improvement. Second, an extra early class-feedback refit round increases local training confidence but overfits the field: with class-hard growth it gives a $53.83\%$ source row, only $54.17\%$ best regional lift, and a $50.92\%$ controller. These smokes close an important loop. The next serious route is not narrower hard-class masks, post-predictive global displacement, or more refit confidence. The higher-leverage target is a validation-selected source generator: mutate the source view, target views, pole family, branch width, growth objective, feedback gate, and consolidation rule together, then keep only source families that improve the calibrated frontier under resampled consolidation.
The next source-generator mini-matrix fixed the class-hard objective and mutated the retinal source/target graph. Local-contrast source $6$ to targets $4,5,7$ is a clear failure: despite nearly perfect local training context, held-out rows are only $42$--$44.5\%$ and the controller is $39.75\%$. DoG/opponent source $7$ to targets $4,5,6$ reaches only a $52.25\%$ target row and a $48.50\%$ controller. High-pass source $3$ to targets $4,5,7$ is closer, with a $54.67\%$ predicted-target row and a $52.00\%$ controller, but it remains below the source-$4$ baseline. Raw RGB source $0$ to targets $4,5,7$ is the only promoted candidate: smoke seed \texttt{20260733} reaches a $55.08\%$ predicted-target-$4$ row and a $53.25\%$ controller, and repeat seed \texttt{20260734} gives a $54.50\%$ regional row and a $53.17\%$ controller. The full clean promotion is competitive but not better than source $4$: raw/source/local rows are $66.47\%/66.64\%/67.61\%$, best regional lift $69.11\%$, and the exported controller $69.87\%$ with $69.59\%$ linear control, versus $69.99\%$ for the source-$4$ class-hard run. The calibrated round-60 audit is flat and strict resampling rejects every addition; top candidate $121$ has $\eta=0.004$, full validation gain $-0.060$ points, lower-tail gain $-0.109$ points, $1.0\%$ win rate, and support $-3$. Thus source $0$ is a real upstream variant but not a frontier-complementary family. The working source graph remains opponent/center-surround source $4$ with edge/contrast/DoG targets $5,6,7$; future generation must mutate more than the source view, for example local reward heads and pole/branch families jointly.
We then ran that joint direction as small controlled smokes, still on source $4\to5,6,7$. Conductance/reversal-potential cells under class-hard growth are worse than the mixed pole bank: best row $51.33\%$, controller $49.58\%$, despite perfect local training context. Increasing mixed capacity from $4\times96$ to $6\times80$ also regresses: best row $53.33\%$, controller $50.08\%$. Pure margin-based synaptogenesis is not the missing reward head: best row $53.67\%$, controller $50.08\%$. Changing class-error selectivity confirms the top-$20\%$ sampler. A sharper top-$10\%$ sampler reaches $54.58\%$ best row but only a $49.25\%$ controller, while top-$30\%$ falls to a $52.50\%$ best row and $49.67\%$ controller. These negative ablations are useful because they narrow the mechanism: the current winner is not simply more biological cell naming, more width, margin-only reward, or arbitrary hard-region sparsity; it is the specific combination of mixed poles, opponent source geometry, class-error regional reward, and moderate top-$20\%$ patch growth.
As a final consolidation check, we made the PGPE logit-architecture search calibrated-aware. Without this correction, the PGPE loader normalized the already calibrated round-60 base and artificially lowered it from $78.36\%$ to $78.12\%$ validation, producing a misleading $77.29\%$ test result at lower validation. The updated runner accepts \texttt{--calibrated\_logits\_npz} and bypasses per-member normalization for those sources, matching the context-gate protocol. Running PGPE over the calibrated round-60 base plus the source-$4$ and source-$0$ class-hard banks with $180$ steps, population $48$, top-$12$, validation-half stability penalty, and an $8192$-sample core reward anchor selects only the base: final validation/test remain $78.36\%/77.15\%$. Thus even population-level nonnegative source weighting does not rescue these candidates once calibration is handled correctly.
The next MPS smokes close the current class-hard patch-growth family. Keeping the grown class-error filters as a separate parallel physical population does not solve the overwrite problem: grown-only reaches $47.33\%$, base-plus-grown concatenation reaches $50.92\%$, the best predicted-target row is $54.92\%$, and the controller is $49.92\%$. A no-backprop confidence route fitted from local class-context solves mostly rejects the grown population; only $1.7\%$ of regions prefer it, and the routed controller falls to $47.75\%$. The sampling-sharpness sweep is also negative. Flattening the class-error sampling power to $1.0$ gives a $54.25\%$ best row but only a $44.08\%$ controller; sharpening to $2.25$ gives $53.08\%$ best and a $45.08\%$ controller. A broad source-$4$ predictive stack to all non-source retinal views reaches only $54.17\%$ best and a $52.25\%$ controller, below the selected target-$5,6,7$ graph. Finally, we tested an explicit forward-only contrastive goodness profile: for each target map, the local state compares the predicted target against the true target and several rolled negative targets. The compact contrastive score is stable but not frontier-moving ($54.25\%$ best contrastive score), while the high-dimensional contrastive residual state is weaker ($53.58\%$ best). The best overall row in that run remains the old regional lift at $54.75\%$, with controller $52.25\%$ and linear control $55.00\%$. The conclusion is that class-hardness top-$20\%$ on source $4\to5,6,7$ is a local optimum for this substrate; the next attempt must create a different source-forming mechanism, not another patch-growth or late routing variant.
We then tested a more explicit cellular architecture in \texttt{apps\_industrial\_breakthrough/bio\_columnar\_predictive\_control\_benchmark.py}. The model has a retinal/V1 front end, $49$ L1 cortical columns with $64$ cable cells each, $16$ L2 association columns with $128$ cells each, a $4096$-cell global field, optional thalamic sensory skip cells, four dendritic branches, and four cable modes per branch. Each branch has stable leak/synapse/diffusion poles, conductance gates, reversal potentials, lateral inhibition, block-local covariance readouts, fusion, and a dopamine-like residual controller. This is closer to the proposed biological architecture than the single global projection, but the first full FashionMNIST run is a negative result: the full-field columnar profile reaches $90.74\%$ on the $60{,}000/10{,}000$ protocol, below the simpler streaming covariance field at $92.32\%$. Matching the streaming ridge and fp32 cache lowers it further to $90.04\%$. The diagnosis is useful. Explicit dendritic geometry alone does not solve credit assignment; the current columnar stack discards or overcompresses class-separable sensory evidence before the closed-form controller. A serious next cellular model must learn or select intermediate predictive targets locally, not merely route fixed cable states into a larger final covariance solve.
The ablations matter. Prototype voting retains useful memory but is weaker than the linear center-state solve. Diagonal Gaussian statistics are too crude. The dense kernel memory is pathological at high capacity, collapsing to about $10$--$12\%$ final accuracy despite excellent early-task performance; the failure is a conditioning/credit-allocation warning against treating every stored center as a dense global kernel atom. The class-subspace attractor reaches only $75.93\%$ and fixed-budget multi-head attention fusion reaches $79.32\%$ on canonical FashionMNIST, so attention-style splitting is not automatically useful without a stable local credit field. The architecture search in \texttt{apps\_industrial\_breakthrough/bio\_plasticity\_architecture\_search.py} adds the knobs the biological thesis actually needs---state degree, local projection depth, neuron count, dendrite count, ridge, and fusion temperature. Degree-$2$ lifts and extra random projection depth increase memory without improving Fashion test accuracy; five-dendrite variants tie validation but cost substantially more memory. The selected three-dendrite row is therefore the current best tradeoff. The working mechanism is more specific: project cellular evidence into independent pole-rich dendritic states, accumulate local eligibility covariances and dopamine cross-covariances, and fuse the resulting quadratic controllers. This is closer to biological plasticity than replayed global-gradient updates, because old information persists as local co-activity statistics rather than as raw examples or weights repeatedly overwritten by backpropagation.
\subsection{Application validation: 2D tensor-product fluid dynamics}
The first application-level validation script, \texttt{apps/01\_fluid\_dynamics/run\_vortex\_street.py}, extends the Hermite trunk from separate 1D passes to a true 2D tensor-product coefficient system. Each grid node stores nine Hermite streams corresponding to value, first derivatives, second derivatives, and mixed derivative channels. The continuous stream function $a_z$ defines the incompressible velocity field by the analytical curl
\[
v_x = \partial_y a_z,
\qquad
v_y = -\partial_x a_z.
\]
Consequently, incompressibility is structural rather than imposed by a penalty. Solid cylinder and wall masks overwrite all nine coefficient streams to zero at masked vertices, enforcing no-slip and zero-flux constraints by coefficient assignment.
The 2D Hermite Gram is assembled as a tensor product of the 1D cross-correlation filters. Applying \texttt{torch.fft.fft2} diagonalizes the spatial part of the block-circulant system, reducing the global solve to independent $9\times9$ complex systems at each frequency coordinate $(\nu_y,\nu_x)$. In the current synthetic unrolled vortex-street validation, the script runs $36$ frames on a $32\times48$ grid with a mean processing duration of $1.2850$ ms per frame. Boundary leakage remains $0.000000\mathrm{e}{+00}$ and the coefficient clamp residual remains $0.000000\mathrm{e}{+00}$ through the unroll. The final momentum residual is $1.876066$ in the script's synthetic nondimensional units.
\subsection{Visual validation: multi-obstacle CFD cinema}
The breakthrough visual script, \texttt{apps\_breakthrough/fluid\_vortex\_cinema.py}, scales the same tensor-product Hermite fluid engine to a $64\times256$ canvas with a multi-obstacle mask inspired by the ``Smiley Face / HI!'' geometry used in the Spline-PINN visual demonstrations. The mask combines disk, capsule, and rectangular primitives to produce a dense nonconvex obstacle field. The solid set includes the obstacle geometry and the domain walls. At every time step, all nine Hermite coefficient streams are overwritten to zero on this set, so no-slip and zero-flux constraints enter as direct coefficient assignments rather than differentiable penalties.
The state variable remains a scalar stream function $a_z$. Velocities are recovered by the analytical curl
\[
(v_x,v_y)=(\partial_y a_z,-\partial_x a_z),
\]
which structurally removes the need for an incompressibility loss. The unrolled update evaluates a synthetic transport-diffusion step
\[
a_z^{n+1}
=
a_z^n
+
\Delta t\left(
\nu \Delta a_z^n
-\eta\,(v^n\cdot\nabla)a_z^n
-\kappa\,\omega^n
+f^n
\right),
\]
where $\omega=\partial_x v_y-\partial_y v_x$ is the vorticity field and $f^n$ is a time-dependent wake forcing. The updated scalar field is mapped into the nine Hermite streams, clamped on the solid set, passed through the block-circulant Hermite Gram, and recovered by the parallel $9\times9$ Fourier solver. A conservative amplitude limiter is applied to keep the synthetic visualization stable over the full $100$-frame unroll; this limiter is a numerical display stabilizer, not a replacement for calibrated Navier--Stokes time integration.
The run exports every fifth step as a high-contrast PNG visualization of speed and signed vorticity. In the verified local run, it produced $20$ frames, maintained boundary leakage $0.000000\mathrm{e}{+00}$ and coefficient clamp residual $0.000000\mathrm{e}{+00}$, and completed with mean latency $13.6396$ ms per frame. The final synthetic momentum residual was $8.212812$ in the script's nondimensional units. The entire run executes under \texttt{torch.no\_grad()} with $0.00$ B autograd graph allocation.
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{@{}p{0.14\linewidth}p{0.1\linewidth}p{0.12\linewidth}p{0.16\linewidth}p{0.17\linewidth}p{0.15\linewidth}p{0.1\linewidth}@{}}
\toprule
Grid & Frames & PNG exports & Mean latency & Boundary leakage & Clamp residual & Autograd \\
\midrule
$64\times256$ & $100$ & $20$ & $13.6396$ ms/frame & $0.000000\mathrm{e}{+00}$ & $0.000000\mathrm{e}{+00}$ & $0.00$ B \\
\bottomrule
\end{tabular}
\caption{Measured execution ledger for the multi-obstacle CFD cinema validation.}
\label{tab:cfd-cinema}
\end{table}
\begin{figure}[h]
\centering
\setlength{\tabcolsep}{2pt}
\begin{tabular}{@{}ccccc@{}}
\includegraphics[width=0.19\linewidth]{figures/frame_000.png} &
\includegraphics[width=0.19\linewidth]{figures/frame_005.png} &
\includegraphics[width=0.19\linewidth]{figures/frame_010.png} &
\includegraphics[width=0.19\linewidth]{figures/frame_015.png} &
\includegraphics[width=0.19\linewidth]{figures/frame_020.png} \\
\scriptsize frame 000 & \scriptsize frame 005 & \scriptsize frame 010 & \scriptsize frame 015 & \scriptsize frame 020 \\
\includegraphics[width=0.19\linewidth]{figures/frame_025.png} &
\includegraphics[width=0.19\linewidth]{figures/frame_030.png} &
\includegraphics[width=0.19\linewidth]{figures/frame_035.png} &
\includegraphics[width=0.19\linewidth]{figures/frame_040.png} &
\includegraphics[width=0.19\linewidth]{figures/frame_045.png} \\
\scriptsize frame 025 & \scriptsize frame 030 & \scriptsize frame 035 & \scriptsize frame 040 & \scriptsize frame 045 \\
\includegraphics[width=0.19\linewidth]{figures/frame_050.png} &
\includegraphics[width=0.19\linewidth]{figures/frame_055.png} &
\includegraphics[width=0.19\linewidth]{figures/frame_060.png} &
\includegraphics[width=0.19\linewidth]{figures/frame_065.png} &
\includegraphics[width=0.19\linewidth]{figures/frame_070.png} \\
\scriptsize frame 050 & \scriptsize frame 055 & \scriptsize frame 060 & \scriptsize frame 065 & \scriptsize frame 070 \\
\includegraphics[width=0.19\linewidth]{figures/frame_075.png} &
\includegraphics[width=0.19\linewidth]{figures/frame_080.png} &
\includegraphics[width=0.19\linewidth]{figures/frame_085.png} &
\includegraphics[width=0.19\linewidth]{figures/frame_090.png} &
\includegraphics[width=0.19\linewidth]{figures/frame_095.png} \\
\scriptsize frame 075 & \scriptsize frame 080 & \scriptsize frame 085 & \scriptsize frame 090 & \scriptsize frame 095
\end{tabular}
\caption{Full exported speed-vorticity rendering sequence from the OSNR multi-obstacle CFD cinema run. Dark regions are hard-clamped obstacle and wall coefficients; color encodes velocity magnitude modulated by signed vorticity.}
\label{fig:cfd-cinema}
\end{figure}
\subsection{Industrial visual validation: sparse graphics super-resolution}
The graphics super-resolution script, \texttt{apps\_industrial\_breakthrough/graphics\_superres\_engine.py}, applies the Tier 2 adaptive sparse core to a high-density geometric rendering problem. The target is a synthetic industrial graphics asset with high-frequency directional contours. Each horizontal scanline is modeled as a finite-rate-of-innovation signal with six discontinuity locations, corresponding to three filled geometric bands. The raw comparison image is produced by evaluating the same asset on a coarse uniform grid and expanding it to the display canvas, which exposes block aliasing at the sub-pixel boundaries.
The OSNR path passes the scanline moments into the TLS matrix-pencil tracker, recovers the fractional transition coordinates, and snaps the sparse step dictionary to those coordinates before reconstruction. The cross-Gram-shielded ADMM sieve suppresses the empty background and uniform interior atoms, while the scale-invariant ridge debiasing pass stabilizes the active discontinuity support. The final continuous field is evaluated on a $512\times512$ canvas and exported as a side-by-side PNG: coarse block-aliased rendering on the left, FRI-snapped OSNR reconstruction on the right.
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{@{}p{0.13\linewidth}p{0.18\linewidth}p{0.13\linewidth}p{0.1\linewidth}p{0.13\linewidth}p{0.15\linewidth}p{0.1\linewidth}@{}}
\toprule
Canvas & Solver layout & Duration & PSNR & Sparsity & Max edge error & Autograd \\
\midrule
$512\times512$ & $B=256$, $N=256$, $K=6$ & $110.5804$ ms & $101.61$ dB & $96.2\%$ & $1.674321\mathrm{e}{-08}$ & $0.00$ B \\
\bottomrule
\end{tabular}
\caption{Measured execution ledger for the adaptive sparse graphics super-resolution validation.}
\label{tab:graphics-superres}
\end{table}
\begin{figure}[h]
\centering
\includegraphics[width=0.95\linewidth]{figures/osnr_graphics_superres.png}
\caption{Graphics super-resolution output. Left: block-aliased uniform-grid rendering. Right: OSNR reconstruction after TLS FRI edge localization and sparse knot snapping.}
\label{fig:graphics-superres}
\end{figure}
\subsection{Industrial aerodynamic validation: high-Reynolds wind tunnel}
The aerodynamic wind-tunnel script, \texttt{apps\_industrial\_breakthrough/aerodynamic\_wind\_tunnel.py}, extends the tensor-product Hermite fluid path to a high-Reynolds engineering surrogate. The obstacle is a multi-element NACA 0012-style body composed of a main airfoil, slat, and deflected flap. The geometry is rasterized into a curved solid mask on a $64\times192$ wind-tunnel grid. The simulated regime uses $\operatorname{Re}=50{,}000$ with reference velocity $U=0.24$, chord $c=0.72$, and kinematic viscosity $\nu=3.45600000\mathrm{e}{-06}$.
As in the CFD cinema experiment, the state variable is a scalar stream function $a_z$ and the velocity field is recovered through
\[
v_x=\partial_y a_z,\qquad v_y=-\partial_x a_z.
\]
This curl parameterization enforces incompressibility structurally. The wall and airfoil masks overwrite all nine tensor-product Hermite coefficient channels to zero at each step, imposing no-slip and zero-flux constraints by assignment. The update evaluates advection, viscous diffusion, and a high-frequency wake forcing through forward finite-difference ladders, then applies the 2D block-circulant Hermite Fourier solve. Every tenth time step is exported as a raw state matrix containing velocity magnitude, vorticity, the solid mask, Reynolds number, and boundary leakage.
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{@{}p{0.12\linewidth}p{0.12\linewidth}p{0.1\linewidth}p{0.14\linewidth}p{0.17\linewidth}p{0.15\linewidth}p{0.1\linewidth}@{}}
\toprule
Grid & Reynolds & Frames & Mean latency & Boundary leakage & Clamp residual & Exports \\
\midrule
$64\times192$ & $50{,}000$ & $50$ & $8.6717$ ms/frame & $0.000000\mathrm{e}{+00}$ & $0.000000\mathrm{e}{+00}$ & $5$ \\
\bottomrule
\end{tabular}
\caption{Measured execution ledger for the high-Reynolds multi-element airfoil wind-tunnel validation. The exported state matrices are stored under \texttt{apps\_industrial\_breakthrough/wind\_tunnel\_states/}.}
\label{tab:wind-tunnel}
\end{table}
\subsection{Industrial benchmark ingestion: Hugging Face video challenger}
The Hugging Face challenger script, \texttt{apps\_industrial\_breakthrough/huggingface\_sota\_challenger.py}, is the first repository path that ingests an external hosted video asset rather than a manufactured field. The script uses the official \texttt{datasets} library and Hugging Face Hub APIs to inspect video metadata, resolves local or HTTPS media references when they are available, and decodes real video containers with a prioritized backend chain: \texttt{decord}, then PyAV, then \texttt{imageio-ffmpeg}. The production execution profile targets $T=30$ frames at $256\times256$ RGB resolution. The public \texttt{APRIL-AIGC/UltraVideo} rows currently expose metadata and YouTube identifiers rather than direct \texttt{.mp4} payloads, so the script requires \texttt{OSNR\_HF\_VIDEO\_FILE} for a local or HTTPS UltraVideo media export. If no decodable media file is provided, it records this condition explicitly and falls back to a real Hugging Face video fixture so the decoding, algebraic compression, and metric path remains executable.
Each RGB scanline is processed as a composite sparse-plus-smooth color track. Before moments are formed, the decoded tensor is passed through a localized separable cubic B-spline prefilter with kernel $[1,4,6,4,1]/16$ along both image axes. This shift-invariant smoothing step suppresses quantization and compression perturbations that otherwise dominate the algebraic roots. Gradient-selected transitions are then de-duplicated by non-maximum suppression, converted into moments, and routed into the packaged \texttt{AdaptiveSparseSolver}. For this noisy-video path, the solver's FRI tracker is replaced by a ridge-regularized TLS matrix pencil: the denoised Hankel coordinate equation is solved as
\[
(H_0^\top H_0+\gamma I)Z=H_0^\top H_1,
\qquad
\gamma=10^{-6}\operatorname{mean}(\operatorname{diag}(H_0^\top H_0)).
\]
This prevents near-null Hankel directions from snapping knots to compression artifacts. The remaining sparse recovery uses the cross-Gram-shielded ADMM sieve with ridge debiasing to suppress inactive background atoms. Because natural video is not a pure step-edge signal, the repaired pipeline adds a pruned smooth residual tier after sparse recovery. The residual is projected onto an orthonormal DCT row dictionary, ridge solved, and hard-pruned to retain only the largest coefficients. This implements the continuous-domain composite model
\[
f(x)=f_{\mathrm{sparse}}(x;\tau_k)+f_{\mathrm{smooth}}(x),
\]
with FRI atoms representing geometry and low-frequency DCT atoms representing illumination, texture, and compression residuals.
\paragraph{Implemented video algorithm.}
The current video path is not a neural training loop. It is a deterministic composite inversion pipeline whose components are tied to the spline theory above. For a decoded video tensor
\[
Y\in[0,1]^{T\times H\times W\times 3},
\qquad T=30,\quad H=W=256,
\]
the implementation proceeds as follows.
\begin{enumerate}[leftmargin=*,itemsep=2pt]
\item \textbf{Decode and normalize.} Load consecutive frames through the prioritized decoder chain \texttt{decord}/PyAV/\texttt{imageio-ffmpeg}; resize to $256\times256$ and normalize RGB values to $[0,1]$.
\item \textbf{Spline prefilter.} Apply the separable cubic B-spline smoothing kernel
\[
b=\frac{1}{16}[1,4,6,4,1]
\]
along $x$ and $y$. This produces a denoised tensor $\widetilde{Y}$ used only for edge moment estimation, not for final metric evaluation.
\item \textbf{Scanline FRI moments.} For every time, row, and color channel, flatten the horizontal trace into a one-dimensional signal $y_{t,h,c}(x)$. Select $K=32$ non-maximum-suppressed gradient transitions and convert them into innovation moments
\[
m_\ell=\sum_{k=1}^{K} a_k\tau_k^\ell,\qquad \ell=0,\ldots,2K+1.
\]
\item \textbf{Ridge TLS matrix pencil.} Build Hankel pairs $(H_0,H_1)$ from the moments, project the concatenated pencil to rank $K$, and solve the stabilized shift equation
\[
(H_0^\top H_0+\gamma I)Z=H_0^\top H_1,
\qquad
\gamma=10^{-6}\operatorname{mean}(\operatorname{diag}(H_0^\top H_0)).
\]
The eigenvalues of $Z$ give the snapped sub-pixel step coordinates $\tau_k$.
\item \textbf{Cross-Gram sparse solve.} Form a smooth sinusoidal dictionary $\A_s$ and a snapped step dictionary $\A_x(\tau_k)$. The ADMM block updates use the cross-Gram shield $\A_s^\top\A_x$ exactly as in the hybrid normal equations:
\[
(\A_s^\top\A_s)\cvec_s=\A_s^\top y-\A_s^\top\A_x z,
\]
\[
(\A_x^\top\A_x+\rho I)\cvec_x
=
\A_x^\top y-\A_x^\top\A_s\cvec_s+\rho z-u.
\]
This is the oblique projection step that prevents smooth illumination atoms and sparse step atoms from absorbing each other's energy.
\item \textbf{Scale-invariant debiasing.} Debias the active FRI atoms with
\[
(\A^\top\A+\epsilon\,\overline{d}\,I)c=\A^\top y,
\qquad
\overline{d}=\operatorname{mean}(\operatorname{diag}(\A^\top\A)),
\quad \epsilon=10^{-6}.
\]
The sparse tensor is stored in a $1024$-slot accounting dictionary; only the FRI-snapped active atoms are solved, while the remaining slots are explicit hard zeros.
\item \textbf{Frame-wise 2D-DCT residual.} Convert the sparse scanline prediction back to a video tensor $\widehat{Y}_{\mathrm{sparse}}$. For each frame and color channel, represent the residual
\[
R_{t,c}=Y_{t,:,:,c}-\widehat{Y}_{\mathrm{sparse},t,:,:,c}
\]
in an orthonormal two-dimensional DCT basis
\[
R_{t,c}(i,j)
\approx
\sum_{p=0}^{P_y-1}\sum_{q=0}^{P_x-1}
d_{t,c,p,q}\,\psi_p(i)\psi_q(j),
\qquad P_y=P_x=256.
\]
Coefficients are computed by the separable projection
\[
D_{t,c}=\Psi_y^\top R_{t,c}\Psi_x.
\]
The practical profile keeps the $32768$ largest coefficients per frame/channel; the ceiling profile keeps all $65536$ coefficients.
\item \textbf{Composite synthesis and export.} The final reconstruction is
\[
\widehat{Y}
=
\widehat{Y}_{\mathrm{sparse}}
+
\Psi_y D \Psi_x^\top,
\]
clipped to $[0,1]$. The script exports target and OSNR frames for both Pareto profiles and evaluates PSNR, SSIM, LPIPS, sparsity, latency, and memory.
\end{enumerate}
This algorithm explains the main empirical observation. The row-DCT variant had no vertical basis functions and therefore generated visible scanline ripple. The 2D-DCT tier restores a true image-plane smooth residual space, eliminating that artifact when enough coefficients are retained. The price is that the ceiling profile becomes a dense transform-codec upper bound rather than a sparse representation claim.
The experimental record is cumulative. We keep the earlier small high-quality run because it provides a reconstructable baseline for the quality ceiling of the current sparse-plus-DCT path: $T=2$ frames at $32\times32$ RGB resolution, $K=8$ transitions per scanline, one B-spline smoothing pass, matrix-pencil ridge scale $10^{-6}$, and a $32$-term DCT residual tier pruned to $30$ coefficients per row. That configuration improved PSNR from the original sparse-only $18.5182$ dB and the ridge-prefiltered $21.1817$ dB result to $60.2304$ dB, with SSIM $0.999686$, LPIPS $0.000004$, $80.21\%$ combined hard-zero parameters, and $95.00\%$ sparse-tier hard-zero parameters.
The overhauled production-profile run used the UltraVideo metadata row as the ingestion target; because that row exposed the non-decodable identifier \texttt{BJRpaBau\_QI}, the script decoded the Hugging Face video fixture while retaining the requested $T=30$, $256\times256$ RGB tensor layout and $K=32$ transitions per scanline. It used two B-spline smoothing passes, matrix-pencil ridge scale $10^{-6}$, a $1024$-slot sparse accounting dictionary, and a $128$-term DCT residual candidate dictionary. The live ADMM solve uses only the FRI-snapped active knot bank; the remaining sparse slots are retained as explicit hard-zero background atoms, avoiding a wasteful $B\times1024\times1024$ Gram expansion. With the strict sparsity setting of $22$ retained DCT coefficients per scanline, the representation reaches $29.1223$ dB PSNR, SSIM $0.805561$, and LPIPS $0.243320$ with $95.31\%$ combined hard-zero parameters. With a quality-prioritized setting of $120$ retained DCT coefficients per scanline, the same decoded tensor reaches $35.0321$ dB PSNR, SSIM $0.958564$, and LPIPS $0.033300$ while retaining $86.81\%$ combined hard-zero parameters and $96.88\%$ sparse-tier hard-zero parameters. The latter is the better production direction because image fidelity is the decisive benchmark; the strict sparsity profile is retained as an ablation, not as the preferred operating point.
Visual inspection of the row-DCT reconstructions revealed coherent vertical ripple artifacts. This is a structural artifact of treating each row independently: the smooth residual tier has no vertical coupling, so natural two-dimensional texture is forced into separable scanline corrections. We therefore added a frame-wise two-dimensional DCT residual tier and unrolled it across the complete $T=30$ decoded sequence. The script exports both the target and OSNR reconstruction at every time step to \texttt{apps\_industrial\_breakthrough/ultravideo\_cinema/}, with separate \texttt{practical\_sparse} and \texttt{quality\_ceiling} directories. The multi-frame run produced $120$ PNG frames: target and reconstruction pairs for both profiles over $t=0,\ldots,29$.
With $32768$ retained 2D coefficients per frame/channel, the full cinema profile reaches $41.3874$ dB and LPIPS $0.021302$, and the visible ripple is largely suppressed. With the full $256\times256$ 2D residual basis retained, the current algebraic framework reaches its quality ceiling across the full sequence: $142.7187$ dB PSNR, SSIM $1.000000$, and LPIPS $0.000000$. This ceiling run is not presented as a compression result; it is an upper-bound diagnostic proving that the sparse FRI geometry plus 2D smooth residual path can reproduce the decoded video exactly when quality is unconstrained. The next meaningful engineering target is therefore the intermediate regime between $32768$ and $65536$ retained 2D residual coefficients, or a more structured perceptual residual dictionary that concentrates the same visual quality into fewer active parameters.
\begin{table}[H]
\centering
\scriptsize
\begin{tabular}{@{}p{0.22\linewidth}p{0.13\linewidth}p{0.1\linewidth}p{0.1\linewidth}p{0.1\linewidth}p{0.13\linewidth}p{0.12\linewidth}@{}}
\toprule
Configuration & Tensor & PSNR & SSIM & LPIPS & Sparsity & Latency \\
\midrule
Small high-quality profile & $T=2$, $32^2$ RGB, $K=8$ & $60.2304$ dB & $0.999686$ & $0.000004$ & $80.21\%$ & $26.4591$ ms/frame \\
Full strict-sparsity profile & $T=30$, $256^2$ RGB, $K=32$ & $29.1223$ dB & $0.805561$ & $0.243320$ & $95.31\%$ & $167.0024$ ms/frame \\
Full quality-prioritized profile & $T=30$, $256^2$ RGB, $K=32$ & $35.0321$ dB & $0.958564$ & $0.033300$ & $86.81\%$ & $167.8295$ ms/frame \\
2D residual practical cinema & $T=30$, $256^2$ RGB, $K=32$ & $41.3874$ dB & $0.950734$ & $0.021302$ & $87.50\%$ & $170.7790$ ms/frame \\
2D residual quality ceiling cinema & $T=30$, $256^2$ RGB, $K=32$ & $142.7187$ dB & $1.000000$ & $0.000000$ & $77.50\%$ & $169.8067$ ms/frame \\
\bottomrule
\end{tabular}
\caption{Cumulative Hugging Face video ingestion benchmark ledger using the B-spline-prefiltered, ridge-regularized, sparse-plus-DCT OSNR codec path. LPIPS is computed with the official \texttt{torchmetrics} AlexNet-backed implementation. Rows are retained as experiment memory rather than overwritten by later ablations.}
\label{tab:hf-challenger}
\end{table}
\begin{figure}[H]
\centering
\begin{tabular}{@{}cccc@{}}
\includegraphics[width=0.22\linewidth]{figures/hf_challenger_quality_frame0_target.png} &
\includegraphics[width=0.22\linewidth]{figures/hf_challenger_quality_frame0_osnr.png} &
\includegraphics[width=0.22\linewidth]{figures/hf_challenger_full256_dct2_ceiling_frame0_target.png} &
\includegraphics[width=0.22\linewidth]{figures/hf_challenger_full256_dct2_ceiling_frame0_osnr.png} \\
\scriptsize small target & \scriptsize small OSNR, $60.23$ dB &
\scriptsize full target & \scriptsize 2D ceiling OSNR, $142.72$ dB
\end{tabular}
\caption{First-frame visual comparisons for the cumulative Hugging Face video challenger ledger. The earlier small high-quality profile is preserved as a reconstructable historical result, while the full $256^2$ 2D-DCT ceiling profile records the current maximum image quality of the algebraic sparse-plus-smooth path.}
\label{fig:hf-challenger}
\end{figure}
\clearpage
\subsection{Frame-by-frame UltraVideo cinema comparison}
Table~\ref{tab:ultravideo-cinema-frames} gives the direct visual audit requested for the multi-frame cinema export. Each row shows the decoded target frame, the full-quality OSNR 2D-DCT ceiling reconstruction, and the practical sparse reconstruction with the same FRI geometry tier but a pruned 2D residual budget. The comparison is intentionally image-first: the ceiling column records the maximum quality attainable by the current algebraic sparse-plus-smooth representation, while the practical column records the visible cost of residual pruning.
\begingroup
\scriptsize
\setlength{\tabcolsep}{2pt}
\renewcommand{\arraystretch}{1.05}
\begin{longtable}{@{}c c c c@{}}
\caption{Frame-by-frame visual comparison for the $T=30$ UltraVideo cinema export. The three image columns are, respectively, the decoded target, the quality-ceiling OSNR reconstruction, and the practical sparse OSNR reconstruction.}
\label{tab:ultravideo-cinema-frames}\\
\toprule
Frame & Target & OSNR quality ceiling & OSNR practical sparse \\
\midrule
\endfirsthead
\toprule
Frame & Target & OSNR quality ceiling & OSNR practical sparse \\
\midrule
\endhead
000 &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/target_frame_000.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/osnr_frame_000.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/practical_sparse/osnr_frame_000.png} \\
001 &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/target_frame_001.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/osnr_frame_001.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/practical_sparse/osnr_frame_001.png} \\
002 &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/target_frame_002.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/osnr_frame_002.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/practical_sparse/osnr_frame_002.png} \\
003 &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/target_frame_003.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/osnr_frame_003.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/practical_sparse/osnr_frame_003.png} \\
004 &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/target_frame_004.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/osnr_frame_004.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/practical_sparse/osnr_frame_004.png} \\
005 &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/target_frame_005.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/osnr_frame_005.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/practical_sparse/osnr_frame_005.png} \\
006 &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/target_frame_006.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/osnr_frame_006.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/practical_sparse/osnr_frame_006.png} \\
007 &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/target_frame_007.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/osnr_frame_007.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/practical_sparse/osnr_frame_007.png} \\
008 &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/target_frame_008.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/osnr_frame_008.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/practical_sparse/osnr_frame_008.png} \\
009 &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/target_frame_009.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/osnr_frame_009.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/practical_sparse/osnr_frame_009.png} \\
010 &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/target_frame_010.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/osnr_frame_010.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/practical_sparse/osnr_frame_010.png} \\
011 &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/target_frame_011.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/osnr_frame_011.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/practical_sparse/osnr_frame_011.png} \\
012 &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/target_frame_012.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/osnr_frame_012.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/practical_sparse/osnr_frame_012.png} \\
013 &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/target_frame_013.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/osnr_frame_013.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/practical_sparse/osnr_frame_013.png} \\
014 &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/target_frame_014.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/osnr_frame_014.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/practical_sparse/osnr_frame_014.png} \\
015 &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/target_frame_015.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/osnr_frame_015.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/practical_sparse/osnr_frame_015.png} \\
016 &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/target_frame_016.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/osnr_frame_016.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/practical_sparse/osnr_frame_016.png} \\
017 &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/target_frame_017.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/osnr_frame_017.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/practical_sparse/osnr_frame_017.png} \\
018 &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/target_frame_018.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/osnr_frame_018.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/practical_sparse/osnr_frame_018.png} \\
019 &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/target_frame_019.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/osnr_frame_019.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/practical_sparse/osnr_frame_019.png} \\
020 &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/target_frame_020.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/osnr_frame_020.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/practical_sparse/osnr_frame_020.png} \\
021 &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/target_frame_021.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/osnr_frame_021.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/practical_sparse/osnr_frame_021.png} \\
022 &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/target_frame_022.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/osnr_frame_022.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/practical_sparse/osnr_frame_022.png} \\
023 &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/target_frame_023.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/osnr_frame_023.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/practical_sparse/osnr_frame_023.png} \\
024 &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/target_frame_024.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/osnr_frame_024.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/practical_sparse/osnr_frame_024.png} \\
025 &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/target_frame_025.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/osnr_frame_025.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/practical_sparse/osnr_frame_025.png} \\
026 &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/target_frame_026.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/osnr_frame_026.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/practical_sparse/osnr_frame_026.png} \\
027 &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/target_frame_027.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/osnr_frame_027.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/practical_sparse/osnr_frame_027.png} \\
028 &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/target_frame_028.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/osnr_frame_028.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/practical_sparse/osnr_frame_028.png} \\
029 &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/target_frame_029.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/quality_ceiling/osnr_frame_029.png} &
\includegraphics[width=0.27\linewidth]{../apps_industrial_breakthrough/ultravideo_cinema/practical_sparse/osnr_frame_029.png} \\
\bottomrule
\end{longtable}
\endgroup
\subsection{Multi-view DL3DV scene reconstruction runner}
The next industrial runner, \texttt{apps\_industrial\_breakthrough/dl3dv\_sota\_challenger.py}, lifts the video pipeline from a regular $x,y,t$ tensor to a multi-view scene tensor governed by camera rays. The runner is designed for the gated \texttt{DL3DV/DL3DV-Benchmark} repository. It deliberately refuses blind full-dataset downloads and requires either a local scene directory via \texttt{OSNR\_DL3DV\_SCENE\_DIR} or an authenticated single-scene Hugging Face prefix via \texttt{OSNR\_DL3DV\_SCENE\_PREFIX}. This is necessary because the public benchmark repository is multi-terabyte scale and requires acceptance of dataset access conditions.
For a selected scene, the runner parses \texttt{transforms.json}, resolves the first $30$ frame images, downsamples them to $256\times256$, and constructs a target tensor
\[
Y\in[0,1]^{30\times256\times256\times3}.
\]
For each view $k$, pixel coordinates are mapped to continuous camera rays by the usual NeRF/COLMAP transformation
\[
\mathbf{x}_{k}(s;i,j)
=
\mathbf{o}_{k}+s\,\mathbf{d}_{k}(i,j),
\]
where $\mathbf{o}_{k}$ is the camera center from the camera-to-world matrix and $\mathbf{d}_{k}$ is obtained by applying the camera rotation to the normalized intrinsic-coordinate direction
\[
\left((i-c_x)/f_x,\;-(j-c_y)/f_y,\;1\right).
\]
The current algebraic solve then uses the same composite partition as the UltraVideo runner: a ridge-regularized TLS matrix pencil estimates scanline discontinuity coordinates, the FRI step dictionary is snapped to those coordinates, and the cross-Gram ADMM shield prevents the smooth sinusoidal tier from absorbing discontinuity energy. The residual is lifted from a frame-wise 2D DCT to a view-volume 3D DCT,
\[
R(v,y,x,c)
\approx
\sum_{p,q,r}
d_{c,p,q,r}\,\psi_p(v)\psi_q(y)\psi_r(x),
\]
with an optional 3D FFT ridge conditioner
\[
\widehat{R}_{\mathrm{cond}}(\omega)
=
\frac{\widehat{R}(\omega)}{1+\gamma}.
\]
This FFT division is the implemented block-circulant identity/ridge solve for the current prototype; a full physically coupled 3D radiance operator remains future work.
The gated live run was executed against the locally cached official scene prefix \texttt{0853979305f7ecb80bd8fc2c8df916410d471ef04ed5f1a64e9651baa41d7695}. The selected path contains the DL3DV \texttt{nerfstudio} layout with \texttt{transforms.json} and RGB source frames. The runner resolved the first image as \texttt{nerfstudio/images/frame\_00001.png}, parsed the camera intrinsics and camera-to-world matrices, constructed per-pixel ray origins and directions of shape $30\times256\times256\times3$, and reconstructed the observed multi-view image stack. This is not yet a novel-view renderer and should not be read as a physical $x,y,z$ radiance-volume solve; the current 3D residual axes are view index, image row, and image column. The experiment is therefore a memory-bounded multi-view stack reconstruction with camera-ray metadata, establishing the ingestion and algebraic reconstruction path before the future physically coupled radiance-field step.
The initial DL3DV ledger measured two profiles. The first retained only $56$ three-dimensional DCT residual coefficients per color channel. It produced $20.4800$ dB PSNR, SSIM $0.448128$, LPIPS $0.718625$, $96.92\%$ combined hard-zero parameters, $96.88\%$ sparse-tier hard-zero parameters, $193{,}302{,}032$ measured peak bytes, and $140.1238$ ms/view. This sparse run is useful as a stress test, but not as the preferred visual-quality profile. The second profile prioritized image quality by retaining the complete $30\times256\times256$ orthonormal DCT support per color channel. It produced $117.2378$ dB PSNR, SSIM $0.99999988$, LPIPS $1.0539\mathrm{e}{-10}$, $77.50\%$ combined hard-zero parameters, $96.88\%$ sparse-tier hard-zero parameters, $215{,}813{,}648$ measured peak bytes, and $135$--$136$ ms/view across repeated runs. Dense NumPy tensor export remained disabled; only metrics and PNG frames were written.
The first failed quality-profile attempt exposed a real implementation bottleneck: the four-operand 3D-DCT projection \texttt{einsum} was killed externally during contraction planning/execution despite the preflight estimate. The corrected implementation now computes the separable DCT projection and synthesis axis-by-axis: $x$ projection, $y$ projection, view projection, followed by view, $y$, and $x$ synthesis. This keeps the DCT stage inside the same memory envelope and makes the quality-profile run reproducible on the local CPU. The TLS matrix-pencil amplitude solve also gained an absolute ridge floor for degenerate rows, so blank or nearly flat scanlines no longer produce singular Gram failures.
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{@{}p{0.22\linewidth}p{0.16\linewidth}p{0.1\linewidth}p{0.1\linewidth}p{0.1\linewidth}p{0.14\linewidth}p{0.13\linewidth}@{}}
\toprule
Configuration & Residual support & PSNR & SSIM & LPIPS & Sparsity & Peak memory \\
\midrule
Sparse residual profile & $30\times64\times64$, $56$ kept/channel & $20.4800$ dB & $0.448128$ & $0.718625$ & $96.92\%$ & $193.30$ MB \\
Quality ceiling profile & $30\times256\times256$, all kept/channel & $117.2378$ dB & $1.000000$ & $0.000000$ & $77.50\%$ & $215.81$ MB \\
\bottomrule
\end{tabular}
\caption{DL3DV multi-view stack reconstruction ledger for scene \texttt{0853979305f7ecb80bd8fc2c8df916410d471ef04ed5f1a64e9651baa41d7695}. Both profiles run under \texttt{torch.no\_grad()} with a $0$ byte autograd graph and dense NPZ tensor export disabled. The quality ceiling is an image-fidelity upper bound, not a sparsity claim.}
\label{tab:dl3dv-stack}
\end{table}
After establishing the quality ceiling, the follow-up Pareto sweep in \texttt{apps\_industrial\_breakthrough/dl3dv\_pareto\_sweep.py} re-ran the same scene while monotonically pruning the 3D-DCT residual support. Each profile used the same $30\times256\times256$ input tensor, $K=32$ sparse knots, $5$ LPIPS views, \texttt{torch.no\_grad()}, disabled dense NPZ export, and the same $2$ GiB preflight guardrail. The sweep deliberately keeps the sparse tier fixed, so the measured curve isolates the visual effect of residual support pruning rather than conflating it with a new boundary locator.
The measured curve is more conservative than the optimistic pre-run hypothesis. The $50\%$ DCT-support profile remains high fidelity at $42.7562$ dB, SSIM $0.981283$, LPIPS $0.002584$, and $87.50\%$ combined hard-zero parameters. The $25\%$ profile reaches $92.50\%$ hard-zero parameters but falls to $35.5073$ dB, and the $10\%$ profile reaches $95.50\%$ hard-zero parameters but falls to $30.4278$ dB. The edge-localization error remains at $0.958244$ for every sweep point, confirming that the present sparse tier is not yet carrying enough of the geometric boundary load; additional gains should come from improving the TLS/Hankel conditioning and boundary model rather than from further blind DCT pruning.
\paragraph{Compression sanity check.}
The same DL3DV quality-ceiling result also motivates a direct compression audit, because an exact orthonormal residual expansion is not automatically a competitive codec. The script \texttt{apps\_industrial\_breakthrough/osnr\_compression\_audit.py} therefore takes the same $30\times256\times256$ RGB target stack, whose raw unsigned-byte footprint is $5{,}898{,}240$ bytes, and measures payload size after scalar quantization and \texttt{np.savez\_compressed} entropy compression. Two OSNR-style transform payloads are tested: a global separable 3D DCT over view, row, and column axes, and an independent per-frame 2D DCT. Both use low-frequency support masks and quantized integer coefficients. The audit then decodes the stored coefficients and evaluates PSNR/SSIM against JPEG and WebP encodings at quality $90$ using the same source images.
The result is intentionally conservative and negative. At similar bits per pixel, the naive OSNR transform payloads are far below mature image codecs: the $5\%$ 3D-DCT profile reaches only $25.2894$ dB at $1.7118$ bpp, and the $5\%$ per-frame 2D-DCT profile reaches $26.4132$ dB at $1.8927$ bpp, while WebP reaches $39.4283$ dB at $1.8132$ bpp. Increasing OSNR support restores quality but destroys payload efficiency: the $50\%$ per-frame 2D-DCT profile reaches $37.4358$ dB, but costs $12.8751$ bpp. This confirms that the $117$ dB observed-stack ceiling is a completeness result, not a compression claim. A serious OSNR codec would need at least perceptual quantization, coefficient ordering, block or geometry-conditioned prediction, motion/view compensation, and a real entropy coder before it should be compared against JPEG, WebP, AV1, or neural codecs. For the present manuscript, compression is therefore recorded as a promising but unfinished direction rather than the next flagship validation target.
A more appropriate compression target is neural-scene distillation rather than still-image coding. The audit \texttt{apps\_industrial\_breakthrough/neural\_scene\_distillation\_audit.py} tests this narrower claim using the frozen nerfstudio even/odd split. It serializes only the fifteen training key views into quantized OSNR 2D-DCT coefficient packages, decodes those key views, and predicts the held-out views by the same deterministic adjacent-view interpolation rule. This is a lightweight scene-streaming proxy: the payload is a compact mathematical scene package rather than a trained radiance field, and the metric is held-out view quality per transmitted byte.
The first result is a foothold, not a SOTA win. At $5\%$ support and $q=0.004$, the OSNR keyview package is $234{,}555$ bytes and reaches $22.3882$ dB held-out PSNR, slightly above the WebP-keyview stream at $255{,}070$ bytes and $21.9900$ dB. An even smaller $2\%$ OSNR package is only $92{,}228$ bytes and still reaches $21.7984$ dB. However, the locally available \texttt{nerfacto} CPU pilot checkpoint is $242{,}859{,}619$ bytes and reaches $27.0830$ dB on the same odd-view protocol. Therefore the current OSNR package is dramatically smaller, but not yet quality-competitive with even a reduced NeRF pilot. The next scene-compression experiment would need a true geometry-aware residual package--for example plane-sweep depth support, sparse COLMAP anchors, and view-dependent residual coefficients--before claiming neural-field model compression.
The follow-up audit \texttt{apps\_industrial\_breakthrough/neural\_scene\_geometry\_package\_audit.py} performs exactly this geometry-aware test. For each package profile, it first decodes the transmitted key views and then runs the deterministic COLMAP-pose plane-sweep renderer from those decoded key views into the held-out cameras. The package byte count includes the serialized keyview payload plus a conservative $17{,}394{,}552$ byte COLMAP camera metadata budget from \texttt{cameras.bin} and \texttt{images.bin}. With uncompressed key views, the geometry package reaches $29.3178$ dB in $20.34$ MB. More importantly, compressed keyview packages still beat the local \texttt{nerfacto} CPU pilot: WebP key views plus geometry reach $29.0865$ dB in $17.62$ MB, and the OSNR $50\%$ keyview package reaches $28.9853$ dB in $18.98$ MB. Compared with the $242.86$ MB \texttt{nerfacto} CPU checkpoint at $27.0830$ dB, this is a concrete local model-compression win: better held-out PSNR with roughly $12$--$14\times$ smaller serialized scene state. The limitation is equally clear. The current OSNR keyview transform is not yet the best keyview codec inside the geometry package--WebP remains slightly better at lower payload--so the next OSNR-specific compression gain must come from geometry-conditioned residual coefficients or a more mature entropy-coded spline payload rather than from naive per-frame DCT pruning alone.
The residual-package audit \texttt{apps\_industrial\_breakthrough/neural\_scene\_residual\_package\_audit.py} tests that next hypothesis directly. It computes leave-one-out plane-sweep residuals on the even training views, projects those residuals onto quantized spatial OSNR DCT packets, interpolates the decoded residual packets to the odd held-out views, and adds them to the held-out geometry render. This is deliberately quality-first: it tests $50\%$ and $25\%$ residual support and does not force extreme sparsity. The result is negative. For raw key views, the geometry-only package remains best at $29.3178$ dB; adding the best residual packet falls to $29.1727$ dB. For WebP key views, geometry-only reaches $29.0865$ dB, while the best residual packet falls to $28.9533$ dB. The residual stream therefore encodes view-specific plane-sweep errors that do not transfer cleanly from even leave-one-out views to odd held-out views. The practical conclusion is that the current quality bottleneck is not residual coefficient capacity; it is visibility/depth correctness. The scene package should next improve geometry--depth maps, occlusion masks, or multi-source visibility confidence--before adding larger residual payloads.
The spline inverse-problem bridge \texttt{apps\_industrial\_breakthrough/dl3dv\_tomographic\_radiance\_bridge.py} then tests whether the McCann--Donati $H^\top H$ convolution idea can already help the real DL3DV held-out split. The script recolors the $81{,}120$ sparse COLMAP points from even training views, deposits them into a $64^3$ compact spline voxel grid, applies a cubic-B-spline FFT normal solve with ridge $0.005$, and renders the regularized radiance grid into the odd held-out cameras. The result is a small but measurable PSNR foothold rather than a finished renderer. Adjacent-view interpolation reaches $22.7803$ dB, SSIM $0.611987$, and LPIPS $0.150600$. The best FFT-tomographic blend uses only $2\%$ of the regularized grid prediction and reaches $22.7955$ dB, but SSIM falls to $0.609815$ and LPIPS rises to $0.155779$. Larger blends degrade quickly. This confirms that the inverse grid contains some held-out radiance signal, while the dominant problem remains visibility-aware measurement construction and occlusion reasoning rather than the speed of the FFT normal solve itself.
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{@{}p{0.34\linewidth}p{0.13\linewidth}p{0.13\linewidth}p{0.13\linewidth}p{0.16\linewidth}@{}}
\toprule
Payload profile & Bytes & bpp & PSNR & SSIM \\
\midrule
OSNR 3D-DCT, $5\%$, $q=0.004$ & $420{,}688$ & $1.7118$ & $25.2894$ dB & $0.950010$ \\
OSNR 2D-DCT, $5\%$, $q=0.004$ & $465{,}151$ & $1.8927$ & $26.4132$ dB & $0.962129$ \\
OSNR 2D-DCT, $50\%$, $q=0.004$ & $3{,}164{,}180$ & $12.8751$ & $37.4358$ dB & $0.997102$ \\
JPEG, quality $90$ & $584{,}387$ & $2.3779$ & $38.0274$ dB & $0.997522$ \\
WebP, quality $90$ & $445{,}608$ & $1.8132$ & $39.4283$ dB & $0.998208$ \\
\bottomrule
\end{tabular}
\caption{Compression sanity check on the same $30$-view DL3DV target stack. The audit measures actual serialized payload bytes after quantization and compression. The observed-stack OSNR reconstruction remains a completeness result; these naive transform payloads are not yet competitive with mature codecs.}
\label{tab:osnr-compression-audit}
\end{table}
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{@{}p{0.36\linewidth}p{0.13\linewidth}p{0.13\linewidth}p{0.13\linewidth}p{0.15\linewidth}@{}}
\toprule
Scene package profile & Bytes & bpp/eval & Held-out PSNR & Held-out SSIM \\
\midrule
Uncompressed key views + interpolation & $2{,}949{,}120$ & $24.0000$ & $21.9836$ dB & $0.889469$ \\
JPEG key views, quality $90$ & $322{,}203$ & $2.6221$ & $21.9584$ dB & $0.888935$ \\
WebP key views, quality $90$ & $255{,}070$ & $2.0758$ & $21.9900$ dB & $0.889673$ \\
OSNR keyview DCT, $5\%$, $q=0.004$ & $234{,}555$ & $1.9088$ & $22.3882$ dB & $0.900308$ \\
OSNR keyview DCT, $2\%$, $q=0.008$ & $92{,}228$ & $0.7506$ & $21.7984$ dB & $0.884918$ \\
\texttt{nerfacto} CPU pilot checkpoint & $242{,}859{,}619$ & n/a & $27.0830$ dB & $0.818526$ \\
\bottomrule
\end{tabular}
\caption{First neural-scene distillation audit on the frozen DL3DV even/odd split. The OSNR keyview package is smaller than JPEG/WebP keyview streams at comparable held-out interpolation quality, but it does not yet match the trained \texttt{nerfacto} pilot's PSNR. This supports scene-streaming potential, not a completed neural-field compression result.}
\label{tab:neural-scene-distillation-audit}
\end{table}
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{@{}p{0.34\linewidth}p{0.13\linewidth}p{0.13\linewidth}p{0.12\linewidth}p{0.12\linewidth}p{0.12\linewidth}@{}}
\toprule
Geometry package profile & Key payload & Total package & Held-out PSNR & Held-out SSIM & LPIPS \\
\midrule
Raw key views + plane sweep & $2{,}949{,}120$ & $20{,}343{,}672$ & $29.3178$ dB & $0.896597$ & $0.102406$ \\
WebP key views + plane sweep & $223{,}258$ & $17{,}617{,}810$ & $29.0865$ dB & $0.889464$ & $0.112057$ \\
JPEG key views + plane sweep & $292{,}277$ & $17{,}686{,}829$ & $29.0429$ dB & $0.888549$ & $0.110024$ \\
OSNR keyview DCT, $50\%$, $q=0.004$ + plane sweep & $1{,}582{,}874$ & $18{,}977{,}426$ & $28.9853$ dB & $0.888095$ & $0.122029$ \\
OSNR keyview DCT, $25\%$, $q=0.004$ + plane sweep & $920{,}438$ & $18{,}314{,}990$ & $28.2040$ dB & $0.862453$ & $0.219444$ \\
\texttt{nerfacto} CPU pilot checkpoint & n/a & $242{,}859{,}619$ & $27.0830$ dB & $0.818526$ & $0.176857$ \\
\bottomrule
\end{tabular}
\caption{Geometry-aware neural-scene package audit. Each compressed-keyview row is decoded before rendering; the deterministic plane-sweep renderer then predicts the odd held-out views from the decoded even views. Under this local CPU-pilot comparison, compact geometry packages are both smaller and higher-PSNR than the available \texttt{nerfacto} checkpoint, while the OSNR-specific keyview transform still trails WebP inside the package.}
\label{tab:neural-scene-geometry-package-audit}
\end{table}
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{@{}p{0.43\linewidth}p{0.14\linewidth}p{0.13\linewidth}p{0.13\linewidth}p{0.12\linewidth}@{}}
\toprule
Residual package profile & Total package & Held-out PSNR & Held-out SSIM & LPIPS \\
\midrule
Raw key views + geometry only & $20{,}343{,}672$ & $29.3178$ dB & $0.896597$ & $0.102406$ \\
Raw key views + best OSNR residual packet & $21{,}444{,}470$ & $29.1727$ dB & $0.893367$ & $0.103023$ \\
WebP key views + geometry only & $17{,}617{,}810$ & $29.0865$ dB & $0.889464$ & $0.112057$ \\
WebP key views + best OSNR residual packet & $18{,}734{,}634$ & $28.9533$ dB & $0.886479$ & $0.111948$ \\
OSNR keyview DCT + geometry only & $18{,}977{,}426$ & $28.9853$ dB & $0.888095$ & $0.122029$ \\
OSNR keyview DCT + best OSNR residual packet & $20{,}079{,}081$ & $28.8556$ dB & $0.884912$ & $0.122517$ \\
\bottomrule
\end{tabular}
\caption{Geometry-conditioned residual package audit. Residual coefficients are fitted only from even-view leave-one-out geometry errors and then evaluated on odd held-out views. The negative result is informative: residual capacity does not solve the current error mode; visibility and depth correctness dominate.}
\label{tab:neural-scene-residual-package-audit}
\end{table}
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{@{}p{0.38\linewidth}p{0.13\linewidth}p{0.13\linewidth}p{0.13\linewidth}p{0.12\linewidth}@{}}
\toprule
DL3DV held-out profile & PSNR & SSIM & LPIPS & Coverage \\
\midrule
Adjacent-view linear interpolation & $22.7803$ dB & $0.611987$ & $0.150600$ & $100.00\%$ \\
Raw recolored COLMAP splat & $21.2616$ dB & $0.537217$ & $0.425973$ & $79.68\%$ \\
FFT-tomographic spline grid, $1\%$ blend & $22.7913$ dB & $0.611576$ & $0.152229$ & $88.34\%$ \\
FFT-tomographic spline grid, $2\%$ blend & $22.7955$ dB & $0.609815$ & $0.155779$ & $88.34\%$ \\
FFT-tomographic spline grid, $4\%$ blend & $22.7841$ dB & $0.602852$ & $0.169208$ & $88.34\%$ \\
\bottomrule
\end{tabular}
\caption{First real-scene spline-tomographic radiance bridge on the frozen DL3DV even/odd split. The $64^3$ FFT-normal grid solve takes $17.2450$ ms; rendering the unoptimized voxel splats takes $848.5788$ ms for the held-out stack. The slight PSNR gain at tiny blend weights is useful evidence, but not yet a NeRF-quality renderer.}
\label{tab:dl3dv-tomographic-radiance-bridge}
\end{table}
The next implementation step adds an explicit world-edge proposal layer without modifying the packaged \texttt{AdaptiveSparseSolver}. The module \texttt{apps\_industrial\_breakthrough/dl3dv\_world\_edge\_atoms.py} computes cubic-B-spline derivative magnitudes in each view, selects non-maximum-suppressed gradient pixels, backprojects them through the parsed camera intrinsics and camera-to-world matrices, triangulates adjacent-view ray pairs by closest point of approach, clusters the resulting candidate points, and reprojects those world atoms back into each camera. The DL3DV runner exposes this through \texttt{--edge\_atom\_mode scanline|world|hybrid}; the hybrid mode merges projected world-edge knots with the original scanline knots while preserving the same no-autograd, RAM-guarded execution path.
This first world-edge ablation is intentionally diagnostic rather than presented as an improvement. With $128$ edge rays per view, $1024$ clustered world atoms, CPA threshold $0.025$, cluster radius $0.015$, and ridge $10^{-6}$, the hybrid projection uses $24.03\%$ projected knots and lowers the scanline-referenced edge delta from $0.958244$ to $0.624169$. However, the multi-view reprojection error is still $13.5410$ pixels, so the injected atoms are not yet selective enough: at $25\%$ DCT support, PSNR falls from $35.5073$ dB to $33.6019$ dB; at $10\%$ support, PSNR falls from $30.4278$ dB to $29.2703$ dB. A stricter CPA run with threshold $0.005$ and cluster radius $0.005$ gives similar quality ($33.6076$ dB at $25\%$ support) and worse reprojection error ($14.3163$ pixels). The conclusion is precise: the world-space atom path is now executable and measurable, but adjacent-view CPA alone must be augmented with epipolar-consistency scoring, depth/COLMAP support, or multi-view consensus pruning before it can replace the DCT cushion.
To separate observed-view reconstruction from genuine view generalization, \texttt{apps\_industrial\_breakthrough/dl3dv\_heldout\_challenger.py} implements an interleaved held-out light-field challenge on the same scene. Even-indexed views $\{0,2,\ldots,28\}$ are the only training/input images; odd-indexed views $\{1,3,\ldots,29\}$ are held out for metrics. The runner compares three deterministic, no-autograd profiles: adjacent-view linear interpolation, DCT interpolation along the camera sequence, and a ray-kernel OSNR model that fits RGB from sampled training rays $(\mathbf{o},\mathbf{d})$ and evaluates the held-out camera rays directly. This experiment is quality-first and does not impose sparsity pruning on the ray model.
The held-out result is a useful boundary marker rather than a new SOTA claim. Linear neighbor interpolation reaches $22.7803$ dB PSNR, SSIM $0.611987$, and LPIPS $0.150600$ on the fifteen held-out views. DCT view interpolation reaches $21.2918$ dB, SSIM $0.534591$, and LPIPS $0.141890$. The first ray-kernel OSNR profile, using $131{,}072$ sampled training rays and $1024$ RBF centers, reaches only $18.0553$ dB, SSIM $0.466200$, and LPIPS $0.904455$; increasing to $262{,}144$ samples and $4096$ centers with a broader kernel worsens PSNR to $14.0241$ dB. This confirms that the earlier $117$ dB DL3DV ceiling is an observed-stack completeness result, not yet a NeRF-style novel-view synthesis result. The next graphics step therefore needs depth-aware or epipolar-consensus geometry rather than a larger ray-only kernel.
The asset audit then found a usable sparse COLMAP reconstruction: \texttt{nerfstudio/colmap/sparse/0/points3D.bin} contains $81{,}120$ points, with matching \texttt{images.bin} and \texttt{cameras.bin}. No depth maps, NumPy geometry arrays, or Gaussian-splat \texttt{.ply} file are cached. The follow-up renderer \texttt{apps\_industrial\_breakthrough/dl3dv\_colmap\_heldout\_renderer.py} therefore tests a deterministic geometry-backed held-out baseline: parse the COLMAP points and image poses, recolor visible points from even training views, z-buffer splat them into the odd held-out cameras, and blend the sparse render with the linear-neighbor fallback. This is still not a trained NeRF or dense 3DGS renderer, but it is the first held-out result in this section that uses actual scene geometry.
The geometry-backed profile beats the interpolation-only baseline. The first refined run used COLMAP poses, training-view recolored points, radius-$1$ splats, and a $45\%$ geometry blend, reaching $23.6510$ dB PSNR. Pushing quality further showed that the limiting artifact is high-frequency splat noise: applying a $3\times3$ smoothing kernel to the sparse geometry before blending raises the held-out score to $23.9984$ dB PSNR and SSIM $0.684186$, compared with $22.7803$ dB and SSIM $0.611987$ for linear neighbor interpolation. The next visibility-aware recoloring pass resolves each training view with a deterministic nearest-depth test before sampling point colors; this removes occluded color assignments and raises the peak to $24.1063$ dB PSNR and SSIM $0.686748$ at a $67\%$ geometry blend. Adaptive surfel splats increased projected coverage from $65.82\%$ to $75.87\%$--$82.73\%$, but did not improve PSNR because the additional coverage carried too much color and visibility noise. Finally, \texttt{apps\_industrial\_breakthrough/dl3dv\_osnr\_residual\_renderer.py} fits residuals only on even training views after subtracting the visibility-clean geometry anchor, then evaluates the residual correction on odd held-out views. A small linear residual correction raises the peak to $24.2145$ dB PSNR, while higher-capacity DCT residuals underperform because they begin to inject view-dependent residual noise.
The stronger Track-A result is the dense, deterministic plane-sweep renderer \texttt{apps\_industrial\_breakthrough/dl3dv\_plane\_sweep\_renderer.py}. It uses the same COLMAP camera convention as the successful sparse renderer and warps the nearest even training views into each odd held-out camera over a bounded depth lattice. A diagnostic depth pass over the COLMAP points gives average visible-depth quantiles $q_{50}\approx 6.31$, $q_{90}\approx 12.06$, and $q_{95}\approx 13.68$, explaining why the $0.4$--$12.0$ range outperforms the initial $0.4$--$8.0$ sweep. The first dense pass, with $64$ depth planes, two source views, and a $5\%$ residual correction, reached $26.3525$ dB PSNR. The v2 pass then replaces hard winner-take-all depth selection with soft plane aggregation, introduces patch-averaged photometric costs, and removes residual correction once it becomes detrimental. This staged refinement raises held-out quality to $29.1731$ dB PSNR at $192$ depths. The v3 ablation shows that confidence-fused multi-pair sources, bilateral edge-aware cost aggregation, and local depth refinement do not improve PSNR on this scene; the best path is still nearest two-view soft aggregation with a uniform patch cost and more depth support. A final high-depth push raises the deterministic ceiling to $29.4187$ dB PSNR, SSIM $0.898331$, and LPIPS $0.099257$ using $512$ linear depth planes, a $13\times13$ patch cost, two source views, and no neural training or autograd. The gain over $320$ planes is measurable but small relative to the added compute, so the $320$-plane profile remains the practical operating point while the $512$-plane profile records the quality ceiling of this deterministic renderer.
The bridge from the controlled multi-ray FRI studies to real COLMAP imagery is \texttt{apps\_industrial\_breakthrough/dl3dv\_multiview\_fri\_geometry\_bridge.py}. The runner keeps the same even/odd held-out split, but augments the depth-lattice score with cubic-B-spline derivative edge coherence from the even training views. At each candidate depth, RGB disagreement between warped source views is combined with a source-edge disagreement term and a small coherent-edge reward; this is a real-image analogue of the controlled multi-ray residual selection loop, but still avoids target-view leakage. A diagnostic no-edge four-source setting reaches only $27.5992$ dB at $160^2$ resolution, while the nearest-two-source edge-coherent setting reaches $30.1786$ dB at $160^2$. At the standard $256^2$ DL3DV size, the $256$-plane edge bridge reaches $29.8085$ dB PSNR, SSIM $0.913584$, and LPIPS $0.070140$ in $21.481$ s for the fifteen held-out views. This exceeds the earlier deterministic $512$-plane photometric ceiling in all three perceptual metrics while using half the number of depth planes, and it also improves over the trained liquid-residual polishing row. The edge bridge is therefore the strongest current geometry result: it is autograd-free, uses only observed training views and camera geometry, and demonstrates that sparse edge evidence transfers from the controlled Haouchat setting into real multi-view reconstruction.
The liquid-neural-network follow-up deliberately tests a different question: whether a very small continuous-time residual corrector can remove systematic plane-sweep artifacts without replacing the deterministic geometry engine. The runner \texttt{apps\_industrial\_breakthrough/dl3dv\_liquid\_residual\_challenger.py} first computes leave-one-out plane-sweep renders on the even training views, fits only the residual image with a tiny exact liquid cell, and evaluates the learned residual on the odd held-out views. Each hidden unit follows the closed-form multi-synapse liquid update
\[
x_{k+1}
=
\gamma_k\left(x_k-\frac{\sum_s A_s f_s(\mathbf{c}_k)}{\omega+\sum_s f_s(\mathbf{c}_k)}\right)
+
\frac{\sum_s A_s f_s(\mathbf{c}_k)}{\omega+\sum_s f_s(\mathbf{c}_k)},
\qquad
\gamma_k=\exp\!\left(-\omega-\sum_s f_s(\mathbf{c}_k)\right),
\]
where the conditioning vector $\mathbf{c}_k$ contains the plane-sweep RGB estimate, the linear-view fallback, confidence, normalized pixel coordinates, normalized view index, and the plane-sweep/fallback discrepancy. This hybrid no longer has the zero-autograd property during fitting, but the learned component is intentionally small and residual-only. At $320$ depth planes it improves the held-out result from $29.3178$ to $29.3994$ dB and reduces LPIPS from $0.102406$ to $0.096436$. At the $512$-plane quality setting, a $64$-state liquid residual improves the plane-sweep ceiling from $29.4187$ to $29.5291$ dB, SSIM from $0.898331$ to $0.901776$, and LPIPS from $0.099257$ to $0.093192$. The gain is consistent but modest; it supports liquid residuals as a polishing layer, not as a substitute for denser visibility-aware geometry.
To make the SOTA comparison falsifiable rather than rhetorical, \texttt{apps\_industrial\_breakthrough/dl3dv\_external\_baseline\_protocol.py} freezes an external baseline protocol for the same scene, view count, resolution, and even/odd split. The exporter writes a \texttt{nerfstudio\_even\_odd\_256} dataset with $30$ resized frames whose basenames explicitly contain \texttt{train} or \texttt{eval}, \texttt{transforms.json}, \texttt{transforms\_train.json}, \texttt{transforms\_eval.json}, and \texttt{split.json}, plus explicit \texttt{nerfacto}, \texttt{splatfacto}, and \texttt{ns-eval} command lines. The regenerated protocol report also audits the local device stack. The machine itself supports Apple Metal: direct shell testing shows that the project \texttt{venv} can allocate \texttt{device='mps'} tensors. However, the current \texttt{.venv\_nerfstudio} Python 3.10 environment reports \texttt{mps=False}, and fresh Homebrew Python 3.11/3.13/3.14 torch-2.12 test environments failed the same MPS runtime gate in this execution context. A targeted follow-up repinned the Python 3.10 nerfstudio environment to torch-2.5.1/torchvision-0.20.1 and then to torch-2.3.1/torchvision-0.18.1; both variants still failed \texttt{torch.ones(1, device='mps')} with the same PyTorch OS-version gate. Python 3.14 cannot be used directly for nerfstudio because Open3D has no compatible wheel. A local CPU-only \texttt{nerfacto} pilot with $1000$ iterations and $1024$ rays per batch reaches $27.0830$ dB PSNR, SSIM $0.818526$, and LPIPS $0.176857$ on the odd held-out views. This remains only a protocol-validation point. The current comparison target for the external run is no longer the older plane-sweep row but the multi-view FRI edge bridge: $29.8085$ dB PSNR, SSIM $0.913584$, and LPIPS $0.070140$. A final trained-NeRF/3DGS comparison therefore requires either a CUDA-capable \texttt{nerfacto}/\texttt{splatfacto} run or a local nerfstudio environment whose PyTorch installation is first verified to allocate MPS tensors.
The consolidation script \texttt{apps\_industrial\_breakthrough/dl3dv\_heldout\_sota\_ledger.py} collects the scattered held-out outputs into one ranked ledger. This makes the current competitive status unambiguous. The naive ray-kernel OSNR row is not competitive, reaching only $18.0553$ dB, so the project cannot claim that coordinate-ray regression alone beats NeRF. The geometry-aware rows are different: the deterministic edge-consistent bridge reaches $29.8085$ dB without neural training, the deterministic $512$-plane sweep reaches $29.4187$ dB, and the tiny residual liquid polishing layer reaches $29.5291$ dB. These exceed the available \texttt{nerfacto} CPU pilot at $27.0830$ dB on the identical split, but the comparison remains a local pilot until a full GPU \texttt{nerfacto}/\texttt{splatfacto} run is executed. The ledger therefore defines the next hard target: retain the held-out quality advantage while cutting the plane-sweep latency and replacing the external CPU pilot with a complete CUDA baseline.
The adaptive-depth follow-up \texttt{apps\_industrial\_breakthrough/dl3dv\_adaptive\_depth\_bridge.py} tests whether this latency can be reduced by replacing the uniform depth lattice with a two-stage proposal scheme. The renderer first runs a coarse edge-aware bridge, then evaluates local per-pixel depth offsets around the selected depth and optional sparse COLMAP point-depth proposals. The first single-depth adaptive profile reduces the held-out-stack latency but loses quality: a $64+17+9$ proposal profile reaches $28.7523$ dB in $6830.4$ ms, while a denser $128+17$ profile reaches $28.8875$ dB in $10598.5$ ms. A top-$K$ proposal-recall upgrade is stronger. Keeping the best three coarse depth hypotheses per pixel and refining each with seven local offsets reaches $29.0585$ dB at $64$ coarse depths, $29.3291$ dB at $128$ coarse depths, and $29.5615$ dB at $192$ coarse depths. The $192$-depth top-$K$ profile is faster than the uniform bridge ($15882.1$ ms versus $21481.0$ ms) and improves perceptual metrics (SSIM $0.918645$, LPIPS $0.064826$), but it still trails the uniform bridge in PSNR. We also tested a sparse COLMAP visibility-consistency penalty that rejects candidates landing behind the source view's nearest sparse point depth. Even a weak penalty (\texttt{visibility\_weight=0.02}, \texttt{visibility\_eps=0.12}) reduces the top-$K$ profile to $29.3382$ dB and increases latency to $21878.8$ ms. Thus sparse COLMAP depth is useful as a proposal hint but too noisy as a direct occlusion veto. Since the adaptive residual variant again reduces PSNR, the dominant error is not a missing smooth residual; it is missed depth/visibility proposal quality. The next DL3DV improvement must therefore use stronger epipolar source-edge intersections or a dense source-depth confidence field before enforcing bidirectional consistency.
The final graphics push in this cycle tests whether the EGGROLL low-rank evolution-strategy idea can serve as a minimal neural visibility selector without turning the method into a full radiance MLP. The script \texttt{apps\_industrial\_breakthrough/dl3dv\_eggro\_visibility\_scorer.py} renders three candidate stacks---linear interpolation, the top-$K$ adaptive bridge, and the uniform FRI edge bridge---then trains a tiny low-rank antithetic ES scorer on even-view leave-one-out pixels. The scorer sees only candidate confidence, candidate disagreement, and local edge features; it blends candidate RGB values but does not synthesize new color. With $100$ ES iterations, population $40$, and rank $8$, the learned blend reaches $29.6504$ dB, SSIM $0.914176$, and LPIPS $0.082802$ on the odd held-out views. This improves over the candidate bridge rendered inside the same joint run, but it still does not beat the frozen deterministic FRI edge bridge at $29.8085$ dB or the top-$K$ profile's LPIPS. A subsequent affine color calibration overfits the leave-one-out training views and falls to $26.4261$ dB. We then expanded the scorer with an OSNR feature lift: local pooled spline-style neighborhoods, candidate-rank channels, luminance/chroma terms, and Fourier coordinate features. The low-resolution smoke improves, but the full $256^2$ held-out run reaches only $29.6202$ dB, below the simpler ES scorer. The conclusion is narrow but useful: a tiny ES visibility scorer is plausible, but raw scorer capacity over finished RGB candidates is not yet a SOTA-grade substitute for better geometry proposals, epipolar edge evidence, or dense visibility/depth reasoning.
To isolate the precise difference from NeRF-style training, \texttt{apps\_industrial\_breakthrough/dl3dv\_osnr\_volume\_renderer.py} implements a shared OSNR volume with actual alpha compositing. Sparse COLMAP points are recolored from even training views, deposited into a compact $64^3$ radiance/density grid, regularized by the FFT B-spline/Laplacian normal solve, and rendered into odd cameras by ray-marching $64$ samples per ray. This borrows NeRF's volumetric visibility equation but not its MLP. The result exposes the missing ingredient. The pure alpha-composited volume reaches only $11.4149$ dB, SSIM $0.376490$, and LPIPS $0.911937$. Blending $1\%$ of the volume render with $99\%$ interpolation gives a tiny PSNR foothold at $22.7898$ dB, but no meaningful view-synthesis gain. Thus the NeRF advantage is not merely alpha compositing; it is direct optimization of a dense occupancy/transmittance field from multi-view ray losses. Sparse COLMAP deposition plus smooth FFT regularization does not provide enough empty-space or surface evidence to create that field.
We then tested the obvious next bridge, \texttt{apps\_industrial\_breakthrough/dl3dv\_osnr\_pinn\_density\_renderer.py}: an OSNR-PINN density volume whose color and density spline-grid coefficients are initialized from the same COLMAP/FFT volume but fitted against even-view rays with a differentiable volume-rendering loss, total-variation regularity, sparsity, an initialization anchor, and an eikonal-style surface-gradient proxy. At $128^2$ resolution with a $48^3$ grid, $48$ samples per ray, and $120$ training iterations, the training photometric loss drops from $0.118027$ to $0.025061$. However, the held-out pure density render reaches only $15.7504$ dB, SSIM $0.327521$, and LPIPS $0.850917$; the best $5\%$ blend with interpolation reaches $24.7899$ dB, slightly below the interpolation baseline at the same resolution ($24.8264$ dB). This negative result is useful: weak physics-style regularity is insufficient. A NeRF-competitive OSNR geometry model needs either dense depth/occupancy supervision, stronger epipolar surface constraints, or an optimizer that directly solves the nonlinear visibility ambiguity rather than only smoothing sparse COLMAP evidence.
The denser initializer \texttt{apps\_industrial\_breakthrough/dl3dv\_dense\_depth\_volume\_renderer.py} then replaces sparse COLMAP deposition with leave-one-out FRI depth estimates for all even training views. At $160^2$ resolution, the runner backprojects $257{,}276$ confidence-weighted dense depth samples into a shared $72^3$ spline/FFT radiance-density volume and renders held-out odd views with $72$ alpha samples per ray. This is a stronger geometry initializer, but the single smoothed volume still fails: the pure dense-depth alpha volume reaches $12.9803$ dB, SSIM $0.325629$, and LPIPS $0.652960$; the best $2\%$ blend reaches $24.1111$ dB, essentially tied with but not better than the same-resolution interpolation baseline ($24.1113$ dB). The failure mode is now clear. The deterministic depth bridge succeeds because it keeps view-conditioned depth hypotheses and local source evidence alive until rendering. Collapsing those hypotheses into one global smoothed density grid discards too much visibility structure. The next NeRF-facing OSNR attempt should therefore preserve surface/depth hypotheses explicitly (for example as layered splines or surfel sheets) or learn opacity with direct multi-view transmittance constraints, not by smoothing depth maps into a volumetric average.
The layered follow-up \texttt{apps\_industrial\_breakthrough/dl3dv\_dense\_surfel\_renderer.py} confirms this diagnosis. Instead of averaging the dense FRI depths into a volume, it keeps the $361{,}290$ backprojected training-view samples as explicit weighted surfels and z-buffers them into held-out views. At $192^2$ resolution with $160$ depth planes, the same-resolution interpolation baseline is $23.5474$ dB. The best layered surfel profile, radius $1$ with a $3\times3$ smoothed $10\%$ blend, reaches $23.6577$ dB and SSIM $0.659134$. Pure surfels remain poor ($18.6729$ dB) because coverage, view-dependent color, and visibility ordering are still imperfect, but this is the first volumetric/surface variant in this sequence to improve over its same-resolution interpolation baseline. The conclusion is narrow: preserving layered surface hypotheses is directionally correct, while collapsing geometry into a single smoothed voxel field is not.
We also tested whether the EGGROLL-style low-rank selector could turn the surfel signal into a stronger visibility model. The scorer in \texttt{apps\_industrial\_breakthrough/dl3dv\_surfel\_visibility\_scorer.py} is trained on even-view leave-one-out pixels over five candidates: interpolation, raw surfels, smoothed surfels, confidence-blended surfels, and the fixed $10\%$ smoothed surfel blend. This did not improve the frontier. The fixed blend remains best at $23.6577$ dB, while the learned ES selector reaches only $23.0595$ dB and the affine-calibrated selector falls to $22.2272$ dB. The current candidate scorer therefore overfits or selects the wrong corrections.
The next quality-first probe replaces the shallow ES selector with a spline-native network in \texttt{apps\_industrial\_breakthrough/dl3dv\_deep\_spline\_surfel\_network.py}. The model uses two dense hidden layers, but replaces standard pointwise activations by learnable per-channel compact spline activations. It receives the same explicit surfel candidate stack, local OSNR feature lift, confidence fields, candidate disagreement, coordinates, and candidate RGB values, and is fitted only on even-view leave-one-out pixels. At $192^2$ resolution, $160$ depth planes, $64$ hidden channels, $25$ activation knots, and $60$ epochs over $131{,}072$ samples, the deep spline selector reaches $23.7305$ dB and SSIM $0.663583$, improving over the same-resolution linear baseline ($23.5474$ dB) and fixed surfel blend ($23.6577$ dB). However, LPIPS worsens to $0.265901$, and the total diagnostic latency rises to $52.433$ s. This is a useful but bounded result: additional spline-network capacity can extract a little more PSNR from the preserved surfel hypotheses, but it does not repair missing visibility, coverage, or view-dependent color evidence. A stronger NeRF/SIREN competitor needs explicit source-view agreement, occlusion ordering, normal/facing estimates, epipolar edge intersections, or dense confidence geometry before adding still larger learned layers.
We therefore added those first-order evidence channels directly in \texttt{apps\_industrial\_breakthrough/dl3dv\_evidence\_spline\_surfel\_network.py}. The renderer records per-pixel surfel confidence, color-consistency variance, depth-coherence variance, projected depth-edge strength, and a crude facing proxy derived from the rendered depth gradient. These measurements define conservative and stronger evidence-gated surfel blends, and the spline network is constrained to choose among these candidates with no free RGB residual by default. This improves the safety of raw surfel corrections but does not change the frontier. At the same $192^2$ diagnostic setting, the fixed $10\%$ smoothed-surface blend remains best at $23.6576$ dB and LPIPS $0.148888$. The evidence-conservative blend reaches $23.6183$ dB but slightly improves LPIPS to $0.146167$, while the evidence-aware spline selector falls to $23.5142$ dB despite a higher SSIM of $0.670806$. This is the strongest negative constraint so far: simple confidence, variance, and facing features are not enough to infer NeRF-grade visibility. The next real push must create better geometry hypotheses themselves, specifically multi-view epipolar edge intersections, source-view agreement at candidate depths before surfel projection, and occlusion-ordered layered surfaces.
To remove ambiguity about benchmark quality scales, the next runner, \texttt{apps\_industrial\_breakthrough/nerf\_synthetic\_spline\_benchmark.py}, switches to the canonical NeRF Synthetic/Blender data layout: \texttt{transforms\_train.json}, \texttt{transforms\_test.json}, RGBA images composited over white, and OpenGL camera-to-world matrices. The script supports real scenes such as \texttt{lego} or \texttt{chair}, but also creates a tiny generated Blender-style sphere scene for camera-convention and memory smoke testing. This benchmark records three profiles: nearest-pose image transfer, deterministic OSNR-style plane sweep, and a small deep spline ray network with Fourier ray features and learnable compact spline activations. On the generated sphere diagnostic ($24$ train views, $8$ held-out views, $96^2$ resolution, $96$ depth planes), nearest-pose transfer reaches $24.0446$ dB, SSIM $0.869139$, and LPIPS $0.016128$. The deterministic plane sweep reaches only $19.1317$ dB but a higher SSIM of $0.887576$, exposing a cost/depth ambiguity despite correct Blender projection. The deep spline ray network fits the training rays down to MSE $2.5344\times10^{-3}$, but collapses on held-out views at $6.2504$ dB and LPIPS $0.694466$. This reproduces the DL3DV lesson under a cleaner benchmark convention: a ray-only coordinate network, even with spline activations, is not a NeRF replacement. The next canonical run should therefore use the actual NeRF volume-rendering transmittance equation with spline-parameterized density/radiance, then evaluate on real \texttt{lego}/\texttt{chair} metrics against published NeRF-family scores.
That volume-rendering step is implemented in \texttt{apps\_industrial\_breakthrough/nerf\_synthetic\_spline\_volume\_renderer.py}. The model keeps NeRF's alpha-compositing equation but replaces the MLP with a compact trilinear spline grid storing density and color coefficients. Training is still by ray loss, so this is not an autograd-free solver; it is a controlled test of whether the missing ingredient is the transmittance model rather than the spline representation. On the same generated sphere benchmark at $96^2$ resolution, a $64^3$ spline volume with $80$ samples per ray and $800$ AdamW iterations drives the training loss to $9.064\times10^{-4}$ and reaches $28.6898$ dB PSNR, SSIM $0.963119$, and LPIPS $0.021464$ on held-out views. This beats nearest-pose transfer by $4.6452$ dB and improves substantially over both plane sweep and the ray-only spline network. The result is the first positive canonical-NeRF-path evidence: OSNR should compete through spline-parameterized density/radiance under the correct volume-rendering operator, not through direct ray-to-RGB regression.
We then moved from the generated smoke scene to real canonical NeRF Synthetic Blender scenes. To keep the diagnostic memory-bounded and comparable to the smoke run, we downloaded only the per-scene \texttt{lego} and \texttt{chair} archives from the NerfBaselines data mirror, evaluated $24$ training views and $8$ held-out test views at $96^2$ resolution, and kept the same $64^3$ trilinear spline grid, $80$ samples per ray, $800$ AdamW iterations, and $87.60$ MiB estimated active footprint. On \texttt{lego}, nearest-pose transfer reaches only $14.6577$ dB, SSIM $0.650337$, and LPIPS $0.180211$, while the spline volume reaches $22.4265$ dB, SSIM $0.891044$, and LPIPS $0.068127$. On \texttt{chair}, nearest-pose transfer reaches $22.6878$ dB, SSIM $0.883803$, and LPIPS $0.128070$, while the spline volume reaches $29.4390$ dB, SSIM $0.954310$, and LPIPS $0.062394$. These real-scene results show that the volume-rendering mechanism transfers beyond the generated sphere and produces large held-out gains over image transfer. They are not yet SOTA: the current model is still a first-order trilinear grid without view-dependent radiance, hierarchical sampling, cubic/exponential spline interpolation, sparse occupancy priors, or closed-form color updates. The next technical bottleneck is therefore not whether to use the NeRF operator, but how to replace the primitive trilinear lattice by a higher-order operator-spline volume with better density localization and view-dependent color.
We next replaced the primitive trilinear sampler by an explicit tensor-product cubic B-spline interpolation path. Each query point now accumulates over a compact $4\times4\times4$ support stencil using the cardinal cubic weights, rather than the $2\times2\times2$ trilinear hat stencil. This tests whether higher-order spline regularity alone improves the NeRF-style volume without changing the loss, grid size, camera model, or radiance parameterization. Because the cubic support is eight times larger, we used a matched reduced ray budget ($64$ samples per ray and $1024$ rays per batch) and ran both cubic and linear controls. On \texttt{chair}, cubic interpolation improves the matched-budget PSNR from $27.4948$ dB to $28.2934$ dB and SSIM from $0.934462$ to $0.947217$, but worsens LPIPS from $0.090736$ to $0.131908$ and is roughly an order of magnitude slower. On \texttt{lego}, cubic improves the matched-budget PSNR from $21.8256$ dB to $22.0616$ dB and SSIM from $0.876625$ to $0.882167$, but again worsens LPIPS from $0.095020$ to $0.148289$. The conclusion is therefore nuanced: higher-order tensor-product splines do help global least-squares-style reconstruction at fixed stochastic ray budget, but regularity alone is not the missing NeRF/SIREN ingredient. The model still needs sharper density localization, hierarchical occupancy sampling, and view-dependent radiance before it can approach published NeRF-family quality.
The strongest geometry result comes from replacing diffuse volumetric density by explicit spline-surface intersections. Inspired by the closed-form convolution/Gram acceleration used in spline snake resampling, we implemented \texttt{apps\_industrial\_breakthrough/nerf\_synthetic\_spline\_surface\_renderer.py}. The diagnostic represents the sphere geometry as a compact tensor-product cubic spline surface
\begin{equation}
\mathbf{s}(u,v)=\sum_{i,j}\mathbf{c}_{ij}\,\beta_3(Mu-i)\,\beta_3(Nv-j)
\end{equation}
and solves ray intersections by local Newton updates on the three unknowns $(u,v,t)$:
\begin{equation}
\mathbf{s}(u,v)-\mathbf{o}-t\mathbf{d}=\mathbf{0}.
\end{equation}
More generally, let $\beta_{\boldsymbol{\alpha}}$ denote a compact exponential or polynomial spline generator with pole vector $\boldsymbol{\alpha}$ and support length equal to the number of poles. A tensor-product spline surface is
\begin{equation}
\mathbf{s}(u,v)
=\sum_{i\in\mathcal{I}}\sum_{j\in\mathcal{J}}
\mathbf{c}_{ij}\,
\beta_{\boldsymbol{\alpha}_u}(Mu-i)\,
\beta_{\boldsymbol{\alpha}_v}(Nv-j),
\qquad
\mathbf{c}_{ij}\in\mathbb{R}^3 .
\label{eq:tensor-product-surface}
\end{equation}
For a camera ray $\mathbf{r}(t)=\mathbf{o}+t\mathbf{d}$, with $\|\mathbf{d}\|_2=1$ and $t>0$, an intersection is a root of
\begin{equation}
\mathbf{F}(u,v,t)
=\mathbf{s}(u,v)-\mathbf{o}-t\mathbf{d}
=\mathbf{0}.
\label{eq:spline-surface-root}
\end{equation}
The local Newton system follows directly from the analytical spline derivative ladder:
\begin{equation}
\begin{bmatrix}
\partial_u\mathbf{s}(u,v) & \partial_v\mathbf{s}(u,v) & -\mathbf{d}
\end{bmatrix}
\begin{bmatrix}
\Delta u\\ \Delta v\\ \Delta t
\end{bmatrix}
=-\mathbf{F}(u,v,t),
\label{eq:spline-surface-newton}
\end{equation}
where
\begin{align}
\partial_u\mathbf{s}(u,v)
&=
M\sum_{i,j}\mathbf{c}_{ij}\,
\beta'_{\boldsymbol{\alpha}_u}(Mu-i)\,
\beta_{\boldsymbol{\alpha}_v}(Nv-j),\\
\partial_v\mathbf{s}(u,v)
&=
N\sum_{i,j}\mathbf{c}_{ij}\,
\beta_{\boldsymbol{\alpha}_u}(Mu-i)\,
\beta'_{\boldsymbol{\alpha}_v}(Nv-j).
\end{align}
Because the support of $\beta_{\boldsymbol{\alpha}}$ is compact, only a small stencil of coefficients contributes to each $(u,v)$ evaluation. For cubic polynomial splines this stencil is $4\times4$ for a surface, while for an order-$P_u$ by order-$P_v$ exponential spline it is $P_u\times P_v$. This is the surface analogue of the convolution/Gram trick used in spline resampling: all repeated products between basis functions and derivative basis functions can be pretabulated as local functions of fractional coordinates, and candidate patches can be culled by compact support before solving \eqref{eq:spline-surface-newton}. The expensive global scene query is therefore reduced to a small number of local $3\times3$ systems.
The boundary conditions are determined by the topology of the parameter domain:
\begin{itemize}
\item \textbf{Rectangular open patches.} For a surface patch over $[0,1]\times[0,1]$, both parameters are non-periodic. Newton updates are clamped or damped to keep $(u,v)$ inside the valid domain, and only basis functions whose support overlaps the rectangle are evaluated. This covers trimmed sheets, local surface charts, and open spline patches used for piecewise object shells.
\item \textbf{Cylindrical topology.} For a cylinder-like surface, one parameter is periodic and the other is open. Typically $u\in\mathbb{R}/\mathbb{Z}$ wraps around the circumference, while $v\in[0,1]$ remains clamped along the height. The coefficient index $i$ is evaluated modulo $M$, but $j$ uses boundary-aware open support. Newton updates wrap $u\leftarrow u\bmod 1$ and clamp $v$.
\item \textbf{Toroidal topology.} For a torus-like surface, both parameters are periodic: $(u,v)\in(\mathbb{R}/\mathbb{Z})^2$. Both coefficient indices are circular, the Gram matrices are block-circulant, and the support search can be diagonalized or accelerated by FFT-style periodic convolution. Newton updates wrap both parameters.
\item \textbf{Spherical topology.} A sphere is periodic in longitude but singular at the poles. The practical implementation used here treats $u$ as periodic and $v\in[0,1]$ as a clamped latitude coordinate, with duplicated/regularized polar control rows. A more invariant construction can use multiple overlapping charts, such as two or six rectangular charts, to avoid polar degeneracy. In either case, the intersection equation remains \eqref{eq:spline-surface-root}; only the index wrapping and chart transition rules change.
\end{itemize}
For closed surfaces the first positive root along the ray is selected, while for multi-layer or self-occluding surfaces the renderer evaluates all candidate local roots and keeps the smallest valid $t$ after residual and normal-facing checks. This gives an explicit alternative to NeRF's volumetric opacity integral: geometry is stored as a low-dimensional spline manifold, visibility is resolved by root ordering, and radiance can be attached to the surface as a second tensor-product field $\boldsymbol{\rho}(u,v,\mathbf{d})$ rather than diffused through a dense 3D volume.
This is not a full unknown-scene method yet: the surface family is known and the initialization uses the sphere's analytic support. It is nevertheless a critical density-localization experiment because the renderer evaluates only compact surface support instead of fitting an opaque 3D density field. On the same generated NeRF Synthetic sphere benchmark, a coarse $32\times17$ surface reaches $24.9397$ dB, SSIM $0.965720$, and LPIPS $0.008263$; a $64\times33$ surface reaches $31.0490$ dB, SSIM $0.991639$, and LPIPS $0.002731$; and a $96\times49$ surface reaches $50.8061$ dB, SSIM $0.999790$, and LPIPS $0.000023$ in $530.8$ ms at only $16.29$ MiB estimated active footprint. This confirms the central geometric hypothesis: when the shape class can be represented explicitly, tensor-product spline surfaces plus direct ray intersection can outperform diffuse volumetric fitting by a large margin. The next research problem is to infer such surfaces from multi-view data, using edge/epipolar evidence and occupancy fields, rather than assuming them.
The first unknown-scene bridge experiment is \texttt{apps\_industrial\_breakthrough/nerf\_synthetic\_occupancy\_shell\_renderer.py}. It trains the same spline density/radiance volume on \texttt{chair}, then attempts to collapse the learned opacity field into a single explicit shell sample per held-out ray. We tested three extraction rules: maximum transmittance weight, first alpha-threshold crossing, and expected-depth projection. This is the direct test of whether a NeRF-style learned density can be turned into an OSNR-style surface renderer without first improving the geometry prior. The result is negative but informative. The alpha-composited spline volume repeats the previous $29.4390$ dB, SSIM $0.954310$, LPIPS $0.062394$ result. The best shell collapse, maximum weight, falls to $24.2233$ dB, SSIM $0.880552$, and LPIPS $0.103145$; first-alpha reaches only $21.5798$ dB, and expected-depth reaches $22.4439$ dB. The density active ratio remains $0.998177$, showing that the learned field is still a diffuse opacity cushion rather than a localized surface shell. This explains why the explicit sphere-surface experiment succeeds while direct shell extraction from the primitive volume fails: explicit spline geometry is powerful, but the current volume training objective does not yet produce extractable geometry. The next step must add an occupancy/surface regularizer, multi-view depth agreement, or an edge-driven shell proposal before collapsing to tensor-product patches.
We therefore recast the Blender held-out problem as a spline-tomographic inverse problem rather than as pure coordinate-network fitting. The implementation is \texttt{apps\_industrial\_breakthrough/nerf\_synthetic\_silhouette\_volume\_solver.py}. It uses the NeRF Synthetic RGBA alpha channel as an explicit silhouette measurement and solves the geometry stage before the radiance stage. Let $\sigma_{\mathbf{c}}(\mathbf{x})\geq0$ be the compact-support spline density volume with coefficients $\mathbf{c}$, and let $\mathbf{r}_{i}(t)=\mathbf{o}_{i}+t\mathbf{d}_{i}$ be a camera ray. The opacity forward operator is
\begin{equation}
\mathcal{H}(\mathbf{c})_i
=1-\exp\left(-\int_{t_{\min}}^{t_{\max}}\sigma_{\mathbf{c}}(\mathbf{r}_i(t))\,dt\right),
\label{eq:silhouette-forward}
\end{equation}
which is the nonlinear analogue of the tomographic projector $H\mathbf{c}$ used in the spline CT and cryo-EM papers. The first stage estimates occupancy by minimizing a balanced foreground/background silhouette data term with positivity built into $\sigma_{\mathbf{c}}=\operatorname{softplus}(\tilde{\mathbf{c}})$:
\begin{equation}
\min_{\tilde{\mathbf{c}}}\;
\operatorname{BCE}\!\left(\mathcal{H}(\operatorname{softplus}(\tilde{\mathbf{c}})),\mathbf{a}\right)
+\lambda_{\mathrm{TV}}\|\nabla \tilde{\mathbf{c}}\|_1
+\lambda_{\mathrm{sp}}\|\operatorname{softplus}(\tilde{\mathbf{c}})\|_1 .
\label{eq:silhouette-inverse}
\end{equation}
This is the NeRF-facing counterpart of the constrained regularized weighted-norm reconstructions of Nilchian and Donati: the unknown is a spline coefficient volume, the data term is a ray projection model, and the priors enforce support, positivity, sparsity, and bounded variation. After the density stage, the density is frozen and a separate compact spline color volume is fitted through the standard alpha compositing integral. This cleanly separates geometry recovery from radiance fitting and prevents the color loss from using diffuse density as an unrestricted numerical cushion.
The result is a useful diagnostic. On \texttt{lego}, $50$ training views, $8$ held-out views, $96^2$ resolution, a $64^3$ grid, $80$ samples per ray, and $500+500$ density/color iterations produce held-out alpha IoU $0.944604$ and reduce the density active ratio from the RGB-only volume's roughly $0.997$ to $0.418766$. The RGB score is $22.6041$ dB, SSIM $0.896264$, and LPIPS $0.076846$. On \texttt{chair}, the same protocol yields alpha IoU $0.937057$, active density $0.341656$, and $28.6128$ dB, SSIM $0.956229$, LPIPS $0.064510$. A quality-relaxed Chair pass with weaker TV/sparsity and a longer color stage reaches $28.9733$ dB, SSIM $0.957624$, LPIPS $0.058525$, alpha IoU $0.937829$, and active density $0.307980$. The interpretation is precise: silhouette tomography fixes the diffuse-geometry failure and creates a compact, visible occupancy field, but it does not yet beat the RGB-only spline volume in PSNR. The remaining bottleneck is surface-aware radiance assignment and visibility, not silhouette geometry. The next NeRF-facing solver should therefore use the recovered occupancy field as an initialization/preconditioner for an adjoint or variable-projection radiance solve, or convert the high-confidence occupancy boundary into explicit tensor-product spline patches.
\begin{figure}[h]
\centering
\begin{minipage}[t]{0.48\linewidth}
\centering
\includegraphics[width=0.48\linewidth]{../apps_industrial_breakthrough/nerf_synthetic_silhouette_volume_lego_g64_i500/test_frame0_target.png}
\includegraphics[width=0.48\linewidth]{../apps_industrial_breakthrough/nerf_synthetic_silhouette_volume_lego_g64_i500/test_frame0_silhouette_volume.png}\\
\includegraphics[width=0.48\linewidth]{../apps_industrial_breakthrough/nerf_synthetic_silhouette_volume_lego_g64_i500/test_frame0_alpha_target.png}
\includegraphics[width=0.48\linewidth]{../apps_industrial_breakthrough/nerf_synthetic_silhouette_volume_lego_g64_i500/test_frame0_alpha_recon.png}\\
\small Lego RGB target/reconstruction and alpha target/reconstruction
\end{minipage}
\hfill
\begin{minipage}[t]{0.48\linewidth}
\centering
\includegraphics[width=0.48\linewidth]{../apps_industrial_breakthrough/nerf_synthetic_silhouette_volume_chair_g64_quality/test_frame0_target.png}
\includegraphics[width=0.48\linewidth]{../apps_industrial_breakthrough/nerf_synthetic_silhouette_volume_chair_g64_quality/test_frame0_silhouette_volume.png}\\
\includegraphics[width=0.48\linewidth]{../apps_industrial_breakthrough/nerf_synthetic_silhouette_volume_chair_g64_quality/test_frame0_alpha_target.png}
\includegraphics[width=0.48\linewidth]{../apps_industrial_breakthrough/nerf_synthetic_silhouette_volume_chair_g64_quality/test_frame0_alpha_recon.png}\\
\small Chair RGB target/reconstruction and alpha target/reconstruction
\end{minipage}
\caption{Silhouette-constrained spline-volume inverse reconstruction on NeRF Synthetic held-out views. The alpha reconstructions show that the occupancy inverse problem localizes the object support accurately (Lego IoU $0.944604$, Chair IoU $0.937829$), while the RGB images expose the remaining surface-radiance and visibility bottleneck.}
\label{fig:nerf-synthetic-silhouette-volume}
\end{figure}
The next NeRF-facing bridge is \texttt{apps\_industrial\_breakthrough/osnr\_nerf\_spline\_mlp.py}, which keeps NeRF's alpha-compositing and hierarchical coarse/fine ray sampling but replaces the plain coordinate MLP by a compact multiresolution spline-grid feature field. The important engineering correction is that compact support must be used as a local-control mechanism, not as a dense expanded positional feature vector. Early variants with dense spline encodings and spline activations improved the matched Fourier/ReLU baseline but were prohibitively slow. The current quality-first configuration samples one shared compact cubic spline grid, feeds the local spline features to split density/color heads, disables dense spline encodings by default, and runs on Apple Metal through \texttt{torch.device='mps'}. On the generated sphere diagnostic, a $4$-level grid with base resolution $6$ and $6$ features per level reaches $22.6470$ dB, SSIM $0.859884$, and LPIPS $0.112952$ at $48^2$ resolution, while the matched Fourier/ReLU profile reaches only $18.9988$ dB. On real NeRF Synthetic \texttt{lego}, the same compact-grid mechanism transfers: at $64^2$ resolution, $24+16$ samples per ray, $24$ training views, $8$ held-out views, and $2000$ MPS iterations, the shared-grid OSNR profile reaches $21.2960$ dB, SSIM $0.858446$, and LPIPS $0.069849$, versus $20.4836$ dB, SSIM $0.830932$, and LPIPS $0.120454$ for the matched Fourier/ReLU model. Increasing MLP width from $64$ to $128$ hidden channels does not help, and raw grid scaling to base resolution $8$ or $5$ levels trades PSNR for perceptual metrics rather than producing a clean improvement. A light ray-geometry concentration prior is more useful: with entropy weight $10^{-4}$ and depth-variance weight $10^{-5}$, the $3000$-step Lego run reaches $21.6605$ dB, SSIM $0.873037$, and LPIPS $0.062367$, while the matched Fourier/ReLU baseline in the same run reaches $21.0730$ dB, SSIM $0.846366$, and LPIPS $0.099627$. Scaling to $96^2$ shows that the $4$-level grid underfits perceptual detail ($21.2913$ dB, SSIM $0.847601$, LPIPS $0.149864$), but adding a fifth compact grid level recovers the high-resolution frontier. With $5000$ MPS iterations, the $96^2$ five-level OSNR profile reaches $21.9754$ dB, SSIM $0.869955$, and LPIPS $0.083635$ without reintroducing dense spline features; extending the same run to $8000$ iterations lowers held-out PSNR to $21.8090$ despite lower training loss, indicating overfitting or stochastic ray-sampling mismatch. Increasing view Fourier frequencies to $10$ worsens the same setting to $21.0993$ dB, doubling the ray batch to $1536$ reaches only $21.5024$ dB, lowering the learning rate to $3\times10^{-4}$ reaches only $21.4863$ dB, and re-enabling compact spline activations reaches only $21.3165$ dB while increasing training time beyond $1000$ s. The stronger quality lever is camera coverage: increasing the Lego training set from $24$ to $50$ views at the same $96^2$ five-level setting raises the OSNR profile to $22.5796$ dB, SSIM $0.883554$, and LPIPS $0.088734$ at $5000$ iterations, $23.0810$ dB, SSIM $0.896646$, and LPIPS $0.077989$ at $8000$ iterations, and $23.7099$ dB, SSIM $0.909780$, and LPIPS $0.066458$ at $12000$ iterations. Using all $100$ training views with only $5000$ iterations improves SSIM to $0.886317$ but lowers PSNR to $22.4800$ and LPIPS to $0.094478$, suggesting that the fixed update budget is then spread too thinly across cameras. A direct $128^2$ scaling run with the $50$-view, five-level configuration reaches $23.1842$ dB, SSIM $0.887382$, and LPIPS $0.132941$. Adding a sixth grid level improves the $128^2$ perceptual score to LPIPS $0.109930$ and SSIM to $0.889882$, but leaves PSNR essentially unchanged at $23.1901$ dB while increasing training time to $834.1$ s. At $96^2$, the six-level Lego profile is a perceptual/detail tradeoff rather than a universal improvement: PSNR drops from $23.7099$ to $23.4193$ and SSIM from $0.909780$ to $0.903694$, but LPIPS improves from $0.066458$ to $0.058955$. A stronger Lego improvement comes from decoupling opacity and radiance support: enabling a separate compact color grid while keeping the five-level density grid raises the $96^2$, $50$-view, $12000$-step result to $24.0426$ dB, SSIM $0.916105$, and LPIPS $0.055512$. Combining separate color support with a sixth level gives the best Lego LPIPS so far, $0.046647$, but drops PSNR to $23.7521$ dB and takes $1629.4$ s to train; this makes it a quality-ceiling/perceptual point, not the efficient frontier. We also implemented edge-weighted ray sampling, mixing CPU-selected high-gradient rays with uniform MPS batches to avoid a large-vector MPS multinomial failure. A $50\%$ edge-biased mixture on the five-level separate-color Lego model improves LPIPS slightly to $0.053491$ but lowers PSNR/SSIM to $23.8910$ dB and $0.910451$, while a gentler $25\%$ mixture falls further to $23.5742$ dB, SSIM $0.909236$, and LPIPS $0.056093$. Static image-gradient sampling is therefore not the correct hard-ray policy; it prioritizes apparent silhouettes before the model has estimated which rays are actually underfit. The successful sampler is residual-driven: after a $3000$-step uniform warmup, replacing $25\%$ of each batch with the highest-error rays from a $2\times$ no-gradient candidate pool raises Lego to a new quality frontier of $24.4613$ dB, SSIM $0.916640$, and LPIPS $0.041176$. A cheaper $1.25\times$ candidate pool reaches $24.2211$ dB, SSIM $0.916619$, and LPIPS $0.044676$ in $1596.4$ s, preserving most of the perceptual gain while reducing runtime. The same residual-mined policy transfers to \texttt{chair}, raising the separate-color frontier from $29.8809$ dB to $30.8454$ dB, SSIM $0.961629$, and LPIPS $0.047816$, which also beats the previous six-level shared-grid Chair LPIPS. These two scenes confirm that model-aware hard-ray selection is a general OSNR-NeRF mechanism; its cost shows that the next engineering problem is a cheaper cached or amortized residual map rather than more model capacity. Chair also benefits from both earlier capacity mechanisms, but in different ways: a sixth shared grid level improves perceptual quality to LPIPS $0.048553$ and raises PSNR to $29.6390$ dB, while a separate five-level color grid gives $29.8809$ dB and $0.961321$ SSIM with LPIPS $0.053649$. The efficient frontier is therefore no longer a single grid-depth setting: decoupled radiance support and residual-driven evidence selection are the strongest general quality levers, while extra local spline scale is a perceptual/detail lever whose value depends on scene content. The limiting factor is not angular encoding bandwidth, stochastic batch noise, learning-rate instability, or pointwise activation expressivity; it is the amount and organization of geometric/radiance evidence available to the compact field. The current lesson is specific: the spline advantage is real when it is implemented as compact local grid control with split radiance/opacity heads, separate radiance support, and model-aware hard-ray selection; larger dense MLPs, dense spline feature expansion, and simply adding more depth samples are not the path forward.
\paragraph{Reproducible residual-mined OSNR-NeRF protocol.}
The residual-mined experiments use the same held-out NeRF Synthetic split throughout: $50$ training views, $8$ test views, $96^2$ render resolution, $24$ coarse samples, $16$ fine hierarchical samples, $12000$ MPS training iterations, $768$ rays per training batch, a $5$-level compact cubic spline grid with base resolution $6$ and $6$ features per level, separate compact radiance support via \texttt{--separate\_color\_grid}, entropy regularization $10^{-4}$, and depth-variance regularization $10^{-5}$. Training starts with a uniform-ray warmup of $3000$ iterations. After warmup, for each step a candidate ray set of size $\kappa B$ is sampled uniformly, rendered under \texttt{torch.no\_grad()}, scored by per-ray RGB MSE, and the top $\rho B$ candidates replace part of the training batch:
\begin{equation}
e_i=\frac{1}{3}\left\|\hat{\mathbf{c}}(\mathbf{r}_i)-\mathbf{c}_i\right\|_2^2,\qquad
\mathcal{H}_t=\operatorname{TopK}_{i\in\mathcal{C}_t}(e_i,\rho B),
\end{equation}
where $B=768$, $\rho=0.25$, and $\kappa\in\{1.25,2.0\}$. The final batch is
\begin{equation}
\mathcal{B}_t=\mathcal{U}_t\cup\mathcal{H}_t,\qquad
|\mathcal{U}_t|=(1-\rho)B,\quad |\mathcal{H}_t|=\rho B.
\end{equation}
This differs from fixed image-gradient sampling: the hard rays are selected by the current model's residual after a warmup, not by an a-priori edge detector. The exact frontier commands are:
\begin{verbatim}
venv/bin/python apps_industrial_breakthrough/osnr_nerf_spline_mlp.py \
--device mps --profiles osnr_spline_wavelet \
--data_dir data/nerf_synthetic --scene lego \
--views_train 50 --views_test 8 --resolution 96 \
--coarse_samples 24 --fine_samples 16 --iters 12000 \
--batch_rays 768 --hidden 64 --layers 4 --skip_layer 2 \
--grid_levels 5 --grid_base 6 --grid_features 6 \
--separate_color_grid --residual_sample_prob 0.25 \
--residual_warmup 3000 --residual_candidate_mult 2 \
--entropy_weight 0.0001 --depth_var_weight 0.00001 \
--lpips_frames 2 \
--output_dir apps_industrial_breakthrough/\
osnr_nerf_spline_mlp_lego96_train50_grid566_colorgrid_residual25_warm3000_mps_geomreg_b_12k
venv/bin/python apps_industrial_breakthrough/osnr_nerf_spline_mlp.py \
--device mps --profiles osnr_spline_wavelet \
--data_dir data/nerf_synthetic --scene chair \
--views_train 50 --views_test 8 --resolution 96 \
--coarse_samples 24 --fine_samples 16 --iters 12000 \
--batch_rays 768 --hidden 64 --layers 4 --skip_layer 2 \
--grid_levels 5 --grid_base 6 --grid_features 6 \
--separate_color_grid --residual_sample_prob 0.25 \
--residual_warmup 3000 --residual_candidate_mult 2 \
--entropy_weight 0.0001 --depth_var_weight 0.00001 \
--lpips_frames 2 \
--output_dir apps_industrial_breakthrough/\
osnr_nerf_spline_mlp_chair96_train50_grid566_colorgrid_residual25_warm3000_mps_geomreg_b_12k
\end{verbatim}
\begin{figure}[h]
\centering
\begin{minipage}[t]{0.48\linewidth}
\centering
\includegraphics[width=0.48\linewidth]{../apps_industrial_breakthrough/osnr_nerf_spline_mlp_lego96_train50_grid566_colorgrid_residual25_warm3000_mps_geomreg_b_12k/test_frame0_target.png}
\includegraphics[width=0.48\linewidth]{../apps_industrial_breakthrough/osnr_nerf_spline_mlp_lego96_train50_grid566_colorgrid_residual25_warm3000_mps_geomreg_b_12k/test_frame0_osnr_spline_wavelet.png}\\
\small Lego target \hfill Lego OSNR
\end{minipage}
\hfill
\begin{minipage}[t]{0.48\linewidth}
\centering
\includegraphics[width=0.48\linewidth]{../apps_industrial_breakthrough/osnr_nerf_spline_mlp_chair96_train50_grid566_colorgrid_residual25_warm3000_mps_geomreg_b_12k/test_frame0_target.png}
\includegraphics[width=0.48\linewidth]{../apps_industrial_breakthrough/osnr_nerf_spline_mlp_chair96_train50_grid566_colorgrid_residual25_warm3000_mps_geomreg_b_12k/test_frame0_osnr_spline_wavelet.png}\\
\small Chair target \hfill Chair OSNR
\end{minipage}
\caption{Held-out NeRF Synthetic visual comparison for the residual-mined compact OSNR-NeRF frontier. Left pair: \texttt{lego} target and OSNR reconstruction ($24.4613$ dB, SSIM $0.916640$, LPIPS $0.041176$). Right pair: \texttt{chair} target and OSNR reconstruction ($30.8454$ dB, SSIM $0.961629$, LPIPS $0.047816$).}
\label{fig:osnr-nerf-residual-lego-chair}
\end{figure}
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{@{}p{0.31\linewidth}p{0.23\linewidth}p{0.1\linewidth}p{0.1\linewidth}p{0.1\linewidth}p{0.11\linewidth}@{}}
\toprule
Held-out profile & Family & PSNR & SSIM & LPIPS & Latency \\
\midrule
Multi-view FRI edge bridge, $256$ depths & Deterministic edge geometry & $29.8085$ dB & $0.913584$ & $0.070140$ & $21481.0$ ms \\
EGGROLL visibility blend & Low-rank ES candidate scorer & $29.6504$ dB & $0.914176$ & $0.082802$ & $45609.7$ ms \\
Top-$K$ adaptive depth bridge, $192$ coarse depths & Adaptive deterministic geometry & $29.5615$ dB & $0.918645$ & $0.064826$ & $15882.1$ ms \\
Plane sweep $+$ liquid residual & Geometry-backed hybrid & $29.5291$ dB & $0.901776$ & $0.093192$ & $35450.9$ ms \\
Plane sweep, $512$ depths & Deterministic geometry renderer & $29.4187$ dB & $0.898331$ & $0.099257$ & $35201.2$ ms \\
Raw geometry keyview package & Compressed scene package & $29.3178$ dB & $0.896597$ & $0.102406$ & $20091.3$ ms \\
\texttt{nerfacto} CPU pilot & External NeRF baseline & $27.0830$ dB & $0.818526$ & $0.176857$ & $10268.9$ ms \\
Deep spline surfel NN, $192^2$ diagnostic & Spline-activation candidate network & $23.7305$ dB & $0.663583$ & $0.265901$ & $52433.2$ ms \\
Evidence-aware surfel blend, $192^2$ diagnostic & Surfel evidence selector & $23.6576$ dB & $0.659127$ & $0.148888$ & $13422.3$ ms \\
NeRF Synthetic sphere smoke, nearest pose & Canonical Blender harness & $24.0446$ dB & $0.869139$ & $0.016128$ & $0.6$ ms \\
NeRF Synthetic sphere smoke, spline volume & Spline density/radiance volume & $28.6898$ dB & $0.963119$ & $0.021464$ & $30687.6$ ms \\
NeRF Synthetic lego, nearest pose & Canonical Blender real scene & $14.6577$ dB & $0.650337$ & $0.180211$ & $0.6$ ms \\
NeRF Synthetic lego, spline volume & Spline density/radiance volume & $22.4265$ dB & $0.891044$ & $0.068127$ & $32628.6$ ms \\
NeRF Synthetic chair, nearest pose & Canonical Blender real scene & $22.6878$ dB & $0.883803$ & $0.128070$ & $0.6$ ms \\
NeRF Synthetic chair, spline volume & Spline density/radiance volume & $29.4390$ dB & $0.954310$ & $0.062394$ & $32494.9$ ms \\
NeRF Synthetic lego, matched linear volume & $64$ samples, $1024$ rays/batch & $21.8256$ dB & $0.876625$ & $0.095020$ & $10753.5$ ms \\
NeRF Synthetic lego, cubic spline volume & Tensor-product cubic grid & $22.0616$ dB & $0.882167$ & $0.148289$ & $126184.4$ ms \\
NeRF Synthetic chair, matched linear volume & $64$ samples, $1024$ rays/batch & $27.4948$ dB & $0.934462$ & $0.090736$ & $10639.7$ ms \\
NeRF Synthetic chair, cubic spline volume & Tensor-product cubic grid & $28.2934$ dB & $0.947217$ & $0.131908$ & $125802.4$ ms \\
NeRF Synthetic sphere, coarse spline surface & Tensor-product surface $32\times17$ & $24.9397$ dB & $0.965720$ & $0.008263$ & $527.6$ ms \\
NeRF Synthetic sphere, mid spline surface & Tensor-product surface $64\times33$ & $31.0490$ dB & $0.991639$ & $0.002731$ & $528.9$ ms \\
NeRF Synthetic sphere, dense spline surface & Tensor-product surface $96\times49$ & $50.8061$ dB & $0.999790$ & $0.000023$ & $530.8$ ms \\
NeRF Synthetic chair, occupancy shell max & Shell from learned density & $24.2233$ dB & $0.880552$ & $0.103145$ & $32033.1$ ms \\
NeRF Synthetic chair, occupancy shell first alpha & Shell from learned density & $21.5798$ dB & $0.888355$ & $0.147828$ & $32034.6$ ms \\
NeRF Synthetic chair, occupancy shell expected depth & Shell from learned density & $22.4439$ dB & $0.871275$ & $0.142029$ & $32036.6$ ms \\
NeRF Synthetic lego, silhouette volume & Occupancy-first inverse volume & $22.6041$ dB & $0.896264$ & $0.076846$ & $26644.9$ ms \\
NeRF Synthetic chair, silhouette volume & Occupancy-first inverse volume & $28.6128$ dB & $0.956229$ & $0.064510$ & $27055.9$ ms \\
NeRF Synthetic chair, silhouette volume quality & Relaxed occupancy-first volume & $28.9733$ dB & $0.957624$ & $0.058525$ & $46943.4$ ms \\
NeRF Synthetic lego, Fourier/ReLU NeRF-MLP & MPS, shared protocol, $3000$ steps & $21.0730$ dB & $0.846366$ & $0.099627$ & $26403.7$ ms \\
NeRF Synthetic lego, compact OSNR-NeRF & MPS, shared spline grid, light geometry prior & $21.6605$ dB & $0.873037$ & $0.062367$ & $79537.3$ ms \\
NeRF Synthetic lego, compact OSNR-NeRF $96^2$ & MPS, $5$-level shared spline grid & $21.9754$ dB & $0.869955$ & $0.083635$ & $183106.9$ ms \\
NeRF Synthetic lego, compact OSNR-NeRF $96^2$, $50$ views & MPS, $5$-level shared spline grid & $23.7099$ dB & $0.909780$ & $0.066458$ & $506179.0$ ms \\
NeRF Synthetic lego, compact OSNR-NeRF $96^2$, $50$ views & MPS, $6$-level shared spline grid & $23.4193$ dB & $0.903694$ & $0.058955$ & $840631.4$ ms \\
NeRF Synthetic lego, compact OSNR-NeRF $96^2$, $50$ views & MPS, separate $5$-level color grid & $24.0426$ dB & $0.916105$ & $0.055512$ & $960367.2$ ms \\
NeRF Synthetic lego, compact OSNR-NeRF $96^2$, $50$ views & MPS, separate $6$-level color grid & $23.7521$ dB & $0.910148$ & $0.046647$ & $1634796.8$ ms \\
NeRF Synthetic lego, compact OSNR-NeRF $96^2$, $50$ views & MPS, separate color grid, $50\%$ edge rays & $23.8910$ dB & $0.910451$ & $0.053491$ & $996875.8$ ms \\
NeRF Synthetic lego, compact OSNR-NeRF $96^2$, $50$ views & MPS, separate color grid, $25\%$ edge rays & $23.5742$ dB & $0.909236$ & $0.056093$ & $1050666.8$ ms \\
NeRF Synthetic lego, compact OSNR-NeRF $96^2$, $50$ views & MPS, separate color grid, residual hard rays & $24.4613$ dB & $0.916640$ & $0.041176$ & $1969545.5$ ms \\
NeRF Synthetic lego, compact OSNR-NeRF $96^2$, $50$ views & MPS, residual hard rays, $1.25\times$ pool & $24.2211$ dB & $0.916619$ & $0.044676$ & $1601067.7$ ms \\
NeRF Synthetic chair, compact OSNR-NeRF $96^2$, $50$ views & MPS, $5$-level shared spline grid & $29.4469$ dB & $0.956702$ & $0.070670$ & $512749.4$ ms \\
NeRF Synthetic chair, compact OSNR-NeRF $96^2$, $50$ views & MPS, $6$-level shared spline grid & $29.6390$ dB & $0.958011$ & $0.048553$ & $848168.5$ ms \\
NeRF Synthetic chair, compact OSNR-NeRF $96^2$, $50$ views & MPS, separate $5$-level color grid & $29.8809$ dB & $0.961321$ & $0.053649$ & $973176.4$ ms \\
NeRF Synthetic chair, compact OSNR-NeRF $96^2$, $50$ views & MPS, separate color grid, residual hard rays & $30.8454$ dB & $0.961629$ & $0.047816$ & $2114642.7$ ms \\
FFT tomographic grid blend & Linearized spline tomography & $22.7957$ dB & $0.609818$ & $0.155683$ & $870.1$ ms \\
Shared OSNR alpha volume & Spline/FFT volume & $22.7898$ dB & $0.612160$ & $0.151110$ & $2485.4$ ms \\
Linear neighbor & View interpolation & $22.7803$ dB & $0.611987$ & $0.150600$ & $2.96$ ms \\
Ray-kernel OSNR & Pure OSNR ray kernel & $18.0553$ dB & $0.466200$ & $0.904455$ & $600.8$ ms \\
\bottomrule
\end{tabular}
\caption{Consolidated held-out DL3DV SOTA ledger generated by \texttt{dl3dv\_heldout\_sota\_ledger.py}. The current win condition is geometry-backed rendering, not naive ray-coordinate regression. The \texttt{nerfacto} row is a local CPU pilot and must be replaced by a full CUDA baseline before making final SOTA claims.}
\label{tab:dl3dv-heldout-sota-ledger}
\end{table}
\begin{figure}[h]
\centering
\includegraphics[width=\linewidth]{../apps_industrial_breakthrough/dl3dv_heldout_sota_ledger_outputs/dl3dv_heldout_sota_ledger.png}
\caption{Ranked DL3DV held-out ledger across interpolation, pure OSNR ray kernels, spline-tomographic radiance, deterministic geometry, hybrid residual correction, and the available \texttt{nerfacto} CPU pilot.}
\label{fig:dl3dv-heldout-sota-ledger}
\end{figure}
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{@{}p{0.15\linewidth}p{0.16\linewidth}p{0.1\linewidth}p{0.1\linewidth}p{0.1\linewidth}p{0.12\linewidth}p{0.12\linewidth}p{0.11\linewidth}@{}}
\toprule
DCT support & Kept/channel & PSNR & SSIM & LPIPS & Sparsity & Max edge error & Time \\
\midrule
$100\%$ & $1{,}966{,}080$ & $117.2378$ dB & $1.000000$ & $0.000000$ & $77.50\%$ & $0.958244$ & $134.5901$ ms/view \\
$50\%$ & $983{,}040$ & $42.7562$ dB & $0.981283$ & $0.002584$ & $87.50\%$ & $0.958244$ & $136.9464$ ms/view \\
$25\%$ & $491{,}520$ & $35.5073$ dB & $0.927023$ & $0.049832$ & $92.50\%$ & $0.958244$ & $136.9298$ ms/view \\
$10\%$ & $196{,}608$ & $30.4278$ dB & $0.824789$ & $0.230503$ & $95.50\%$ & $0.958244$ & $135.8118$ ms/view \\
$5\%$ & $98{,}304$ & $28.0242$ dB & $0.746550$ & $0.391810$ & $96.50\%$ & $0.958244$ & $136.1801$ ms/view \\
$1\%$ & $19{,}661$ & $24.6923$ dB & $0.604307$ & $0.591024$ & $97.30\%$ & $0.958244$ & $136.6344$ ms/view \\
\bottomrule
\end{tabular}
\caption{DL3DV 3D-DCT residual Pareto sweep for the same locally cached scene. The preflight estimator reports $552{,}895{,}644$ bytes and measured peak memory stays at $215{,}813{,}648$ bytes for all sweep points. The constant edge-error column is a diagnostic: this sweep changes residual capacity only, not the sparse-tier Hankel locator.}
\label{tab:dl3dv-pareto}
\end{table}
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{@{}p{0.13\linewidth}p{0.11\linewidth}p{0.09\linewidth}p{0.09\linewidth}p{0.09\linewidth}p{0.11\linewidth}p{0.12\linewidth}p{0.11\linewidth}p{0.11\linewidth}@{}}
\toprule
Profile & Edge mode & PSNR & SSIM & LPIPS & Sparsity & Reproj. error & Edge delta & Time \\
\midrule
$25\%$ & scanline & $35.5073$ dB & $0.927023$ & $0.049832$ & $92.50\%$ & n/a & $0.958244$ & $138.6768$ ms/view \\
$25\%$ & hybrid & $33.6019$ dB & $0.892827$ & $0.064125$ & $92.50\%$ & $13.5410$ px & $0.624169$ & $67.8946$ ms/view \\
$10\%$ & scanline & $30.4278$ dB & $0.824789$ & $0.230503$ & $95.50\%$ & n/a & $0.958244$ & $139.8724$ ms/view \\
$10\%$ & hybrid & $29.2703$ dB & $0.792317$ & $0.198944$ & $95.50\%$ & $13.5410$ px & $0.624169$ & $67.3256$ ms/view \\
$25\%$ & hybrid strict CPA & $33.6076$ dB & $0.893035$ & $0.062962$ & $92.50\%$ & $14.3163$ px & $0.622570$ & $45.3286$ ms/view \\
\bottomrule
\end{tabular}
\caption{First DL3DV world-edge atom ablation. The hybrid mode activates ray-consistent projected knots and improves the scanline-referenced edge-delta diagnostic, but the reprojection error remains too high for a visual-quality gain. This table is included to document the structural progression and the current bottleneck, not as a final compression result.}
\label{tab:dl3dv-world-edge-ablation}
\end{table}
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{@{}p{0.24\linewidth}p{0.18\linewidth}p{0.12\linewidth}p{0.12\linewidth}p{0.12\linewidth}p{0.13\linewidth}@{}}
\toprule
Held-out profile & Training signal & PSNR & SSIM & LPIPS & Latency \\
\midrule
Linear neighbor & even views only & $22.7803$ dB & $0.611987$ & $0.150600$ & $2.9613$ ms \\
DCT view interpolation & even views only & $21.2918$ dB & $0.534591$ & $0.141890$ & $1.1113$ ms \\
Ray-kernel OSNR, $1024$ centers & even-view rays only & $18.0553$ dB & $0.466200$ & $0.904455$ & $600.8342$ ms \\
Ray-kernel OSNR, $4096$ centers & even-view rays only & $14.0241$ dB & $0.409003$ & $0.923252$ & $11142.5733$ ms \\
\bottomrule
\end{tabular}
\caption{Held-out DL3DV light-field generalization challenge. Metrics are computed only on odd-indexed views that were excluded from the fitting/input set. The ray-kernel rows evaluate held-out camera rays directly from $(\mathbf{o},\mathbf{d})$ coordinates, but do not yet include depth-aware visibility or epipolar-consensus geometry.}
\label{tab:dl3dv-heldout}
\end{table}
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{@{}p{0.28\linewidth}p{0.12\linewidth}p{0.12\linewidth}p{0.12\linewidth}p{0.12\linewidth}p{0.13\linewidth}@{}}
\toprule
Held-out geometry profile & PSNR & SSIM & LPIPS & Coverage & Latency \\
\midrule
Linear neighbor & $22.7803$ dB & $0.611987$ & $0.150600$ & n/a & n/a \\
Raw COLMAP colors, radius $1$ & $11.7702$ dB & $0.168650$ & $0.958861$ & $65.82\%$ & $242.5846$ ms \\
Training recolor, radius $1$ & $22.4994$ dB & $0.604704$ & $0.355278$ & $65.82\%$ & $242.3686$ ms \\
Training recolor, radius $1$, $35\%$ blend & $23.5950$ dB & $0.646699$ & $0.207789$ & $65.82\%$ & $242.3686$ ms \\
Training recolor, radius $1$, $45\%$ blend & $23.6510$ dB & $0.648680$ & $0.238721$ & $65.82\%$ & $276.9294$ ms \\
Training recolor, radius $1$, $55\%$ blend & $23.6146$ dB & $0.647230$ & $0.268017$ & $65.82\%$ & $242.3686$ ms \\
Training recolor, radius $1$, $3\times3$ smooth, $64\%$ blend & $23.9984$ dB & $0.684186$ & $0.285195$ & $65.82\%$ & $277.1769$ ms \\
Visible recolor, radius $1$, $3\times3$ smooth, $67\%$ blend & $24.1063$ dB & $0.686748$ & $0.293838$ & $65.82\%$ & $274.9366$ ms \\
Visible geometry + $20\%$ linear residual & $24.2137$ dB & $0.683271$ & $0.243624$ & $65.82\%$ & $293.4945$ ms \\
Plane sweep, $64$ depths, $0.4$--$12.0$, $5\%$ residual & $26.3525$ dB & $0.830533$ & $0.144778$ & dense & $4591.8559$ ms \\
Plane sweep v2, $192$ depths, soft, $13\times13$ patch & $29.1731$ dB & $0.893662$ & $0.109423$ & dense & $11749.4877$ ms \\
Plane sweep v3, $320$ depths, soft, $13\times13$ patch & $29.3178$ dB & $0.896597$ & $0.102406$ & dense & $20545.6590$ ms \\
Plane sweep final, $512$ depths, soft, $13\times13$ patch & $29.4187$ dB & $0.898331$ & $0.099257$ & dense & $35201.2051$ ms \\
Multi-view FRI edge bridge, $256$ depths & $29.8085$ dB & $0.913584$ & $0.070140$ & dense edge-consistent & $21481.0390$ ms \\
Plane sweep $+$ exact liquid residual, $320$ depths & $29.3994$ dB & $0.899609$ & $0.096436$ & dense + $32$ liquid states & $37285.2$ ms \\
Plane sweep $+$ exact liquid residual, $512$ depths & $29.5291$ dB & $0.901776$ & $0.093192$ & dense + $64$ liquid states & $106447.5$ ms \\
Nerfstudio \texttt{nerfacto} CPU pilot, $1000$ iters & $27.0830$ dB & $0.818526$ & $0.176857$ & trained MLP/hash grid & CPU pilot \\
\bottomrule
\end{tabular}
\caption{Geometry-backed held-out renderers on the same DL3DV split. Sparse COLMAP splats improve PSNR over interpolation but remain coverage-limited; the dense COLMAP-pose plane sweep is the first Track-A renderer to deliver a large held-out quality gain without neural training or autograd. The multi-view FRI edge bridge adds real-image derivative coherence to the deterministic depth score and becomes the strongest autograd-free geometry row. The liquid rows add a tiny trained residual corrector on top of the deterministic renderer and are therefore reported as hybrid quality-polishing experiments.}
\label{tab:dl3dv-colmap-heldout}
\end{table}
\begin{figure}[h]
\centering
\includegraphics[width=0.31\linewidth]{../apps_industrial_breakthrough/dl3dv_multiview_fri_bridge_outputs_256_d256/heldout_frame1_target.png}
\includegraphics[width=0.31\linewidth]{../apps_industrial_breakthrough/dl3dv_multiview_fri_bridge_outputs_256_d256/heldout_frame1_linear.png}
\includegraphics[width=0.31\linewidth]{../apps_industrial_breakthrough/dl3dv_multiview_fri_bridge_outputs_256_d256/heldout_frame1_multiview_edge_bridge.png}
\caption{Held-out DL3DV view-one comparison for the multi-view FRI edge bridge. Left: target odd view excluded from fitting. Middle: adjacent even-view interpolation. Right: deterministic edge-consistent geometry render using only even training views and COLMAP camera poses.}
\label{fig:dl3dv-multiview-fri-bridge}
\end{figure}
\begin{figure}[h]
\centering
\includegraphics[width=0.31\linewidth]{../apps_industrial_breakthrough/dl3dv_liquid_residual_outputs_512/heldout_frame1_target.png}
\includegraphics[width=0.31\linewidth]{../apps_industrial_breakthrough/dl3dv_liquid_residual_outputs_512/heldout_frame1_plane_sweep.png}
\includegraphics[width=0.31\linewidth]{../apps_industrial_breakthrough/dl3dv_liquid_residual_outputs_512/heldout_frame1_liquid.png}
\caption{Held-out DL3DV view-one comparison for the quality-first liquid residual run. Left: target odd view excluded from fitting. Middle: deterministic $512$-plane sweep. Right: plane sweep plus exact liquid residual correction trained only from even-view leave-one-out residuals.}
\label{fig:dl3dv-liquid-residual}
\end{figure}
\begin{figure}[h]
\centering
\includegraphics[width=0.31\linewidth]{../apps_industrial_breakthrough/dl3dv_pareto_outputs/profile_100p0pct/frame0_osnr.png}
\includegraphics[width=0.31\linewidth]{../apps_industrial_breakthrough/dl3dv_pareto_outputs/profile_050p0pct/frame0_osnr.png}
\includegraphics[width=0.31\linewidth]{../apps_industrial_breakthrough/dl3dv_pareto_outputs/profile_025p0pct/frame0_osnr.png}\\[-2pt]
\includegraphics[width=0.31\linewidth]{../apps_industrial_breakthrough/dl3dv_pareto_outputs/profile_010p0pct/frame0_osnr.png}
\includegraphics[width=0.31\linewidth]{../apps_industrial_breakthrough/dl3dv_pareto_outputs/profile_005p0pct/frame0_osnr.png}
\includegraphics[width=0.31\linewidth]{../apps_industrial_breakthrough/dl3dv_pareto_outputs/profile_001p0pct/frame0_osnr.png}
\caption{Representative view-zero reconstructions from the DL3DV Pareto sweep, ordered left-to-right and top-to-bottom by retained 3D-DCT support: $100\%$, $50\%$, $25\%$, $10\%$, $5\%$, and $1\%$. The corresponding target view is identical to the target in Table~\ref{tab:dl3dv-stack-frames}.}
\label{fig:dl3dv-pareto-view0}
\end{figure}
\begingroup
\scriptsize
\setlength{\tabcolsep}{2pt}
\renewcommand{\arraystretch}{1.05}
\begin{longtable}{@{}c c c@{}}
\caption{Frame-by-frame visual comparison for the $30$-view DL3DV stack quality-ceiling export. The columns are the observed target view and the OSNR reconstruction of the same view.}
\label{tab:dl3dv-stack-frames}\\
\toprule
View & Target & OSNR quality ceiling \\
\midrule
\endfirsthead
\toprule
View & Target & OSNR quality ceiling \\
\midrule
\endhead
000 &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/target_frame_000.png} &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/osnr_frame_000.png} \\
001 &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/target_frame_001.png} &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/osnr_frame_001.png} \\
002 &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/target_frame_002.png} &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/osnr_frame_002.png} \\
003 &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/target_frame_003.png} &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/osnr_frame_003.png} \\
004 &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/target_frame_004.png} &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/osnr_frame_004.png} \\
005 &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/target_frame_005.png} &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/osnr_frame_005.png} \\
006 &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/target_frame_006.png} &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/osnr_frame_006.png} \\
007 &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/target_frame_007.png} &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/osnr_frame_007.png} \\
008 &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/target_frame_008.png} &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/osnr_frame_008.png} \\
009 &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/target_frame_009.png} &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/osnr_frame_009.png} \\
010 &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/target_frame_010.png} &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/osnr_frame_010.png} \\
011 &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/target_frame_011.png} &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/osnr_frame_011.png} \\
012 &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/target_frame_012.png} &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/osnr_frame_012.png} \\
013 &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/target_frame_013.png} &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/osnr_frame_013.png} \\
014 &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/target_frame_014.png} &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/osnr_frame_014.png} \\
015 &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/target_frame_015.png} &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/osnr_frame_015.png} \\
016 &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/target_frame_016.png} &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/osnr_frame_016.png} \\
017 &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/target_frame_017.png} &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/osnr_frame_017.png} \\
018 &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/target_frame_018.png} &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/osnr_frame_018.png} \\
019 &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/target_frame_019.png} &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/osnr_frame_019.png} \\
020 &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/target_frame_020.png} &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/osnr_frame_020.png} \\
021 &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/target_frame_021.png} &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/osnr_frame_021.png} \\
022 &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/target_frame_022.png} &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/osnr_frame_022.png} \\
023 &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/target_frame_023.png} &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/osnr_frame_023.png} \\
024 &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/target_frame_024.png} &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/osnr_frame_024.png} \\
025 &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/target_frame_025.png} &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/osnr_frame_025.png} \\
026 &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/target_frame_026.png} &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/osnr_frame_026.png} \\
027 &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/target_frame_027.png} &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/osnr_frame_027.png} \\
028 &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/target_frame_028.png} &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/osnr_frame_028.png} \\
029 &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/target_frame_029.png} &
\includegraphics[width=0.42\linewidth]{../apps_industrial_breakthrough/dl3dv_sota_outputs_quality_all/frames/osnr_frame_029.png} \\
\bottomrule
\end{longtable}
\endgroup
\subsection{Application validation: biharmonic structural mechanics}
The structural-shell validation script, \texttt{apps/03\_structural\_shells/biharmonic\_plate.py}, uses the same 2D tensor-product Hermite machinery for a fourth-order clamped-plate operator
\[
\Delta^2 \Phi
=
\partial_{xxxx}\Phi
+2\partial_{xxyy}\Phi
+\partial_{yyyy}\Phi
=
f(x,y).
\]
The validation constructs a manufactured clamped deflection field, computes the load by an autograd-free fourth-order finite-difference ladder, maps the deflection into nine Hermite streams, applies the 2D block-circulant Gram, and recovers the coefficient tensor through parallel $9\times9$ Fourier-domain solves. Rigid plate edges are enforced by overwriting all boundary coefficient streams to zero.
On a $48\times48$ grid, the current run completes in $42.9853$ ms, reports boundary clamping residual $0.000000\mathrm{e}{+00}$, interior deflection RMS error $4.697917\mathrm{e}{-08}$, and relative biharmonic operator residual $1.065298\mathrm{e}{-04}$. The high maximum frequency-system condition number, approximately $1.8757\mathrm{e}{09}$, identifies the expected low-frequency stiffness of fourth-order tensor-product Gram systems and motivates more specialized biharmonic preconditioning before external structural-mechanics comparisons.
\subsection{SOTA comparison: SIREN versus adaptive sparse OSNR}
The first comparative benchmark, \texttt{benchmarks\_sota/compare\_siren\_sdf.py}, evaluates a standard sinusoidal representation network against the Tier 2 adaptive sparse OSNR solver on an identical non-bandlimited geometric target. The target is a two-dimensional silhouette with high-frequency wavy boundaries and step discontinuities along each scanline. It is intentionally hostile to smooth coordinate MLPs because the field is not bandlimited and its boundary locations fall between grid samples.
The SIREN baseline follows the Sitzmann et al. implicit representation pattern: a fully parameterized multilayer perceptron maps coordinates $(x,y)$ to occupancy values through sinusoidal hidden layers. The benchmark uses a $64\times192$ coordinate grid, a hidden width of $64$, three hidden sine layers, $\omega_0=30$, and Adam optimization for $1000$ full-batch epochs. This produces a dense model with $12{,}737$ trainable weights. Its final prediction is thresholded to estimate boundary locations, yielding $24.32$ dB PSNR and edge blurring error $4.803569\mathrm{e}{-03}$ after $5{,}622.11$ ms of optimization in the current rerun.
The OSNR path uses the same target samples but does not optimize a coordinate network. Each scanline is encoded as a finite-rate-of-innovation signal with two step horizons. The TLS matrix-pencil pre-filter recovers those continuous edge coordinates from moments, the sparse knot frame is snapped to the recovered horizons, and the cross-Gram-shielded sparse solver debiases the active shock atoms with scale-invariant Tikhonov stabilization. The OSNR pass runs under \texttt{torch.no\_grad()}, uses $128$ active sparse knots across the batch, hard-zeros $97.9\%$ of the sparse parameter tensor, and achieves $111.89$ dB PSNR with edge localization error $1.443290\mathrm{e}{-15}$ in $19.23$ ms.
A third path, \texttt{apps\_industrial\_breakthrough/osnr\_operator\_atlas\_rank\_probe.py}, tests the broader no-backprop idea inspired by local random-feature FBPINNs and rank-revealing feature filtering. It does not use the exact FRI edge moments. Instead, it covers the same field with $12\times12$ and $24\times24$ partition-of-unity charts, evaluates frozen local polynomial, DCT, and SIREN probe features, greedily keeps locally independent rank directions, orthogonalizes the retained chart directions, and solves one global ridge system. The best fixed-grid pilot keeps $10{,}080$ of $32{,}400$ local candidate directions, drops $68.9\%$ of the atlas, and reaches $35.318555$ dB PSNR with edge error $2.642476\mathrm{e}{-03}$ in $49{,}629.10$ ms for the unoptimized Python prototype. This row is not a substitute for the exact FRI solver; it is evidence for the more general operator-atlas thesis: local frozen features plus rank-revealed algebra can beat a trained global SIREN even when the exact sparse innovation coordinates are not supplied.
The next adaptive variant, \texttt{apps\_industrial\_breakthrough/osnr\_operator\_atlas\_adaptive\_refine.py}, removes the hand-picked dense fine grid. A coarse $8\times8$ pilot atlas scores a $40\times40$ candidate chart lattice by residual energy, target-gradient energy, and rank density. The solver then keeps only $641$ refined charts after one-cell dilation, evaluates frozen polynomial, DCT, SIREN, and curved local edge-step atoms, applies target-aware local rank filtering, and solves one global ridge system. The condition-diagnostic run keeps $8{,}820$ of $131{,}130$ local candidate atoms, drops $93.27\%$ of the candidate atlas, and reaches $86.360935$ dB PSNR with edge error $2.615928\mathrm{e}{-03}$. Two seed repeats reach $89.188964$ dB and $84.082762$ dB, respectively. Thus the result is not a one-seed random-feature accident. It also improves the fixed rank-revealed atlas by $51.04$ dB while retaining fewer directions. The remaining gap to the exact FRI row is expected: the FRI row is given the exact sparse innovation model, whereas the adaptive atlas only receives samples and a frozen local operator dictionary.
The same adaptive runner now includes a global eigentruncated right-preconditioned solve. Instead of trusting the full ridge normal system, it diagonalizes the global atlas Gram matrix, removes directions below a relative eigenthreshold, and solves in the retained eigenspace. On the silhouette target, the publication-friendly threshold $3\times10^{-11}$ keeps $86.093140$ dB while reducing the effective condition from $5.389\times10^{12}$ to $3.235\times10^{10}$, a $166\times$ reduction, with retained eigenspace rank $6{,}778$. A more aggressive $10^{-8}$ threshold still reaches $84.601181$ dB while reducing the effective condition to $9.841\times10^7$, about $5.48\times10^4$ lower than the raw system. Thus the adaptive atlas result survives explicit right-preconditioned rank truncation rather than depending on hidden nearly-null directions.
The dense eigentruncation is now cross-checked by a randomized projected eigensolve. With Gaussian sketching, one subspace iteration, and projected dimension $6{,}714$, the randomized solve reaches $84.601150$ dB with the same edge error $2.615928\mathrm{e}{-03}$ and retained Ritz rank $6{,}669$, matching the dense aggressive row to within $3.1\times10^{-5}$ dB. A projected dimension of $6{,}592$ still reaches $82.817051$ dB. More aggressive compression is not free: projected dimensions $6{,}464$, $6{,}208$, and $4{,}160$ reach $75.836651$, $63.725327$, and $20.082664$ dB, respectively. This negative boundary is useful: the atlas win is broad-rank and rank-revealed, not a tiny hidden low-rank shortcut.
The global solve can also avoid dense normal-matrix formation. A row-block Jacobi-preconditioned CG mode applies the field/operator design only through matrix-vector products, with optional weighted row sampling. On the silhouette target, full-row PCG reaches $84.500732$ dB with edge error $2.615928\mathrm{e}{-03}$ after $6{,}400$ iterations, within $0.10$ dB of the dense aggressive eigentruncated row while never forming the dense Gram matrix. Fixed-policy replay separates solver effects from adaptive chart drift: replayed row-norm sampling at $10{,}000$ of $12{,}288$ rows keeps $80.457415$ dB and edge error $2.615923\mathrm{e}{-03}$, whereas coarse spatial block and stratified schedules collapse on this discontinuity target. Thus row scheduling must respect the operator/objective geometry; it is not a generic block-dropping problem.
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{@{}p{0.18\linewidth}p{0.22\linewidth}p{0.18\linewidth}p{0.16\linewidth}p{0.2\linewidth}@{}}
\toprule
Architecture & Duration & Parameters & Sparsity & PSNR / edge error \\
\midrule
SIREN MLP & $5{,}622.11$ ms / $1000$ epochs & $12{,}737$ dense weights & $0.0\%$ hard zeros & $24.32$ dB / $4.803569\mathrm{e}{-03}$ \\
Rank-revealed OSNR atlas & $49{,}629.10$ ms / prototype solve & $10{,}080$ retained directions & $68.9\%$ directions dropped & $35.32$ dB / $2.642476\mathrm{e}{-03}$ \\
Adaptive OSNR atlas & $34{,}841.52$ ms / diagnostic solve & $8{,}820$ retained directions & $93.27\%$ atoms dropped & $86.36$ dB / $2.615928\mathrm{e}{-03}$ \\
OSNR Tier 2 & $19.23$ ms / single pass & $128$ active knots & $97.9\%$ hard zeros & $111.89$ dB / $1.443290\mathrm{e}{-15}$ \\
\bottomrule
\end{tabular}
\caption{First SOTA-style comparison on a non-bandlimited multi-edge silhouette. The SIREN row reports trained coordinate-network performance after $1000$ Adam epochs. The rank-revealed atlas rows are no-backprop local-feature prototypes that do not use the exact FRI edge moments; the adaptive row selects charts by pilot residual/rank maps and uses local discontinuity atoms. The OSNR Tier 2 row reports the matched FRI-snapped sparse solver, which remains the exact-structure oracle for this target.}
\label{tab:siren-osnr}
\end{table}
As a first non-silhouette cross-check, the same adaptive runner now supports a mixed SPDE target generated by a smooth, Gaussian, and sparse-event innovation passed through a periodic advection-diffusion-reaction inverse. In this setting edge atoms are actively harmful, which is the expected guardrail for a smooth operator field. A field-only adaptive atlas with polynomial/DCT/SIREN atoms reaches $56.786560$ dB versus a $512$-feature global DCT baseline at $46.317344$ dB, but its operator relative RMSE remains $0.982192$. Adding operator rows to the algebraic normal equation,
\[
\min_c \|A c-u\|_2^2+\lambda^2\|\mathcal{L}A c-\mathcal{L}u\|_2^2+\gamma\|c\|_2^2,
\]
improves the best mixed-SPDE point to $57.140339$ dB with operator relative RMSE $0.209998$, compared with global DCT operator relative RMSE $2.773887$. This is a positive second validation of the adaptive atlas idea outside the silhouette benchmark, while also exposing the next numerical issue: the operator-augmented normal system is ill-conditioned, with diagnostic condition estimate $4.805\times10^{13}$, so RRQR/right-preconditioned block solves are the next required improvement before making external PDE benchmark claims.
The eigentruncated solve materially improves that numerical story. With relative eigenthreshold $10^{-8}$, the mixed-SPDE atlas keeps $56.927949$ dB and operator relative RMSE $0.212386$ while reducing the effective condition from $4.805\times10^{13}$ to $9.860\times10^7$. Only $2.24\%$ of global eigendirections are removed, indicating that the SPDE instability is concentrated in a small global null-like subspace. This is still a dense diagnostic solve, not yet a scalable PDE production method, but it validates the intended RRQR/right-preconditioning direction.
The randomized projected solve also preserves the operator-aware SPDE result. At projected dimension $7{,}404$, it reaches $56.911218$ dB with operator relative RMSE $0.212675$, essentially matching the dense eigentruncated row. At projected dimension $7{,}164$, it still reaches $56.343715$ dB and operator relative RMSE $0.217228$, while projected dimension $6{,}208$ collapses to $13.261593$ dB and operator relative RMSE $3.790479$. Thus the next scaling target is not smaller global rank alone; it is matrix-free or block-randomized least squares that avoids full Gram formation while preserving the broad well-conditioned Ritz subspace.
The no-dense-normal PCG path gives the same conclusion on the SPDE target. With all $24{,}576$ augmented field/operator rows, PCG reaches $56.984770$ dB and operator relative RMSE $0.210942$ after $6{,}400$ iterations, slightly stronger field PSNR than the dense eigentruncated row. Replayed row schedules then identify the correct compression geometry. Independent row-norm sampling at $20{,}000$ rows preserves field PSNR, $56.967060$ dB, but degrades operator relative RMSE to $0.709043$; uniform, spatial-stratified, equal field/operator quota, and field-full/operator-sampled controls also fail to preserve both objectives. The positive schedule keeps all operator rows exactly and samples only the field rows by row norm. At $20{,}000$ of $24{,}576$ rows it reaches $57.600055$ dB and operator relative RMSE $0.209357$, slightly beating the full-row PCG anchor while using $18.6\%$ fewer augmented rows. The same operator-shell schedule remains strong at $18{,}000$ rows ($57.153551$ dB, $0.209805$), $16{,}000$ rows ($56.930231$ dB, $0.215080$), $14{,}000$ rows ($55.953736$ dB, $0.210489$), and $13{,}000$ rows ($55.315051$ dB, $0.210262$). The $12{,}288$-row operator-only cliff collapses to $7.165436$ dB and operator relative RMSE $1.656706$, showing that the sampled field equations are nullspace anchors for the operator shell rather than expendable data rows.
The same result now survives streamed atlas-column construction. In \texttt{streamed\_row\_sketch\_pcg} mode, retained atlas directions are stored as chart-local support blocks, and the solver evaluates $A v$, $A^\top v$, and $A^\top \mathcal{L}^\ast\mathcal{L} A v$ without materializing either the dense field design $A$ or the dense operator design $\mathcal{L}A$. A smoke run matches materialized PCG to within $3.91\times10^{-8}$ dB PSNR and $7.56\times10^{-9}$ operator relative RMSE. On the replayed SPDE policy, streamed full-row PCG reaches $56.991617$ dB and operator relative RMSE $0.210990$ in $17.109$ s, compared with $56.984770$ dB and $0.210942$ in $77.045$ s for the materialized full-row run. Streamed operator-shell PCG keeps the compressed frontier: $20{,}000$ rows reaches $57.598052$ dB and $0.209398$ in $16.830$ s, while $13{,}000$ rows reaches $55.321231$ dB and $0.210299$ in $17.831$ s; the streamed $12{,}288$-row operator-only cliff still collapses to $7.165883$ dB and $1.656653$. Thus the SPDE atlas result is now no-dense-Gram and no-dense-design in the global solve. The next scaling step is larger streamed operator-atlas validation, not blind row dropping.
\subsection{External PDEBench Darcy sparse OSNR assimilation}
The first external Darcy result is deliberately framed as sparse-observation assimilation rather than a blind coefficient-to-solution solver claim. On the real \texttt{PDEBench\_2D\_DarcyFlow\_beta0.01} shard, a hand-coded finite-volume elliptic bridge maps each coefficient field to a solution shape, but high low-conductivity inclusion regimes expose a hidden amplitude/interface convention. The previous scalar sparse-sensor audit showed that a few target observations can calibrate the dominant amplitude. The new runner, \texttt{pdebench\_darcy\_osnr\_sparse\_residual\_ladder.py}, asks a stronger question: can sparse observations also identify a compact OSNR residual dictionary on top of the PDE-shaped solution,
\[
\widehat u(x)=\alpha u_{\mathrm{CG}}(x)+\beta+\sum_{k=1}^{K} c_k \phi_k(x),
\]
where $\phi_k$ are low-frequency DCT residual modes. All ranks, ridges, and sensor policies are selected on the training split and reported once on $300$ held-out Darcy fields.
The first ladder uses $1000$ samples, $700$ for training, $240$ of those for hyperparameter selection, sensor budgets from $1$ to $128$, residual ranks $0,2,4,8,16,32$, and random, grid, training-residual-variance, and energy sensor policies. The blind finite-volume CG bridge has held-out mean nRMSE $0.3262416$; the full-field scalar oracle, which uses all target pixels only to choose one amplitude, has mean nRMSE $0.0585205$. The first deployable sparse row uses only $64$ fixed grid sensors out of $128^2$ pixels ($0.390625\%$ of the field), selects the PDE shape plus a rank-$32$ residual dictionary with ridge $10^{-4}$, and reaches mean nRMSE $0.0153122$, median $0.0122784$, p90 $0.0325075$, and max $0.0501816$.
The active-design follow-up keeps the same train/held-out split but chooses additional sensor locations by the leverage geometry of the PDE+DCT feature system. Policies such as \texttt{grid\_dopt128} first allocate a coarse grid prefix, then greedily add points by target-independent D-optimal posterior leverage. They use the coefficient field, the CG solution, and the frozen residual dictionary, but not unobserved target residuals. With residual ranks $96$ and $128$, the fixed grid already breaks the $0.01$ barrier at $256$ sensors, reaching mean nRMSE $0.0097030$. The best active row, \texttt{grid\_dopt128}, reaches $0.0088075$ at $256$ sensors and $0.0086891$ at $384$ sensors. Thus the external sparse-assimilation frontier moves from ``below the scalar oracle'' to a sub-$10^{-2}$ held-out PDEBench error with no neural retraining.
The neural follow-up then performs the comparison that this result demands. The runner \texttt{pdebench\_darcy\_neural\_sparse\_assimilation\_baseline.py} trains a U-Net sparse assimilator on the same $700$ training fields. Its inputs are the coefficient field, CG base, same-sensor scalar-calibrated base, sparse observed target values, sparse residual values, the binary observation mask, and coordinate channels. After the neural prediction, the same OSNR residual adapter is fitted from the same sparse observations, but now around the neural field rather than the CG field:
\[
\widehat u(x)=\alpha u_{\mathrm{neural}}(x)+\beta+\sum_{k=1}^{K} c_k \phi_k(x).
\]
This makes the claim harder: OSNR must improve an already trained sparse neural assimilator rather than only beat a scalar or DCT control.
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{@{}p{0.28\linewidth}p{0.16\linewidth}p{0.18\linewidth}p{0.15\linewidth}p{0.15\linewidth}@{}}
\toprule
Method & Sensors & Model & Mean nRMSE & Median nRMSE \\
\midrule
Blind finite-volume CG & $0$ & Darcy PDE bridge & $0.3262416$ & $0.1820685$ \\
Full-field scalar oracle & all pixels & scalar amplitude only & $0.0585205$ & $0.0313853$ \\
Grid sparse scalar & $64$ & scalar amplitude only & $0.0585245$ & $0.0313892$ \\
Grid plain DCT & $64$ & rank-$32$ DCT only & $0.1174098$ & $0.1227564$ \\
Grid sparse OSNR & $64$ & PDE + rank-$32$ DCT residual & $0.0153122$ & $0.0122784$ \\
Grid sparse OSNR & $32$ & PDE + rank-$16$ DCT residual & $0.0294030$ & $0.0196162$ \\
Grid sparse OSNR & $16$ & PDE + rank-$8$ DCT residual & $0.0387592$ & $0.0245881$ \\
Random sparse OSNR & $128$ & PDE + rank-$32$ DCT residual & $0.0202490$ & $0.0161723$ \\
Variance sparse OSNR & $128$ & PDE + rank-$32$ DCT residual & $0.0164891$ & $0.0125584$ \\
Grid sparse OSNR & $256$ & PDE + rank-$128$ DCT residual & $0.0097030$ & $0.0069747$ \\
Grid-D-opt sparse OSNR & $256$ & PDE + rank-$128$ DCT residual & $0.0088075$ & $0.0062498$ \\
Grid-D-opt sparse OSNR & $384$ & PDE + rank-$128$ DCT residual & $0.0086891$ & $0.0061179$ \\
Grid-D-opt sparse U-Net & $256$ & trained neural assimilator & $0.0080866$ & $0.0058787$ \\
Grid-D-opt U-Net + OSNR & $256$ & neural + rank-$32$ adapter & $0.0063245$ & $0.0045115$ \\
Grid-D-opt U-Net + OSNR & $384$ & neural + rank-$64$ adapter & $0.0072188$ & $0.0059974$ \\
Budgeted Grid-D-opt U-Net + OSNR & $256$ & wider neural + rank-$32$ adapter & $\mathbf{0.0059962}$ & $\mathbf{0.0047928}$ \\
\bottomrule
\end{tabular}
\caption{External PDEBench Darcy sparse-observation assimilation. The PDE-shaped residual model is selected on the training split. Plain DCT-only interpolation is a negative mechanism control; it is much worse than the PDE-shaped residual row, showing that the elliptic operator bridge supplies the dominant field prior. Grid-D-opt policies are target-independent active measurement designs based on PDE+DCT feature leverage. The U-Net rows are trained sparse assimilators under the same train/held-out split; the OSNR adapter is then fitted from the same sparse observations at test time. The final budgeted row uses a separately budgeted target-independent $256$-sensor design rather than the nested $256$-prefix from the joint $256,384$ sweep.}
\label{tab:pdebench-darcy-sparse-osnr}
\end{table}
The improvement is strongest exactly where the blind bridge was weakest. In the high-inclusion bin, $\mathrm{low\_fraction}\geq0.75$ ($44$ held-out fields), blind CG has mean nRMSE $0.8641549$ and the full-field scalar oracle has $0.1287972$. The $64$-sensor rank-$32$ PDE+DCT residual row reduces this to $0.0274448$, and the $384$-sensor grid-D-opt rank-$128$ row reduces it further to $0.0149138$. The neural adapter pushes the same hard bin to $0.0107779$ at $384$ sensors and $0.0101835$ in the focused $256$-sensor run. Mid/high bins show the same pattern: for low-fraction $0.50$--$0.75$, the rank-$128$ CG+OSNR row improves $0.4736753$ blind and $0.0900559$ scalar-oracle nRMSE to $0.0120798$, while the focused neural+OSNR row reaches $0.0076549$; for $0.25$--$0.50$, it reaches $0.0037485$.
This result is the first external PDEBench row in the manuscript where OSNR beats the scalar-oracle ceiling rather than only calibrating amplitude, and the neural follow-up changes the status of the claim. OSNR is no longer only a standalone no-retraining sparse assimilator; it is also a test-time correction layer that improves a trained neural sparse assimilator. In the main MPS run, the $256$-sensor sparse U-Net reaches mean nRMSE $0.0080866$, while U-Net+OSNR reaches $0.0063245$. The focused $256$-sensor run reaches $0.0059962$ mean nRMSE, median $0.0047928$, p90 $0.0119379$, and max $0.0161753$, using only $1.5625\%$ of pixels. The scientific claim remains sparse-observation assimilation rather than blind coefficient-to-solution neural-operator SOTA, but the mechanism is now more general: OSNR can operate both as the primary PDE-shaped residual solver and as a plug-in residual adapter on top of a learned neural prior.
\subsection{SOTA comparison: Spline-PINN regime versus tensor-product OSNR CFD}
The second comparative benchmark, \texttt{benchmarks\_sota/compare\_spline\_pinn\_cfd.py}, targets the fluid-surrogate regime studied by Wandel et al. for Spline-PINN. The script uses a DFG-style cylinder domain with a $41\times220$ spatial layout and compares the reported Spline-PINN training regime against the measured Tier 3 OSNR tensor-product Hermite path. The Spline-PINN row is therefore not a rerun of the authors' training code; it is an explicit reported-regime reference capturing the relevant structural cost: a spline-interpolated U-Net update model trained with physics-informed losses and data recycling over one to two days, with real-time inference reported at approximately $30$ updates per second.
The OSNR path uses the same grid and obstacle geometry but does not train a time-step network. A stream function $a_z$ generates the velocity field by the hard incompressible curl map
\[
v_x = \partial_y a_z,\qquad
v_y = -\partial_x a_z,
\]
and the cylinder no-slip boundary is enforced by overwriting the nine tensor-product Hermite coefficient channels at masked vertices. The nonlinear advection and viscous diffusion terms are evaluated through the forward Hermite derivative ladder, and the coefficient update is resolved through the 2D block-circulant Fourier solver. The measured OSNR rows run under \texttt{torch.no\_grad()} with zero autograd graph allocation.
For Reynolds-number calibration, the benchmark uses
\[
\operatorname{Re}=\frac{\rho U D}{\mu},
\qquad
\rho=1,\quad U=0.20,\quad D=10\ \text{grid cells}.
\]
This gives $\mu=0.100000$ for $\operatorname{Re}=20$ and $\mu=0.020000$ for $\operatorname{Re}=100$. Boundary leakage is measured directly on the obstacle mask after coefficient overwriting. The divergence residual is the root-mean-square divergence of the reconstructed velocity field on the discrete validation grid.
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{@{}p{0.22\linewidth}p{0.19\linewidth}p{0.15\linewidth}p{0.15\linewidth}p{0.14\linewidth}p{0.1\linewidth}@{}}
\toprule
Architecture & Initialization/training & Frame latency & Boundary leakage & $\nabla\cdot v$ RMS & Autograd memory \\
\midrule
Spline-PINN mock, $\operatorname{Re}=20$ & $1$--$2$ days reported & $\approx 33.33$ ms reported & boundary-loss dependent & vector-potential hard constraint & training graph required \\
OSNR Tier 3, $\operatorname{Re}=20$ & $0.00$ ms & $6.7893$ ms & $0.000000\mathrm{e}{+00}$ & $1.774261\mathrm{e}{-02}$ & $0.00$ B \\
Spline-PINN mock, $\operatorname{Re}=100$ & $1$--$2$ days reported & $\approx 33.33$ ms reported & boundary-loss dependent & vector-potential hard constraint & training graph required \\
OSNR Tier 3, $\operatorname{Re}=100$ & $0.00$ ms & $6.9435$ ms & $0.000000\mathrm{e}{+00}$ & $1.825696\mathrm{e}{-02}$ & $0.00$ B \\
\bottomrule
\end{tabular}
\caption{SOTA-style CFD comparison on a DFG-style cylinder grid. The Spline-PINN rows summarize the reported training and inference regime; the OSNR rows are measured package outputs from the tensor-product Hermite CFD benchmark.}
\label{tab:splinepinn-osnr-cfd}
\end{table}
\subsection{SOTA comparison: classic PINN versus operator-spline boundary solver}
The third comparative benchmark, \texttt{benchmarks\_sota/compare\_classic\_pinn.py}, isolates the cost of high-order automatic differentiation in the classical physics-informed neural-network formulation. The validation problem is the fourth-order boundary-value system
\[
u^{(4)}(x)=(2\pi)^4\sin(2\pi x),
\qquad x\in[0,1],
\]
with strict Dirichlet and curvature constraints
\[
u(0)=u(1)=0,\qquad u''(0)=u''(1)=0.
\]
The exact interior solution is $u(x)=\sin(2\pi x)$, which satisfies both the boundary values and the curvature clamps.
The PINN baseline follows the Raissi et al. pattern: a deep fully connected tanh network is trained with Adam on a joint physics-plus-boundary objective. The physics loss is evaluated by repeated backward-mode automatic differentiation through the network to obtain $u^{(4)}(x)$, while the boundary loss separately differentiates the boundary predictions to obtain $u''(0)$ and $u''(1)$. The benchmark uses $2000$ optimization epochs. On the CPU validation run, where PyTorch does not expose a global peak autograd allocator analogous to CUDA peak memory, the script reports a conservative graph-footprint estimate built from the derivative tapes, layer activations, parameters, gradients, and Adam state tensors.
The OSNR path uses the same spatial dimension but replaces the learned function with a calibrated knot grid and a Fourier biharmonic symbol inversion. In the periodized operator basis, the fourth derivative is diagonalized by the Fourier symbol $(2\pi\nu)^4$, so the coefficient recovery is a single element-wise division in the frequency domain. Boundary constraints are then represented as hard coefficient-layer constraints rather than soft penalties. The complete OSNR segment runs under \texttt{torch.no\_grad()} and allocates no autograd graph.
\begin{table}[h]
\centering
\scriptsize
\begin{tabular}{@{}p{0.2\linewidth}p{0.24\linewidth}p{0.2\linewidth}p{0.16\linewidth}p{0.12\linewidth}@{}}
\toprule
Architecture & Training/solve duration & Peak autograd graph memory & Boundary leakage & Interior PSNR \\
\midrule
Classic PINN & $8.0585$ s / $2000$ epochs & $1{,}430{,}560$ B & $9.771605\mathrm{e}{-02}$ & $25.76$ dB \\
OSNR Tier 1+3 & $0.000336$ s / single pass & $0.00$ B & $0.000000\mathrm{e}{+00}$ & $313.02$ dB \\
\bottomrule
\end{tabular}
\caption{Classic PINN comparison on a fourth-order boundary-value problem. The PINN row measures iterative tanh-network training with repeated fourth-derivative autograd; the OSNR row measures the calibrated operator-spline Fourier inversion.}
\label{tab:classic-pinn-osnr}
\end{table}
The absolute PSNR values are extremely high because these are controlled algebraic verification problems with exact synthetic data and matched model assumptions. Future external comparisons must include noisy measurements, non-exact operators, multidimensional fields, and standardized PINN/SIREN baselines.
\subsection{Operator-compiled constitutive KANs}
\label{sec:operator-compiled-kan}
The edge-function idea of a Kolmogorov--Arnold network becomes more useful for
operator learning when the learnable function is placed at the constitutive
uncertainty, rather than at the entire PDE right-hand side. Consider
\[
u_t=\nu u_{xx}-\partial_x F(u),
\qquad
F_c(u)=\sum_{j=1}^{J}c_j\phi_j(u).
\]
The chain rule compiles the unknown flux into a linear coefficient problem,
\[
u_t-\nu u_{xx}
=-F_c'(u)u_x
=\sum_{j=1}^{J}c_j[-\phi_j'(u)u_x].
\]
Consequently, coefficient identification is one regularized least-squares
solve even though the resulting PDE is nonlinear in $u$. At rollout we
evaluate $F_c(u)$ and apply the discrete spectral derivative to the complete
flux, rather than separately sampling the chain-rule factors. On a periodic
grid this gives
\[
\frac{\dd}{\dd t}\sum_n u_n=0
\]
up to floating-point roundoff for every learned coefficient vector. The
resolution-sensitive differential operator is never approximated by the
network.
Uniform cubic cardinal B-splines provide local adaptation and a matrix--vector
evaluation path, but compact support creates an unavoidable amplitude-
extrapolation ambiguity. If training states occupy only an interval
$I_{\rm tr}$, coefficients whose supports lie outside $I_{\rm tr}$ are not
identified; minimum-norm fitting makes the represented flux flatten outside
the observed interval. A global carrier is therefore not an implementation
detail but an identifiability requirement. We test polynomial carriers and a
sparse exponential-polynomial atlas
\[
\mathcal A=\{u,u^2,u^3,\sin(\omega u),\cos(\omega u):
\omega=1,\ldots,6\}.
\]
These atoms are the null-space functions associated with repeated zero poles
and conjugate imaginary poles. Cardinal exponential splines reproduce the
same spaces; the present experiment operates directly in the reproduction
space and is therefore a pole-discovery/compiler test, not yet a compact
E-spline implementation.
The matrix--vector qualification is operational, not merely asymptotic. If
$z=(u-u_0)/h=j+t$ with $t\in[0,1)$, a centered cardinal cubic edge is
evaluated from only four adjacent coefficients by
\[
F_c(u)=
\begin{bmatrix}1&t&t^2&t^3\end{bmatrix}
\frac{1}{6}
\begin{bmatrix}
1&4&1&0\\[-1mm]
-3&0&3&0\\
3&-6&3&0\\
-1&3&-3&1
\end{bmatrix}
\begin{bmatrix}c_{j-1}&c_j&c_{j+1}&c_{j+2}\end{bmatrix}^{\!\top}.
\]
Thus a layer is a batched gather followed by a fixed small matrix contraction;
neither Cox--de Boor recursion nor a dense all-knot basis tensor is needed.
The same compilation extends to objectives. Badoual, Schmitter, and Unser's
periodic inner-product calculus gives
\[
\langle F_c,F_d\rangle_{L_2}=c^\top A d,
\qquad
A_{k\ell}=\langle\phi(\cdot-k),\phi(\cdot-\ell)\rangle,
\]
and derivative energies replace $A$ by precomputed derivative cross-Grams.
For an equal periodic cardinal grid these matrices are circulant. A compact
cubic mass-matrix application is therefore a seven-tap $O(J)$ stencil, while
regularized inversion and broader composite operators are diagonalized by the
DFT. The Hermite construction of Appendix~\ref{app:hermite-block-gram} is the
multichannel version of exactly this identity. Consequently cardinality
provides two separate accelerators: a fixed local evaluation kernel for neural
edges and exact coefficient-space calculus for training losses and operator
solves.
A dedicated CPU benchmark tests both claims against independent paths. For
8,192--32,768 edge outputs and 16, 32, and 64 knots, the local cubic matrix
kernel agrees with generic vectorized Cox--de Boor evaluation to at worst
$1.58\times10^{-16}$ and is approximately $4.8\times$--$12.2\times$ faster
across the sweep. Its gathered payload remains four values per edge,
whereas the recursive dense basis workspace grows with the knot count. For
64--4096 periodic coefficients, compact-stencil and FFT Gram products agree
to at worst $1.09\times10^{-15}$, the coefficient bilinear form agrees with
32-point-per-cell continuous quadrature to at worst $2.67\times10^{-8}$, and
FFT ridge solves have relative residual at most $1.03\times10^{-15}$. At
$J=4096$, a dense Gram alone is 128 MiB, compared with approximately 0.063 MiB
for its real kernel and complex half-spectrum. The timing boundary is useful:
the seven-tap direct stencil is faster than FFT at the largest tested compact
case, so FFT should be reserved for inversion or noncompact/composite
circulant operators rather than applied dogmatically. These are NumPy CPU
prototype timings. An unfused PyTorch CPU control is negative: although the dense explicit-cardinal
and local paths agree to $2.33\times10^{-6}$ in float32, composing gather,
mask, and reduction primitives makes the local path $1.95\times$--$4.31\times$
slower in the forward pass and $3.44\times$--$7.68\times$ slower for
forward--backward. The direct KAN therefore keeps the dense explicit-cardinal
contraction as its CPU default and exposes the local matrix path as an opt-in
reference. Apple MPS exposes the predicted resolution crossover even without
a custom kernel. In a synchronized sweep with batch 1024, 32 inputs, 64
outputs, and 16--512 knots per edge, a memory-bounded implementation streams
the four taps instead of materializing a batch--input--tap--output tensor.
Forward evaluation first beats the dense explicit-cardinal path at 128 knots;
forward--backward first wins at 256 knots. At 512 knots, streamed evaluation
is $4.23\times$ faster forward and $2.95\times$ faster forward--backward,
while its explicit workspace is 8 MiB rather than the dense basis's 64 MiB.
The two paths agree to relative $1.44\times10^{-5}$ in float32 at this finest
grid. The direct KAN now selects the streamed path automatically on MPS from
256 knots and retains the dense path below the crossover. A truly fused
Metal/Triton kernel remains valuable because it should lower the crossover
into the small-grid regime; the present result already establishes a measured
high-resolution GPU training-throughput advantage.
The same inner-product compiler removes another KAN-specific failure mode:
changing grid resolution need not change the represented edge. For source
basis $\widetilde\Phi$ and target basis $\Phi$, the continuous $L_2$ projection
is the precomputable map
\[
c=A^{-1}\widetilde A\widetilde c,
\qquad
A=\langle\Phi,\Phi\rangle,
\quad
\widetilde A=\langle\Phi,\widetilde\Phi\rangle.
\]
This is the resampling identity in Badoual et al. and in the accompanying
E-snake derivation, which further gives a short characteristic-polynomial
description of a special exponential-spline circulant inverse. We test the
projection identity for periodic cubic cardinal edges using four-point Gauss
integration on the union knot partition, which is exact for the piecewise
degree-six cross-products. Nested refinements $24\to48$ and $24\to96$
preserve the continuous edge to worst relative $L_2$ error
$1.71\times10^{-15}$, whereas periodic coefficient interpolation and using
new-knot values as coefficients incur $0.95\%$--$2.66\%$ error. The
$24\to96\to24$ coefficient round trip has relative error
$1.18\times10^{-15}$. A nonnested $48\to72$ projection has $0.060\%$
median error versus $0.53\%/0.77\%$ for the two heuristics. For
$64\to32,24,16$ coarsening, the exact projection reduces median continuous
error by $1.8\times$--$2.6\times$ and preserves the integral to
$1.39\times10^{-16}$, while heuristic integral drift reaches $2.76\%$.
After matrix precomputation, 32-edge projection solves take 15--67 $\mu$s on
CPU. Thus cardinal grid extension can be exactly function preserving when
the spaces are nested and $L_2$ optimal otherwise. Implementing the
E-spline-specific reciprocal-root inverse as an IIR/parallel-scan kernel is a
separate gate, which we close next.
The E-snake Gram has the more specific real symmetric circulant form
\[
A=pI+q(S+S^{-1})+r(S^2+S^{-2}).
\]
Consequently $A^{-1}$ must itself be symmetric circulant: if $g$ is its first
row, $g_k=g_{M-k}$. Let $z_1,z_2$ be the two roots inside the unit disk of
$rz^4+qz^3+pz^2+qz+r$, and set $\gamma=r/(z_1z_2)$. Reciprocal pairing gives
\[
A=\gamma\prod_{i=1}^2(I-z_iS)(I-z_iS^{-1}).
\]
Partial fractions therefore produce the explicitly symmetric periodic Green
kernels
\[
H_z[k]=\frac{z^k+z^{M-k}}{(1-z^2)(1-z^M)},\qquad
g_k=a_1H_{z_1}[k]+a_2H_{z_2}[k],
\]
where
\[
a_1=\frac{z_1}{\gamma(z_1-z_2)(1-z_1z_2)},\qquad
a_2=\frac{z_2}{\gamma(z_2-z_1)(1-z_1z_2)}.
\]
This exposes a consequential correction. Under the circulant convention
stated in the E-snake note, its one-sided equations (33)--(34), transcribed as
printed, have inverse residual $0.474$--$0.492$ and symmetry defect
$0.659$--$0.666$ for $M=8$--128. The expression above has exactly zero
measured symmetry defect and worst inverse residual $1.28\times10^{-15}$.
The equivalent four cyclic first-order filters agree with dense inversion to
worst residual $1.18\times10^{-15}$ in both sequential and logarithmic-depth
parallel implementations.
The factorization is also a stable learnable parameterization. For this
real-root family, $z_i=-\operatorname{sigmoid}(\theta_i)$ and
$\gamma=\exp(\eta)$ put the roots strictly inside the unit disk and induce the
positive spectrum
\[
\widehat A(\omega)=\gamma\prod_{i=1}^2
(1-2z_i\cos\omega+z_i^2)>0.
\]
Autograd root/scale derivatives agree with centered differences to at worst
$1.52\times10^{-10}$. Both FFT and associative-scan realizations propagate
finite MPS gradients through the right-hand side, roots, and scale. The
systems control is negative for an unfused scan: across two synchronized
batch-256 sweeps over $M=64$--4096, the root-spectrum FFT is
$2.35$--$9.08\times$ faster forward and $2.20$--$7.66\times$ faster
forward--backward, with scan/FFT float32 disagreement at most
$7.95\times10^{-7}$. At $M=4096$, FFT takes $0.305$--$0.307$ ms forward
versus $2.77$ ms for the scan and avoids a 64 MiB dense float32 inverse.
Thus symmetric-circulant Fourier diagonalization is the preferred compiled
path on FFT-capable hardware; the exact bidirectional filters are retained for
streaming or no-FFT targets rather than claimed as an MPS speed improvement.
We then insert this transfer into a trainable six-edge additive KAN. Every arm
starts from the same 24-knot checkpoint, moves to 96 knots, resets the Adam
optimizer to isolate parameter transport, and continues training. On a target
with deliberately unresolved frequency-15 and localized detail, exact transfer
has median/worst immediate function drift
$1.97\times10^{-15}/1.98\times10^{-15}$ across five seeds, loss-jump ratio
one, and integral drift at roundoff. Coefficient interpolation changes the
function by $2.63\%$ and raises loss by $5.67\%$; a zero restart loses the
entire function and raises loss by $70.98\times$. At 1\% training noise the
coarse median test MSE is $2.0122\times10^{-2}$. Fine training after exact
transport reaches $2.2615\times10^{-5}$ without regularization. The exact
curvature Gram
\[
R_{k\ell}=\langle\phi_k'',\phi_\ell''\rangle,
\qquad \lambda\sum_e c_e^\top R c_e,
\]
with $\lambda=10^{-9}$ selected on development seeds, improves the median to
$2.0896\times10^{-5}$ on five untouched seeds. It wins all five paired
comparisons with a $1.100\times$ geometric improvement factor, giving roughly
$963\times$ improvement over the coarse unresolved model. In the clean
control, exact initialization reaches $3.37\times10^{-8}$ after 250 steps,
where a zero restart remains at $6.20\times10^{-6}$.
Refinement is not universally beneficial. On the earlier smooth target, the
24-knot model already has median test MSE $5.02\times10^{-6}$ at 1\% noise;
unregularized refinement worsens it to $2.24\times10^{-5}$, and the best tested
curvature setting only returns to approximately $5.94\times10^{-6}$. Thus the
complete rule is guarded: project exactly so a proposed refinement cannot
damage the current function, train new fine modes under the continuous
Sobolev Gram, and accept the enlarged grid only if held-out evidence improves.
This separates stability of grid transport from necessity of grid growth.
The controlled benchmark uses three periodic viscous conservation laws
(quadratic, quadratic plus sinusoidal, and saturating flux), centered temporal-
difference targets, six training trajectories at 64 points, and held-out
long-horizon rollouts at unseen amplitude and at 128 points. Across three
seeds the pole atlas obtains median/worst amplitude-OOD nRMSE
$9.287\times10^{-7}/8.475\times10^{-6}$ on the oscillatory flux. Median
errors for the compact cardinal edge, degree-five polynomial/SINDy, a
345-parameter direct cardinal KAN, and a 337-parameter MLP are respectively
$0.1870$, $0.06477$, $0.3300$, and $0.2977$. The selected reference law is
$0.499993u^2+0.079997\sin(4u)$, recovering the true
$0.5u^2+0.08\sin(4u)$. On the unmatched saturating flux, median amplitude-OOD
error is $0.005199$, versus $0.04113/0.1839/0.1165$ for
SINDy/direct-KAN/MLP. Conservative models preserve mass to approximately
$10^{-16}$; the direct networks drift by $10^{-3}$--$10^{-2}$.
The stronger mixed-constitutive gate uses
\[
u_t=\nu u_{xx}-\partial_xF(u)+R(u)
\]
and assigns a separate copy of the atlas to the flux and reaction edges. The
compiled columns are $-\phi_j'(u)u_x$ for the former and $\phi_j(u)$ for the
latter. Across three seeds the mixed atlas reaches median/worst amplitude-OOD
nRMSE $2.205\times10^{-5}/2.979\times10^{-5}$, compared with median
$0.1531$, $0.02890$, $0.2544$, and $0.2719$ for mixed cardinal,
polynomial/SINDy, direct KAN, and MLP models. Median resolution-OOD error is
$5.762\times10^{-7}$. The flux
$0.5u^2+0.06\sin(4u)$ is selected consistently to at least six coefficient
digits; the reaction function is recovered to median $7.91\times10^{-4}$
nRMSE.
Functional identifiability does not imply symbolic identifiability. The
reaction atom lists vary across seeds because $u$ and $\sin u$ are nearly
collinear over the narrow training amplitude interval. They synthesize nearly
the same reaction on the tested range, but their individual coefficients
cannot be interpreted uniquely. Broader excitation, group sparsity by pole
family, or annihilator-based incoherence constraints are required before the
mixed model can support a symbolic reaction-law claim.
\paragraph{Weak compilation under observation noise.}
For a separable test $\eta(t)\psi(x)$, periodic integration by parts gives
\[
-\iint \eta_t\psi u
=\nu\iint\eta\psi_{xx}u
+\iint\eta\psi_xF(u)
+\iint\eta\psi R(u).
\]
No derivative acts on the noisy observation. In the implementation, the
temporal test uses the exact discrete adjoint of the centered time-difference
operator, while sine/cosine spatial tests supply analytic first and second
derivatives. Across three seeds at $1\%$ additive observation noise, median
amplitude-OOD nRMSE is $0.001401$ for weak pole identification, $0.03844$ for
pointwise pole identification, and $0.003194$ for weak polynomial/SINDy. At
$2\%$, weak pole remains far stronger than pointwise
($0.006581$ versus $0.1464$) but slightly trails weak polynomial
($0.005207$). The flux function remains below $10^{-3}$ median error through
$2\%$ noise; the reaction function loses identifiability first. Thus weak
compilation delays, but does not remove, the variance cost of the larger pole
dictionary.
\paragraph{Annihilator-designed excitation.}
If $u(x,t)=a(t)$ is spatially constant, then
$\partial_xF(u)=0$ for every flux law. Such trajectories are therefore
operator-null-space probes of the reaction edge. Merely replacing half the
generic trajectories by four constant probes is a negative: it never recovers
the exact support in three seeds. Randomly broadening generic amplitudes is
also unreliable (one of three exact recoveries). The successful design uses
four generic flux trajectories plus a balanced 24-level constant sweep over
$[-1.5,1.5]$. It recovers the exact support
\[
\{F:u^2,\sin(4u);\quad R:u,u^3\}
\]
in all three seeds. Relative to eight narrow generic trajectories, median
reaction-function nRMSE falls from $0.002197$ to
$6.569\times10^{-7}$ and amplitude-OOD rollout nRMSE from
$2.689\times10^{-5}$ to $2.049\times10^{-7}$. This converts the earlier
functional identifiability into repeatable symbolic identifiability by
experimental design, without supervising either edge separately.
\paragraph{Nonseparable closure and quotient dictionaries.}
For the additional closure
\[
C(u,u_x)=\gamma\sin(2u)\sin(u_x),
\]
scalar flux and reaction edges are structurally misspecified. Adding sparse
tensor products of pole atoms lowers three-seed median amplitude-OOD nRMSE at
$\gamma=0.05$ from $0.005192$ for the scalar compiler, $0.1012$ for the
direct KAN, and $0.07287$ for the MLP to $0.0002971$. The unrestricted
dictionary is not identifiable, however, because
\[
a(u)u_x=\partial_x A(u),\qquad A'(u)=a(u).
\]
It can move a conservative flux term into the generic interaction edge without
changing the PDE residual. The naive fit does exactly this and has median
interaction-function nRMSE $23.71$. Defining the interaction dictionary on
the quotient that removes all atoms linear in $u_x$ restores flux attribution
and reduces interaction error to $0.2246$, with unchanged rollout.
\paragraph{Guarded scalar-to-tensor growth.}
A validation router selects the quotient tensor atlas only when at least one
interaction survives and held-out residual error falls. Across three seeds
and five mismatch strengths, it retains the scalar model in all three
$\gamma=0$ cases and grows the tensor model in all 12 nonzero cases. At zero
mismatch it prevents a worst-case tensor amplitude-OOD error of $0.02601$; at
$\gamma=0.05$ it lowers median error from $0.006218$ to
$6.197\times10^{-6}$. This supplies a concrete KAN-like growth rule: add
product structure only after operator residuals demonstrate scalar-edge
mismatch. Occasional worst-seed errors up to $0.004567$ show that sparse
selection inside the enlarged tensor space still needs a stability penalty.
\paragraph{Validation-selected tensor order.}
The surplus-interaction failure is not intrinsic to the product atlas. At
$\gamma=0.05$, a six-term quotient fit selects exactly one interaction atom
in all three seeds. Relative to the loose eight-term fit, median/worst
amplitude-OOD nRMSE improves from
$2.971\times10^{-4}/3.204\times10^{-4}$ to
$2.625\times10^{-5}/7.994\times10^{-5}$, while median interaction-function
nRMSE falls from $0.2246$ to $0.006863$. A five-term control is worse
($4.221\times10^{-4}/1.740\times10^{-3}$ median/worst rollout error), because
one small atom is needed to absorb the centered-difference bias in addition to
the generating support. A non-oracle selector fits budgets on six
trajectories and chooses the sparsest model within $10\%$ of the minimum
residual on two held-out trajectories. It selects six in all three
strong-interaction seeds and reproduces the oracle result after a full refit.
For $\gamma=0.005$--$0.02$, accurate prediction does not guarantee interaction
recovery: in some seeds the closure signal is comparable to discretization
bias, so functional attribution remains unidentifiable.
\paragraph{Weak-residual routing is not rollout routing.}
A final control uses held-out weak moments to choose between pole and
degree-five polynomial weak models across six noise levels and three seeds.
It selects the amplitude-OOD rollout winner in only 9 of 18 cases, fails all
three seeds at both $1\%$ and $4\%$ noise, and has worst rollout regret
$4.95\times$. The two libraries can have nearly identical integrated
residuals but different errors after recursive deployment. Thus the weak form
is an effective estimator, but its regression residual is not by itself a
safe architecture-selection objective; the validation functional must include
rollout stability or a provable surrogate for it.
Short low-pass rollout validation raises the development decision accuracy to
14 of 18, including all three $1\%$ cases, but does not solve the problem:
near-tied scores cause a worst regret of $5.32\times$. A robust alternative is
to average models when their identity is below the noise-resolution limit. We
estimate relative white-noise amplitude from the upper spatial half-band, keep
the pole law below $1.5\%$, and above that threshold deploy the convex law
\[
F_{\mathrm{avg}}=0.75F_{\mathrm{pole}}+0.25F_{\mathrm{poly}},\qquad
R_{\mathrm{avg}}=0.75R_{\mathrm{pole}}+0.25R_{\mathrm{poly}}.
\]
Because averaging occurs before the known outer operators, conservation and
the compiled calculus are unchanged. After freezing the rule on ten
development seeds, ten new seeds confirm the result. At $2\%$ noise,
mean/worst amplitude-OOD nRMSE is $0.006288/0.01107$ for the blend, versus
$0.007666/0.01424$ for pole and $0.007699/0.02382$ for polynomial. At
$4\%$, blend median/worst is $0.005424/0.009171$, versus
$0.008142/0.01449$ and $0.007678/0.02368$. The practical rule is therefore
continuous as well as operator-aware: use hard architectural growth when the
validation gap is resolved, but use an operator-compatible model average when
the data cannot support that decision.
\paragraph{Two-dimensional jet-space experimental design.}
Consider the anisotropic extension
\[
u_t=\nu\Delta u-\partial_xF_x(u)-\partial_yF_y(u)+R(u).
\]
Constants annihilate both conservative edges, x-only fields annihilate the
y-flux, and y-only fields annihilate the x-flux. These probes make the inverse
problem block triangular: estimate $R$ from constants, subtract it from each
directional balance, and then estimate the corresponding flux. There is an
additional jet-space requirement. A zero-mean wave does not adequately span
$(u,u_x)$ because its largest state values occur where $u_x=0$. We therefore
use offset directional waves $u(x)=c+v(x)$ and sweep $c$, independently varying
the constitutive argument and its multiplying derivative.
This design converts the 2D experiment from approximate prediction to exact
law recovery. The hierarchical atlas uses 10,200 compressed scalar rows and
recovers
\[
F_x(u)=0.5u^2+0.05\sin(3u),\quad
F_y(u)=-0.3u^2+0.04\sin(5u),\quad
R(u)=0.2u-0.15u^3
\]
to seven--eight coefficient digits in every one of three seeds. On unseen 2D
fields of amplitude $1.35$, beyond the designed value range, median/worst
amplitude-OOD nRMSE is $5.695\times10^{-9}/5.739\times10^{-9}$, while median
resolution- and joint-OOD errors are $7.404\times10^{-9}$ and
$5.702\times10^{-9}$. The generic 2D control consumes 204,800 rows---a
$20.1\times$ larger scalar data matrix---yet has median/worst amplitude-OOD
error $0.03235/0.06712$; randomly broadening its amplitudes gives
$0.04093/0.09698$. Earlier zero-offset directional probes remain at
$10^{-3}$--$10^{-2}$ rollout error. The resulting design rule is more precise
than ``excite broadly'': span the jet variables that multiply each compiled
operator and choose null-space probes that triangularize the unknown edges.
\paragraph{Noisy 2D derivative compilation.}
The exact clean recovery does not survive raw temporal differencing under
measurement noise. We therefore test two theoretically matched preprocessing
steps: projection onto the transverse-invariant subspace of each directional
probe, and a Savitzky--Golay local-polynomial filter along time before applying
the fourth-order difference. Window selection itself requires a held-out
audit. A seven-sample window looks best at $1\%$ on three development seeds,
but on five untouched seeds a 21-sample window is more robust and also remains
stable at $2\%$.
With width 21, temporal filtering alone lowers held-out median/worst
amplitude-OOD nRMSE from $0.06918/0.07640$ to
$0.003980/0.006052$ at $1\%$ noise, and from $0.1183/0.1355$ to
$0.005420/0.02143$ at $2\%$. The symmetry hypothesis is only partly useful.
Projection alone has median errors $0.07136$ and $0.09785$ and therefore does
not repair differentiated noise. Combining projection and temporal filtering
gives median $0.002678$ at $1\%$ but $0.007249$ at $2\%$, so it is not
uniformly better than temporal filtering alone. The supported rule is to put
the reproducing regularizer on the coordinate that will be differentiated;
null-space symmetry averaging may reduce variance but is not the causal
robustness mechanism here.
The same construction reduces the inverse problem itself. A generic
$N\times N$ trajectory contributes $O(N^2)$ scalar rows to a joint 45-column
design. A transverse-invariant directional trajectory contributes only
$O(N)$ distinct rows, and triangularization means that only one 15-column edge
block is materialized at a time. At $N=64$, the benchmark therefore uses
786,432 versus 19,008 rows ($41.4\times$ fewer) and 270 MiB versus 1.05 MiB of
logical peak design storage ($256\times$ lower). Identification is
$2.7$--$2.9\times$ faster across three seeds even before fused or GPU kernels.
At $N=16$--$24$, separate block-fit overhead dominates, with the measured
runtime crossover near $N=32$; the asymptotic storage gain is already present.
\paragraph{Coupled vector systems.}
The triangular principle also applies across field channels. We test
\[
u_t=\nu_u u_{xx}-\partial_xF(u)+C(v),\qquad
v_t=\nu_v v_{xx}-\partial_xG(v)+D(u),
\]
where all four constitutive edges are unknown. Constant paired states identify
the two cross-reactions. Their contributions are then subtracted from offset
directional-wave balances before fitting the two conservative fluxes. This
remains valid while the fields evolve and drive each other; the separation is
by typed operator action, not by freezing the other channel.
Across three seeds, the resulting atlas recovers all seven generating atoms and
their edge functions to $10^{-8}$--$1.5\times10^{-7}$. Median/worst
amplitude-OOD rollout nRMSE is
$2.419\times10^{-9}/2.964\times10^{-9}$; median resolution- and joint-OOD
errors are $1.650\times10^{-9}$ and $2.202\times10^{-9}$. The generic joint
fit has median/worst amplitude-OOD error $0.002016/0.004428$, selects between
7 and 12 atoms, and has median functional error $0.4453$ on the cubic
cross-reaction. Thus the relevant object is a typed graph of operator-compiled
edges: null-space probes can order that graph into identifiable blocks even
when its state dynamics remain coupled.
The exact result has two jointly necessary causes. Keeping the jet probes and
triangular solve fixed while deleting only the true frequency-five atom raises
hierarchical median/worst amplitude-OOD nRMSE from
$5.695\times10^{-9}/5.739\times10^{-9}$ to $0.01853/0.02243$ and median
y-flux function error to $0.1303$. Removing every sinusoidal pole raises
median rollout and y-flux errors to $0.05825$ and $0.3471$. The probes make the
restricted design full rank, but cannot synthesize an absent exponential mode;
conversely, a correct pole library is not identifiable without jet coverage.
This is the experimental form of the reproduction/identifiability factorization.
\paragraph{Finite-dimensional identifiability criterion.}
Let $S_e$ denote the active reproduction atoms for edge $e$, and let
$X_{qe}(P)$ be the compiled response of those atoms in equation $q$ under a
probe family $P$. If the probes can be ordered so that
\[
X_S(P)=
\begin{bmatrix}
X_{11} & 0 & \cdots & 0\\
X_{21} & X_{22} & \ddots & \vdots\\
\vdots & \ddots & \ddots & 0\\
X_{m1} & \cdots & X_{m,m-1} & X_{mm}
\end{bmatrix},
\]
then the active coefficients are uniquely identifiable exactly when every
diagonal block $X_{ee}$ has full column rank, modulo the explicit gauge quotient.
This follows directly by block forward substitution; if a diagonal block is
rank deficient, a nonzero coefficient perturbation in its null space produces
the same measurements. Pole inclusion guarantees that the truth lies in the
column span, annihilators create the zero blocks, and offset jet probes supply
rank to the diagonal blocks. The three experimental ablations separately
remove each condition: missing poles create approximation bias, generic probes
destroy triangular attribution, and zero-offset waves lose jet rank near state
extrema.
\paragraph{Continuous pole profiling.}
The same design makes learnable exponential-spline poles tractable. For
\[
F(u)=a u^2+b\sin(\omega u),
\]
fixing $\omega$ leaves a two-column compiled design, so $(a,b)$ are eliminated
by a ridge solve. We evaluate the resulting scalar profile objective on a
coarse frequency grid and refine its best interval by bounded scalar
minimization. This is variable projection: the difficult pole is not optimized
jointly with the linear edge coefficients.
Across five noninteger frequencies from $1.7$ to $5.6$ and three seeds, six
offset jet probes recover the pole with median/worst absolute error
$2.158\times10^{-9}/1.355\times10^{-8}$ and median/worst amplitude-OOD rollout
nRMSE $2.581\times10^{-10}/9.446\times10^{-10}$. The same profiled model on
generic trajectories has worst pole error $0.02548$ and worst rollout
$0.001606$. Fixed integer poles give median/worst rollout
$0.001574/0.01527$, and degree-five polynomial gives $0.03169/0.1191$.
Pole estimation is still statistically fragile. With a 21-sample temporal
polynomial prefilter and 24 replicated offset probes, median rollout errors are
$0.000709$, $0.002612$, and $0.008512$ at $0.1\%$, $0.5\%$, and $1\%$
observation noise. These improve strongly over six noisy probes and over the
fixed/polynomial offset controls, but generic continuous-pole trajectories are
competitive or better; median pole error reaches $0.1052$ at $1\%$. Variable
projection removes coefficient--pole optimization coupling, while weak-form or
probabilistic inference is still required to control pole variance.
For multiple poles, sequential pursuit is not reliable: after profiling the
linear coefficients, a greedy second-pole insertion can enter a wrong basin
even as the true separation increases. The operator calculus supplies a
finite-dimensional repair. Evaluate every candidate pole column once, form
its Gram matrix and target correlations, and score every coarse pole pair by a
constant-size profiled solve. Joint continuous refinement is then initialized
from several distinct low-residual pairs. Offset jet experiments recover
grid-aligned separations from $0.6$ down to $0.025$ to machine precision across
three seeds. The normalized active-design condition number increases from
$7.36$ to $178.2$ over that range. Off-grid tests separate prediction from
symbol recovery: at gap $0.058$ rollout remains near $10^{-7}$ with pole error
near $10^{-2}$, whereas at gap $0.025$ individual poles become unstable even
though their combined flux is accurate. The relevant resolution certificate
is therefore the conditioned active Gram, not convergence of a nonconvex
optimizer alone.
Noise requires compiling the profile in weak form, not differentiating a
prefiltered trajectory. For a separable test $\eta(t)\psi(x)$, conservation
gives the pole column directly as
\[
\int\!\!\int \eta(t)\psi_x(x)\sin(\omega u(x,t))\,\mathrm dx\,\mathrm dt,
\]
while temporal and diffusive derivatives act only on $\eta$ and $\psi$.
With offset jets, a frozen 40-sample window and four spatial modes, median
amplitude-OOD nRMSE at $0.1\%$, $0.5\%$, and $1\%$ noise is respectively
$1.266\times10^{-4}$, $5.232\times10^{-4}$, and $7.639\times10^{-4}$,
versus $3.548\times10^{-3}$, $3.502\times10^{-2}$, and
$4.426\times10^{-2}$ for matched pointwise profiling. Weak fixed-integer and
degree-seven polynomial controls are also worse. However, the median maximum
pole error rises to $0.0658$, $0.2687$, and $2.7$. Consequently the weak
profile is a robust predictive estimator beyond the pole-identification
regime; noisy symbolic claims require uncertainty sets, replicated excitation,
or an explicit minimum-separation prior.
The same Gram geometry suggests D-optimal probes: differentiate compiled
columns with respect to their coefficients and poles, then maximize a nominal
sensitivity log determinant. A strict control shows why this apparently
natural rule must not be accepted without matching excitation support. Its
initial advantage over offsets in $[-0.75,0.75]$ disappears when both arms use
the same $[-1.2,1.2]$ range and identical wave phases. We also compile a
second D-optimal rule from the exact weak sensitivity Gram under a nominal
simulator. Across ten deterministic-phase seeds at $0.5\%$ noise, broad
uniform offsets attain median rollout/flux/pole errors
$2.875\times10^{-4}/6.984\times10^{-4}/0.05319$, versus
$3.418\times10^{-4}/9.788\times10^{-4}/0.1052$ for narrow uniform.
Geometric improvement factors are $1.43$, $1.59$, and $2.27$, with bootstrap
95\% intervals above one. Neither D-optimal rule significantly improves on
broad uniform. Thus broad jet coverage is causal; the tested local Fisher
surrogates are not.
A one-versus-two profiled BIC is likewise only a conservative symbolic flag,
not a rollout selector. Measurement design, attribution, and deployment
require separate validation criteria.
A weak-form cardinal edge provides the direct KAN-inspired control. Compact
B-spline columns with negligible weak sensitivity must first be pruned;
otherwise their standardized coefficients diverge. Even after pruning and
ridge tuning, a 33-knot edge has median amplitude-OOD error about $0.105$ at
$0.5\%$ noise. Supplying the correct quadratic conservative base and using
the spline only as an innovation lowers this to $0.0371$, still approximately
$71\times$ above the continuous-pole result. Local support reduces parameter
interference, but cannot substitute for the correct global reproduction space
when extrapolation and weak observability are decisive.
The profiled estimator provides a local resolution certificate without ground
truth. Append to the weak linear design the pole-sensitivity columns
\[
b_j\,\frac{\partial}{\partial\omega_j}
\int\!\!\int\eta\psi_x\sin(\omega_j u)
=b_j\int\!\!\int\eta\psi_x u\cos(\omega_j u),
\]
and estimate covariance from the weak residual variance and the pseudoinverse
of this full Jacobian Gram. Across 20 independent broad-jet trials per noise
level, nominal 95\% intervals jointly cover both true poles in 20/20, 20/20,
and 18/20 cases at $0.1\%$, $0.5\%$, and $1\%$ noise; marginal coverage at
$1\%$ is 95\% for each pole. Median half-widths grow from
$[0.0555,0.0697]$ to $[0.2706,0.3970]$ and $[0.6464,0.6921]$.
The certificate requiring both half-widths below $0.1$ accepts every low-noise
case and rejects every medium/high-noise case. It therefore distinguishes
resolved symbolic poles from an accurate but non-identifiable combined flux.
The same compiler supports genuine complex exponential-spline roots. For
\[
F(u)=0.5u^2+0.04\exp(\sigma u)\sin(\omega u),
\]
fixing $(\sigma,\omega)$ again leaves only a small linear least-squares
problem. A two-dimensional Gram profile followed by local refinement recovers
both positive and negative real parts. Across nine clean offset-jet cases,
median/worst amplitude-OOD nRMSE is
$1.486\times10^{-10}/2.827\times10^{-10}$, and the maximum error in either
pole component is $2.23\times10^{-11}$. Imaginary-only and degree-seven
polynomial controls have median rollout $0.01020$ and $0.002269$.
In weak form the new feature is compiled without differentiating the noisy
state,
\[
\int\!\!\int \eta\psi_x
\exp(\sigma u)\sin(\omega u)\,\mathrm dx\,\mathrm dt.
\]
Three-seed median amplitude-OOD errors are $5.458\times10^{-5}$,
$2.059\times10^{-4}$, and $5.167\times10^{-4}$ at $0.1\%$, $0.5\%$, and
$1\%$ noise. The imaginary-only weak model remains near $1.15\times10^{-2}$
and weak polynomial closure near $4.2\times10^{-3}$. At $1\%$ noise, median
absolute errors in $(\sigma,\omega)$ are $(0.00717,0.00122)$. This establishes
that the learned carrier need not be a Fourier atom: its real and imaginary
pole parts can both be recovered from noisy conservation-law data.
For a complex pole the profile Jacobian appends both carrier sensitivities,
\[
b\int\!\!\int\eta\psi_x u\exp(\sigma u)\sin(\omega u),\qquad
b\int\!\!\int\eta\psi_x u\exp(\sigma u)\cos(\omega u).
\]
Across 20 new trials at every noise level, the resulting joint 95\% intervals
cover $(\sigma,\omega)$ in 20/20 cases at $0.1\%$, $0.5\%$, and $1\%$.
Median half-widths grow from $[0.00325,0.00199]$ to
$[0.01664,0.01011]$ and $[0.03251,0.01992]$; median component errors remain
smaller at $[0.000666,0.000373]$, $[0.00458,0.00156]$, and
$[0.00581,0.00338]$. Median/worst rollout at $1\%$ is
$4.264\times10^{-4}/1.120\times10^{-3}$. Thus the local covariance is
conservative on the nominal-amplitude task; a maximum-half-width threshold of
$0.05$ accepts all 60 trials, and must next be challenged by weakening the
carrier rather than by retroactively changing the threshold.
The amplitude sweep verifies that behavior. At $1\%$ noise, amplitude $0.03$
has median half-widths $[0.0441,0.0271]$ and an 8/10 acceptance rate. At
amplitudes $0.02$, $0.01$, and $0.005$, median real-part half-widths become
$0.0661$, $0.1331$, and $0.2638$, and the fixed certificate accepts 0/5 in
each group. Joint coverage is 100\% throughout and median rollout remains
below $5.5\times10^{-4}$, explicitly separating prediction from symbol
resolution. Five-seed tests at $\sigma=-0.75$ and $+0.75$ retain 100\%
coverage and about $5\times10^{-4}$ median rollout, excluding proximity to the
profile bounds as the cause.
Local cardinal support becomes useful when the law contains a localized
constitutive defect, but only with an attribution constraint. In a controlled
test the true flux is the complex carrier plus one compact cardinal cubic.
Simultaneously fitting the pole and a dense local dictionary reduces prediction
error yet lets the nuisance dictionary absorb pole perturbations. We instead
compute each probe's carrier weak residual, refit the carrier on the
least-mismatched half of the offset jets, freeze it, and estimate the local
innovation from all probes. At $0.1\%$ noise over five seeds this residual-
routed hybrid has median amplitude-OOD error $1.927\times10^{-4}$, compared
with $6.902\times10^{-3}$ for the global carrier, $6.292\times10^{-4}$ for a
cardinal-only edge, and $1.025\times10^{-2}$ for degree seven. Its median
$(\sigma,\omega)$ errors are $(2.86\times10^{-4},1.14\times10^{-3})$.
At $0.5\%$ noise the routed median is $7.737\times10^{-4}$, compared with
$6.986\times10^{-3}$ and $8.747\times10^{-4}$ for carrier-only and
cardinal-only, while pole errors remain $(1.13\times10^{-3},2.78\times10^{-3})$.
This triangular scheme gives a precise role to the KAN idea: a local edge is a
routed nuisance correction around an identified operator carrier, not a free
competitor for the same signal.
For deployment we sparsify the nuisance step. Each fixed-grid cardinal atom
is scored after residual routing, and an extended BIC charges for both its
linear amplitude and the search over centers. In five defect-present and five
defect-absent trials at each of $0.1\%$ and $0.5\%$ noise, this gate accepts all
10 true defects, rejects all 10 nulls, and selects the exact center $u=0.4$ in
every accepted run. Present-case median rollout changes from
$5.652\times10^{-3}$ to $9.386\times10^{-5}$ at $0.1\%$ and from
$5.634\times10^{-3}$ to $3.293\times10^{-4}$ at $0.5\%$. In null cases the
gate returns the original all-probe carrier, so median rollout is exactly
unchanged at $3.965\times10^{-5}$ and $2.039\times10^{-4}$. This is the
operative synthesis: exponential-spline reproduction supplies global
extrapolation, residual routing protects pole attribution, and a sparse
cardinal edge supplies conditional local repair with no null-case accuracy tax.
Fixed cardinal centers introduce a separate quantization boundary. Defects at
grid centers $-0.6$, $0$, and $0.8$ are accepted and localized exactly in 9/9
trials at $0.5\%$ noise. With true center $0.35$, however, fixed pursuit
chooses $0.3$ or $0.4$ and obtains median/worst rollout
$2.034\times10^{-3}/7.140\times10^{-3}$. Retaining cardinal screening but
refining only the winning center by bounded scalar variable projection recovers
$0.34774$, $0.35014$, and $0.34989$ and reduces median/worst error to
$2.715\times10^{-4}/4.632\times10^{-4}$. The extra center degree of freedom
is charged in the extended BIC; 5/5 new null cases are rejected with exact
carrier fallback.
The offset-jet design remains essential. Under generic zero-centered probes,
the same off-grid task has median carrier rollout $0.02291$ and adaptive-hybrid
rollout $0.02420$, with large pole errors and scattered selected centers. The
evidence criterion can identify a residual feature in the observed state band,
but cannot manufacture the missing state-space separation. Local refinement
and annihilator/jet excitation solve different parts of the inverse problem.
The same independence issue controls local order. Applying ordinary EBIC to
every overlapping weak row can add a spurious boundary atom in a one-defect
case. A probe-block criterion counts non-overlapping temporal windows times
orthogonal spatial tests and then adds the combinatorial support penalty. On
five null, five one-defect, and five two-defect trials at $0.5\%$ noise, it
selects order 0, 1, and 2 correctly in all 15 cases, always recovering the
exact nonempty support $\{0.4\}$ or $\{-0.6,0.4\}$. Median rollout is
$1.754\times10^{-4}$, $5.048\times10^{-4}$, and
$4.330\times10^{-4}$, compared with carrier-only medians
$7.106\times10^{-3}$ and $8.368\times10^{-3}$ in the nonnull groups. Thus
hierarchical local growth is feasible, but its likelihood must be defined on
independent experiment blocks.
The remaining bridge is discretization bias. In a matched three-seed audit at
$0.5\%$ noise, spectral-data training yields sparse-hybrid median/worst rollout
$4.238\times10^{-4}/1.005\times10^{-3}$. A separately implemented centered
conservative finite-volume generator transfers partially at
$1.405\times10^{-3}/1.745\times10^{-3}$, still below its carrier-only median
$6.206\times10^{-3}$. A Rusanov generator fails: median/worst error becomes
$7.414\times10^{-3}/3.533\times10^{-2}$ and pole estimates are biased because
the compiler interprets numerical viscosity as physical constitutive signal.
Mesh doubling improves matched Rusanov and centered cases only to
$2.733\times10^{-3}$ and $1.464\times10^{-3}$. Hence weak differentiation
removes observation-noise amplification, but it does not remove misspecification
of the discrete diffusion operator.
A scalar calibration profile partly repairs this mismatch without using
rollout labels. Scanning the viscosity inside the compiler and selecting by
weak residual chooses $\nu_{\mathrm{eff}}=0.08$ in 3/3 Rusanov trials. Median
rollout falls from $7.414\times10^{-3}$ to $2.100\times10^{-3}$, a
$3.5\times$ reduction, and the selected value matches the rollout-oracle grid
choice in two trials. Median real-pole error remains about $0.08$, however.
Scalar operator calibration captures an average modified-equation viscosity;
state-dependent numerical diffusion remains a nuisance operator and prevents
symbolic attribution.
For centered spatial discretization the continuous--discrete bridge can be
exact on every retained Fourier test. Replace the continuum derivative
symbols by
\[
\widetilde k_1={\sin(k\Delta x)\over\Delta x},\qquad
\widetilde k_2={4\sin^2(k\Delta x/2)\over\Delta x^2}.
\]
\paragraph{Discrete-adjoint exactness.}
Let $D_h$ be any periodic translation-invariant spatial stencil with discrete
Fourier symbol $d_h(k)$, and let $\langle\cdot,\cdot\rangle_h$ denote the grid
inner product. For every retained Fourier test $\psi_k$ and sampled field
$v$, one has exactly
\[
\langle \psi_k,D_hv\rangle_h
=\langle D_h^*\psi_k,v\rangle_h
=\overline{d_h(k)}\,\langle\psi_k,v\rangle_h.
\]
Consequently a weak constitutive design compiled with $\overline{d_h(k)}$
matches the semidiscrete generator on those modes without differentiating the
observations. The statement follows because circulant stencils are diagonal
in the discrete Fourier basis and their adjoints conjugate the eigenvalues.
Continuum weak forms are recovered as $d_h(k)\to(ik)^r$; using continuum
symbols at finite $h$ instead introduces a deterministic bridge bias that no
increase in the constitutive dictionary can remove.
Equivalently, rescale $\psi_x$ by $\sin(k\Delta x)/(k\Delta x)$ and
$\psi_{xx}$ by $[\sin(k\Delta x/2)/(k\Delta x/2)]^2$. This acts only on the
analytic tests. Across three centered-volume trials it reduces median/worst
rollout from $1.405\times10^{-3}/1.745\times10^{-3}$ to
$4.197\times10^{-4}/1.003\times10^{-3}$, matching the spectral-data median
$4.238\times10^{-4}$.
For constant-speed Rusanov, the modified equation adds exactly
$\lambda\Delta x/2$ to viscosity. The resulting $\nu_{\mathrm{eff}}=0.07927$
is selected by weak residual in 3/3 paired trials; combined with the discrete
symbols it gives median/worst rollout
$6.332\times10^{-4}/1.607\times10^{-3}$ and median pole errors
$(0.00127,0.00253)$. Without symbol correction the median is
$1.939\times10^{-3}$. For state-dependent Rusanov, exact centered symbols
improve the calibrated median only from $2.100\times10^{-3}$ to
$1.849\times10^{-3}$ and real-pole error remains about $0.081$. This isolates
the unresolved term: not observation derivatives or stencil mismatch, but the
state-dependent numerical-viscosity operator itself.
That nuisance can be removed by one operator fixed point. Initialize from the
scalar-calibrated flux, evaluate its state derivative to obtain Rusanov's local
face speed, construct the corresponding numerical-viscosity right-hand side,
and subtract its projection using the already stored weak test weights. Refit
the physical-viscosity carrier and local edge afterward; observations are never
differentiated. Across three paired state-dependent Rusanov trials, median/
worst rollout becomes $3.774\times10^{-4}/8.934\times10^{-4}$, versus
$1.849\times10^{-3}/1.852\times10^{-3}$ after scalar and exact-symbol
calibration and $7.414\times10^{-3}/3.533\times10^{-2}$ without correction.
All three recover the true local center, with median $(\sigma,\omega)$ errors
$(0.00317,0.00890)$. The discrete nuisance operator can therefore be inferred,
compiled out, and separated from the physical constitutive law. The expanded
audit selects local order one in all 10 defect cases, with median/worst rollout
$4.025\times10^{-4}/1.259\times10^{-3}$ and median pole errors
$(0.00558,0.00667)$. All five fresh no-defect controls select order zero and
have median/worst rollout $2.210\times10^{-4}/2.581\times10^{-4}$.
At $1\%$ observation noise, 5/5 further trials recover the exact one-atom
support with median/worst rollout
$6.248\times10^{-4}/1.173\times10^{-3}$ and median pole errors
$(0.00235,0.00459)$. With two local defects at $0.5\%$ noise, 5/5 trials
select exact order two and support $\{-0.6,0.4\}$, with median/worst rollout
$6.143\times10^{-4}/1.100\times10^{-3}$ and pole errors
$(0.00700,0.00363)$. Hence the discrete nuisance projection is stable to both
statistical and sparse structural scaling in this controlled solver-transfer
test.
The data stencil need not be supplied as an oracle, but it should not be
selected by the constitutive residual that it also changes. Full-fit residual
selection chooses the correct continuous or exact-centered compiler in only
18/20 cases. A two-probe holdout remains unstable: a frozen 2\% preference
margin gets all five untouched spectral trials but only one of five centered
trials. Raw holdout is 10/10 on those new trials but already missed one pilot.
This is a useful negative---operator and constitutive selection are coupled at
this sample size.
An independent dispersion fingerprint resolves the coupling. Apply small-
amplitude single-mode probes, estimate each mode's exponential amplitude decay
and unwrapped phase rate, and profile the unknown diffusion and transport
coefficients under either $(k,k^2)$ or the centered symbols above. With 64
cells and modes 1--12, the lower normalized modal residual identifies the
generator in 80/80 trials at 0.5--1\% noise. The identifiability boundary is
set by Fourier-symbol separation: at 1\% noise, maximum modes 2, 3, and 4 give
only 20/40, 23/40, and 28/40 pooled correct decisions, while modes 6 and 8
give 40/40. Modes through 12 give 40/40 at 128 cells and 38/40 at 256 cells;
extending the 256-cell probe to mode 20 restores 40/40. Across these cases the
reliable design has approximately $k_{\max}\Delta x\geq0.5$; below it, the
continuous and discrete symbols converge faster than noise permits their
separation.
The resulting two-stage protocol is non-oracle: fingerprint the numerical
measurement operator, compile its exact adjoint, then identify the nonlinear
constitutive edge. On ten seeds per generator, this selection yields median
amplitude-OOD rollout $3.053\times10^{-4}$ for spectral data and
$3.074\times10^{-4}$ for centered data, versus $1.137\times10^{-3}$ and
$1.331\times10^{-3}$ with the mismatched bridges. The short calibration is
therefore not a preprocessing convenience; it is an identifiability experiment
that prevents discretization artifacts from being reported as physical poles.
State offsets provide the second excitation axis needed to recognize nonlinear
artificial diffusion. Under the Rusanov hypothesis we constrain modal decay
to a shared physical viscosity plus $|F'(u_0)|\Delta x/2$, taking $F'(u_0)$
from the phase speed fitted independently at each small-amplitude offset jet.
Across spectral, centered, and state-dependent Rusanov generators, five offsets
and modes 1--12 give 120/120 correct three-way classifications at 0.5--1\%
noise. At 1\% noise the Rusanov fit recovers physical viscosity with median
absolute error $2.72\times10^{-4}$. A single offset is genuinely
nonidentifiable: centered versus Rusanov selection is correct in only 19/40
pooled trials because either model can absorb one effective decay rate into
viscosity. Two separated offsets restore 40/40 decisions and median Rusanov
viscosity error $2.12\times10^{-4}$; three offsets also give 40/40. Hence
Fourier-mode diversity identifies the derivative stencil, while state-offset
diversity identifies its nonlinear viscosity. Nuisance separation comes from
orthogonal axes of excitation, not from enlarging the constitutive dictionary.
\paragraph{Offset identifiability.}
Linearize a conservative flux about constant states $u_j$ and write
$c_j=F'(u_j)$. For a centered semidiscretization, the mode-$k$ eigenvalue is
\[
\lambda^{\mathrm C}_{jk}=-\nu\widetilde k_2(k)
-\mathrm{i}c_j\widetilde k_1(k),
\]
whereas local Lax--Friedrichs/Rusanov flux adds, to first order in probe
amplitude,
\[
\lambda^{\mathrm R}_{jk}
=-\bigl(\nu+|c_j|\Delta x/2\bigr)\widetilde k_2(k)
-\mathrm{i}c_j\widetilde k_1(k).
\]
At one offset the added term is exactly confounded with an unknown $\nu$. At
two offsets with $|c_1|\ne|c_2|$, the centered hypothesis requires equal
decay divided by $\widetilde k_2$, while the Rusanov hypothesis requires their
difference to equal $(|c_1|-|c_2|)\Delta x/2$. Phase identifies the $c_j$
independently through $\widetilde k_1$. Thus, given one nonzero retained mode
and noiseless linearized rates, the two hypotheses and physical viscosity are
identifiable from two such offsets. Multiple modes provide noise averaging
and distinguish the continuous from discrete symbols; the preceding sweeps
quantify the finite-noise bandwidth required for that second distinction.
Finite amplitude exposes the expected bias--variance compromise. Three-way
selection remains 30/30 through 5\% relative noise and for amplitudes from
0.025 to 0.4, but median Rusanov physical-viscosity error increases from
$2.72\times10^{-4}$ at amplitude 0.025 to $4.88\times10^{-3}$ at 0.4 because
the local linearization is no longer exact. With fixed absolute noise
$5\times10^{-4}$, amplitudes 0.005, 0.01, and 0.025 give 23/30, 29/30, and
30/30 correct decisions; their Rusanov viscosity errors are respectively
$1.04\times10^{-4}$, $1.09\times10^{-4}$, and $2.73\times10^{-4}$. Hence an
adaptive calibration should increase amplitude only until symbol separation is
statistically decisive, then stop before nonlinear bias dominates.
The calibration can replace the final oracle in the nonlinear pipeline. We
freeze the independent Rusanov estimate $\widehat\nu=0.040272$ and use it in
the exact-symbol fixed-point nuisance compiler, rather than resetting to the
simulator's true $\nu=0.04$. Across ten new 0.5\%-noise constitutive trials,
the resulting pipeline selects the one-atom support in 10/10 and has
median/worst rollout $5.222\times10^{-4}/9.080\times10^{-4}$. The paired
true-viscosity oracle gives $5.252\times10^{-4}/9.296\times10^{-4}$; the median
paired error ratio is 1.018. Median pole errors are
$(0.00462,0.00383)$ with estimated viscosity and $(0.00464,0.00384)$ with the
oracle. Thus the measured modal/offset response supplies all numerical-
operator quantities required by the fixed-point constitutive discovery stage.
The candidate stencil can itself be removed. Represent an unknown periodic
translation-invariant first derivative by the odd symbol
\[
d_{\mathbf a}(k)={1\over\Delta x}\sum_{r=1}^{R}a_r
\sin(kr\Delta x)
\]
and its even diffusive part by
\[
\ell_{\mathbf b}(k)={2\over\Delta x^2}\sum_{r=1}^{R}b_r
\{1-\cos(kr\Delta x)\}.
\]
The offset-by-mode phase matrix is rank one in the linearized regime, so its
right singular vector supplies the shape of $d_{\mathbf a}$. Phase speed and
symbol scale have a gauge; first-derivative consistency removes it through
$\sum_r r a_r=1$. Modal decay identifies the physical coefficients
$\mathbf b$ directly. Stencil radius is then an ordinary held-out model-
selection problem.
The companion experiment fits ten calibration records and validates on ten
untouched records. Radius one is selected with score
$2.641\times10^{-5}$, versus $2.650\times10^{-5}$,
$2.663\times10^{-5}$, and $2.668\times10^{-5}$ for radii two through four;
the learned odd and even physical coefficients are 1.0 and 0.0400019. On ten
new centered-data constitutive trials, compiling these learned symbols gives
median/worst amplitude-OOD rollout
$5.6696\times10^{-4}/1.1341\times10^{-3}$, numerically identical to the
analytic centered-stencil oracle
$5.6689\times10^{-4}/1.1342\times10^{-3}$. The continuum compiler gives
$1.446\times10^{-3}/2.107\times10^{-3}$. If the rank-one symbol is instead
normalized arbitrarily by $d(1)=1$, median/worst error is only
$9.553\times10^{-4}/1.406\times10^{-3}$. This negative control establishes
that continuum consistency, not curve fitting alone, closes the unknown-
stencil bridge.
To exclude a three-point coincidence, we repeat the procedure with an
independently implemented fourth-order five-point generator. The one-standard-
error rule selects radius two at both 0.5\% and 1\% calibration noise. At 32
cells the recovered odd coefficients are $(1.333332,-0.166666)$, compared with
$(4/3,-1/6)$, and the even physical coefficients are
$(0.0533288,-0.00333008)$, compared with $(0.0533333,-0.00333333)$. On ten new
coarse-grid, eight-mode constitutive trials at 0.1\% noise, the learned-symbol
compiler yields median/worst rollout
$6.480\times10^{-5}/2.974\times10^{-4}$, the analytic-symbol oracle yields
$6.349\times10^{-5}/2.945\times10^{-4}$, and the continuum compiler yields
$4.323\times10^{-4}/7.364\times10^{-4}$. At 0.5\% noise, learned and oracle
remain matched at about $5.168\times10^{-4}/1.525\times10^{-3}$, whereas the
continuum approximation has a slightly smaller median
$4.503\times10^{-4}$ but a larger maximum $1.927\times10^{-3}$ and worse pole
errors. This is the honest statistical boundary: exact adjoints are required
for attribution and control deterministic bias, but a biased low-bandwidth
model can occasionally regularize a noisy finite sample.
Calibration cost can be compressed spectrally. With four candidates
(continuous, three-point centered, fourth-order centered, and Rusanov), two
offsets, 12 separate modes, and 80 steps give 80/80 correct decisions at 1\%
noise, requiring 24 trajectories. A single multisine can carry all 12 modes:
one amplitude-0.005 trajectory at each of two offsets, observed for 160 steps,
also gives 80/80. At 0.5\% noise and 80 steps it gives 79/80. The shorter
1\%-noise, 80-step audit falls to 69/80, demonstrating a temporal-aperture
boundary. Replacing log-amplitude/phase regression by a forward--backward
complex AR(1) estimate is a negative at 43/80 despite one successful pilot;
noise in both consecutive Fourier coefficients invalidates that shortcut.
Thus broadband excitation reduces trajectory count by a factor of 12, while
time duration remains governed by modal-rate signal to noise.
The unknown-stencil fit can use the same compression. One two-offset
multisine record estimates coefficients and two further records validate
radius, for six trajectories total. The one-standard-error rule selects the
true radius two in 6/6 disjoint three-record groups. Freezing one such
calibration and applying it to ten new 32-cell/eight-mode constitutive trials at
0.1\% noise gives median/worst rollout
$1.215\times10^{-4}/3.308\times10^{-4}$, compared with
$4.323\times10^{-4}/7.364\times10^{-4}$ for continuum compilation,
$6.480\times10^{-5}/2.974\times10^{-4}$ for the larger learned-symbol
calibration, and $6.349\times10^{-5}/2.945\times10^{-4}$ for analytic symbols.
Thus six broadband trials already obtain a $3.6\times$ median gain without a
named stencil; additional calibration reduces coefficient variance toward the
oracle limit.
Sparse instrumentation introduces a third, purely sampling-theoretic boundary.
For $K=12$ multisine modes, we recover modal coefficients by least squares from
fixed spatial sensors. Twenty-four uniformly distributed sensors are rank
deficient and produce only 29/80 correct four-way decisions. At the exact
$2K+1=25$ real-sample threshold, the audit gives 80/80 at 0.5\% noise and
77/80 at 1\%; extending the latter from 160 to 240 time steps restores 80/80.
Random placement does not inherit the same conditioning: at 0.5\% noise,
25, 32, 40, 48, and 56 random sensors yield respectively 20, 53, 75, 77, and
79 decisions out of 80. This is the cardinal sampling requirement in an
experimental-design role: spatial geometry must make the trigonometric frame
stable before temporal modal rates can identify the discrete operator.
Uniform time sampling is not required. With full spatial readout, 40 random
time stamps over a 240-step aperture retain 80/80 four-way decisions at 1\%
noise; 40 samples over only 160 steps give 73/80. Under joint sparsity,
25 uniform sensors and 40 random times give 74/80, while 80 random times restore
80/80. Raising the sensor count to 32 but keeping only 40 times gives 75/80,
confirming that the temporal rate estimate is then limiting. The successful
joint protocol consumes $2\times25\times80=4000$ scalar observations, versus
$2\times64\times241=30848$ under full space--time sampling. Above the spatial
sampling threshold, identifiability depends on temporal aperture and count,
not on a uniform clock.
The final composition learns the stencil from the jointly sparse records and
then transfers it across resolution. With 25 uniform sensors and 80 irregular
times, a ten-record fit and ten-record validation split selects radius two and
recovers odd coefficients $(1.33540,-0.16770)$ and even physical coefficients
$(0.053235,-0.003306)$. A stencil coefficient vector is grid independent,
whereas its sampled Fourier symbol is not. We therefore evaluate the learned
trigonometric polynomial anew on the 32-cell constitutive grid. Across ten
new 0.1\%-noise nonlinear trials, this compiler selects the exact one-atom
support in 10/10 and attains median/worst amplitude-OOD rollout
$1.510\times10^{-4}/2.707\times10^{-4}$, versus
$1.738\times10^{-4}/2.505\times10^{-4}$ for the analytic-stencil oracle and
$4.331\times10^{-4}/8.078\times10^{-4}$ for the continuum compiler. Its
paired median error ratio to the oracle is 0.996. Directly reusing the
64-cell sampled factors at 32 cells instead gives median
$4.081\times10^{-4}$, a $2.7\times$ regression. The learned operator must
therefore cross resolutions as coefficients and be recompiled at the target
mesh; this is the discrete counterpart of transporting a cardinal spline by
its generator rather than by samples tied to one grid.
This sparse design also exposes a replication and experimental-design
boundary. The initial symmetric-offset panel yielded 80/80 classifications,
but a disjoint matched panel yields only 35/40 (38/40 with twice as many
retained times). At offsets $(-0.4,+0.4)$ the two values of $|F'(u_0)|$ are
too similar, so state-dependent numerical viscosity is nearly coherent with
the shared physical-viscosity column. Changing only the offsets to
$(-0.4,+0.8)$ raises a larger validation panel from 72/80 to 77/80 at the
same measurement budget, while four wide/asymmetric pairs each give 40/40 in
the pilot. If perturbed sensor coordinates are supplied to the Fourier frame,
even one-cell jitter remains 74/80; if a 0.025-cell perturbation is unmodelled,
accuracy falls to 58/80. We therefore define the normalized selection margin
$(S_{(2)}-S_{(1)})/S_{(2)}$ and freeze a threshold 0.15 from the pilot. It
accepts 210 of 320 cases in the subsequent multi-regime validation and all 210
accepted decisions are correct. The resulting rule is operator-theoretic:
choose offsets that reduce column coherence, include sensor geometry in the
analysis operator, and abstain whenever the candidate quotient is not
separated.
The same excitation can identify the analysis geometry. Let
$q_j(x)$ denote the known initial multisine of probe $j$ and let $\bar x_m$ be
the nominal sensor position. We estimate its displacement independently by
\[
\widehat x_m=\arg\min_{|x-\bar x_m|\leq\rho\Delta x}
\sum_j\{y_{jm}(0)-q_j(x)\}^2 ,
\]
then build the trigonometric sampling matrix at $\widehat x_m$. Independent
probe phases make the local code increasingly injective as $j$ grows. At 1\%
noise and unknown 0.1-cell jitter, two, three, and five probes give 33/40,
37/40, and 38/40 classifications, compared with 19/40 under nominal
coordinates. In an 80-case five-probe validation, nominal, self-calibrated,
and known coordinates yield 31/80, 75/80, and 79/80. Median coordinate error
is 0.00370 cells and the median trial-wise maximum is 0.01493; median Rusanov
viscosity error improves from $3.74\times10^{-4}$ nominal to
$1.14\times10^{-4}$, with known-position error $6.50\times10^{-5}$. The
frozen margin gate accepts 55 self-calibrated cases and all 55 are correct.
Thus the excitation first self-surveys the sampling operator and only then
identifies the discrete differential operator.
For harmonic generators the shift law eliminates even the nonlinear local
search. Two static fields give
$(y_m^{s},y_m^{c})=A(\sin Kx_m,\cos Kx_m)+\boldsymbol\epsilon_m$, so
\[
\widehat x_m={1\over K}\operatorname{atan2}(y_m^s,y_m^c)
\pmod {2\pi/K},
\]
with the branch nearest $\bar x_m$ selected. Its small-noise position variance
scales as $K^{-2}$, motivating a locator mode above the dynamics band. Raising
$K$ from 12 to 28 reduces pilot median coordinate error from 0.00406 to
0.00170 cells. On 80 new trials with 0.1-cell unknown jitter, nominal,
quadrature-calibrated, and known coordinates give 24/80, 78/80, and 79/80
classifications. Quadrature calibration has median coordinate error 0.00178
cells and median trial-wise maximum 0.00602; its median Rusanov viscosity error
is $8.82\times10^{-5}$, between the known-position
$6.56\times10^{-5}$ and nominal-position $1.11\times10^{-3}$. The frozen
margin gate accepts 68 cases and all 68 are correct. Only two 25-value static
snapshots are added to two $25\times80$ dynamics records, for 4050 scalar
measurements total. This is a direct algorithmic consequence of exponential
reproduction: translation becomes phase, so the sampling operator can be
self-calibrated before the differential operator is learned.
The fine phase has branch radius $n_x/(2K)$ in grid-cell units. For $K=28$
this is 1.14 cells: the single-frequency locator changes from 36/40 correct at
one-cell jitter to only 8/40 at 1.25 cells, with errors of one phase period
(2.29 cells). A coarse-to-fine construction first decodes mode 8 and chooses
the mode-28 branch nearest that estimate. It maintains median and median-
maximum coordinate errors of approximately 0.0017 and 0.0057 cells through
two-cell jitter, and every confidence-gated downstream decision is correct.
The remaining loss follows conditioning of the irregular spatial frame, even
when positions are known. At two-cell jitter, 64 sensors and 80 times give
40/40 known and 39/40 self-calibrated classifications, while 40 sensors and
160 times give 39/40 and 38/40. The former lowers the median sampling-frame
condition number to about 3.06. Thus multi-frequency exponential reproduction
sets the coordinate range, and spatial oversampling controls the subsequent
modal inversion variance.
The composed experiment removes both geometry and named-stencil oracles.
Twenty quadrature-self-calibrated five-point records, each with 25 sensors and
80 irregular times under hidden 0.1-cell jitter, select radius two and recover
odd coefficients $(1.33978,-0.16989)$ and even physical coefficients
$(0.053134,-0.003260)$. We transfer these coefficients and re-evaluate their
symbols on a 32-cell grid. Across ten new nonlinear discovery trials, exact
local support is selected in 10/10. Median/worst amplitude-OOD rollout is
$2.091\times10^{-4}/4.544\times10^{-4}$ for the learned compiler,
$1.498\times10^{-4}/4.012\times10^{-4}$ for the analytic oracle, and
$4.603\times10^{-4}/9.046\times10^{-4}$ for continuum compilation. Learned
beats continuum in every paired trial and by a factor 2.2 in median; its median
penalty relative to the oracle is 1.37. Hence the sampling geometry, discrete
adjoint, target-grid symbol, and sparse constitutive innovation can all be
identified sequentially from designed measurements, with a quantified final
calibration-variance cost.
The clean nonseparable result has a stricter noise boundary. Pointwise tensor
selection, temporal plus spatial smoothing, ridge variation over seven orders,
and increasing the generic trajectory count from 8 to 64 all leave the
interaction-function error near one. Prescribed initial jets remove regressor
noise but require a boundary time derivative; even 256 averaged bursts do not
make that route symbolic. We therefore compile a space--time weak tensor
design in which the temporal, diffusion, and conservative derivatives act on
analytic tests, while only the irreducible nonlinear $\sin(u_x)$ factor is
evaluated from the state. Although one pilot recovers
$0.04725\sin(2u)\sin(u_x)$ for a true coefficient 0.05, a 45-case audit does
not reproduce atom identity reliably. The robust result is predictive: a
development-frozen mixture with 25\% weak tensor and 75\% weak scalar beats the
scalar in 10/10 new $\gamma=0.05$, 0.5\%-noise trials, reducing median/worst
amplitude-OOD rollout from
$3.776\times10^{-3}/8.016\times10^{-3}$ to
$3.203\times10^{-3}/5.558\times10^{-3}$. Thus the operator-compatible tensor
edge contributes below the symbolic resolution threshold, but must be averaged
rather than interpreted.
Repeated observation reveals a coherence rather than a variance floor. Over
five fixed clean ensembles, increasing independent noisy replicates from one
to 32 decreases median interaction nRMSE from 0.1299 to 0.0204 and median
rollout error from $7.20\times10^{-4}$ to $1.79\times10^{-4}$, but exact atom
recovery saturates at four of five; 64--512 replicates do not remove the
failure. Increasing random trajectories from 8 to 64 is non-monotone
(2/5, 3/5, 4/5, and 3/5 exact). More importantly, a coefficient-frequency
certificate developed over repeated noise is externally falsified: one of
five new clean ensembles stably certifies two false surrogate interactions.
Stability to measurement noise is not identifiability of the physical law.
Weak-feature excitation design nevertheless provides a strong intermediate
result. Selecting eight of sixteen candidate trajectories by greedy
D-optimality improves exact recovery from 3/10 to 7/10, reduces median/worst
interaction nRMSE from 0.997/19.16 to 0.0385/1.009, and reduces median/worst
rollout from $2.730\times10^{-3}/1.342\times10^{-1}$ to
$1.125\times10^{-3}/1.532\times10^{-2}$, with eight paired rollout wins.
This does not close the symbolic gate: selecting 8 or 12 from 32 candidates
retains a shared catastrophic seed, and quotienting scalar features before
D-optimal selection worsens exact recovery from 4/5 to 3/5. The conclusion is
that designed excitation is the right control variable, but determinant
volume and replicate confidence alone do not resolve nuisance coherence.
The failure atoms identify the missing design coordinate: coverage of the
nonlinear argument. On the original amplitude-0.65 trajectories,
$\sin(2u)$ remains coherent after weak projection with higher trigonometric
surrogates. Holding the weak compiler and sparse selector fixed, amplitude
0.95 yields 4/5 exact recoveries, whereas amplitude 1.20 with D-optimal
selection yields 5/5. A frozen ten-seed validation then gives 9/10 exact for
random wide-amplitude trajectories and 10/10 for D-optimal wide-amplitude
trajectories. The D-optimal coefficient lies in $[0.04945,0.05045]$ for truth
0.05, with median/worst interaction nRMSE 0.00400/0.01096. Its median rollout
is neutral relative to random ($8.62\times10^{-4}$ versus
$8.13\times10^{-4}$, five paired wins), but its worst rollout improves from
$4.03\times10^{-3}$ to $1.42\times10^{-3}$. A no-averaging audit is stronger:
D-optimal wide-amplitude design recovers the exact atom in 5/5 pilot and 10/10
untouched single-observation seeds, compared with 5/5 and 9/10 for random
wide-amplitude design. Its validation median/worst interaction nRMSE is
0.00674/0.04524 versus 0.01620/0.66393, and median/worst rollout is
$6.77\times10^{-4}/8.18\times10^{-4}$ versus
$7.54\times10^{-4}/2.64\times10^{-3}$. This closes the noisy symbolic gate:
operator compilation removes derivative noise, argument-range coverage
separates nonlinear atoms, and D-optimal selection controls the remaining
support tail. Replication refines coefficients but is not the source of
identifiability.
Without replicate averaging, a five-seed-per-level noise sweep gives exact
support in 5/5 cases at 1\%, 2\%, and 3\% noise. The corresponding
median/worst interaction nRMSE values are 0.0308/0.0568, 0.0730/0.1143, and
0.1967/0.2720, so quantitative accuracy degrades before atom identity. At
5\%, 7.5\%, and 10\%, support drops to 4/5, 3/5, and 1/5, median interaction
nRMSE rises to 0.661, 1.663, and 4.060, and median rollout rises to
$7.55\times10^{-3}$, $3.82\times10^{-2}$, and $6.88\times10^{-2}$.
Selected clean trajectories cover approximately $u\in[-1.24,1.22]$. Hence
the useful quantitative operating regime is about 2\% noise or less; support
alone at 3\% must not be read as an accurate recovered law.
At fixed 0.5\% raw noise, lowering the true interaction coefficient to 0.02,
0.01, and 0.005 gives exact support in 5/5, 4/5, and 2/5 cases, with
median/worst interaction nRMSE 0.0387/0.1093, 0.1246/1.0, and 1.0/1.392.
Rollout medians nevertheless stay near $8\times10^{-4}$ because the omitted
term becomes dynamically small. This separates a discovery threshold near
coefficient 0.02 from a much weaker prediction threshold and shows why rollout
agreement alone cannot certify a learned physical edge.
Configurable target tests establish both transfer and a parity boundary. With
the same wide-amplitude D-optimal single-observation protocol,
$\sin(3u)\sin(u_x)$ and $\cos(2u)\sin(u_x)$ are each recovered exactly in 5/5
new seeds, with median/worst interaction nRMSE 0.00329/0.00875 and
0.00991/0.01346. In contrast, $\sin(2u)\cos(u_x)$ is exact in only 2/5 with
median nRMSE 0.436: near zero gradient, its even factor has a reaction-like
constant component. IC frequency scaling by two or three yields only 0/5
and 2/5; at scale two, shortening weak windows to 8, 16, or 24 steps yields
0/3 at every setting. Hence the method transfers over state-side poles and
odd-gradient factors, but even-gradient terms need an explicit gauge quotient
or controlled gradient offset rather than indiscriminate high-frequency
excitation.
The appropriate repair is an operator quotient. Replacing each even-gradient
interaction by $a(u)[\cos(q u_x)-1]$ assigns its null-gradient component to the
reaction edge and leaves only irreducible gradient dependence in the tensor
edge. This centered basis is exact in 5/5 pilot seeds. On ten untouched
paired seeds, the ordinary basis is exact in only 4/10 with median/worst
interaction nRMSE 0.441/0.453; the quotient basis is exact in 10/10 with
0.00805/0.0408. Median/worst rollout improves from
$5.64\times10^{-4}/7.95\times10^{-4}$ to
$2.48\times10^{-4}/3.88\times10^{-4}$. Thus hierarchical edge ownership can
be compiled algebraically: annihilate each higher-order atom at the reference
jet of every lower-order edge before selection. This is an identifiability
operation rather than a numerical preconditioner.
With this quotient fixed, random excitation is likewise exact in 10/10
(median/worst interaction nRMSE 0.0133/0.0303), so the algebra itself closes
the support gate. D-optimal selection chiefly controls the dynamic tail,
reducing median/worst rollout from
$3.19\times10^{-4}/1.12\times10^{-3}$ to
$2.48\times10^{-4}/3.88\times10^{-4}$.
The quotient transfers to $\cos(2u)[\cos(u_x)-1]$ and
$\sin(3u)[\cos(2u_x)-1]$, each exact in 5/5 seeds, with median/worst
interaction nRMSE 0.00994/0.0301 and 0.0467/0.0685. The higher-gradient case
has median/worst rollout $8.09\times10^{-4}/1.26\times10^{-3}$. The
state-cosine case reveals a residual hierarchy boundary: two fits misidentify
the lower-order reaction component, producing worst rollout 0.1156 despite
exact tensor attribution. Thus the quotient is reusable, but its receiving
lower-order edge also requires robust staged identification.
A one-pass lower-first scheme is a negative control: selecting five base atoms
before one interaction fails in 0/5, with median/worst interaction nRMSE
0.898/0.909 and rollout $8.25\times10^{-3}/8.60\times10^{-2}$. Freezing an
early surrogate makes the hierarchy irreversible. The next solver must use
alternation or hierarchical group constraints within a joint objective.
The negative results determine the architecture. The plain cardinal model is
excellent in interpolation and resolution transfer but has median amplitude-
OOD errors of order $0.16$--$0.19$. A cubic carrier only partly repairs the
non-polynomial cases. Once the correct pole carrier is selected, fitting a
local spline innovation is neutral or harmful. Hence the current hypothesis
is narrower and stronger than generic KAN substitution: discover a sparse
operator reproduction space for the unknown constitutive edge, compile exact
outer calculus, and introduce local spline innovations only when held-out
evidence demonstrates residual mismatch.
\section{Limitations}
The current implementation establishes the algebraic path but is not yet a complete scientific library. Important limitations remain:
\begin{enumerate}[leftmargin=2em]
\item E-spline basis functions are represented in simplified benchmark-specific forms.
\item FRI moment inputs are exact in the current verification scripts; noisy measurement pipelines remain future work.
\item Boundary conditions must be formalized per operator: periodic, finite interval, causal, or corrected by null-space constraints.
\item The first multidimensional tensor-product Hermite solver is implemented, but it remains a synthetic periodic validation rather than a calibrated external CFD benchmark.
\item Fourth-order structural mechanics is validated on a manufactured clamped plate, but it still requires dedicated biharmonic preconditioning and external benchmark comparison.
\item Benchmark comparisons against external PINN/SIREN baselines must be standardized and repeated under controlled conditions.
\item The operator-compiled constitutive result now includes controlled 2D anisotropic and coupled two-field laws, noisy weak multipole prediction, and identified noisy bivariate interactions, but noisy pole-level identifiability, broader continuous interaction dictionaries, high-dimensional interaction selection, and external trajectories remain open validation gates.
\end{enumerate}
\section{Conclusion}
OSNR is best understood as a continuous-domain signal-processing architecture that adopts the interface of neural representations while rejecting their blind optimization core. For operator-bound fields, the correct spline dictionary collapses training into stable coefficient recovery. For sparse non-Gaussian fields, FRI localization and matched sparse dictionaries avoid the grid leakage and coherence traps of uniform frames. The external sparse-assimilation results add a practical systems role: OSNR can act as a deterministic test-time correction layer on top of physical or neural priors, with the Darcy U-Net adapter, the PDEBench Test~28 temporal FNO rescue, and the station-gated Test~28 neural-prior DST high-pass ladder showing the same mechanism in static elliptic and time-dependent vorticity settings. The engineering rule is strict: continuous-domain exactness only survives when the discrete bridge is correct. That bridge consists of calibrated knot maps, stable bases, proper inner products, cross-Gram coupling, boundary-aware solvers, and regularized inverses where identifiability fails.
\appendix
\section{Hermite Block-Circulant Inner Products}
\label{app:hermite-block-gram}
This appendix records the coefficient-domain inner-product calculus for the second-order Hermite tier. It is included explicitly because the OSNR use of the full autocorrelation tensor is a library-level construction rather than a theorem that can be cited as a pre-existing implementation recipe.
\subsection{Multichannel synthesis}
Let the Hermite generator be
\[
\Phi(t)=
\begin{bmatrix}
\phi_0(t) & \phi_1(t) & \phi_2(t)
\end{bmatrix}^{\top},
\]
with $\operatorname{supp}\phi_p\subset[-1,1]$. A coefficient sequence is
\[
\mathbf{c}[k]=
\begin{bmatrix}
c_0[k] & c_1[k] & c_2[k]
\end{bmatrix}^{\top}.
\]
The continuous field is
\[
f(t)=\sum_{k\in\Z}\mathbf{c}[k]^{\top}\Phi(t-k)
=\sum_{p=0}^{2}\sum_{k\in\Z}c_p[k]\phi_p(t-k).
\]
\subsection{Autocorrelation tensor}
The continuous $L_2$ energy expands into nine generator-pair channels:
\[
\|f\|_{L_2}^2
=
\int_{\R}f(t)^2\,\dd t
=
\sum_{p=0}^{2}\sum_{q=0}^{2}
\sum_{k\in\Z}\sum_{m\in\Z}
c_p[k]c_q[m]
\int_{\R}\phi_p(t-k)\phi_q(t-m)\,\dd t.
\]
Define
\[
\Gamma_{pq}[n]
=
\int_{\R}\phi_p(t)\phi_q(t+n)\,\dd t.
\]
Changing variables gives
\[
\int_{\R}\phi_p(t-k)\phi_q(t-m)\,\dd t
=
\Gamma_{pq}[k-m].
\]
Therefore
\[
\|f\|_{L_2}^2
=
\sum_{p=0}^{2}\sum_{q=0}^{2}
\sum_{k,m} c_p[k]\Gamma_{pq}[k-m]c_q[m].
\]
The tensor $\Gamma_{pq}$ is the Hermite analogue of the scalar spline autocorrelation filter used in periodic spline inner-product calculus \cite{badoual2016inner,badoual2018periodic}.
Because every $\phi_p$ is supported on $[-1,1]$, $\Gamma_{pq}[n]=0$ for $|n|\geq 2$ in the non-periodic infinite-line case. Only the shifts $n\in\{-1,0,1\}$ can contribute. This local support is the algebraic reason the spatial-domain block Gram is sparse before periodization.
\subsection{Block Toeplitz and block-circulant matrices}
For a finite coefficient vector with $M$ knots, define the $M\times M$ block
\[
[\mathbf{\Gamma}_{pq}]_{k,m}
=
\Gamma_{pq}[k-m].
\]
On the infinite line or with finite non-periodic truncation, these blocks are Toeplitz away from boundary corrections. Under periodic boundary conditions, indices are taken modulo $M$, and each block becomes circulant:
\[
[\mathbf{\Gamma}^{\mathrm{per}}_{pq}]_{k,m}
=
\Gamma^{\mathrm{per}}_{pq}[(k-m)\bmod M],
\]
where
\[
\Gamma^{\mathrm{per}}_{pq}[n]
=
\sum_{\ell\in\Z}\Gamma_{pq}[n+\ell M].
\]
The full Hermite Gram matrix is the $3M\times 3M$ block matrix
\[
\mathbf{\Gamma}_{\mathrm{H}}
=
\begin{bmatrix}
\mathbf{\Gamma}_{00} & \mathbf{\Gamma}_{01} & \mathbf{\Gamma}_{02}\\
\mathbf{\Gamma}_{10} & \mathbf{\Gamma}_{11} & \mathbf{\Gamma}_{12}\\
\mathbf{\Gamma}_{20} & \mathbf{\Gamma}_{21} & \mathbf{\Gamma}_{22}
\end{bmatrix}.
\]
The transpose symmetry follows directly from the definition:
\[
\Gamma_{pq}[n]
=
\Gamma_{qp}[-n],
\qquad
\mathbf{\Gamma}_{pq}
=
\mathbf{\Gamma}_{qp}^{\top}.
\]
Thus the diagonal blocks are symmetric, while the off-diagonal blocks need not be symmetric individually. In particular, the slope channel is odd/asymmetric for the standard Hermite construction, so value-slope and curvature-slope blocks encode directional cross-talk.
\subsection{Fourier block diagonalization}
Let $\widehat{\Gamma}_{pq}[\ell]$ be the $M$-point DFT of the first column of the circulant block $\mathbf{\Gamma}^{\mathrm{per}}_{pq}$. The DFT simultaneously diagonalizes all nine circulant blocks:
\[
\mathbf{\Gamma}^{\mathrm{per}}_{pq}
=
\mathbf{F}^{-1}
\operatorname{diag}(\widehat{\Gamma}_{pq}[0],\ldots,\widehat{\Gamma}_{pq}[M-1])
\mathbf{F}.
\]
After applying the DFT to the knot dimension of each Hermite channel, the large $3M\times 3M$ system decouples into $M$ independent $3\times 3$ Hermite channel systems:
\[
\widehat{\mathbf{\Gamma}}[\ell]
=
\begin{bmatrix}
\widehat{\Gamma}_{00}[\ell] & \widehat{\Gamma}_{01}[\ell] & \widehat{\Gamma}_{02}[\ell]\\
\widehat{\Gamma}_{10}[\ell] & \widehat{\Gamma}_{11}[\ell] & \widehat{\Gamma}_{12}[\ell]\\
\widehat{\Gamma}_{20}[\ell] & \widehat{\Gamma}_{21}[\ell] & \widehat{\Gamma}_{22}[\ell]
\end{bmatrix}.
\]
For a right-hand side with Hermite-channel DFT coefficients $\widehat{\mathbf{b}}[\ell]\in\C^3$, the exact periodic normal-equation solve is
\[
\widehat{\mathbf{c}}[\ell]
=
\left(\widehat{\mathbf{\Gamma}}[\ell]+\gamma\mathbf{I}_3\right)^{-1}
\widehat{\mathbf{b}}[\ell],
\qquad \ell=0,\ldots,M-1.
\]
The ridge $\gamma$ is optional for strictly Riesz-stable settings but mandatory in finite precision whenever boundary constraints, redundant channels, or composite dictionaries create near-null directions. The complexity is $O(3M\log M)$ for the channel FFTs plus $O(27M)$ for the $M$ dense $3\times 3$ solves, instead of $O((3M)^3)$ for a dense inversion.
\subsection{Physical meaning for OSNR Tier 3}
The block matrix is not a bookkeeping artifact. It is the exact continuous $L_2$ metric for value, slope, and curvature streams. For a Hermite neural operator, a branch encoder may output coefficient tensors, but the comparison of predicted and target fields should be performed through
\[
\langle f,g\rangle_{L_2}
=
\sum_{p,q=0}^{2}\mathbf{c}_{f,p}^{\top}\mathbf{\Gamma}_{pq}\mathbf{c}_{g,q},
\]
not through a sampled coordinate loss unless sampling is required by the measurement model. Boundary clamping is similarly direct: at a boundary knot $k_b$, Dirichlet, Neumann, and curvature data are imposed by assigning $c_0[k_b]$, $c_1[k_b]$, and $c_2[k_b]$. The Hermite tier therefore converts soft boundary penalties into coefficient constraints and converts continuous PDE energies into block-circulant linear algebra.
\subsection{2D tensor-product block-circulant calculus}
For the 2D tensor-product Hermite generator, index the nine channels by
\[
a=(p_x,p_y),\qquad b=(q_x,q_y),
\qquad p_x,p_y,q_x,q_y\in\{0,1,2\}.
\]
The generator pair is
\[
H_a(x,y)
=
\phi_{p_x}^{x}(x)\phi_{p_y}^{y}(y).
\]
The 2D cross-correlation filter between channels $a$ and $b$ is
\[
\Gamma^{2D}_{ab}[n_x,n_y]
=
\iint_{\R^2}
H_a(x,y)
H_b(x+n_x,y+n_y)\,
\dd x\,\dd y.
\]
By separability,
\[
\Gamma^{2D}_{ab}[n_x,n_y]
=
\Gamma^x_{p_xq_x}[n_x]\,
\Gamma^y_{p_yq_y}[n_y].
\]
Since each one-dimensional Hermite generator is supported on $[-1,1]$, the nonzero spatial shifts satisfy
\[
n_x,n_y\in\{-1,0,1\}.
\]
For a periodic $M_x\times M_y$ grid, every channel pair defines a block-circulant-with-circulant-blocks matrix. The full 2D Hermite Gram is
\[
\mathbf{\Gamma}_{\mathrm{H},2D}
=
\begin{bmatrix}
\mathbf{\Gamma}_{00} & \cdots & \mathbf{\Gamma}_{08}\\
\vdots & \ddots & \vdots\\
\mathbf{\Gamma}_{80} & \cdots & \mathbf{\Gamma}_{88}
\end{bmatrix},
\]
where each $\mathbf{\Gamma}_{ab}$ is the circulant 2D convolution operator associated with $\Gamma^{2D}_{ab}$.
Let $\widehat{\Gamma}^{2D}_{ab}[\nu_y,\nu_x]$ be the 2D DFT of the first column/filter of the $(a,b)$ block. Applying a 2D DFT to the spatial dimensions of all coefficient channels yields, for every frequency coordinate $(\nu_y,\nu_x)$, the local $9\times9$ system
\[
\widehat{\mathbf{\Gamma}}_{2D}[\nu_y,\nu_x]
\widehat{\mathbf{c}}[\nu_y,\nu_x]
=
\widehat{\mathbf{b}}[\nu_y,\nu_x],
\]
with
\[
\widehat{\mathbf{\Gamma}}_{2D}[\nu_y,\nu_x]
=
\begin{bmatrix}
\widehat{\Gamma}^{2D}_{00}[\nu_y,\nu_x] & \cdots & \widehat{\Gamma}^{2D}_{08}[\nu_y,\nu_x]\\
\vdots & \ddots & \vdots\\
\widehat{\Gamma}^{2D}_{80}[\nu_y,\nu_x] & \cdots & \widehat{\Gamma}^{2D}_{88}[\nu_y,\nu_x]
\end{bmatrix}.
\]
Thus the global dense inverse is replaced by
\[
\widehat{\mathbf{c}}[\nu_y,\nu_x]
=
\left(
\widehat{\mathbf{\Gamma}}_{2D}[\nu_y,\nu_x]
+\gamma\mathbf{I}_9
\right)^{-1}
\widehat{\mathbf{b}}[\nu_y,\nu_x].
\]
In code, this is the \texttt{torch.fft.fft2} path implemented by the 2D block-Fourier solver. The spatial complexity is governed by the FFTs, $O(9M_xM_y\log(M_xM_y))$, while each frequency bin performs a constant-size $9\times9$ complex solve. This is the algebraic mechanism behind the measured $1.28$ ms per-frame 2D tensor-Hermite fluid validation and the fourth-order biharmonic structural-shell validation.
\nocite{khalidov2006differential,forster2006complex,unser2003wavelet,unser2007selfsimilarity1,blu2007selfsimilarity2,pad2015operator,pad2017optimized,parhi2023cycle,dadi2020matched,blu2003kernels,blu2004linear,schmitter2015shape,schmitter2018landmark,vandeville2004hex,vandeville2005polyharmonic,wandel2022splinepinn,schmitter2016hermite,lu2021deeponet,jin2022mionet}
\input{theory_sensor_composition}
\input{theory_policy_transfer}
\input{colony_flagship_20260914}
\input{inspection_memory_20260914}
\input{question_inspection_20260914}
\bibliographystyle{plain}
\bibliography{references}
\end{document}