Research record

Policy-transfer theory and experiments

Historical source. Some claims in older records were subsequently corrected. The associated article states the adopted interpretation. This record preserves the original source alongside its rendered reading view.

Rendered archival TeX

This is an HTML reading rendition of the local TeX record. Mathematical notation is rendered with KaTeX; archived figures are included when their source assets are part of this collection.

From compact physical models to useful policies: an empirical boundary

Exact operator calculus and compact coefficient estimation solve a model construction problem. Useful control additionally depends on observation coverage, model errors along the policy's trajectories, policy optimization and selection. A low prediction residual is not a bound on recovered task return without further assumptions. This distinction motivates a controlled full-policy refinement test instead of another inverse-evaluation microbenchmark.

Scope of the joint quadratic identification step.

Write the static coefficients as c=μ+Da/(1−ρ)c=\mu+Da/(1-\rho), where μ,D\mu,D are fixed source quantities and 0≤ρ≤0.850\leq\rho\leq0.85. For the unclipped one-step model, the predicted response is Bμ+[p−BμBD][ρa],B\mu+\begin{bmatrix}p-B\mu&BD\end{bmatrix} \begin{bmatrix}\rho\\a\end{bmatrix}, so squared fitting error and coefficient-monotonicity constraints are quadratic and linear, respectively, in (ρ,a)(\rho,a). The implemented regularizers act on the scaled correction DaDa and scaled drive (1−ρ)μ+Da(1-\rho)\mu+Da. In static-law coordinates, their penalties are λ(1−ρ)2∥F(c−μ)∥2\lambda(1-\rho)^2\|F(c-\mu)\|^2 and κ(1−ρ)2∥F2c∥2\kappa(1-\rho)^2\|F_2c\|^2, not pole-independent static-law penalties. The small pole ridge is also retained. Furthermore, deployment uses ρp+(1−ρ)clip⁡(Bc,−1,1)\rho p+(1-\rho)\operatorname{clip}(Bc,-1,1); this clipped loss is not the quadratic objective being minimized. Convexity therefore describes the specified identification surrogate, not arbitrary pole learning or globally optimal fitting of the clipped physical simulator. Executable objective identities record this distinction without changing any frozen fit.

Finite-region treatment of the clipped objective.

For sorted commands and coefficient-monotone scalar functions, observations can only form a lower-saturated prefix, an interior block and an upper-saturated suffix. There are (n+1)(n+2)/2(n+1)(n+2)/2 assignments, including empty blocks and threshold ties. With gi=(1−ρ)Biμ+BiDag_i=(1-\rho)B_i\mu+B_iDa, deployment is ρpi+clip⁡(gi,−(1−ρ),1−ρ)\rho p_i+\operatorname{clip}(g_i,-(1-\rho),1-\rho). Its three region formulas are respectively ρ(pi+1)−1,Biμ+ρ(pi−Biμ)+BiDa,ρ(pi−1)+1.\rho(p_i+1)-1,\qquad B_i\mu+\rho(p_i-B_i\mu)+B_iDa,\qquad \rho(p_i-1)+1. Each is affine in (ρ,a)(\rho,a), and region consistency is linear in these parameters. Thus enumeration of the assignments reduces the specified clipped loss, with the unchanged scaled regularizers, to finitely many convex quadratic subproblems. Monotonicity permits checking only the boundary observations of each block. This uses a fixed basis and one stable response pole; it is not a convex formulation for arbitrary learned exponential generators or higher-order pole sets.

A lower bound for any region follows by dropping the nonnegative interior loss and function penalties, retaining saturated observations and pole ridge: Lsat=min⁡0≤ρ≤0.85{12∑i∈S(αiρ+bi−yi)2+10−82ρ2}.L_{\rm sat}=\min_{0\leq\rho\leq0.85} \left\{\frac12\sum_{i\in{\cal S}} (\alpha_i\rho+b_i-y_i)^2+\frac{10^{-8}}2\rho^2\right\}. Here (αi,bi)=(pi+1,−1)(\alpha_i,b_i)=(p_i+1,-1) or (pi−1,1)(p_i-1,1). Its minimizer is the box projection of −∑iαi(bi−yi)/(∑iαi2+10−8)-\sum_i\alpha_i(b_i-y_i)/(\sum_i\alpha_i^2+10^{-8}). Prefix sums supply bounds for every clipping assignment. A bound exceeding the current feasible objective excludes that region without an LP or QP. This is standard lower-bound pruning specialized to the operator model, not a new general optimization theorem. Floating-point implementations retain rounding guards, LP statuses and QP dual checks; these are numerical accounting, not exact-arithmetic certificates or statistical uncertainty.

The exhaustive reference tests 24 scalar fits on two exposed laws and two representations: 3,672 regions. Six-actuator time is 1.76–2.00 seconds. One high-loss QP in each asymmetric case reports an unresolved optimization status; the frozen complete-accounting rule excludes both from control. Both compound cases resolve all regions. Despite improved aggregate later-record prediction and pole estimates, replacing the inverse beneath unchanged actors changes shared restoration from 0.8894 to 0.8865 and polynomial restoration from 0.8518 to 0.8528. All 128k evaluator transitions and original bitwise replay checks pass. Better identification does not automatically improve a policy trained with a different inverse.

Saturated-row pruning is 7–13×\times faster than exhaustive enumeration in a three-repeat exposed-case CPU comparison, but six-actuator times remain 136–281 ms. A stronger bound drops the linear region/monotonicity constraints from each complete quadratic while retaining the pole box. For H≻0H\succ0, let z0=H−1rz_0=H^{-1}r and v=H−1e0v=H^{-1}e_0; its relaxed box minimizer is zbox=z0+v clip⁡((z0)0,0,0.85)−(z0)0v0.z_{\rm box}=z_0+ v\,\frac{\operatorname{clip}((z_0)_0,0,0.85)-(z_0)_0}{v_0}. This follows by eliminating the remaining coordinates and minimizing the scalar Schur complement. Batched matrix products construct the region Hessians; batched solves and residual-corrected dual values then screen them before any generic constrained optimization. These small data Hessians are dense and are not assumed circulant.

In a separate paired comparison, matrix bounds reduce six-actuator time to 39–67 ms with identical preceding objective values and 13–33 final QPs out of 918 candidate regions. A broader exposed-archive check completes all 576 scalar fits. Ninety of 96 six-actuator cases are below 100 ms, but the worst is 133.39 ms; the all-cases latency condition fails. Retained scalar bound arrays occupy 226,440 bytes, excluding other temporaries; peak process RSS is 303.80 MiB. No control outcome is changed by this compiler comparison.

In Gate 333, four fixed world models—shared six-coordinate with sixteen target observations, cold MLP with sixteen, cardinal with 128, and privileged exact law—each support standard SAC training on four exposed lag families and two seeds. The shared representation inherits 768 prior observations; every actor/critic inherits a 20M-transition pretrained policy. Mechanics and six force sensors are supplied. Each run uses 200,000 virtual transitions, three virtual validation resets and eleven eligible checkpoints. Training is backed up before evaluation on eight held-out changed-physics resets. All 32 runs complete with immutable worlds and finite trajectories.

WorldDeadzoneSmoothAsymmetricCompound
Shared six, 160.86880.86840.84780.8011
Cold MLP, 160.88370.37080.21470.7398
Cardinal, 1280.87100.85440.84380.8116
Exact law0.86470.88060.79030.7724

Entries are mean paired reset-normalized returns within a training run, then median across two training seeds. They are not independent-law confidence estimates. Primary gains over its initial policy are 0.0021/0.0206/0.2190/0.1030 nominal units. The last two clear the five-point gain clause, but every family misses 90% nominal restoration. All primary training runs remain within 600 seconds on one memory-capped MPS learner. The complete frozen panel therefore fails.

An exact-law learner is a diagnostic of the specified learning procedure, budget and reset-selection protocol, not an optimal-control upper bound. A learned-world policy exceeding it does not establish more accurate physics. Likewise, failure of a cold small-data MLP does not compare against all neural models or a neural shared prior. This is offline simulated-policy refinement for later episodes, not uninterrupted hardware repair or a gradient-free end-to-end method. The next frozen test triples the learning budget and includes a nested budget control and matched shared-prior polynomial directions, without replacing these failed thresholds.

Budget and selection follow-up.

Gate 337 completes all twelve 600k runs and their audits. Shared-six median restoration is 0.9228/0.8894 on asymmetric/compound lag, versus polynomial 0.8719/0.8518 and exact-law learning 0.9051/0.9130. The primary gains 0.0763/0.0754 over nested 200k and 0.05085/0.03763 over polynomial: both predeclared extra-budget and coordinate-specific clauses pass. The full gate fails the compound 0.90 restoration clause. All primary runs meet 1,800 seconds. This is two-seed evidence on exposed laws with new resets, not independent-law generalization or immediate repair. Both fits use the same cardinal carrier. The functional-PCA versus fixed polynomial comparison does not isolate exact-Gram geometry from source-learned subspace selection; a matched learned coefficient-metric prior is absent.

Gate 339 instead retains the older eight 200k archives and evaluates their 88 checkpoints on sixteen new virtual reset states before physical testing. Expanded selection restores only 0.8359/0.8010 for shared-six asymmetric/compound policies, versus 0.8472/0.8007 originally: the gate fails. The separate post-panel maximum of mean paired-normalized return over each finite archive is below 0.8554 in all eight cases. Thus no single fixed checkpoint from these archives can meet 90% on this test panel. This is neither a control-theoretic ceiling nor permission to select using future physical outcomes. It motivates testing added capability rather than further selection tuning on already exposed responses.

Descriptor equivalence and saturation boundaries.

For fitted models c=cˉ+Dac=\bar c+Da with D⊤GD=ID^\top GD=I, the shared descriptor is aa and the polynomial descriptor is E⊤GD aE^\top GD\,a. For the frozen source geometry, the latter six-by-six map is invertible with condition number 1.66186. Standardization and the common pole preserve an affine relation, which an unrestricted first affine policy layer can absorb. Thus these encodings do not add target-model information or neural expressivity on the fitted shared subspace. Their source-bank projections are not globally equivalent, and finite-budget optimization need not agree. Three source-only tests verify this distinction. In contrast, Gate 337 fits different subspaces; its completed shared-six/poly medians are 0.9228/0.8719 on asymmetric and 0.8894/0.8518 on compound lag.

Known static clipping is another boundary: coefficient functions can differ only inside saturated tails and induce the same bounded response. Partitioning at cardinal knots and clipping crossings makes every bounded piece constant or cubic. Four-point Gaussian integration then produces exact pairwise products, up to floating point, as a factor Gram. Four tests include an invisible-tail witness and independent integration. The 36 source profiles require 102 intervals and 408 nodes; the resulting Gram is positive semidefinite, not necessarily invertible or circulant. This is uniform-command geometry, not task-risk certification.

A separate 220-byte six-channel cache selects a nearer-to-zero command on each outer saturation plateau. On twelve archived-model validation pairs, all actions, forces, observations and velocities remain bitwise identical, while squared-command cost falls by mean 208.86–215.49 return units across the four policies. Four additional tests verify model-equivalent outputs and immutable caches. A fitted saturation threshold can be wrong on the true plant: these are source-world efficiency results, not physical recovery, electrical-energy savings or amendments to the frozen policy protocols.

Exact stable response geometry: an identity, not a new kernel theorem.

Let ϕi(u)\phi_i(u) be the bounded cardinal response, ρi∈[0,1)\rho_i\in[0,1) its causal pole and bi=ϕi(0)b_i=\phi_i(0) its zero-command equilibrium. Start at force pp, apply uu once and then zero commands. For k≥0k\geq0, direct substitution in the first-order recurrence gives ri(k,u,p)=yk+1−bi=ρik{ρip+(1−ρi)ϕi(u)−bi}.r_i(k,u,p)=y_{k+1}-b_i =\rho_i^k\{\rho_i p+(1-\rho_i)\phi_i(u)-b_i\}. Use the equilibrium once and the infinite transient as a feature, with independent uniform reference measures on p,u∈[−1,1]p,u\in[-1,1]. Writing δi(u)=(1−ρi)ϕi(u)−bi\delta_i(u)=(1-\rho_i)\phi_i(u)-b_i yields Kij=bibj+ρiρj/3+Eu[δi(u)δj(u)]1−ρiρj.K_{ij}=b_i b_j+ \frac{\rho_i\rho_j/3+\mathbb E_u[\delta_i(u)\delta_j(u)]} {1-\rho_i\rho_j}. The proof is the convergent geometric series, E[p]=0\mathbb E[p]=0 and E[p2]=1/3\mathbb E[p^2]=1/3. Static products use the clipped-polynomial partition; no time unrolling is required. This is an inner-product Gram and hence positive semidefinite, without a general invertibility or circulant claim. The equilibrium has unit finite weight: its nonzero constant trajectory is not summed over infinite time. The reference measure is not a task-risk distribution, and stability is essential.

Four tests compare this formula with direct 300-step causal recurrences, check saturation aliases and pole distinctions, verify centered kernel-PCA projection and reject unstable/rank-deficient inputs. A six-mode source encoder plus explicit pole retains 13,128 numeric bytes, including the source curves needed for cross-products. Its response-metric reconstruction error is not a policy-performance bound. The frozen next experiment tests the same source worlds, 600k policy budget, sixteen-row fit and same-episode continuation against static and zero-context controls on new laws; its completed negative capability result is reported below.

Kernels on dynamical systems, including initial conditions and efficient linear-system calculations, have established prior art [vishwanathan2007binet]. Kernel-based Hammerstein identification likewise predates this work [risuleo2016hammerstein]. The scalar response identity and centered kernel PCA are not novelty claims; the open question is useful matched physical-skill transfer.

Fresh-law source-policy reuse: conditional benefit, failed recovery.

Gate 340 completes three source learners (1.8M new virtual transitions plus 378k validation transitions) before evaluating eight new laws and eight paired reset states per law. Every context uses the same sixteen-row fitted shared inverse. Asymmetric/compound restoration is 0.8202/0.8915 for shared context, 0.8287/0.8643 for polynomial and 0.8655/0.8767 for zero context. Primary within-law paired gains over polynomial are 0.02707/0.02671, but gains over zero are −0.04091/+0.01477-0.04091/+0.01477. Both families fail the recovery and zero-context gain clauses; all remaining clauses and resource/trace checks pass. The paired contrast is not the difference of marginal medians. Invertible target encodings do not guarantee equal finite-budget training, but these comparisons also do not establish added target-model information.

Maximum fitting plus encoding is 8.981 ms, primary numerical state is 654,084 bytes, and no target policy-gradient update occurs. All 448 controller episodes and 1,024 separately acquired prefix transitions are preserved. Identical simulated prefixes are replayed with extra fallback steps charging standalone computation; Git/orchestrator delay is not plant time. The changed law is present from reset. This is causal startup adaptation in a known-mechanics, force-sensed simulator, not unknown-change detection, hardware real-time control or a zero-forgetting result. A true inverse under the same actor and inferred context gives only 0.8325/0.8941 restoration; it is a limited intervention, not a fully informed-policy upper bound.

Exact response geometry does not guarantee task-useful reuse.

Gate 341 adds one source learner using the response kernel above, with 600k training and 126k validation transitions and a bitwise matched source world schedule. On eight new laws and eight resets each, asymmetric/compound restoration is 0.8993/0.8582, versus 0.8664/0.8620 for static shared, 0.8273/0.8861 for static polynomial and 0.9021/0.8913 for zero context. Paired response gains over both static encodings exceed two points only on asymmetric laws. Both families fail restoration and the zero-context gain clause; the full operator-specific and reusable-recovery gates fail. Exact uniform-reference response geometry is not the task's trajectory geometry or a control-value guarantee.

All 512k controller and 1,024 prefix transitions pass the trace, immutable actor and charged-delay audit. The maximum fit plus four descriptors is 11.006 ms, numerical primary state 665,028 bytes, evaluator RSS 311.609 MiB. The source learner takes 836.920 s with 790.516 MiB peak RSS and 28,098,560 Metal driver bytes. No target gradient occurs; inherited experience, ongoing force sensing and known mechanics remain essential qualifications. The conditional calibration-in-the-loop source-training gate follows its predeclared failure trigger, not a kernel-selection sweep on these targets.

A physically motivated descriptor can be actively unhelpful.

Fixed-actor interventions after Gate 340 preserve the inverse and all mechanical/force observations. Replacing the inferred normalized descriptor with its source centroid changes shared restoration from 0.8202 to 0.8736 on asymmetric laws and from 0.8915 to 0.8851 on compound laws. The paired original-minus-centroid contrasts are −0.05343/+0.00259-0.05343/+0.00259; all four asymmetric laws improve under the centroid. Polynomial-context contrasts are −0.03413/−0.00906-0.03413/-0.00906. Cyclic channel shifting generally hurts the aggregate, showing sensitivity without establishing uniformly useful information. All 384k diagnostic transitions pass immutable-actor, startup-prefix and bitwise original-replay checks. Zero normalized context is not a zero plant law or the separately trained zero-context policy. Interventions can create out-of-distribution combinations; none modifies the primary experiment. The subsequently completed Gate-341 diagnostic adds 576k transitions. For its response actor, original/centroid restoration is 0.8993/0.9195 and 0.8582/0.8966, with paired original-minus-centroid contrasts −0.02158/−0.01103-0.02158/-0.01103. All original replays remain bitwise identical. The polynomial actor instead benefits from its inferred descriptor on the compound panel (paired gain 0.02471). These results reject a universal context-benefit interpretation without proving context is unnecessary.

Prior art limits the compiler novelty claim.

Cardinal-spline input lifting for Hammerstein identification predates this work [chan2006cardinalhammerstein]; those natural interpolating splines are not identical to the compact uniform-grid carrier used here. Saturation-data partitioning [pupeikis2006saturation], mixed-integer global piecewise-affine identification [roll2004piecewiseidentification], and branching methods for bounded monotone smoothing [sasane2019monotone] also precede our finite-region specialization. The measured acceleration compares our own successive implementations, not those prior solvers. No global-method priority, statistically unbiased noisy-feedback fit or algorithmic SOTA claim follows.

Clipped-loss memory needs more than the ordinary empirical Gram.

On one cardinal cubic cell, Gram entries depend on input moments through degree six. At ui=i/112u_i=i/112, i=0,…,7i=0,\ldots,7, put weights (7i)\binom{7}{i} on even indices in one zero-target dataset and odd indices in another. Their sample counts, full cardinal data Grams and target statistics agree, since the seventh finite difference annihilates every polynomial of degree at most six. Yet the weighted squared clipped losses for the reproduced monotone affine drive 64u−164u-1, with zero pole, differ by 16/716/7. Two tests check the exact rational identity and the actual 37-coordinate basis. Thus the unclipped pooled-Gram contract does not extend automatically to a parameter-dependent clipped objective. Additional region/order information or retained observations are needed; this elementary counterexample does not invalidate immutable model archives.

Clipping-aware inverse and descriptor correction is not sufficient.

The final fresh-law Gate 343 holds Gate-340 source actors fixed and crosses old/new clipped-objective fits independently through the inverse and descriptor. Asymmetric/compound restoration is 0.8166/0.8851 for old/old, 0.8181/0.8932 for new-inverse/old-context, 0.7750/0.8739 for old-inverse/new-context and 0.8039/0.8832 for new/new. The primary new/new paired gains over old/old are −0.01267/−0.00500-0.01267/-0.00500; both recovery gates fail. Positive factorial interactions (0.01624/0.000910.01624/0.00091) do not imply a positive primary effect.

All 384 scalar fits complete numerical region accounting. Both pipelines and encodings take 36.65–133.57 ms, below this separately frozen 200-ms paired-pipeline cap, without erasing earlier 100-ms misses. Common activation occurs at steps 17–19; all rewards and delay steps count. All 641,024 evaluator transitions are retained with unchanged source actors and zero new gradients. Primary numerical state is 651,060 bytes and evaluator RSS 406.047 MiB. Reaggregation is compared in memory by a separate audit adapter after the original exclusive-output writer rejects an existing metrics file; the frozen experiment and its artifacts are not changed. More faithful fitting and faster numerical compilation do not establish useful policy compatibility.

Calibration-consistent source training also fails the recovery gate.

The predeclared conditional Gate 342 trains new response and zero-context actors through the actual sixteen-row fitted inverse. The two source-world schedules, calibration observations, fitted models and startup traces match bitwise; all 1,484 source/validation fits reproduce from their permitted rows. Each actor uses 600k learner transitions plus 11,088 startup transitions; combined source validation adds 252k transitions. The two training times are 879.199/876.969 s, with maximum source RSS 813.625 MiB.

On eight fresh laws, asymmetric/compound restoration is 0.8514/0.8595 for calibrated response and 0.8953/0.8974 for its matched zero-context actor. Previous response gives 0.8654/0.8681. Primary paired gains over matched zero are −0.04580/−0.03713-0.04580/-0.03713, and over previous response −0.02409/−0.00602-0.02409/-0.00602. Both families fail restoration, matched-zero gain and the all-law floor; neither training-consistency clause passes. All 641,024 evaluator transitions and resource checks pass. Maximum fit plus four descriptors is 11.686 ms, primary numerical state 665,020 bytes and evaluator RSS 316.031 MiB. No target policy gradient occurs.

This retires the overnight descriptor/training-recipe branch under its predeclared stopping rule. The previous-response comparison changes a whole training recipe, not only one matched causal factor; only the two new arms are fully matched. The zero-context control also misses 90% in both family medians and is not a post-hoc passing primary. Known mechanics, ongoing force sensing, inherited pretraining, startup changes and one source seed limit the interpretation. A true inverse with the primary's inferred descriptor remains a limited intervention, not an optimal-policy bound.

Same-policy feasibility limits the remaining inversion hypothesis.

Gate 344 fixes the existing calibrated-zero actor and removes descriptor differences across all methods. Sixteen fresh laws with eight resets each give shared-spline restoration 0.9084/0.8966 (asymmetric/compound), frozen-RLS 0.8850/0.8722, online-RLS 0.8766/0.8568, polynomial-prior 0.8999/0.8901, and true-inverse 0.9123/0.8977. Paired gains over frozen RLS are 0.02124/0.02292 and over online RLS 0.02900/0.05269; descriptive law-bootstrap lower bounds are positive. The polynomial advantage remains below the required two points in both families, while compound misses recovery and the all-law floor. Both the full operator-capability gate and the two-family true-inverse feasibility criterion fail.

This is a modest matched-controller identification result, not universal operator superiority. Shared and polynomial models have the same cardinal carrier; learned subspace and geometry are not independently isolated. True inversion's paired advantage over the fitted spline is small and not uniformly positive. It does not bound an optimized policy, but weakens the case for more fitting refinements as the route to a large gain here. All 898,048 transitions, 1,536 scalar-fit reproductions and actor/inverse/ sensor/force recurrences pass. Maximum common fitting is 15.488 ms, every activation is step 17, numerical state is 651,900 bytes and RSS 317.672 MiB. No new policy training occurs; the selected actor inherits 611,088 source training transitions, 126k validation transitions and earlier 20M pretraining. These are startup-change, continuously force-sensed simulations, not a demonstration of unknown-change detection, safe probing or useful memory.

Direct forward planning: valid algebra, unsuccessful development task.

Gate 345 changes the operational task to reduced suspended-load motion, with known linear mechanics, position/angle sensing and an unknown monotone drive law with stable lag. In the oracle-model pilot, augmenting the state by force gives zk+1=Aρzk+Bρg(uk)z_{k+1}=A_\rho z_k+B_\rho g(u_k); lifting a finite horizon gives Z=Sρz0+TρVZ=S_\rho z_0+T_\rho V. A quadratic state/static-drive objective with input bounds is therefore a condensed convex quadratic program. This is established Hammerstein/MPC structure, not a new consequence unique to cardinal splines. Filtering basis rows through the same operator also makes position observations linear in constitutive coefficients for fixed pole, with separate linear initial-state nuisance columns. General finite-horizon Grams are not circulant.

Five numerical tests and four full pilot audits validate these identities, observer recurrences, commands and constrained-plan dual-gap witnesses. Nevertheless, the true-forward whole cost is 302.9361 versus true-inverse 304.8170: only 0.617% improvement. Settled-position RMS 0.1260 m fails the proposed 0.035 m target, as does nominal 0.1251 m. The reference preview allows early departures that conflict with the declared settling criterion. We retain this failed task contract and do not retune it on its exposed trace. All 1,600 development transitions complete, with primary p95 1.970 ms and 263,920 bytes of numeric controller state. No proposed fresh-panel laws are sampled, no learner is trained, and no fresh-panel or safety claim follows.

Adverse evidence must survive a retrieval decision, but rejection does not repair.

The 192 archived historical/current-only Gate 332 traces reproduce every alarm and reuse score. At long dwell, 510/566 low-noise and 1,742/1,778 high-noise returning historical reuses fail the selected model's own alarm threshold on the four triggering observations. The subsequent sixteen probes are evaluated by a different RMS/floor criterion and can miss that local discrepancy. Current-only behaves similarly. These empirical threshold failures are not uniform noise or physical-model certificates.

In a fixed exposed-case intervention, reuse is additionally checked against the retained alarm rows. All 270,400 physical steps complete and untreated trajectories replay bitwise. Low-noise historical RMS rises from 0.002084 to 0.004193, probe actions from 1,504 to 7,856, and the full bank refuses 54 fits. High-noise probes fall from 2,272 to 800 but RMS rises from 0.003655 to 0.003869. Thus consistency filtering alone does not establish useful memory. The result motivates a separately tested local correction and subsequent- observation verifier, not relaxed thresholds or a retroactive passing score.

Compact local correction composes the calculus without global task guarantees.

Let a constitutive parent be s(u)s(u) and propose sa(u)=s(u)+∑iaihi(u)s_a(u)=s(u)+\sum_i a_i h_i(u), where cardinal linear hats hih_i are supported inside two parent cells and the endpoint corrections are zero. The continuous penalties a⊤Maa^\top M a and a⊤Kaa^\top K a use exact hat mass and derivative Grams. Parent cubic derivative minima on each subcell turn monotonicity into linear slope inequalities. For each retained observation block bb, the squared-loss change under a scaled proposal is 2η⟨eb,Δb⟩+η2∥Δb∥22\eta\langle e_b,\Delta_b\rangle+\eta^2\|\Delta_b\|^2. Choosing a common η∈[0,1]\eta\in[0,1] below all favorable quadratic roots preserves the observed block losses and the convex monotonicity constraints. These are standard compact-support and quadratic identities, here combined with a sixteen-subsequent-observation deployment verifier. They are not a future-risk or physical stability theorem.

The exposed twelve-trajectory pilot completes all 405,600 evaluator steps and full causal replay/fit/validation checks. Low-noise repaired-bank RMS is 0.001357, versus original 0.002084 and matched repaired current-only 0.001383; probe actions fall from 1,504 to 560. High-noise RMS 0.003876 is worse than original 0.003655 and no correction is accepted. Thus the operational gain is conditional and unique historical value remains limited. An audit recovers an overwritten line-search metadata field from unchanged saved coefficients; no experimental outputs are replaced. Static complement preservation must also not be confused with trajectory preservation: for identical commands outside the support, a remaining force-state difference obeys Δfk+m=ρmΔfk\Delta f_{k+m}=\rho^m\Delta f_k, rather than becoming identically zero. Different closed-loop commands require a separate analysis.

The unchanged mechanism subsequently passes its frozen eight-law, two-noise, six-control replication. All 3,244,800 evaluator transitions and complete independent replays pass. Low-noise whole-error/original and consistent- no-repair median ratios are 0.5539 and 0.3666; their descriptive paired-law intervals lie below one. High-noise ratio to consistent no-repair is 1.0002, so the high-noise result does not isolate a patching benefit. Against matched repaired current-only, returning early-error ratios 0.7641/0.7086 and probe ratios 0.3700/0.2933 pass both declared memory screens. This operational improvement is stronger than unchanged archive coefficients, but not uniform: one high-noise law has 17.03% worse whole error. Primary numeric state peaks at 52,176 bytes. A fixed gain-one integral-feedback control fails to match the learned controllers on the two exposed pilot cases despite zero probes and 32-byte state; all its action/physical replays pass. Neither this narrow classical falsifier nor the eight-law median screen is a general stability, no-forgetting or robotics-superiority theorem.

Alarm attribution requires separating observation and model errors.

For a frozen fitted law on a recorded command sequence, write the force observation as f~k=fk+ϵk\tilde f_k=f_k+\epsilon_k. Then ρ^f~k−1+(1−ρ^)s^(uk)−f~k=ekmodel+ρ^ϵk−1−ϵk,ekmodel=ρ^fk−1+(1−ρ^)s^(uk)−fk.\hat\rho\tilde f_{k-1}+(1-\hat\rho)\hat s(u_k)-\tilde f_k =e_k^{\rm model}+\hat\rho\epsilon_{k-1}-\epsilon_k, \qquad e_k^{\rm model}=\hat\rho f_{k-1}+(1-\hat\rho)\hat s(u_k)-f_k. Squared-error attribution must retain the cross term; pooled energy and counts of threshold events answer different questions. In 64 old long-dwell traces this decomposition reproduces all observed alarms. For returning historical runs, model-only errors exceed the four-row threshold in 630/682 low-noise alarms, with zero noise-only events. At high noise, 51/1,791 meet the model-only predicate while 1,733 meet the noise-only predicate under the valid identity law. The component predicates are not generally disjoint. The nominal empirical threshold is below the high-noise observation RMS. Neither a local approximation correction nor a retrieval mechanism remedies that statistical mismatch by itself. Actual commands and previously fitted coefficients remain noise-dependent in this diagnostic, so removing the explicit observation term is not a counterfactual noise-free experiment or a new uniform validity guarantee.

Predictable paired validation avoids a Gaussian inverse-force assumption.

Let parent and candidate forecasts pt,ctp_t,c_t be measurable before observing yt=μt+ϵty_t=\mu_t+\epsilon_t, with conditionally Gaussian noise of known covariance Σt\Sigma_t. No membership of μt\mu_t in either model class is required. For dt=ct−ptd_t=c_t-p_t, observed gain Dt=∥pt−yt∥2−∥ct−yt∥2D_t=\|p_t-y_t\|^2-\|c_t-y_t\|^2 and conditional mean gain At=∥pt−μt∥2−∥ct−μt∥2A_t=\|p_t-\mu_t\|^2-\|c_t-\mu_t\|^2, Dt−At=2dt⊤ϵt,Vn=4∑t≤ndt⊤Σtdt.D_t-A_t=2d_t^\top\epsilon_t,\qquad V_n=4\sum_{t\le n}d_t^\top\Sigma_t d_t. The classical two-sided normal-mixture boundary[howard2021confidence] therefore yields, for fixed η>0\eta>0, Pr⁡ ⁣(∀n:∣∑t≤n(Dt−At)∣<(Vn+η)log⁡Vn+ηηα2)≥1−α.\Pr\!\left(\forall n:\left|\sum_{t\le n}(D_t-A_t)\right| <\sqrt{(V_n+\eta)\log\frac{V_n+\eta}{\eta\alpha^2}}\right)\ge1-\alpha. Summable error allocations over adaptively initiated comparisons protect multiple proposals, provided each forecast remains predictable and the noise proxy is valid. A positive cumulative lower bound certifies improvement on those conditional forecast losses, not every state or a different control trajectory. If an observation map is linear in an edit, dt=Ftδd_t=F_t\delta, gain and variance are coefficient products with residual moments and Grams; an arbitrary nonlinear simulator does not inherit that linearity. Identical predictions contribute exactly zero gain and variance.

Using known next-velocity noise rather than assuming Gaussian output from the bounded nonlinear force observer, a frozen retrospective audit finds positive final lower bounds on all 66 existing sixteen-row validation blocks, including all 61 accepted repairs. All 1,056 prefix intervals cover evaluator truth; the five extra positive comparisons had been rejected by a different force- adequacy predicate. This is not a new-controller false-update estimate or a retroactive prospective trial. The sequential boundary is established prior work; the application connects local program repair to its stated observation contract without promoting prediction confidence to physical safety.

The subsequent frozen prospective-supervisor pilot makes this separation operationally concrete. All 415 selected comparisons have positive true cumulative forecast gain, and 405,600 physical steps plus an equal independent replay pass. Yet the primary exhausts fitting before returning regimes in both noise conditions and fails the useful-memory and repair-attribution screens. Low-noise whole RMS is 2.618 times the old repaired bank's; probe counts are 3.886/3.163 times old low/high counts. All thirteen low-noise rejected primary global proposals are genuinely worse on their validation states. A correct rejection rule does not make continued acquisition useful. The statistical statement concerns predictable losses at issued actions; it neither identifies a regime switch nor establishes value of further experimentation or performance at an edited model's inverse actions.

Ending rejected transactions is a causal lifecycle intervention, not a change to that statistical theorem. In the next frozen pilot, 88,064 old-prefix steps match bitwise before the first rejection, and 270,400 new steps plus equal physical replay pass. Primary error improves 24.60%/17.07% over the failed supervisor without primary exhaustion, but full useful-memory and repair-attribution screens still fail. High-noise historical memory has 1.1785 times matched current-only error. All 593 selected comparisons have positive true prediction gain, reinforcing rather than closing the gap between comparison validity, experiment value and useful behavior.

A cardinal counterexample at the inverse-action interface.

Let the true static plant be f(u)=uf(u)=u, parent p(u)=2up(u)=2u, and request r=1/2r=1/2. Define a continuous piecewise-linear cc through (0,0),(1/4,1/4),(7/8,3/10),(1,2)(0,0),(1/4,1/4),(7/8,3/10),(1,2), with c=pc=p outside [0,1][0,1]. Its correction c−pc-p is an exact cardinal B1B_1 expansion on spacing 1/81/8 with zero endpoint coefficients. For two nonempty blocks of observations at the parent action up=p−1(r)=1/4u_p=p^{-1}(r)=1/4, the existing quadratic guard admits q=p+(19/20)(c−p)q=p+(19/20)(c-p). Each residual falls from 1/41/4 to 1/801/80, a 99.75% SSE reduction. The minimum slope of qq is 22/125>022/125>0, its off-support values remain exactly those of pp, and its correction has finite mass and derivative energy given by the usual coefficient Grams. Yet uq=q−1(1/2)=192217,∣f(uq)−1/2∣∣f(up)−1/2∣=334217>1.u_q=q^{-1}(1/2)=\frac{192}{217},\qquad \frac{|f(u_q)-1/2|}{|f(u_p)-1/2|}=\frac{334}{217}>1. The squared-error ratio is 2.3690. This disproves control-improvement implications based only on monotonicity, protected support and blockwise observed-loss improvement. It does not assert that the regularized fitter would construct this deliberately chosen edit. The new action lies inside the edit support but outside the validation set; these are different domains. Established safe-learning approaches instead require appropriate model bounds and verify the proposed policy's state-action region, with explicit regularity, initial-policy and Lyapunov assumptions[berkenkamp2017safe].

Behavioral model assistance is different from trusting a model.

For fixed potential parent/candidate costs L0,L1L_0,L_1, pre-outcome forecasts m0,m1m_0,m_1, and a fair randomized policy choice AA, consider the standard model-assisted contrast[dudik2011doubly] Z=m0−m1+21{A=0}(L0−m0)−21{A=1}(L1−m1).Z=m_0-m_1+2\mathbf{1}_{\{A=0\}}(L_0-m_0) -2\mathbf{1}_{\{A=1\}}(L_1-m_1). Enumerating the two choices gives EAZ=L0−L1,Var⁡AZ=(L0+L1−m0−m1)2.\mathbb{E}_A Z=L_0-L_1,\qquad \operatorname{Var}_A Z=(L_0+L_1-m_0-m_1)^2. The conditional variance depends on the predicted cost sum, not the predicted ordering: true costs (1,3)(1,3) and forecasts (3,1)(3,1) have zero conditional randomization variance but the wrong direct ordering. Neither this identity nor a measured variance ratio establishes total uncertainty over future noise and contexts, fewer physical trials, or safe deployment. Time-valid off-policy evaluation with reward predictors and deployment gates is already established under its contextual-bandit assumptions[karampatziakis2021offpolicy]; those assumptions must not be silently extended to persistent physical control.

The common-state audit of all 66 actual repair-validation proposals gives 58/61 behavioral improvements among accepted edits and none among the five rejected edits, but three accepted fitter outputs regress in every one of four shared-noise continuations. No horizon crosses a regime boundary. The frozen all-proposal behavioral screen fails at high noise; fixed parent/candidate-world cost averaging gives conditional variance ratios 0.628/0.484 to IPS, failing both large-gain screens. These are finite-panel counterfactual diagnostics, not fresh physical laws or deployable oracle data.

Fine-grid evidence memory does not imply a useful retention guard.

For a fixed parent and pole, let FF be the scaled cardinal-linear design and ee its observed residual on an old calibration block. A future patch cc in that declared fine space changes the block loss by 2c⊤b+c⊤Gc2c^\top b+c^\top Gc, where b=F⊤eb=F^\top e and G=F⊤FG=F^\top F. Two-tap support makes GG tridiagonal for each independent output channel. The two-block, two-channel 65-knot implementation occupies 6,288 numeric bytes plus its shared grid, and matches direct and additive evaluation on all 66 archived calibrations. The fine moments are accumulated from those observations, not inferred from insufficient coarse statistics. Requiring non-increase on both old blocks nevertheless vetoes only two of three harmful accepted repairs and retains 37/58 useful ones. This failed guard illustrates the distinction between exact retained empirical objectives, physical-law compatibility and future control performance; it does not establish a zero-forgetting theorem.

Discounted continuous learning is an exact objective, not timely adaptation.

For the fixed cardinal basis with affine tails, parameterize observed force as f^t=ρpt+ϵut+B(ut)v,0≤ρ≤.85,vi+1≥vi,ϵ=.02.\hat f_t=\rho p_t+\epsilon u_t+B(u_t)v,\qquad 0\leq\rho\leq .85,\quad v_{i+1}\geq v_i,\quad \epsilon=.02. The static coefficients are c=(ϵ ξ+v)/(1−ρ)c=(\epsilon\,\xi+v)/(1-\rho), where the centers ξ\xi reproduce the identity. This scaled slope floor makes the fitting tails consistent with the positive-floor compiled inverse, but excludes an exactly flat dead zone. With fixed discount γ=.995\gamma=.995, the empirical quadratic obeys Ht=γHt−1+xtxt⊤H_t=\gamma H_{t-1}+x_tx_t^\top and bt=γbt−1+xt(ft−ϵut)b_t=\gamma b_{t-1}+x_t(f_t-\epsilon u_t) for xt=(pt,B(ut))x_t=(p_t,B(u_t)). Its cardinal block remains banded with a causal border. The undiscounted prior integrates [ρu+B(u)v−(1−ϵ)u]2[\rho u+B(u)v-(1-\epsilon)u]^2 over the uniform interior; four Gaussian nodes per cell exactly integrate each degree-six product in real arithmetic. Restricting vv to an affine sequence gives the matched three-parameter control with the same functional prior.

Solving this objective every sixteen rows is numerically verified on all 16,896 channel fits, with a 6,072-byte cardinal learner state. Nevertheless the resulting probe-free controller loses to RLS and the repair bank under both noise conditions; return-early error exceeds the bank by factors 2.584/3.352. Discounted sufficient statistics preserve the declared weighted objective, not knowledge of which hidden physical regime is currently active. Online spline identification and inverse control are established precedents[hong2012inverse,folgheraiter2016bsnn]; this negative capability test does not establish novelty or rule out their broader class.

Addressable retained behavior is distinct from hindsight selection.

For each immutable past bank program mm, four existing ordinary alarm rows define an empirical score Sm=∑t=k−3k∥vt+1obs−v^t+1(m)∥2S_m=\sum_{t=k-3}^{k}\|v_{t+1}^{\rm obs}-\widehat v_{t+1}^{(m)}\|^2. The predictor uses recorded current state, issued command, preceding observed force and known nominal mechanics. All programs predate these rows; the minimizing slot is sealed before evaluator-only physical continuations. The hindsight minimum of realized 160-step bank costs is a separate finite- bank diagnostic, not an optimal-control bound or a selectable causal program.

On sixty first-return alarm windows, this empirical addressing captures 99.93%/98.58% of low/high-noise aggregate adaptive-source-to-hindsight squared-cost headroom, passing the fixed coverage and behavioral screens. Force-SSE ranking gives identical choices. Thus these data support useful early retrieval but no distinct gain from velocity-domain prediction. Alarm timing remains inherited and empirically calibrated; neither selecting four stopped observations nor ranking their SSE creates a confidence sequence. Full lifetime operation at unfamiliar as well as returning dynamics is still untested. Multiple-model selection is established prior art[narendra2003multiple]; compact executable memory and this conditional behavioral result do not by themselves establish switching stability or zero future forgetting.

A fixed-prefix bound prevents misplaced architectural optimization.

If a proposed intervention leaves squared errors et2e_t^2 unchanged for t<T0t<T_0, its best possible whole-run RMS relative to a fixed baseline with squared cost CbC_b obeys RMS⁡newRMS⁡b≥∑t<T0et2Cb.\frac{\operatorname{RMS}_{\rm new}}{\operatorname{RMS}_b} \geq \sqrt{\frac{\sum_{t<T_0}e_t^2}{C_b}}. This follows only from nonnegative future costs and equal horizon length; equality need not be physically achievable. In the full causal early-reuse pilot, both 14,600-step pre-return traces match the old bank bitwise. The high-noise bound is 0.97426, so a target ratio of 0.80 cannot be reached by changing returns alone. The primary indeed eliminates returning probes and improves median early-return RMS, but fails the whole high-noise screen. The bound identifies first-time learning as the next task-level bottleneck, not a reason to weaken the frozen criterion or assert universal zero forgetting.

Causal-history propagation keeps a conditional linear fit.

For a fixed stable scalar pole ρ\rho, let q=(1−ρ)cq=(1-\rho)c and form Φt=ρΦt−1+B(ut),Φ−1=0,f^t=ρt+1f−1obs+Φtq.\Phi_t=\rho\Phi_{t-1}+B(u_t),\quad \Phi_{-1}=0,\qquad \widehat f_t=\rho^{t+1}f_{-1}^{\rm obs}+\Phi_tq. Monotonicity is imposed by q=Tzq=Tz, with cumulative-sum matrix TT, free initial coefficient and nonnegative increments. A second-difference penalty and ridge on qq give a bounded linear least-squares problem conditional on ρ\rho. Its empirical Gram is generally dense. A finite numerical pole search does not prove global optimality. Using propagated force history avoids repeated noisy lagged-force regressors, but the noisy initial condition and adaptive inputs remain; this is not raw-velocity maximum likelihood.

The corresponding sixteen-row physical diagnostic fails the distinct behavioral advantage over the existing jointly linear ARX cardinal fit. An initial TRF numerical failure is retained; a separately frozen BVLS implementation verifies KKT conditions, independent convolution-matrix assembly and full physical replay, but yields primary/ARX RMS ratios 1.00009/0.99274. Both methods can reduce source continuation cost, without establishing reliable unrestricted short acquisition. Exact operator algebra and correct optimization therefore do not establish new practical learning efficiency, and this fitting intervention closes without a parameter sweep.

Sealed predictive validation does not certify an acquisition lifecycle.

A candidate fitted to an initial probe block can be independent of later validation targets while still lacking any closed-loop improvement guarantee. The relevant distribution changes when its inverse chooses actions; a finite validation inequality does not bound that change, archive selection, or future reacquisition cost. Gate 357 implements genuine before-target forecasts for a sixteen-probe candidate and sixteen ordinary RLS-controlled validation actions, without refitting on validation. Complete independent replay confirms this causality, yet whole tracking error increases by factors 1.73054/3.07420 relative to unchanged early reuse at low/high noise. High-noise probing grows to 16,864 actions. All arms remain healthy and the stored programs immutable. Thus empirical predictive improvement, exact program preservation and operational adaptation are separate properties; none of the first two is a substitute for the third. The failed fixed intervention does not rule out all short-data adaptation or all memory-based controllers.

Cardinal weak motion balance compiles to a bounded quadratic.

For a known single-joint load and an unknown residual, write (J0+a)q¨=u+Φ(u,q˙,q)c−g0sin⁡q(J_0+a)\ddot q=u+\Phi(u,\dot q,q)c-g_0\sin q, with 0≤a≤10\leq a\leq1. A compact test function with zero value and slope at its endpoints gives a⟨q,ϕ′′⟩−⟨Φ,ϕ⟩c=⟨u−g0sin⁡q,ϕ⟩−J0⟨q,ϕ′′⟩.a\langle q,\phi''\rangle-\langle\Phi,\phi\rangle c =\langle u-g_0\sin q,\phi\rangle-J_0\langle q,\phi''\rangle. Integrating sample hats against a cardinal ϕ\phi and ϕ′′\phi'' yields fixed filters H0,H2H_0,H_2 that are exact for those interpolants, not for unknown between-sample physics. Uniform decimation evaluates only retained test centers. In Gate 358, overlapping endpoint hats make H2H2⊤H_2H_2^\top pentadiagonal, not circulant or tridiagonal. Banded Cholesky whitens the independent-position-noise contribution; nonlinear regressor noise and model error remain. This is an application of established weak/GLS identification, not a new estimation theorem[messenger2020weak].

For the regularized normal NN and right-hand side bb, partition the bounded scalar from cc. Eliminating cc gives a positive scalar Schur quadratic; clamp its unconstrained minimizer to [0,1][0,1] and re-solve cc for that value. The resulting global bounded-quadratic solution is checked against independently assembled BVLS. The empirical coefficient normal is dense despite local test covariance and continuous tensor-Gram structure. The corresponding real-motion screen reduces mean RMS by 34.5% relative to the same basis with smoothed pointwise acceleration, but fails its absolute trajectory requirement. Whitening adds no measured advantage. Exact calculus and a correct convex solve therefore remain distinct from a complete latent-state physical model.

Separate the weak operator from the chosen spatial representation.

The same integrated quadratic accepts polynomial features. For monomials za,zbz^a,z^b on the normalized cube, their mean-square Gram is Mab=∏im(ai+bi)M_{ab}=\prod_i m(a_i+b_i), where m(k)=0m(k)=0 for odd kk and 1/(k+1)1/(k+1) otherwise. Pure-coordinate curvature products follow by reducing the corresponding exponents by four and multiplying the two second-derivative factors. These exact moments share the temporal filter, whitening and Schur solve of the cardinal model; they do not require spline evaluation. Matched and full-cubic controls compile into a 4×4×44\times4\times4 power tensor with value/slope-matched coordinatewise affine tails. The real-motion cardinal/control RMS ratios are 0.49887 and 0.88202. The richer control thus rejects the fixed 20% representation-advantage requirement, despite three of four individual cardinal wins. Exact inner-product calculus is a reusable advantage, but this does not establish that a cardinal basis is necessary or that its extra parameters buy a new physical capability.

Known diffusion and cardinal-linear forcing need no learned propagator.

For spherical shell volumes W=diag⁡(wi)W=\operatorname{diag}(w_i), the finite-volume diffusion matrix has WA=K=K⊤⪯0WA=K=K^\top\preceq0. Thus W1/2AW−1/2=QΛQ⊤W^{1/2}AW^{-1/2}=Q\Lambda Q^\top and the modal state z=Q⊤W1/2cz=Q^\top W^{1/2}c has diagonal dynamics. This is weighted symmetry, not circulant structure; the zero eigenvalue represents conserved mass and must not be divided out. Across constant diffusivity and radius, poles scale by D/R2D/R^2 while normalized geometry is reusable. For a time step hh with linear current interpolation, each mode satisfies zk+1=ehλzk+hg{[φ1(hλ)−φ2(hλ)]Ik+φ2(hλ)Ik+1},φ1(x)=ex−1x,φ2(x)=ex−1−xx2.z_{k+1}=e^{h\lambda}z_k+h g\{[\varphi_1(h\lambda)-\varphi_2(h\lambda)]I_k +\varphi_2(h\lambda)I_{k+1}\}, \qquad \varphi_1(x)=\frac{e^x-1}{x},\quad \varphi_2(x)=\frac{e^x-1-x}{x^2}. Analytic continuation gives φ1(0)=1\varphi_1(0)=1, φ2(0)=1/2\varphi_2(0)=1/2; stable series avoid cancellation. Gate 360 checks this exact interpolant calculus against quadrature, independent augmented-matrix exponentials and PyBaMM. It retains all forty particle states and a nonlinear voltage readout, but neither learns a generator nor proves an exponential-spline representation necessary. The nearly equal classical-control timing limits the novelty claim. Known constant coefficients and a fixed finite-volume model are critical assumptions; concentration-dependent diffusion loses this fixed diagonal propagation. Accurate voltage also cannot remove parameter non-identifiability from grouped parameters or flat OCV slopes.

Exact response features remove rollout backpropagation in a restricted class.

For known ak=exp⁡(−hk/τ)a_k=\exp(-h_k/\tau) and exogenous feature rows ϕk\phi_k, consider Vk+1=akVk+(1−ak)(gk+ϕkθ)V_{k+1}=a_k V_k+(1-a_k)(g_k+\phi_k\theta). Let bk+1=akbk+(1−ak)gkb_{k+1}=a_kb_k+(1-a_k)g_k and Hk+1=akHk+(1−ak)ϕkH_{k+1}=a_kH_k+(1-a_k)\phi_k, with b0=V0,H0=0b_0=V_0,H_0=0. Induction gives Vk=bk+HkθV_k=b_k+H_k\theta. Thus a whole-trajectory weighted least-squares objective plus exact mass/curvature penalties is quadratic in θ\theta, with no differentiation through learned rollouts. This standard linear-response identity works for polynomial and cardinal features alike. It assumes a fixed generator and exogenous features: dependence of ϕk\phi_k on the unknown predicted state generally destroys the affine reduction. On real battery-aging development data[wang2026ani], the corresponding 35-coefficient cardinal fit is inexpensive but loses to an equally small polynomial and misses the absolute target. Continuous calculus and correct deployment replay do not alone establish missing-physics identification.

Composition requires cross errors and correct nonlinear calculus.

For component coefficient errors collected in EE and a function Gram MM, the exact mixture error is ∥∑jγjej∥2=γ⊤E⊤MEγ\|\sum_j\gamma_j e_j\|^2= \gamma^\top E^\top M E\gamma. The cross-Gram is PSD, but individual off-diagonal entries can be negative. In particular e2=−e1e_2=-e_1 gives zero error at mixture (1,1)(1,1) without identifying either component. A displayed lower-bound proof in the physical-operator literature[gopakumar2026physical] reverses the triangle inequality; its dependent universal capacity conclusion does not follow from that proof. A fixed affine-network counterexample and fifteen independent quadratures are archived in the composition audit. Likewise, for constant weights and ut+uux=0u_t+uu_x=0, vt−νvxx=0v_t-\nu v_{xx}=0, the residual of w=au+bvw=au+bv is (a2−a)uux+b2vvx+ab(uvx+vux)−aνuxx.(a^2-a)uu_x+b^2vv_x+ab(uv_x+vu_x)-a\nu u_{xx}. The illustrative CompNO expansion[hmida2026compositional] omits the −auux-a uu_x term. Symbolic differentiation verifies the correction. Neither audit is a reproduced refutation of the authors' neural experiments: a nonlinear trained aggregator can learn effects absent from an illustration. Known commuting flows also compose exactly without learning, but this does not resolve identifiability or transfer of genuinely unknown components.

Certify the assembled physical law, not the potential's label.

For constant M≻0M\succ0, the unforced equation Mq¨+∇U(q)+∇F(q˙)=0M\ddot q+\nabla U(q)+\nabla F(\dot q)=0 gives E˙=−v⊤∇F(v)\dot E=-v^\top\nabla F(v), where v=q˙v=\dot q and E=12v⊤Mv+U(q)E=\tfrac12v^\top Mv+U(q). Nonnegative FF alone does not certify dissipation: F(v)=v2(v−1)2≥0F(v)=v^2(v-1)^2\geq0, F(0)=0F(0)=0, but at v=3/4v=3/4, vF′(v)=−9/64vF'(v)=-9/64. The displayed nonnegative-potential penalty in LOpInf-SpML[sharma2024lagrangian] is therefore not by itself a global passivity certificate. Likewise scaling (M,U,F)(M,U,F) by ϵ>0\epsilon>0 shrinks the unnormalized unforced residual without changing the represented ODE. The paper's fixed-mass rod and membrane examples exclude that scaling; the criticism concerns its displayed flexible objective, not an observed collapse of its beam fit or uninspected implementation.

A constructive polynomial-cardinal alternative sets g(s)=η+∑iciBi(s)g(s)=\eta+\sum_i c_i B_i(s) with η>0,ci≥0\eta>0,c_i\geq0, and F(v)=12∫0v2g(s) ds,F′(v)=vg(v2),vF′(v)≥ηv2.F(v)=\frac12\int_0^{v^2}g(s)\,ds,\qquad F'(v)=v g(v^2),\qquad vF'(v)\geq\eta v^2. The force is coefficient-linear at observed velocities. With fixed physical normalization, quadratic fitting under c≥0c\geq0 remains convex; nonlinear rollout fitting is not thereby made quadratic. Some permitted potentials are nonconvex: g=.1+B3g=.1+B_3 with a centered unit-scale cubic generator gives F′′=−61/240F''=-61/240 and positive power 29/16029/160 at v2=3/2v^2=3/2. The correct finite-velocity force Gram is Gij=∫−SSv2Bi(v2)Bj(v2) dv=∫0Ss Bi(s)Bj(s) ds.G_{ij}=\int_{-\sqrt S}^{\sqrt S}v^2 B_i(v^2)B_j(v^2)\,dv =\int_0^S\sqrt{s}\,B_i(s)B_j(s)\,ds. Piecewise-polynomial weighted moments are analytic, but the weight destroys translation invariance. Sixteen exact entries agree with independent velocity-space quadrature within 4.44×10−164.44\times10^{-16}; one noiseless 121-observation nonnegative fit recovers four coefficients within 1.67×10−161.67\times10^{-16}. This is an elementary calculus construction and numerical audit, not a new dissipative-systems theorem or a real-data result. Directional sums extend the safeguard, but do not cover every anisotropic or history-dependent constitutive law.

Normalization also matters for nonlinear solves. For G(u)=ϵKu+D⊤F(u)−bG(u)=\epsilon Ku+D^\top F(u)-b, a sufficient contraction condition is ∥K−1D⊤∥L/ϵ<1\|K^{-1}D^\top\|L/\epsilon<1 when FF is LL-Lipschitz. The dependence printed in Geo-NeW Eqs. 9 and 12[shaffer2026geonew] cannot generally replace this by ϵ∥K−1∥L<1\epsilon\|K^{-1}\|L<1. For K=D=1,ϵ=.1,F(u)=−.2tanh⁡(u),b=0K=D=1,\epsilon=.1,F(u)=-.2\tanh(u),b=0, the latter product is .02, yet the equation has roots approximately −1.915008048,0,1.915008048-1.915008048,0,1.915008048. The correctly scaled bound is 2 and makes no uniqueness assertion. These symbolic and numerical checks concern the displayed premises, not reproduced failures of the authors' empirical neural models.

Compile scalar maps before filtering their outputs.

For fixed coefficients, a causal finite-history force predictor of the form y^t=b+∑r,j,icrji∑ℓ=08wrℓβi(xt−ℓ,j)=b+∑r∑ℓ=08wrℓ∑j,icrjiβi(xt−ℓ,j)⏟vr,t−ℓ\hat y_t=b+\sum_{r,j,i}c_{rji}\sum_{\ell=0}^{8}w_{r\ell} \beta_i(x_{t-\ell,j}) =b+\sum_r\sum_{\ell=0}^{8}w_{r\ell} \underbrace{\sum_{j,i}c_{rji}\beta_i(x_{t-\ell,j})}_{v_{r,t-\ell}} can filter the projected output vectors vrv_r instead of the full basis vectors. Zero pre-trajectory basis history is used on both sides. For wrℓ=ρrℓ(1−ρr)/(1−ρr9)w_{r\ell}=\rho_r^\ell(1-\rho_r)/(1-\rho_r^9), the exact finite convolution has the subtractive recurrence zr,t=ρrzr,t−1+arvr,t−arρr9vr,t−9z_{r,t}=\rho_r z_{r,t-1}+a_r v_{r,t}-a_r\rho_r^9v_{r,t-9}, where ar=(1−ρr)/(1−ρr9)a_r=(1-\rho_r)/(1-\rho_r^9); ρr=0\rho_r=0 is instantaneous. Cardinal-to-power matrices compile each scalar map, including value/slope- matched affine tails. This is linear filtering algebra, not a new theorem or a learned exponential B-spline generator. Numerical preflight checks verify full design/compiled/streaming equivalence and trajectory resets.

Exact continuous mass and curvature penalties remain quadratic in the coefficients. Cross-feature and temporal mixing generally make the empirical normal dense, however; compact deployed state must not be conflated with small learning statistics or a circulant inverse. Coefficient updates also require recomputing the retained projected history under the new map. Reusing the old filter state would not implement the updated finite-history model. The corresponding real-telemetry development comparison subsequently rejects both additive cardinal candidates. Primary all-task force MAE is 0.27052 N, versus a small MLP's 0.21411 N, and no-contact error violates the frozen limit. Although fixed-weight state is 125,816 bytes and every validation stream replays correctly, learning normals require 27,040,648 bytes and CPU inference is slower than the small MLP. These numerical identities alone therefore do not establish a practical neural advantage; all test traces remain unopened. Known geometric mixing of joint-response maps is a distinct structural hypothesis, whereas additional knots do not remove the additive class's general inability to represent mixed-coordinate interactions.

Known time-varying mixing preserves linear fitting, not filter commutation.

Let ϕjp(xt)\phi_{jp}(x_t) be a fixed joint-local feature vector, cjpc_{jp} its coefficients, wpℓw_{p\ell} a finite-history filter and W(qt)W(q_t) a known current-pose mixer. Then f^t=b+W(qt)at,(at)j=∑p,ℓwpℓϕjp(xt−ℓ)⊤cjp+g(qt)⊤dj\widehat f_t=b+W(q_t)a_t,\qquad (a_t)_j=\sum_{p,\ell}w_{p\ell} \phi_{jp}(x_{t-\ell})^\top c_{jp}+g(q_t)^\top d_j is linear in (b,c,d)(b,c,d) for recorded inputs and poses. Consequently its trajectory-weighted least-squares objective has exact additive normal statistics, although the resulting matrix is generally dense. Fixed coefficient projection may precede each filter, leaving a ring of joint outputs. In contrast, replacing W(qt)∑ℓwℓat−ℓW(q_t)\sum_\ell w_\ell a_{t-\ell} by ∑ℓwℓW(qt−ℓ)at−ℓ\sum_\ell w_\ell W(q_{t-\ell})a_{t-\ell} changes the model unless the relevant mixers coincide or a special cancellation holds.

The concrete geometric screen uses W(q)=0.1 m (J(q)J(q)⊤+(0.01 m)2I)−1J(q)W(q)=0.1\,\mathrm m\,(J(q)J(q)^\top+(0.01\,\mathrm m)^2I)^{-1}J(q), 643 scalar coefficients and 15,305 fixed streaming numeric bytes. All 36,400 validation predictions and the three classical full training-statistic reconstructions pass independent audits. Nevertheless primary MAE 0.27495 N is worse than constant mixing 0.26795, matched-input neural 0.23753 and the unchanged full-input neural 0.21411 N; no-contact error also fails its limit. The primary's 3,312,744-byte learning statistics and 0.612 s fit do not establish an application advantage. The damped inverse, base-frame labels and latent joint-response semantics are modeling assumptions, not an identified physical torque law. This development-only geometric branch is rejected before test data are opened, preserving the algebra without promoting the capability claim.

Feedback attribution is not established by a nonzero input weight.

A frozen retrospective diagnostic reproduces all six asymmetric selected policies on sixteen exposed physical resets and clamps only their six policy-force coordinates. The inverse retains force sensing. Full-policy traces replay bitwise; all 192k evaluator transitions are finite and complete. Full/clamped restoration is 0.9228/0.9326 for shared-six, 0.8719/0.8716 for polynomial and 0.9051/0.8463 for exact-law policies. The extra policy input is therefore not a demonstrated positive mechanism for the primary recovery, even though its learned weights are nonzero. Clamping can be out of distribution and does not substitute for a matched training comparison or remove the inverse's sensing requirement. Larger-budget shared-six control makes fewer model-infeasible requests than its nested 200k policy (37.02% to 25.12%), lowers mean command cost (1.1208 to 0.9152) and increases velocity (13.1383 to 14.0158). Its on-policy one-step model RMSE changes from 0.01866 to 0.02060: better skill is not the same as better identification, and these trajectories do not share a fixed error-sampling distribution. After the compound primary panel completes, the same intervention gives full/clamped restoration 0.8894/0.8570 for shared-six, 0.8518/0.8621 for polynomial and 0.9130/0.8945 for exact-law learning. The force-input contribution is therefore family- and policy-dependent. Across both families all 384k evaluator transitions and bitwise original replay identities pass; none changes the frozen primary score.

Original: paper/theory_policy_transfer.tex · Raw source file

View raw TEX source
\section{From compact physical models to useful policies: an empirical boundary}
\label{sec:policy-transfer-boundary}

Exact operator calculus and compact coefficient estimation solve a model
construction problem. Useful control additionally depends on observation
coverage, model errors along the policy's trajectories, policy optimization
and selection. A low prediction residual is not a bound on recovered task
return without further assumptions. This distinction motivates a controlled
full-policy refinement test instead of another inverse-evaluation microbenchmark.

\paragraph{Scope of the joint quadratic identification step.}
Write the static coefficients as $c=\mu+Da/(1-\rho)$, where $\mu,D$
are fixed source quantities and $0\leq\rho\leq0.85$.
For the \emph{unclipped} one-step model, the predicted response is
\[
 B\mu+\begin{bmatrix}p-B\mu&BD\end{bmatrix}
 \begin{bmatrix}\rho\\a\end{bmatrix},
\]
so squared fitting error and coefficient-monotonicity constraints are
quadratic and linear, respectively, in $(\rho,a)$.
The implemented regularizers act on the scaled correction $Da$ and scaled
drive $(1-\rho)\mu+Da$. In static-law coordinates, their penalties are
$\lambda(1-\rho)^2\|F(c-\mu)\|^2$ and
$\kappa(1-\rho)^2\|F_2c\|^2$, not pole-independent static-law penalties.
The small pole ridge is also retained. Furthermore, deployment uses
$\rho p+(1-\rho)\operatorname{clip}(Bc,-1,1)$; this clipped loss is not
the quadratic objective being minimized. Convexity therefore describes the
specified identification surrogate, not arbitrary pole learning or globally
optimal fitting of the clipped physical simulator. Executable objective
identities record this distinction without changing any frozen fit.

\paragraph{Finite-region treatment of the clipped objective.}
For sorted commands and coefficient-monotone scalar functions, observations
can only form a lower-saturated prefix, an interior block and an
upper-saturated suffix. There are $(n+1)(n+2)/2$ assignments, including empty
blocks and threshold ties. With $g_i=(1-\rho)B_i\mu+B_iDa$, deployment is
$\rho p_i+\operatorname{clip}(g_i,-(1-\rho),1-\rho)$.
Its three region formulas are respectively
\[
 \rho(p_i+1)-1,\qquad
 B_i\mu+\rho(p_i-B_i\mu)+B_iDa,\qquad
 \rho(p_i-1)+1.
\]
Each is affine in $(\rho,a)$, and region consistency is linear in these
parameters. Thus enumeration of the assignments reduces the specified
clipped loss, with the unchanged scaled regularizers, to finitely many
convex quadratic subproblems. Monotonicity permits checking only the
boundary observations of each block. This uses a fixed basis and one stable
response pole; it is not a convex formulation for arbitrary learned
exponential generators or higher-order pole sets.

A lower bound for any region follows by dropping the nonnegative interior
loss and function penalties, retaining saturated observations and pole ridge:
\[
 L_{\rm sat}=\min_{0\leq\rho\leq0.85}
 \left\{\frac12\sum_{i\in{\cal S}}
       (\alpha_i\rho+b_i-y_i)^2+\frac{10^{-8}}2\rho^2\right\}.
\]
Here $(\alpha_i,b_i)=(p_i+1,-1)$ or $(p_i-1,1)$.
Its minimizer is the box projection of
$-\sum_i\alpha_i(b_i-y_i)/(\sum_i\alpha_i^2+10^{-8})$.
Prefix sums supply bounds for every clipping assignment. A bound exceeding
the current feasible objective excludes that region without an LP or QP.
This is standard lower-bound pruning specialized to the operator model,
not a new general optimization theorem. Floating-point implementations retain
rounding guards, LP statuses and QP dual checks; these are numerical
accounting, not exact-arithmetic certificates or statistical uncertainty.

The exhaustive reference tests 24 scalar fits on two exposed laws and two
representations: 3,672 regions. Six-actuator time is 1.76--2.00 seconds.
One high-loss QP in each asymmetric case reports an unresolved optimization
status; the frozen complete-accounting rule excludes both from control.
Both compound cases resolve all regions. Despite improved aggregate
later-record prediction and pole estimates, replacing the inverse beneath
unchanged actors changes shared restoration from 0.8894 to 0.8865 and
polynomial restoration from 0.8518 to 0.8528. All 128k evaluator transitions
and original bitwise replay checks pass. Better identification does not
automatically improve a policy trained with a different inverse.

Saturated-row pruning is 7--13$\times$ faster than exhaustive enumeration in
a three-repeat exposed-case CPU comparison, but six-actuator times remain
136--281 ms. A stronger bound drops the linear region/monotonicity constraints
from each complete quadratic while retaining the pole box. For
$H\succ0$, let $z_0=H^{-1}r$ and $v=H^{-1}e_0$; its relaxed box minimizer is
\[
 z_{\rm box}=z_0+
 v\,\frac{\operatorname{clip}((z_0)_0,0,0.85)-(z_0)_0}{v_0}.
\]
This follows by eliminating the remaining coordinates and minimizing the
scalar Schur complement. Batched matrix products construct the region
Hessians; batched solves and residual-corrected dual values then screen them
before any generic constrained optimization. These small data Hessians are
dense and are not assumed circulant.

In a separate paired comparison, matrix bounds reduce six-actuator time
to 39--67 ms with identical preceding objective values and 13--33 final QPs
out of 918 candidate regions. A broader exposed-archive check completes all
576 scalar fits. Ninety of 96 six-actuator cases are below 100 ms, but the
worst is 133.39 ms; the all-cases latency condition fails. Retained scalar
bound arrays occupy 226,440 bytes, excluding other temporaries; peak process
RSS is 303.80 MiB. No control outcome is changed by this compiler comparison.

In Gate 333, four fixed world models---shared six-coordinate with sixteen
target observations, cold MLP with sixteen, cardinal with 128, and privileged
exact law---each support standard SAC training on four exposed lag families
and two seeds. The shared representation inherits 768 prior observations;
every actor/critic inherits a 20M-transition pretrained policy. Mechanics and
six force sensors are supplied. Each run uses 200,000 virtual transitions,
three virtual validation resets and eleven eligible checkpoints. Training is
backed up before evaluation on eight held-out changed-physics resets.
All 32 runs complete with immutable worlds and finite trajectories.

\begin{center}
\begin{tabular}{lrrrr}
\toprule
World & Deadzone & Smooth & Asymmetric & Compound\\
\midrule
Shared six, 16 & 0.8688 & 0.8684 & 0.8478 & 0.8011\\
Cold MLP, 16 & 0.8837 & 0.3708 & 0.2147 & 0.7398\\
Cardinal, 128 & 0.8710 & 0.8544 & 0.8438 & 0.8116\\
Exact law & 0.8647 & 0.8806 & 0.7903 & 0.7724\\
\bottomrule
\end{tabular}
\end{center}
Entries are mean paired reset-normalized returns within a training run,
then median across two training seeds. They are not independent-law
confidence estimates. Primary gains over its initial policy are
0.0021/0.0206/0.2190/0.1030 nominal units. The last two clear the five-point
gain clause, but every family misses 90\% nominal restoration. All primary
training runs remain within 600 seconds on one memory-capped MPS learner.
The complete frozen panel therefore fails.

An exact-law learner is a diagnostic of the specified learning procedure,
budget and reset-selection protocol, not an optimal-control upper bound.
A learned-world policy exceeding it does not establish more accurate
physics. Likewise, failure of a cold small-data MLP does not compare against
all neural models or a neural shared prior. This is offline simulated-policy
refinement for later episodes, not uninterrupted hardware repair or a
gradient-free end-to-end method. The next frozen test triples the learning
budget and includes a nested budget control and matched shared-prior
polynomial directions, without replacing these failed thresholds.

\paragraph{Budget and selection follow-up.}
Gate 337 completes all twelve 600k runs and their audits. Shared-six median
restoration is 0.9228/0.8894 on asymmetric/compound lag, versus polynomial
0.8719/0.8518 and exact-law learning 0.9051/0.9130. The primary gains
0.0763/0.0754 over nested 200k and 0.05085/0.03763 over polynomial:
both predeclared extra-budget and coordinate-specific clauses pass.
The full gate fails the compound 0.90 restoration clause. All primary runs
meet 1,800 seconds. This is two-seed evidence on exposed laws with new resets,
not independent-law generalization or immediate repair.
Both fits use the same cardinal carrier. The functional-PCA versus fixed
polynomial comparison does not isolate exact-Gram geometry from source-learned
subspace selection; a matched learned coefficient-metric prior is absent.

Gate 339 instead retains the older eight 200k archives and evaluates their
88 checkpoints on sixteen new virtual reset states before physical testing.
Expanded selection restores only 0.8359/0.8010 for shared-six
asymmetric/compound policies, versus 0.8472/0.8007 originally: the gate fails.
The separate post-panel maximum of mean paired-normalized return over each
finite archive is below 0.8554 in all eight cases. Thus no single fixed
checkpoint from these archives can meet 90\% on this test panel.
This is neither a control-theoretic ceiling nor permission to select using
future physical outcomes. It motivates testing added capability rather than
further selection tuning on already exposed responses.

\paragraph{Descriptor equivalence and saturation boundaries.}
For fitted models $c=\bar c+Da$ with $D^\top GD=I$, the shared descriptor
is $a$ and the polynomial descriptor is $E^\top GD\,a$.
For the frozen source geometry, the latter six-by-six map is invertible
with condition number 1.66186. Standardization and the common pole preserve
an affine relation, which an unrestricted first affine policy layer can
absorb. Thus these encodings do not add target-model information or neural
expressivity on the fitted shared subspace. Their source-bank projections
are not globally equivalent, and finite-budget optimization need not agree.
Three source-only tests verify this distinction. In contrast, Gate 337
fits different subspaces; its completed shared-six/poly medians are
0.9228/0.8719 on asymmetric and 0.8894/0.8518 on compound lag.

Known static clipping is another boundary: coefficient functions can differ
only inside saturated tails and induce the same bounded response.
Partitioning at cardinal knots and clipping crossings makes every bounded
piece constant or cubic. Four-point Gaussian integration then produces exact
pairwise products, up to floating point, as a factor Gram.
Four tests include an invisible-tail witness and independent integration.
The 36 source profiles require 102 intervals and 408 nodes; the resulting
Gram is positive semidefinite, not necessarily invertible or circulant.
This is uniform-command geometry, not task-risk certification.

A separate 220-byte six-channel cache selects a nearer-to-zero command on
each outer saturation plateau. On twelve archived-model validation pairs,
all actions, forces, observations and velocities remain bitwise identical,
while squared-command cost falls by mean 208.86--215.49 return units across
the four policies. Four additional tests verify model-equivalent outputs and
immutable caches. A fitted saturation threshold can be wrong on the true
plant: these are source-world efficiency results, not physical recovery,
electrical-energy savings or amendments to the frozen policy protocols.

\paragraph{Exact stable response geometry: an identity, not a new kernel theorem.}
Let $\phi_i(u)$ be the bounded cardinal response, $\rho_i\in[0,1)$ its
causal pole and $b_i=\phi_i(0)$ its zero-command equilibrium.
Start at force $p$, apply $u$ once and then zero commands.
For $k\geq0$, direct substitution in the first-order recurrence gives
\[
r_i(k,u,p)=y_{k+1}-b_i
=\rho_i^k\{\rho_i p+(1-\rho_i)\phi_i(u)-b_i\}.
\]
Use the equilibrium once and the infinite transient as a feature, with
independent uniform reference measures on $p,u\in[-1,1]$.
Writing $\delta_i(u)=(1-\rho_i)\phi_i(u)-b_i$ yields
\[
K_{ij}=b_i b_j+
\frac{\rho_i\rho_j/3+\mathbb E_u[\delta_i(u)\delta_j(u)]}
{1-\rho_i\rho_j}.
\]
The proof is the convergent geometric series, $\mathbb E[p]=0$ and
$\mathbb E[p^2]=1/3$. Static products use the clipped-polynomial partition;
no time unrolling is required. This is an inner-product Gram and hence
positive semidefinite, without a general invertibility or circulant claim.
The equilibrium has unit finite weight: its nonzero constant trajectory
is not summed over infinite time. The reference measure is not a task-risk
distribution, and stability is essential.

Four tests compare this formula with direct 300-step causal recurrences,
check saturation aliases and pole distinctions, verify centered kernel-PCA
projection and reject unstable/rank-deficient inputs.
A six-mode source encoder plus explicit pole retains 13,128 numeric bytes,
including the source curves needed for cross-products.
Its response-metric reconstruction error is not a policy-performance bound.
The frozen next experiment tests the same source worlds, 600k policy budget,
sixteen-row fit and same-episode continuation against static and zero-context
controls on new laws; its completed negative capability result is reported below.

Kernels on dynamical systems, including initial conditions and efficient
linear-system calculations, have established prior art
\cite{vishwanathan2007binet}. Kernel-based Hammerstein identification likewise
predates this work \cite{risuleo2016hammerstein}.
The scalar response identity and centered kernel PCA are not novelty claims;
the open question is useful matched physical-skill transfer.

\paragraph{Fresh-law source-policy reuse: conditional benefit, failed recovery.}
Gate 340 completes three source learners (1.8M new virtual transitions plus
378k validation transitions) before evaluating eight new laws and eight
paired reset states per law. Every context uses the same sixteen-row fitted
shared inverse. Asymmetric/compound restoration is 0.8202/0.8915 for shared
context, 0.8287/0.8643 for polynomial and 0.8655/0.8767 for zero context.
Primary within-law paired gains over polynomial are 0.02707/0.02671,
but gains over zero are $-0.04091/+0.01477$. Both families fail the recovery
and zero-context gain clauses; all remaining clauses and resource/trace
checks pass. The paired contrast is not the difference of marginal medians.
Invertible target encodings do not guarantee equal finite-budget training,
but these comparisons also do not establish added target-model information.

Maximum fitting plus encoding is 8.981 ms, primary numerical state is
654,084 bytes, and no target policy-gradient update occurs. All 448
controller episodes and 1,024 separately acquired prefix transitions are
preserved. Identical simulated prefixes are replayed with extra fallback
steps charging standalone computation; Git/orchestrator delay is not plant
time. The changed law is present from reset. This is causal startup adaptation
in a known-mechanics, force-sensed simulator, not unknown-change detection,
hardware real-time control or a zero-forgetting result. A true inverse under
the same actor and inferred context gives only 0.8325/0.8941 restoration;
it is a limited intervention, not a fully informed-policy upper bound.

\paragraph{Exact response geometry does not guarantee task-useful reuse.}
Gate 341 adds one source learner using the response kernel above, with
600k training and 126k validation transitions and a bitwise matched source
world schedule. On eight new laws and eight resets each, asymmetric/compound
restoration is 0.8993/0.8582, versus 0.8664/0.8620 for static shared,
0.8273/0.8861 for static polynomial and 0.9021/0.8913 for zero context.
Paired response gains over both static encodings exceed two points only
on asymmetric laws. Both families fail restoration and the zero-context
gain clause; the full operator-specific and reusable-recovery gates fail.
Exact uniform-reference response geometry is not the task's trajectory
geometry or a control-value guarantee.

All 512k controller and 1,024 prefix transitions pass the trace, immutable
actor and charged-delay audit. The maximum fit plus four descriptors is
11.006 ms, numerical primary state 665,028 bytes, evaluator RSS 311.609 MiB.
The source learner takes 836.920 s with 790.516 MiB peak RSS and 28,098,560
Metal driver bytes. No target gradient occurs; inherited experience, ongoing
force sensing and known mechanics remain essential qualifications.
The conditional calibration-in-the-loop source-training gate follows its
predeclared failure trigger, not a kernel-selection sweep on these targets.

\paragraph{A physically motivated descriptor can be actively unhelpful.}
Fixed-actor interventions after Gate 340 preserve the inverse and all
mechanical/force observations. Replacing the inferred normalized descriptor
with its source centroid changes shared restoration from 0.8202 to 0.8736
on asymmetric laws and from 0.8915 to 0.8851 on compound laws. The paired
original-minus-centroid contrasts are $-0.05343/+0.00259$; all four asymmetric
laws improve under the centroid. Polynomial-context contrasts are
$-0.03413/-0.00906$. Cyclic channel shifting generally hurts the aggregate,
showing sensitivity without establishing uniformly useful information.
All 384k diagnostic transitions pass immutable-actor, startup-prefix and
bitwise original-replay checks. Zero normalized context is not a zero plant
law or the separately trained zero-context policy. Interventions can create
out-of-distribution combinations; none modifies the primary experiment.
The subsequently completed Gate-341 diagnostic adds 576k transitions.
For its response actor, original/centroid restoration is
0.8993/0.9195 and 0.8582/0.8966, with paired original-minus-centroid
contrasts $-0.02158/-0.01103$. All original replays remain bitwise identical.
The polynomial actor instead benefits from its inferred descriptor on the
compound panel (paired gain 0.02471). These results reject a universal
context-benefit interpretation without proving context is unnecessary.

\paragraph{Prior art limits the compiler novelty claim.}
Cardinal-spline input lifting for Hammerstein identification predates this
work \cite{chan2006cardinalhammerstein}; those natural interpolating splines
are not identical to the compact uniform-grid carrier used here.
Saturation-data partitioning \cite{pupeikis2006saturation}, mixed-integer
global piecewise-affine identification \cite{roll2004piecewiseidentification},
and branching methods for bounded monotone smoothing \cite{sasane2019monotone}
also precede our finite-region specialization. The measured acceleration
compares our own successive implementations, not those prior solvers.
No global-method priority, statistically unbiased noisy-feedback fit or
algorithmic SOTA claim follows.

\paragraph{Clipped-loss memory needs more than the ordinary empirical Gram.}
On one cardinal cubic cell, Gram entries depend on input moments through
degree six. At $u_i=i/112$, $i=0,\ldots,7$, put weights
$\binom{7}{i}$ on even indices in one zero-target dataset and odd indices
in another. Their sample counts, full cardinal data Grams and target
statistics agree, since the seventh finite difference annihilates every
polynomial of degree at most six. Yet the weighted squared clipped losses
for the reproduced monotone affine drive $64u-1$, with zero pole, differ
by $16/7$. Two tests check the exact rational identity and the actual
37-coordinate basis. Thus the unclipped pooled-Gram contract does not
extend automatically to a parameter-dependent clipped objective.
Additional region/order information or retained observations are needed;
this elementary counterexample does not invalidate immutable model archives.

\paragraph{Clipping-aware inverse and descriptor correction is not sufficient.}
The final fresh-law Gate 343 holds Gate-340 source actors fixed and crosses
old/new clipped-objective fits independently through the inverse and descriptor.
Asymmetric/compound restoration is 0.8166/0.8851 for old/old,
0.8181/0.8932 for new-inverse/old-context,
0.7750/0.8739 for old-inverse/new-context and 0.8039/0.8832 for new/new.
The primary new/new paired gains over old/old are $-0.01267/-0.00500$;
both recovery gates fail. Positive factorial interactions
($0.01624/0.00091$) do not imply a positive primary effect.

All 384 scalar fits complete numerical region accounting. Both pipelines
and encodings take 36.65--133.57 ms, below this separately frozen 200-ms
paired-pipeline cap, without erasing earlier 100-ms misses. Common activation
occurs at steps 17--19; all rewards and delay steps count. All 641,024
evaluator transitions are retained with unchanged source actors and zero
new gradients. Primary numerical state is 651,060 bytes and evaluator RSS
406.047 MiB. Reaggregation is compared in memory by a separate audit adapter
after the original exclusive-output writer rejects an existing metrics file;
the frozen experiment and its artifacts are not changed. More faithful fitting
and faster numerical compilation do not establish useful policy compatibility.

\paragraph{Calibration-consistent source training also fails the recovery gate.}
The predeclared conditional Gate 342 trains new response and zero-context
actors through the actual sixteen-row fitted inverse. The two source-world
schedules, calibration observations, fitted models and startup traces match
bitwise; all 1,484 source/validation fits reproduce from their permitted rows.
Each actor uses 600k learner transitions plus 11,088 startup transitions;
combined source validation adds 252k transitions. The two training times are
879.199/876.969 s, with maximum source RSS 813.625 MiB.

On eight fresh laws, asymmetric/compound restoration is 0.8514/0.8595
for calibrated response and 0.8953/0.8974 for its matched zero-context actor.
Previous response gives 0.8654/0.8681. Primary paired gains over matched
zero are $-0.04580/-0.03713$, and over previous response
$-0.02409/-0.00602$. Both families fail restoration, matched-zero gain
and the all-law floor; neither training-consistency clause passes.
All 641,024 evaluator transitions and resource checks pass. Maximum fit
plus four descriptors is 11.686 ms, primary numerical state 665,020 bytes
and evaluator RSS 316.031 MiB. No target policy gradient occurs.

This retires the overnight descriptor/training-recipe branch under its
predeclared stopping rule. The previous-response comparison changes a whole
training recipe, not only one matched causal factor; only the two new arms
are fully matched. The zero-context control also misses 90\% in both
family medians and is not a post-hoc passing primary. Known mechanics,
ongoing force sensing, inherited pretraining, startup changes and one
source seed limit the interpretation. A true inverse with the primary's
inferred descriptor remains a limited intervention, not an optimal-policy bound.

\paragraph{Same-policy feasibility limits the remaining inversion hypothesis.}
Gate 344 fixes the existing calibrated-zero actor and removes descriptor
differences across all methods. Sixteen fresh laws with eight resets each
give shared-spline restoration 0.9084/0.8966 (asymmetric/compound), frozen-RLS
0.8850/0.8722, online-RLS 0.8766/0.8568, polynomial-prior 0.8999/0.8901, and
true-inverse 0.9123/0.8977. Paired gains over frozen RLS are
0.02124/0.02292 and over online RLS 0.02900/0.05269; descriptive
law-bootstrap lower bounds are positive. The polynomial advantage remains
below the required two points in both families, while compound misses
recovery and the all-law floor. Both the full operator-capability gate and
the two-family true-inverse feasibility criterion fail.

This is a modest matched-controller identification result, not universal
operator superiority. Shared and polynomial models have the same cardinal
carrier; learned subspace and geometry are not independently isolated.
True inversion's paired advantage over the fitted spline is small and not
uniformly positive. It does not bound an optimized policy, but weakens the
case for more fitting refinements as the route to a large gain here.
All 898,048 transitions, 1,536 scalar-fit reproductions and actor/inverse/
sensor/force recurrences pass. Maximum common fitting is 15.488 ms, every
activation is step 17, numerical state is 651,900 bytes and RSS 317.672 MiB.
No new policy training occurs; the selected actor inherits 611,088 source
training transitions, 126k validation transitions and earlier 20M pretraining.
These are startup-change, continuously force-sensed simulations, not a
demonstration of unknown-change detection, safe probing or useful memory.

\paragraph{Direct forward planning: valid algebra, unsuccessful development task.}
Gate 345 changes the operational task to reduced suspended-load motion, with
known linear mechanics, position/angle sensing and an unknown monotone drive
law with stable lag. In the oracle-model pilot, augmenting the state by force
gives $z_{k+1}=A_\rho z_k+B_\rho g(u_k)$; lifting a finite horizon gives
$Z=S_\rho z_0+T_\rho V$. A quadratic state/static-drive objective with input
bounds is therefore a condensed convex quadratic program. This is established
Hammerstein/MPC structure, not a new consequence unique to cardinal splines.
Filtering basis rows through the same operator also makes position observations
linear in constitutive coefficients for fixed pole, with separate linear
initial-state nuisance columns. General finite-horizon Grams are not circulant.

Five numerical tests and four full pilot audits validate these identities,
observer recurrences, commands and constrained-plan dual-gap witnesses.
Nevertheless, the true-forward whole cost is 302.9361 versus true-inverse
304.8170: only 0.617\% improvement. Settled-position RMS 0.1260 m fails the
proposed 0.035 m target, as does nominal 0.1251 m. The reference preview allows
early departures that conflict with the declared settling criterion. We
retain this failed task contract and do not retune it on its exposed trace.
All 1,600 development transitions complete, with primary p95 1.970 ms and
263,920 bytes of numeric controller state. No proposed fresh-panel laws are
sampled, no learner is trained, and no fresh-panel or safety claim follows.

\paragraph{Adverse evidence must survive a retrieval decision, but rejection does not repair.}
The 192 archived historical/current-only Gate 332 traces reproduce every alarm
and reuse score. At long dwell, 510/566 low-noise and 1,742/1,778 high-noise
returning historical reuses fail the selected model's own alarm threshold on
the four triggering observations. The subsequent sixteen probes are evaluated
by a different RMS/floor criterion and can miss that local discrepancy.
Current-only behaves similarly. These empirical threshold failures are not
uniform noise or physical-model certificates.

In a fixed exposed-case intervention, reuse is additionally checked against
the retained alarm rows. All 270,400 physical steps complete and untreated
trajectories replay bitwise. Low-noise historical RMS rises from 0.002084 to
0.004193, probe actions from 1,504 to 7,856, and the full bank refuses 54 fits.
High-noise probes fall from 2,272 to 800 but RMS rises from 0.003655 to
0.003869. Thus consistency filtering alone does not establish useful memory.
The result motivates a separately tested local correction and subsequent-
observation verifier, not relaxed thresholds or a retroactive passing score.

\paragraph{Compact local correction composes the calculus without global task guarantees.}
Let a constitutive parent be $s(u)$ and propose
$s_a(u)=s(u)+\sum_i a_i h_i(u)$, where cardinal linear hats $h_i$ are supported
inside two parent cells and the endpoint corrections are zero. The continuous
penalties $a^\top M a$ and $a^\top K a$ use exact hat mass and derivative
Grams. Parent cubic derivative minima on each subcell turn monotonicity into
linear slope inequalities. For each retained observation block $b$, the
squared-loss change under a scaled proposal is
$2\eta\langle e_b,\Delta_b\rangle+\eta^2\|\Delta_b\|^2$.
Choosing a common $\eta\in[0,1]$ below all favorable quadratic roots preserves
the observed block losses and the convex monotonicity constraints. These are
standard compact-support and quadratic identities, here combined with a
sixteen-subsequent-observation deployment verifier. They are not a future-risk
or physical stability theorem.

The exposed twelve-trajectory pilot completes all 405,600 evaluator steps
and full causal replay/fit/validation checks. Low-noise repaired-bank RMS is
0.001357, versus original 0.002084 and matched repaired current-only 0.001383;
probe actions fall from 1,504 to 560. High-noise RMS 0.003876 is worse than
original 0.003655 and no correction is accepted. Thus the operational gain
is conditional and unique historical value remains limited. An audit recovers
an overwritten line-search metadata field from unchanged saved coefficients;
no experimental outputs are replaced. Static complement preservation must
also not be confused with trajectory preservation: for identical commands
outside the support, a remaining force-state difference obeys
$\Delta f_{k+m}=\rho^m\Delta f_k$, rather than becoming identically zero.
Different closed-loop commands require a separate analysis.

The unchanged mechanism subsequently passes its frozen eight-law, two-noise,
six-control replication. All 3,244,800 evaluator transitions and complete
independent replays pass. Low-noise whole-error/original and consistent-
no-repair median ratios are 0.5539 and 0.3666; their descriptive paired-law
intervals lie below one. High-noise ratio to consistent no-repair is 1.0002,
so the high-noise result does not isolate a patching benefit. Against matched
repaired current-only, returning early-error ratios 0.7641/0.7086 and probe
ratios 0.3700/0.2933 pass both declared memory screens. This operational
improvement is stronger than unchanged archive coefficients, but not uniform:
one high-noise law has 17.03\% worse whole error. Primary numeric state peaks
at 52,176 bytes. A fixed gain-one integral-feedback control fails to match the
learned controllers on the two exposed pilot cases despite zero probes and
32-byte state; all its action/physical replays pass. Neither this narrow
classical falsifier nor the eight-law median screen is a general stability,
no-forgetting or robotics-superiority theorem.

\paragraph{Alarm attribution requires separating observation and model errors.}
For a frozen fitted law on a recorded command sequence, write the force
observation as $\tilde f_k=f_k+\epsilon_k$. Then
\[
 \hat\rho\tilde f_{k-1}+(1-\hat\rho)\hat s(u_k)-\tilde f_k
 =e_k^{\rm model}+\hat\rho\epsilon_{k-1}-\epsilon_k,
 \qquad
 e_k^{\rm model}=\hat\rho f_{k-1}+(1-\hat\rho)\hat s(u_k)-f_k.
\]
Squared-error attribution must retain the cross term; pooled energy and
counts of threshold events answer different questions. In 64 old long-dwell
traces this decomposition reproduces all observed alarms. For returning
historical runs, model-only errors exceed the four-row threshold in 630/682
low-noise alarms, with zero noise-only events. At high noise, 51/1,791 meet
the model-only predicate while 1,733 meet the noise-only predicate under the
valid identity law. The component predicates are not generally disjoint.
The nominal empirical threshold is below the high-noise observation RMS.
Neither a local approximation correction nor a retrieval mechanism remedies
that statistical mismatch by itself. Actual commands and previously fitted
coefficients remain noise-dependent in this diagnostic, so removing the
explicit observation term is not a counterfactual noise-free experiment or
a new uniform validity guarantee.

\paragraph{Predictable paired validation avoids a Gaussian inverse-force assumption.}
Let parent and candidate forecasts $p_t,c_t$ be measurable before observing
$y_t=\mu_t+\epsilon_t$, with conditionally Gaussian noise of known covariance
$\Sigma_t$. No membership of $\mu_t$ in either model class is required. For
$d_t=c_t-p_t$, observed gain $D_t=\|p_t-y_t\|^2-\|c_t-y_t\|^2$ and conditional
mean gain $A_t=\|p_t-\mu_t\|^2-\|c_t-\mu_t\|^2$,
\[
 D_t-A_t=2d_t^\top\epsilon_t,\qquad
 V_n=4\sum_{t\le n}d_t^\top\Sigma_t d_t.
\]
The classical two-sided normal-mixture boundary\cite{howard2021confidence}
therefore yields, for fixed $\eta>0$,
\[
 \Pr\!\left(\forall n:\left|\sum_{t\le n}(D_t-A_t)\right|
 <\sqrt{(V_n+\eta)\log\frac{V_n+\eta}{\eta\alpha^2}}\right)\ge1-\alpha.
\]
Summable error allocations over adaptively initiated comparisons protect
multiple proposals, provided each forecast remains predictable and the noise
proxy is valid. A positive cumulative lower bound certifies improvement on
those conditional forecast losses, not every state or a different control
trajectory. If an observation map is linear in an edit, $d_t=F_t\delta$,
gain and variance are coefficient products with residual moments and Grams;
an arbitrary nonlinear simulator does not inherit that linearity. Identical
predictions contribute exactly zero gain and variance.

Using known next-velocity noise rather than assuming Gaussian output from the
bounded nonlinear force observer, a frozen retrospective audit finds positive
final lower bounds on all 66 existing sixteen-row validation blocks, including
all 61 accepted repairs. All 1,056 prefix intervals cover evaluator truth;
the five extra positive comparisons had been rejected by a different force-
adequacy predicate. This is not a new-controller false-update estimate or a
retroactive prospective trial. The sequential boundary is established prior
work; the application connects local program repair to its stated observation
contract without promoting prediction confidence to physical safety.

The subsequent frozen prospective-supervisor pilot makes this separation
operationally concrete. All 415 selected comparisons have positive true
cumulative forecast gain, and 405,600 physical steps plus an equal independent
replay pass. Yet the primary exhausts fitting before returning regimes in
both noise conditions and fails the useful-memory and repair-attribution
screens. Low-noise whole RMS is 2.618 times the old repaired bank's; probe
counts are 3.886/3.163 times old low/high counts. All thirteen low-noise
rejected primary global proposals are genuinely worse on their validation
states. A correct rejection rule does not make continued acquisition useful.
The statistical statement concerns predictable losses at issued actions;
it neither identifies a regime switch nor establishes value of further
experimentation or performance at an edited model's inverse actions.

Ending rejected transactions is a causal lifecycle intervention, not a change
to that statistical theorem. In the next frozen pilot, 88,064 old-prefix
steps match bitwise before the first rejection, and 270,400 new steps plus
equal physical replay pass. Primary error improves 24.60\%/17.07\% over the
failed supervisor without primary exhaustion, but full useful-memory and
repair-attribution screens still fail. High-noise historical memory has
1.1785 times matched current-only error. All 593 selected comparisons have
positive true prediction gain, reinforcing rather than closing the gap
between comparison validity, experiment value and useful behavior.

\paragraph{A cardinal counterexample at the inverse-action interface.}
Let the true static plant be $f(u)=u$, parent $p(u)=2u$, and request $r=1/2$.
Define a continuous piecewise-linear $c$ through
$(0,0),(1/4,1/4),(7/8,3/10),(1,2)$, with $c=p$ outside $[0,1]$.
Its correction $c-p$ is an exact cardinal $B_1$ expansion on spacing $1/8$
with zero endpoint coefficients. For two nonempty blocks of observations
at the parent action $u_p=p^{-1}(r)=1/4$, the existing quadratic guard admits
$q=p+(19/20)(c-p)$. Each residual falls from $1/4$ to $1/80$, a 99.75\%
SSE reduction. The minimum slope of $q$ is $22/125>0$, its off-support
values remain exactly those of $p$, and its correction has finite mass and
derivative energy given by the usual coefficient Grams. Yet
\[
 u_q=q^{-1}(1/2)=\frac{192}{217},\qquad
 \frac{|f(u_q)-1/2|}{|f(u_p)-1/2|}=\frac{334}{217}>1.
\]
The squared-error ratio is 2.3690. This disproves control-improvement
implications based only on monotonicity, protected support and blockwise
observed-loss improvement. It does not assert that the regularized fitter
would construct this deliberately chosen edit. The new action lies inside
the edit support but outside the validation set; these are different domains.
Established safe-learning approaches instead require appropriate model bounds
and verify the proposed policy's state-action region, with explicit regularity,
initial-policy and Lyapunov assumptions\cite{berkenkamp2017safe}.

\paragraph{Behavioral model assistance is different from trusting a model.}
For fixed potential parent/candidate costs $L_0,L_1$, pre-outcome forecasts
$m_0,m_1$, and a fair randomized policy choice $A$, consider the standard
model-assisted contrast\cite{dudik2011doubly}
\[
 Z=m_0-m_1+2\mathbf{1}_{\{A=0\}}(L_0-m_0)
          -2\mathbf{1}_{\{A=1\}}(L_1-m_1).
\]
Enumerating the two choices gives
\[
 \mathbb{E}_A Z=L_0-L_1,\qquad
 \operatorname{Var}_A Z=(L_0+L_1-m_0-m_1)^2.
\]
The conditional variance depends on the predicted cost sum, not the predicted
ordering: true costs $(1,3)$ and forecasts $(3,1)$ have zero conditional
randomization variance but the wrong direct ordering. Neither this identity
nor a measured variance ratio establishes total uncertainty over future
noise and contexts, fewer physical trials, or safe deployment. Time-valid
off-policy evaluation with reward predictors and deployment gates is already
established under its contextual-bandit assumptions\cite{karampatziakis2021offpolicy};
those assumptions must not be silently extended to persistent physical control.

The common-state audit of all 66 actual repair-validation proposals gives
58/61 behavioral improvements among accepted edits and none among the five
rejected edits, but three accepted fitter outputs regress in every one of
four shared-noise continuations. No horizon crosses a regime boundary.
The frozen all-proposal behavioral screen fails at high noise; fixed
parent/candidate-world cost averaging gives conditional variance ratios
0.628/0.484 to IPS, failing both large-gain screens. These are finite-panel
counterfactual diagnostics, not fresh physical laws or deployable oracle data.

\paragraph{Fine-grid evidence memory does not imply a useful retention guard.}
For a fixed parent and pole, let $F$ be the scaled cardinal-linear design and
$e$ its observed residual on an old calibration block. A future patch $c$ in
that declared fine space changes the block loss by
$2c^\top b+c^\top Gc$, where $b=F^\top e$ and $G=F^\top F$.
Two-tap support makes $G$ tridiagonal for each independent output channel.
The two-block, two-channel 65-knot implementation occupies 6,288 numeric bytes
plus its shared grid, and matches direct and additive evaluation on all 66
archived calibrations. The fine moments are accumulated from those observations,
not inferred from insufficient coarse statistics. Requiring non-increase on
both old blocks nevertheless vetoes only two of three harmful accepted repairs
and retains 37/58 useful ones. This failed guard illustrates the distinction
between exact retained empirical objectives, physical-law compatibility and
future control performance; it does not establish a zero-forgetting theorem.

\paragraph{Discounted continuous learning is an exact objective, not timely adaptation.}
For the fixed cardinal basis with affine tails, parameterize observed force as
\[
 \hat f_t=\rho p_t+\epsilon u_t+B(u_t)v,\qquad
 0\leq\rho\leq .85,\quad v_{i+1}\geq v_i,\quad \epsilon=.02.
\]
The static coefficients are $c=(\epsilon\,\xi+v)/(1-\rho)$, where the
centers $\xi$ reproduce the identity. This scaled slope floor makes the
fitting tails consistent with the positive-floor compiled inverse, but
excludes an exactly flat dead zone. With fixed discount $\gamma=.995$,
the empirical quadratic obeys $H_t=\gamma H_{t-1}+x_tx_t^\top$ and
$b_t=\gamma b_{t-1}+x_t(f_t-\epsilon u_t)$ for $x_t=(p_t,B(u_t))$.
Its cardinal block remains banded with a causal border. The undiscounted
prior integrates $[\rho u+B(u)v-(1-\epsilon)u]^2$ over the uniform interior;
four Gaussian nodes per cell exactly integrate each degree-six product in
real arithmetic. Restricting $v$ to an affine sequence gives the matched
three-parameter control with the same functional prior.

Solving this objective every sixteen rows is numerically verified on all
16,896 channel fits, with a 6,072-byte cardinal learner state. Nevertheless
the resulting probe-free controller loses to RLS and the repair bank under
both noise conditions; return-early error exceeds the bank by factors
2.584/3.352. Discounted sufficient statistics preserve the declared weighted
objective, not knowledge of which hidden physical regime is currently active.
Online spline identification and inverse control are established
precedents\cite{hong2012inverse,folgheraiter2016bsnn}; this negative
capability test does not establish novelty or rule out their broader class.

\paragraph{Addressable retained behavior is distinct from hindsight selection.}
For each immutable past bank program $m$, four existing ordinary alarm rows
define an empirical score
$S_m=\sum_{t=k-3}^{k}\|v_{t+1}^{\rm obs}-\widehat v_{t+1}^{(m)}\|^2$.
The predictor uses recorded current state, issued command, preceding observed
force and known nominal mechanics. All programs predate these rows; the
minimizing slot is sealed before evaluator-only physical continuations.
The hindsight minimum of realized 160-step bank costs is a separate finite-
bank diagnostic, not an optimal-control bound or a selectable causal program.

On sixty first-return alarm windows, this empirical addressing captures
99.93\%/98.58\% of low/high-noise aggregate adaptive-source-to-hindsight
squared-cost headroom, passing the fixed coverage and behavioral screens.
Force-SSE ranking gives identical choices. Thus these data support useful
early retrieval but no distinct gain from velocity-domain prediction.
Alarm timing remains inherited and empirically calibrated; neither selecting
four stopped observations nor ranking their SSE creates a confidence sequence.
Full lifetime operation at unfamiliar as well as returning dynamics is still
untested. Multiple-model selection is established prior art\cite{narendra2003multiple};
compact executable memory and this conditional behavioral result do not by
themselves establish switching stability or zero future forgetting.

\paragraph{A fixed-prefix bound prevents misplaced architectural optimization.}
If a proposed intervention leaves squared errors $e_t^2$ unchanged for
$t<T_0$, its best possible whole-run RMS relative to a fixed baseline with
squared cost $C_b$ obeys
\[
 \frac{\operatorname{RMS}_{\rm new}}{\operatorname{RMS}_b}
 \geq \sqrt{\frac{\sum_{t<T_0}e_t^2}{C_b}}.
\]
This follows only from nonnegative future costs and equal horizon length;
equality need not be physically achievable. In the full causal early-reuse
pilot, both 14,600-step pre-return traces match the old bank bitwise. The
high-noise bound is 0.97426, so a target ratio of 0.80 cannot be reached by
changing returns alone. The primary indeed eliminates returning probes and
improves median early-return RMS, but fails the whole high-noise screen.
The bound identifies first-time learning as the next task-level bottleneck,
not a reason to weaken the frozen criterion or assert universal zero forgetting.

\paragraph{Causal-history propagation keeps a conditional linear fit.}
For a fixed stable scalar pole $\rho$, let $q=(1-\rho)c$ and form
\[
 \Phi_t=\rho\Phi_{t-1}+B(u_t),\quad \Phi_{-1}=0,\qquad
 \widehat f_t=\rho^{t+1}f_{-1}^{\rm obs}+\Phi_tq.
\]
Monotonicity is imposed by $q=Tz$, with cumulative-sum matrix $T$, free
initial coefficient and nonnegative increments. A second-difference penalty
and ridge on $q$ give a bounded linear least-squares problem conditional on
$\rho$. Its empirical Gram is generally dense. A finite numerical pole search
does not prove global optimality. Using propagated force history avoids
repeated noisy lagged-force regressors, but the noisy initial condition and
adaptive inputs remain; this is not raw-velocity maximum likelihood.

The corresponding sixteen-row physical diagnostic fails the distinct
behavioral advantage over the existing jointly linear ARX cardinal fit.
An initial TRF numerical failure is retained; a separately frozen BVLS
implementation verifies KKT conditions, independent convolution-matrix
assembly and full physical replay, but yields primary/ARX RMS ratios
1.00009/0.99274. Both methods can reduce source continuation cost, without
establishing reliable unrestricted short acquisition. Exact operator algebra
and correct optimization therefore do not establish new practical learning
efficiency, and this fitting intervention closes without a parameter sweep.

\paragraph{Sealed predictive validation does not certify an acquisition lifecycle.}
A candidate fitted to an initial probe block can be independent of later
validation targets while still lacking any closed-loop improvement guarantee.
The relevant distribution changes when its inverse chooses actions; a finite
validation inequality does not bound that change, archive selection, or
future reacquisition cost. Gate 357 implements genuine before-target
forecasts for a sixteen-probe candidate and sixteen ordinary RLS-controlled
validation actions, without refitting on validation. Complete independent
replay confirms this causality, yet whole tracking error increases by factors
1.73054/3.07420 relative to unchanged early reuse at low/high noise. High-noise
probing grows to 16,864 actions. All arms remain healthy and the stored
programs immutable. Thus empirical predictive improvement, exact program
preservation and operational adaptation are separate properties; none of
the first two is a substitute for the third. The failed fixed intervention
does not rule out all short-data adaptation or all memory-based controllers.

\paragraph{Cardinal weak motion balance compiles to a bounded quadratic.}
For a known single-joint load and an unknown residual, write
$(J_0+a)\ddot q=u+\Phi(u,\dot q,q)c-g_0\sin q$, with $0\leq a\leq1$.
A compact test function with zero value and slope at its endpoints gives
\[
 a\langle q,\phi''\rangle-\langle\Phi,\phi\rangle c
 =\langle u-g_0\sin q,\phi\rangle-J_0\langle q,\phi''\rangle.
\]
Integrating sample hats against a cardinal $\phi$ and $\phi''$ yields fixed
filters $H_0,H_2$ that are exact for those interpolants, not for unknown
between-sample physics. Uniform decimation evaluates only retained test
centers. In Gate 358, overlapping endpoint hats make $H_2H_2^\top$
pentadiagonal, not circulant or tridiagonal. Banded Cholesky whitens the
independent-position-noise contribution; nonlinear regressor noise and model
error remain. This is an application of established weak/GLS identification,
not a new estimation theorem\cite{messenger2020weak}.

For the regularized normal $N$ and right-hand side $b$, partition the bounded
scalar from $c$. Eliminating $c$ gives a positive scalar Schur quadratic;
clamp its unconstrained minimizer to $[0,1]$ and re-solve $c$ for that value.
The resulting global bounded-quadratic solution is checked against independently
assembled BVLS. The empirical coefficient normal is dense despite local
test covariance and continuous tensor-Gram structure. The corresponding
real-motion screen reduces mean RMS by 34.5\% relative to the same basis with
smoothed pointwise acceleration, but fails its absolute trajectory requirement.
Whitening adds no measured advantage. Exact calculus and a correct convex
solve therefore remain distinct from a complete latent-state physical model.

\paragraph{Separate the weak operator from the chosen spatial representation.}
The same integrated quadratic accepts polynomial features. For monomials
$z^a,z^b$ on the normalized cube, their mean-square Gram is
$M_{ab}=\prod_i m(a_i+b_i)$, where $m(k)=0$ for odd $k$ and $1/(k+1)$
otherwise. Pure-coordinate curvature products follow by reducing the
corresponding exponents by four and multiplying the two second-derivative
factors. These exact moments share the temporal filter, whitening and Schur
solve of the cardinal model; they do not require spline evaluation.
Matched and full-cubic controls compile into a $4\times4\times4$ power tensor
with value/slope-matched coordinatewise affine tails. The real-motion
cardinal/control RMS ratios are 0.49887 and 0.88202. The richer control thus
rejects the fixed 20\% representation-advantage requirement, despite three
of four individual cardinal wins. Exact inner-product calculus is a reusable
advantage, but this does not establish that a cardinal basis is necessary
or that its extra parameters buy a new physical capability.

\paragraph{Known diffusion and cardinal-linear forcing need no learned propagator.}
For spherical shell volumes $W=\operatorname{diag}(w_i)$, the finite-volume
diffusion matrix has $WA=K=K^\top\preceq0$. Thus
$W^{1/2}AW^{-1/2}=Q\Lambda Q^\top$ and the modal state
$z=Q^\top W^{1/2}c$ has diagonal dynamics. This is weighted symmetry, not
circulant structure; the zero eigenvalue represents conserved mass and must
not be divided out. Across constant diffusivity and radius, poles scale by
$D/R^2$ while normalized geometry is reusable. For a time step $h$ with
linear current interpolation, each mode satisfies
\[
 z_{k+1}=e^{h\lambda}z_k+h g\{[\varphi_1(h\lambda)-\varphi_2(h\lambda)]I_k
                +\varphi_2(h\lambda)I_{k+1}\},
 \qquad \varphi_1(x)=\frac{e^x-1}{x},\quad
 \varphi_2(x)=\frac{e^x-1-x}{x^2}.
\]
Analytic continuation gives $\varphi_1(0)=1$, $\varphi_2(0)=1/2$; stable
series avoid cancellation. Gate 360 checks this exact interpolant calculus
against quadrature, independent augmented-matrix exponentials and PyBaMM.
It retains all forty particle states and a nonlinear voltage readout, but
neither learns a generator nor proves an exponential-spline representation
necessary. The nearly equal classical-control timing limits the novelty
claim. Known constant coefficients and a fixed finite-volume model are
critical assumptions; concentration-dependent diffusion loses this fixed
diagonal propagation. Accurate voltage also cannot remove parameter
non-identifiability from grouped parameters or flat OCV slopes.

\paragraph{Exact response features remove rollout backpropagation in a restricted class.}
For known $a_k=\exp(-h_k/\tau)$ and exogenous feature rows $\phi_k$, consider
$V_{k+1}=a_k V_k+(1-a_k)(g_k+\phi_k\theta)$.
Let $b_{k+1}=a_kb_k+(1-a_k)g_k$ and
$H_{k+1}=a_kH_k+(1-a_k)\phi_k$, with $b_0=V_0,H_0=0$.
Induction gives $V_k=b_k+H_k\theta$. Thus a whole-trajectory weighted
least-squares objective plus exact mass/curvature penalties is quadratic in
$\theta$, with no differentiation through learned rollouts. This standard
linear-response identity works for polynomial and cardinal features alike.
It assumes a fixed generator and exogenous features: dependence of $\phi_k$
on the unknown predicted state generally destroys the affine reduction.
On real battery-aging development data\cite{wang2026ani}, the corresponding
35-coefficient cardinal fit is inexpensive but loses to an equally small
polynomial and misses the absolute target. Continuous calculus and correct
deployment replay do not alone establish missing-physics identification.

\paragraph{Composition requires cross errors and correct nonlinear calculus.}
For component coefficient errors collected in $E$ and a function Gram $M$,
the exact mixture error is $\|\sum_j\gamma_j e_j\|^2=
\gamma^\top E^\top M E\gamma$. The cross-Gram is PSD, but individual
off-diagonal entries can be negative. In particular $e_2=-e_1$ gives zero
error at mixture $(1,1)$ without identifying either component. A displayed
lower-bound proof in the physical-operator literature\cite{gopakumar2026physical}
reverses the triangle inequality; its dependent universal capacity conclusion
does not follow from that proof. A fixed affine-network counterexample and
fifteen independent quadratures are archived in the composition audit.
Likewise, for constant weights and $u_t+uu_x=0$, $v_t-\nu v_{xx}=0$,
the residual of $w=au+bv$ is
\[
 (a^2-a)uu_x+b^2vv_x+ab(uv_x+vu_x)-a\nu u_{xx}.
\]
The illustrative CompNO expansion\cite{hmida2026compositional} omits the
$-a uu_x$ term. Symbolic differentiation verifies the correction. Neither
audit is a reproduced refutation of the authors' neural experiments: a
nonlinear trained aggregator can learn effects absent from an illustration.
Known commuting flows also compose exactly without learning, but this does
not resolve identifiability or transfer of genuinely unknown components.

\paragraph{Certify the assembled physical law, not the potential's label.}
For constant $M\succ0$, the unforced equation
$M\ddot q+\nabla U(q)+\nabla F(\dot q)=0$ gives
$\dot E=-v^\top\nabla F(v)$, where $v=\dot q$ and
$E=\tfrac12v^\top Mv+U(q)$. Nonnegative $F$ alone does not certify
dissipation: $F(v)=v^2(v-1)^2\geq0$, $F(0)=0$, but at $v=3/4$,
$vF'(v)=-9/64$. The displayed nonnegative-potential penalty in
LOpInf-SpML\cite{sharma2024lagrangian} is therefore not by itself a global
passivity certificate. Likewise scaling $(M,U,F)$ by $\epsilon>0$
shrinks the unnormalized unforced residual without changing the represented
ODE. The paper's fixed-mass rod and membrane examples exclude that scaling;
the criticism concerns its displayed flexible objective, not an observed
collapse of its beam fit or uninspected implementation.

A constructive polynomial-cardinal alternative sets
$g(s)=\eta+\sum_i c_i B_i(s)$ with $\eta>0,c_i\geq0$, and
\[
 F(v)=\frac12\int_0^{v^2}g(s)\,ds,\qquad
 F'(v)=v g(v^2),\qquad vF'(v)\geq\eta v^2.
\]
The force is coefficient-linear at observed velocities. With fixed physical
normalization, quadratic fitting under $c\geq0$ remains convex; nonlinear
rollout fitting is not thereby made quadratic. Some permitted potentials
are nonconvex: $g=.1+B_3$ with a centered unit-scale cubic generator gives
$F''=-61/240$ and positive power $29/160$ at $v^2=3/2$.
The correct finite-velocity force Gram is
\[
 G_{ij}=\int_{-\sqrt S}^{\sqrt S}v^2 B_i(v^2)B_j(v^2)\,dv
       =\int_0^S\sqrt{s}\,B_i(s)B_j(s)\,ds.
\]
Piecewise-polynomial weighted moments are analytic, but the weight destroys
translation invariance. Sixteen exact entries agree with independent
velocity-space quadrature within $4.44\times10^{-16}$; one noiseless
121-observation nonnegative fit recovers four coefficients within
$1.67\times10^{-16}$. This is an elementary calculus construction and
numerical audit, not a new dissipative-systems theorem or a real-data result.
Directional sums extend the safeguard, but do not cover every anisotropic
or history-dependent constitutive law.

Normalization also matters for nonlinear solves. For
$G(u)=\epsilon Ku+D^\top F(u)-b$, a sufficient contraction condition is
$\|K^{-1}D^\top\|L/\epsilon<1$ when $F$ is $L$-Lipschitz.
The dependence printed in Geo-NeW Eqs.~9 and 12\cite{shaffer2026geonew}
cannot generally replace this by $\epsilon\|K^{-1}\|L<1$.
For $K=D=1,\epsilon=.1,F(u)=-.2\tanh(u),b=0$, the latter product is .02,
yet the equation has roots approximately $-1.915008048,0,1.915008048$.
The correctly scaled bound is 2 and makes no uniqueness assertion.
These symbolic and numerical checks concern the displayed premises, not
reproduced failures of the authors' empirical neural models.

\paragraph{Compile scalar maps before filtering their outputs.}
For fixed coefficients, a causal finite-history force predictor of the form
\[
 \hat y_t=b+\sum_{r,j,i}c_{rji}\sum_{\ell=0}^{8}w_{r\ell}
       \beta_i(x_{t-\ell,j})
 =b+\sum_r\sum_{\ell=0}^{8}w_{r\ell}
       \underbrace{\sum_{j,i}c_{rji}\beta_i(x_{t-\ell,j})}_{v_{r,t-\ell}}
\]
can filter the projected output vectors $v_r$ instead of the full basis
vectors. Zero pre-trajectory basis history is used on both sides. For
$w_{r\ell}=\rho_r^\ell(1-\rho_r)/(1-\rho_r^9)$, the exact finite convolution
has the subtractive recurrence
$z_{r,t}=\rho_r z_{r,t-1}+a_r v_{r,t}-a_r\rho_r^9v_{r,t-9}$,
where $a_r=(1-\rho_r)/(1-\rho_r^9)$; $\rho_r=0$ is instantaneous.
Cardinal-to-power matrices compile each scalar map, including value/slope-
matched affine tails. This is linear filtering algebra, not a new theorem
or a learned exponential B-spline generator. Numerical preflight checks
verify full design/compiled/streaming equivalence and trajectory resets.

Exact continuous mass and curvature penalties remain quadratic in the
coefficients. Cross-feature and temporal mixing generally make the empirical
normal dense, however; compact deployed state must not be conflated with
small learning statistics or a circulant inverse. Coefficient updates also
require recomputing the retained projected history under the new map. Reusing
the old filter state would not implement the updated finite-history model.
The corresponding real-telemetry development comparison subsequently rejects
both additive cardinal candidates. Primary all-task force MAE is 0.27052 N,
versus a small MLP's 0.21411 N, and no-contact error violates the frozen limit.
Although fixed-weight state is 125,816 bytes and every validation stream
replays correctly, learning normals require 27,040,648 bytes and CPU inference
is slower than the small MLP. These numerical identities alone therefore do
not establish a practical neural advantage; all test traces remain unopened.
Known geometric mixing of joint-response maps is a distinct structural
hypothesis, whereas additional knots do not remove the additive class's
general inability to represent mixed-coordinate interactions.

\paragraph{Known time-varying mixing preserves linear fitting, not filter commutation.}
Let $\phi_{jp}(x_t)$ be a fixed joint-local feature vector, $c_{jp}$ its
coefficients, $w_{p\ell}$ a finite-history filter and $W(q_t)$ a known
current-pose mixer. Then
\[
 \widehat f_t=b+W(q_t)a_t,\qquad
 (a_t)_j=\sum_{p,\ell}w_{p\ell}
                 \phi_{jp}(x_{t-\ell})^\top c_{jp}+g(q_t)^\top d_j
\]
is linear in $(b,c,d)$ for recorded inputs and poses. Consequently its
trajectory-weighted least-squares objective has exact additive normal
statistics, although the resulting matrix is generally dense. Fixed
coefficient projection may precede each filter, leaving a ring of joint
outputs. In contrast, replacing $W(q_t)\sum_\ell w_\ell a_{t-\ell}$ by
$\sum_\ell w_\ell W(q_{t-\ell})a_{t-\ell}$ changes the model unless the
relevant mixers coincide or a special cancellation holds.

The concrete geometric screen uses
$W(q)=0.1\,\mathrm m\,(J(q)J(q)^\top+(0.01\,\mathrm m)^2I)^{-1}J(q)$,
643 scalar coefficients and 15,305 fixed streaming numeric bytes. All
36,400 validation predictions and the three classical full training-statistic
reconstructions pass independent audits. Nevertheless primary MAE 0.27495 N
is worse than constant mixing 0.26795, matched-input neural 0.23753 and the
unchanged full-input neural 0.21411 N; no-contact error also fails its limit.
The primary's 3,312,744-byte learning statistics and 0.612 s fit do not establish
an application advantage. The damped inverse, base-frame labels and latent
joint-response semantics are modeling assumptions, not an identified physical
torque law. This development-only geometric branch is rejected before test
data are opened, preserving the algebra without promoting the capability claim.

\paragraph{Feedback attribution is not established by a nonzero input weight.}
A frozen retrospective diagnostic reproduces all six asymmetric selected
policies on sixteen exposed physical resets and clamps only their six
policy-force coordinates. The inverse retains force sensing. Full-policy
traces replay bitwise; all 192k evaluator transitions are finite and complete.
Full/clamped restoration is 0.9228/0.9326 for shared-six,
0.8719/0.8716 for polynomial and 0.9051/0.8463 for exact-law policies.
The extra policy input is therefore not a demonstrated positive mechanism
for the primary recovery, even though its learned weights are nonzero.
Clamping can be out of distribution and does not substitute for a matched
training comparison or remove the inverse's sensing requirement.
Larger-budget shared-six control makes fewer model-infeasible requests than
its nested 200k policy (37.02\% to 25.12\%), lowers mean command cost
(1.1208 to 0.9152) and increases velocity (13.1383 to 14.0158).
Its on-policy one-step model RMSE changes from 0.01866 to 0.02060:
better skill is not the same as better identification, and these trajectories
do not share a fixed error-sampling distribution.
After the compound primary panel completes, the same intervention gives
full/clamped restoration 0.8894/0.8570 for shared-six,
0.8518/0.8621 for polynomial and 0.9130/0.8945 for exact-law learning.
The force-input contribution is therefore family- and policy-dependent.
Across both families all 384k evaluator transitions and bitwise original
replay identities pass; none changes the frozen primary score.