Physical systems · Research & Algorithms

Update a robot’s dynamics model without rewriting its entire memory

A changed actuator may need a small model update. The useful unit of adaptation is a checked proposal, not an unconditional replacement of everything the system knows.

Editorial 3D rendering of a mechanically plausible quadruped robot in a laboratory; the near hip has a copper actuator casingLOCAL MODEL MAINTENANCEAn actuator response changes
REPAIR A MODEL, NOT ITS HISTORY

Let one part change.
Keep the rest intact.

A local coefficient update changes a chosen response region. Protected behavior stays on the original curve.

protectedinput →response

— proposed response · ··· original

Mechanism illustration, not a measured actuator trace. Preserving this curve is not a robot-safety guarantee.

65%
AI-generated editorial robot rendering with an analytically computed, synthetic local-update overlay. The archived control experiment used HalfCheetah-v5 simulation—not the pictured quadruped or a physical laboratory trial. The curve illustrates local response preservation; the separate admission calculation and closed-loop results are explained below.

Follow the information

From input to outcome

Identification produces an actuator response model; the admitted inverse changes the command interface, not the actor weights. Clipping-region calculations handle the deployed response rather than silently replacing it by an unclipped quadratic.

Scroll the diagram horizontally to follow the route. Keyboard: focus the diagram, then use the arrow keys.

Sensed command–force → Fit actuator correction → Admit candidate → Inverse adapter → Changed actuator. Identification produces an actuator response model; the admitted inverse changes the command interface, not the actor weights. Clipping-region calculations handle the deployed response rather than silently replacing it by an unclipped quadratic.
Information-flow map. Simulated control with sensing and source priors; no real-robot safety claim. Original vector schematic based on the method and evidence discussed in this article; signal shapes and icons are illustrative, not additional measurements. Open full-size diagram ↗

Read the main route from left to right; labelled side branches show additional inputs, checks or feedback. The sections below explain the operations and their experimental limits.

Repair the part that changed

An actuator develops a deadzone or a delayed response while its surrounding mechanics remain known. A source-learned response space can make a new fit small: identify a few correction coordinates and a stable lag from force observations. This is attractive because it separates existing knowledge from new evidence. It does not identify an unknown change without sensing it.

Where the adapter sits in the control loop

A sensed command–force prefix fits a low-dimensional actuator response. An inverse adapter then maps the unchanged actor’s requested action into a command for the changed actuator. The actor does not receive a new policy-gradient update. Identification, admission, and behavior are separate modules, with the fitting delay charged to the rollout.

The response includes clipping, so its objective is not globally the same quadratic that would describe an unclipped linear-in-coefficient model. Enumerating consistent saturation regions recovers smaller quadratic subproblems. Ordinary stored moments can omit the information needed to choose those regions.

A model repair still has to help the actor
A model repair still has to help the actor. Original scientific diagram; the stated component and information flow, not an additional experiment. Open full-size figure ↗

Make the update a transaction

For a fixed feature map, the loss along a proposed coefficient update is a quadratic in the update fraction. Retained Grams and right-hand sides can evaluate that entire curve without replaying every row. A candidate can be admitted on separate observations, rejected, or scaled back. The guarantee applies to that measured objective, not every future robot trajectory.

ΔL(η)=2bη+aη2,a=∥ΦΔc∥2≥0\begin{gathered}\Delta L(\eta)=2b\eta+a\eta^2,\\ a=\|\Phi\Delta c\|^2\ge0\end{gathered}
For a fixed quadratic measurement objective, every interpolation step between old and proposed coefficients can be evaluated from retained statistics.

The deployed equation matters

A fitted actuator clips its output. An ordinary unclipped least-squares objective is then only a surrogate. For a monotone scalar response, sorted observations permit a finite clipping-region decomposition into constrained quadratic problems. The implementation accelerates those problems through lower bounds and batched algebra. It does not make arbitrary learned dynamics globally convex.

The small update rests on a larger system

The control studies supply mechanics, continuing force sensing, a source prior, and a heavily pretrained actor. In the matched-policy panel, sixteen new sensed rows fit the inverse, but the actor inherits millions of earlier transitions. Those resources explain the benefit and must stay in the story. A polynomial-prior fit with the same carrier is an important control.

Complete compound-lag family medians under the same frozen actor. The true inverse is an intervention, not an optimal-policy upper bound. The asymmetric family remains in the table.
Complete compound-lag family medians under the same frozen actor. The true inverse is an intervention, not an optimal-policy upper bound. The asymmetric family remains in the table. Open full-size figure ↗

Repair an interface, then test the behavior

The illustrative robot is not the experimental platform: the control evidence comes from simulation. The practical idea is modular repair of a changed actuator interface. The measured benefit is modest, and even a true inverse fails the larger two-family restoration requirement under the fixed actor. That boundary belongs in the main story.

A useful engineering boundary

Versioned model maintenance can provide reproducible updates and rollback without being a complete autonomous learner. Its interface should bind the coordinates, observation semantics, permitted changes, and consuming controller. Exact preservation of an old model is not a proof that the repaired machine is safe. The companion paper reports both the algebra and the failed full recovery gate.

Evidence & further reading

The links below distinguish the project record from foundational literature. This revised story does not add a new application-validation experiment.

  1. Policy-transfer theory and experiments. Daniel Schmitter (2026). Local archive snapshot.
  2. Consolidated research results, including constitutive edges and continual memory. Daniel Schmitter (2026). Local archive snapshot.
  3. Consolidated limitations and research boundaries. Daniel Schmitter (2026). Local archive snapshot.
  4. Correction: what exact Gram memory does and does not establish. Spline research archive (2026). Local archive snapshot.