Repair the part that changed
An actuator develops a deadzone or a delayed response while its surrounding mechanics remain known. A source-learned response space can make a new fit small: identify a few correction coordinates and a stable lag from force observations. This is attractive because it separates existing knowledge from new evidence. It does not identify an unknown change without sensing it.
Where the adapter sits in the control loop
A sensed command–force prefix fits a low-dimensional actuator response. An inverse adapter then maps the unchanged actor’s requested action into a command for the changed actuator. The actor does not receive a new policy-gradient update. Identification, admission, and behavior are separate modules, with the fitting delay charged to the rollout.
The response includes clipping, so its objective is not globally the same quadratic that would describe an unclipped linear-in-coefficient model. Enumerating consistent saturation regions recovers smaller quadratic subproblems. Ordinary stored moments can omit the information needed to choose those regions.
Make the update a transaction
For a fixed feature map, the loss along a proposed coefficient update is a quadratic in the update fraction. Retained Grams and right-hand sides can evaluate that entire curve without replaying every row. A candidate can be admitted on separate observations, rejected, or scaled back. The guarantee applies to that measured objective, not every future robot trajectory.
The deployed equation matters
A fitted actuator clips its output. An ordinary unclipped least-squares objective is then only a surrogate. For a monotone scalar response, sorted observations permit a finite clipping-region decomposition into constrained quadratic problems. The implementation accelerates those problems through lower bounds and batched algebra. It does not make arbitrary learned dynamics globally convex.
The small update rests on a larger system
The control studies supply mechanics, continuing force sensing, a source prior, and a heavily pretrained actor. In the matched-policy panel, sixteen new sensed rows fit the inverse, but the actor inherits millions of earlier transitions. Those resources explain the benefit and must stay in the story. A polynomial-prior fit with the same carrier is an important control.
Repair an interface, then test the behavior
The illustrative robot is not the experimental platform: the control evidence comes from simulation. The practical idea is modular repair of a changed actuator interface. The measured benefit is modest, and even a true inverse fails the larger two-family restoration requirement under the fixed actor. That boundary belongs in the main story.
A useful engineering boundary
Versioned model maintenance can provide reproducible updates and rollback without being a complete autonomous learner. Its interface should bind the coordinates, observation semantics, permitted changes, and consuming controller. Exact preservation of an old model is not a proof that the repaired machine is safe. The companion paper reports both the algebra and the failed full recovery gate.
Evidence & further reading
The links below distinguish the project record from foundational literature. This revised story does not add a new application-validation experiment.
- Policy-transfer theory and experiments. Daniel Schmitter (2026). Local archive snapshot.
- Consolidated research results, including constitutive edges and continual memory. Daniel Schmitter (2026). Local archive snapshot.
- Consolidated limitations and research boundaries. Daniel Schmitter (2026). Local archive snapshot.
- Correction: what exact Gram memory does and does not establish. Spline research archive (2026). Local archive snapshot.
