A controller creates its own test distribution
A regression model predicts a recorded next state under a known input. A controller chooses the next input using states created by its previous choices. Small model errors can send it into unfamiliar regions, while a policy trained around one inverse may behave differently after that inverse improves. Average prediction error is not the whole causal chain.
Use an intervention to locate the bottleneck
Suppose a model predicts the actuator almost perfectly but policy performance remains poor. Improving the model further may be the wrong experiment. The true-inverse arm makes this question concrete: keep the actor fixed and replace estimated inversion with supplied true inversion. It is an intervention on the adapter, not an optimal-policy upper bound.
Its failure to restore the required capability shows that inverse-fit accuracy is not enough for this actor and disturbance family. The visited command distribution, delayed force, and actor behavior remain part of the task. A response metric averaged under uniform commands measures a different distribution.
Hold the policy fixed
To isolate the remaining compensation question, the archived panel uses one frozen actor across sixteen fresh actuator laws. It compares shared-prior and polynomial-prior fits, frozen and online RLS, and a supplied true inverse. Mechanics and force sensing remain available. This is a more informative comparison than changing the policy, observations, and model simultaneously.
There is a modest benefit, not the desired recovery
Shared-prior restoration reaches 0.9084 for asymmetric lag and 0.8966 for compound lag. It improves over the RLS controls, but the paired advantage over the polynomial prior is below the required two points in both families. Compound lag also misses the 0.90 restoration threshold and the all-law floor. The complete capability gate fails.
Even the true inverse does not close this gap
Under the same actor, true inversion reaches 0.9123 and 0.8977. That is not an optimal-control ceiling, but it weakens the case that another small fitting improvement will unlock the intended behavior here. The policy’s competence and compatibility deserve attention alongside the model. The archived branch was stopped rather than repeatedly tuning the exposed laws.
Find the bottleneck before refining the model
This is broadly relevant to learned world models. A better simulator metric earns practical meaning only when the policy or planner can exploit it. The experiment does not prove model learning is useless; it identifies why further small improvements to this fitted component did not justify a recovery claim.
The general lesson for world models
Evaluate a model on the behavior it is meant to support. One-step loss, rollout loss, physical residual, and task return each answer a different question. A small fast model can still be a useful component, but its component metrics cannot stand in for the assembled system. This is a boundary worth discovering before deploying a model-maintenance method on a real robot.
Evidence & further reading
The links below distinguish the project record from foundational literature. This revised story does not add a new application-validation experiment.
- Policy-transfer theory and experiments. Daniel Schmitter (2026). Local archive snapshot.
- Consolidated research results, including constitutive edges and continual memory. Daniel Schmitter (2026). Local archive snapshot.
- Consolidated limitations and research boundaries. Daniel Schmitter (2026). Local archive snapshot.


