Physical systems · Research & Algorithms

Why a better dynamics model can still produce a worse controller

Prediction error is easy to measure. Useful action is harder—and can remain poor even when the actuator model is nearly right.

EXPLORE THE IDEA

The feedback loop is the test

A nearly correct inverse did not deliver full recovery.

COMPOUND-LAG RESTORATIONSame frozen actor, fresh actuator lawsShared-prior fit0.8966Supplied true inverse0.8977PREDECLARED TARGETThe full recovery gate still failed0.90restoration thresholdA better model is notautomatically a better controller.The true inverse is not an optimal-policy ceiling
50%
Actual compound-lag restoration summaries: shared prior 0.8966; supplied true inverse 0.8977; predeclared target 0.90. These are aggregate archived results with the same frozen actor, not a simulated rollout video or an optimal-control bound.

Follow the information

From input to outcome

The actor consumes states created by its own earlier decisions. Replacing the adapter is an intervention inside that loop. A good inverse fit and good closed-loop behavior are different outputs, so both must be measured.

Scroll the diagram horizontally to follow the route. Keyboard: focus the diagram, then use the arrow keys.

Observed control state → Frozen actor → Chosen inverse adapter → Actuator + mechanics → Control outcome. The actor consumes states created by its own earlier decisions. Replacing the adapter is an intervention inside that loop. A good inverse fit and good closed-loop behavior are different outputs, so both must be measured.
Information-flow map. True inversion is a diagnostic arm, not an optimal-policy upper bound. Original vector schematic based on the method and evidence discussed in this article; signal shapes and icons are illustrative, not additional measurements. Open full-size diagram ↗

Read the main route from left to right; labelled side branches show additional inputs, checks or feedback. The sections below explain the operations and their experimental limits.

A controller creates its own test distribution

A regression model predicts a recorded next state under a known input. A controller chooses the next input using states created by its previous choices. Small model errors can send it into unfamiliar regions, while a policy trained around one inverse may behave differently after that inverse improves. Average prediction error is not the whole causal chain.

at=π(s^t),s^t+1=f^(s^t,at)\begin{gathered}a_t=\pi(\widehat s_t),\\ \widehat s_{t+1}=\widehat f(\widehat s_t,a_t)\end{gathered}
Model errors change the states the policy sees, and the policy changes the inputs the model must handle. The relevant test is the complete feedback loop.

Use an intervention to locate the bottleneck

Suppose a model predicts the actuator almost perfectly but policy performance remains poor. Improving the model further may be the wrong experiment. The true-inverse arm makes this question concrete: keep the actor fixed and replace estimated inversion with supplied true inversion. It is an intervention on the adapter, not an optimal-policy upper bound.

Its failure to restore the required capability shows that inverse-fit accuracy is not enough for this actor and disturbance family. The visited command distribution, delayed force, and actor behavior remain part of the task. A response metric averaged under uniform commands measures a different distribution.

A model repair still has to help the actor
A model repair still has to help the actor. Original scientific diagram; the stated component and information flow, not an additional experiment. Open full-size figure ↗

Hold the policy fixed

To isolate the remaining compensation question, the archived panel uses one frozen actor across sixteen fresh actuator laws. It compares shared-prior and polynomial-prior fits, frozen and online RLS, and a supplied true inverse. Mechanics and force sensing remain available. This is a more informative comparison than changing the policy, observations, and model simultaneously.

There is a modest benefit, not the desired recovery

Shared-prior restoration reaches 0.9084 for asymmetric lag and 0.8966 for compound lag. It improves over the RLS controls, but the paired advantage over the polynomial prior is below the required two points in both families. Compound lag also misses the 0.90 restoration threshold and the all-law floor. The complete capability gate fails.

Even the true inverse does not close this gap

Under the same actor, true inversion reaches 0.9123 and 0.8977. That is not an optimal-control ceiling, but it weakens the case that another small fitting improvement will unlock the intended behavior here. The policy’s competence and compatibility deserve attention alongside the model. The archived branch was stopped rather than repeatedly tuning the exposed laws.

Complete compound-lag family medians under the same frozen actor. The true inverse is an intervention, not an optimal-policy upper bound. The asymmetric family remains in the table.
Complete compound-lag family medians under the same frozen actor. The true inverse is an intervention, not an optimal-policy upper bound. The asymmetric family remains in the table. Open full-size figure ↗

Find the bottleneck before refining the model

This is broadly relevant to learned world models. A better simulator metric earns practical meaning only when the policy or planner can exploit it. The experiment does not prove model learning is useless; it identifies why further small improvements to this fitted component did not justify a recovery claim.

The general lesson for world models

Evaluate a model on the behavior it is meant to support. One-step loss, rollout loss, physical residual, and task return each answer a different question. A small fast model can still be a useful component, but its component metrics cannot stand in for the assembled system. This is a boundary worth discovering before deploying a model-maintenance method on a real robot.

Evidence & further reading

The links below distinguish the project record from foundational literature. This revised story does not add a new application-validation experiment.

  1. Policy-transfer theory and experiments. Daniel Schmitter (2026). Local archive snapshot.
  2. Consolidated research results, including constitutive edges and continual memory. Daniel Schmitter (2026). Local archive snapshot.
  3. Consolidated limitations and research boundaries. Daniel Schmitter (2026). Local archive snapshot.