The same snapshot can have two futures
A mass passing a point from the left and the same mass passing from the right share a position but not a state. No deterministic position-only predictor can give both correct next positions. Adding more units to that predictor does not recover the missing velocity. History, a measurement, or a justified state estimator must supply the distinction.
What the readout cannot infer from a snapshot
A compact head receives only the information its state exposes. If two command histories produce the same instantaneous input but different actuator conditions, a memoryless head cannot distinguish them regardless of how many spline coefficients it stores. An exponential trace gives the head a coordinate describing the recent command history at a particular time scale. Several traces expose several scales.
The drone program uses six held-input recurrences with time constants from 0.02 to 1 second. Their state joins physical rotor mixtures and kinematic features before direct horizon readout. This is a supplied representation of possible memory, followed by fitted coefficients; it is not a claim that the exact physical time constants were discovered.
Memory can be a physical coordinate
A first-order actuator state summarizes how recent commands are still affecting force. Several stable exponential states offer different time scales. Their recurrence has an exact solution for held input; learning can focus on how the state relates to the measured output. This is not a proof that a small bank is sufficient for every physical system. It is a way to make the memory assumption inspectable.
The ablation tells us what mattered
In the nano-drone archive, removing exponential-memory columns worsens the score from 0.41192 to 0.52886. Removing a much larger set of generic cardinal horizon/state/input columns instead improves it to 0.39438. Within that exposed panel, physical memory contributes more than broad spline capacity. This is a specific ablation, not a universal argument against flexible networks.
A teacher can help choose the coordinates
The aircraft compiler inherits response directions from a neural teacher but replaces its feedback history with a stable carrier. This combination approaches the teacher’s measured accuracy using fewer deployment arrays. The teacher’s computation remains part of acquisition cost. Distillation can be valuable precisely because learning and deployment have different budgets.
Give the learner a state it can use
This offers a useful design test for small ML models: before enlarging the readout, ask whether the hidden physical state is observable in its inputs. The ablation makes that question concrete. Removing memory hurt; removing broad generic spline features helped in the named panel. The right state can be more valuable than a larger function approximator.
The architectural question to ask first
What histories are indistinguishable to this model, and can they require different answers? That question precedes optimizer tuning. If the representation merges decision-relevant states, the experiment needs more information or a different state construction. If the representation is adequate, exact calculus and efficient fitting can then make its use cheaper.
Evidence & further reading
The links below distinguish the project record from foundational literature. This revised story does not add a new application-validation experiment.
- Consolidated research results, including constitutive edges and continual memory. Daniel Schmitter (2026). Local archive snapshot.
- Experiment-family evidence map. Spline research archive (2026). Local archive snapshot.
- Consolidated limitations and research boundaries. Daniel Schmitter (2026). Local archive snapshot.

