Evaluation practice · E39 · Negative result

Prediction gains do not automatically become control gains

A multi-arm locomotion gate tests whether an improved model enables a useful policy—not just whether a fitted curve looks better.

Frozen NumPy policiesSimulation harnessProtocol/report files
A better fit to actuator observations and a better controlled trajectory are different outcomes that require separate evaluation.
Figure 1. Prediction is not control. A better fit to actuator observations and a better controlled trajectory are different outcomes that require separate evaluation. Conceptual evaluation gap. Original vector illustration.

Follow the information

From input to outcome

Model fitting, policy selection and final control evaluation are distinct stages. The final gate consumes held-out control measurements, not merely a lower actuator prediction loss.

Model fitting, policy selection and final control evaluation are distinct stages. The final gate consumes held-out control measurements, not merely a lower actuator prediction loss.
Figure 2. Information flow. Solid arrows carry observations, tensors or artifacts; other routes are explicitly labelled. Signal shapes, matrices and network icons are schematic, not measured samples or literal neuron counts. Open full-size SVG ↗ On narrow screens, scroll the diagram horizontally.

Read this alongside Figure 1: A better fit to actuator observations and a better controlled trajectory are different outcomes that require separate evaluation. The module map and layer-level figures below expand the operations in this route.

Prediction gains do not automatically become control gains: system and evaluation mapFrozen actuator arms: Same declared conditions → SAC training seeds: Checkpoint validation → Selected policies: Frozen before final test → Held-out resets: Physical-model evaluation → Joint capability gate: All required conditions. A high-level module map; comparison branches and training details are explained in the article.EVALUATION PRACTICE / E39 / MODULE MAP01 INPUTFrozen actuator armsSame declared conditions02 MODULESAC training seedsCheckpoint validation03 MODULESelected policiesFrozen before final test04 MODULEHeld-out resetsPhysical-model evaluation05 OUTPUTJoint capability gateAll required conditions
Source-grounded module map. Boxes summarize operations, not individual neurons; comparison arms and training paths are detailed below. On a small screen, scroll the diagram horizontally.
Frozen actuator arms — Same declared conditions

The architecture in context

What this comparison asks

The experiment moves beyond measuring a model in isolation. Policies are trained through different actuator representations and evaluated under the declared physical model. The question is whether the representation changes downstream control capability. The saved report includes 32 runs across four families and rejects the primary gate in every family.

Who does what in the stack

Frozen NumPy policies
Support independent evaluation from saved tensors.
Simulation harness
Pairs actuator representations with physical transitions.
Protocol/report files
Retain all required comparisons and the failed gate.

The custom protocol combines training receipts, checkpoint selection and a separate final evaluator. The code excerpt comes from that evaluator; the primary evidence is the report. This is an evaluation case study linked to the SAC architecture, not a new controller invented from the results.

Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph.
Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph. Open full-size SVG ↗

From module map to executable structure

Inside SAC policy with added sensor channels

This results or evaluation article shares the implementation in E37. The architecture below describes that companion, not a newly trained model.

Observation grows 17→23; action dimension 6; actor and twin critics have two 256-unit ReLU layers.

Layer-level implementation. B denotes batch size; parameter and shape conventions are expanded in the table.
Layer-level implementation. B denotes batch size; parameter and shape conventions are expanded in the table. Open full-size SVG ↗
Layer / tensor / operation ledger
Layer or branchOutput shapeImplementation detail
Observation + sensorsB × 23Original 17 coordinates plus 6 added sensor values.
Actor trunkB × 256Linear 23→256 ReLU → Linear 256→256 ReLU.
Action distributionB × 6SAC policy distribution; gSDE is enabled in the configured library policy.
Twin critic trunksEach B × 256Each receives 23 observations +6 actions =29; two 256-unit ReLU layers.
Two Q outputsEach B × 1Critic losses and entropy-regularized actor objective; target critics are separate.

Zero initialization preserves the first-layer output at the instant of expansion. For critics, action columns must move past the new observation columns; simply appending zeros would mix sensor and action meaning. Actor and critic branches are parallel consumers of the observation, not one sequential network.

The equation and the update

Wactor+=[Wactor  0],Wcritic+=[Wobs  0  Waction]W_{\rm actor}^{+}=[W_{\rm actor}\;0],\qquad W_{\rm critic}^{+}=[W_{\rm obs}\;0\;W_{\rm action}]

SAC configured learning rate 1e-4, replay 100,000, batch 256, learning_starts 1,000, gamma .99, tau .005, train_freq 1 / gradient_steps 1, automatic entropy initialized .01; gSDE frequency 4 and warm-up enabled.

Learning or solution path. A parameter-update path is different from the forward inference path; see text for target-network, frozen-feature and local-loss boundaries.
Learning or solution path. A parameter-update path is different from the forward inference path; see text for target-network, frozen-feature and local-loss boundaries. Open full-size SVG ↗

Implementation card / no invented benchmarks

Capacity, budget and execution evidence

Parameters / retained state
Expansion adds 1,536 input weights to actor and each critic first layer; policy-distribution parameters, both critics and target critics must also be counted. No unsupported whole-policy total is supplied.
Duration and hardware evidence
The gate/report, not generic M4 hardware specifications, defines the recorded training budgets. No physical-robot timing is claimed.
Source coordinates
E37 lines 83–111; E38 execution gate; E39 report
Current reproduction context
Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.

Counts above are calculated from the stated layer shapes unless identified as saved measurements. They exclude optimizer state and nontrainable buffers. No archived training was rerun for this revision.

What these design choices change

A function-preserving initialization is useful when loading a pretrained controller, but it is not a guarantee after optimizer updates. New sensor columns start with zero influence, yet can acquire gradients from the first update. Optimizer state and target-network mapping must be reconciled independently of parameter shape.

Reproduction and measurement protocol

Compare old and expanded actor outputs with added sensors set to zero. Repeat for critics using arbitrary actions, then nonzero added sensors: initial equality should still hold because those columns are zero. The simulation study’s failed acceptance gates remain failures even if this checkpoint transformation is algebraically correct.

For a new run, save the resolved Python/framework versions, backend, dtype, seed, input shapes, batch size and exact source revision. Start with one batch and one update. Log training steps separately from epochs or environment steps. Do not equate the configured maximum with a completed budget or convergence.

Measure initialization/compilation, data preparation, warmed forward pass, training updates and evaluation separately. Synchronize accelerator work around timed regions using the chosen framework’s supported mechanism. Report peak process memory and framework allocation separately; parameter bytes exclude activations, gradients, optimizer state and input buffers. On a shared machine, begin with a single CPU worker and a small batch rather than claiming all available resources.

A closer look at the implementation

The code that carries the idea

The evaluation path restores frozen policy arrays and evaluates declared conditions without further optimization. The report’s two SAC seeds and eleven checkpoint levels describe the experiment’s structure; checkpoint count is not an independent-seed count.

Python · file · lines 143–166
def physical_evaluation(condition,kind,seed):
    root=case_path(condition,kind,seed); training=json.loads((root/'trained.json').read_text()); assert training['dependencies']==dependencies()
    actuator,_=law(condition,kind); true=OracleActuator(source.sample(condition,SOURCE_CASE)); results={}
    for name,step in (('selected',training['selected_step']),('initial',0),('nominal',0)):
        folder=root/'evaluation'/name; folder.mkdir(parents=True,exist_ok=False)
        path=root/'checkpoints'/f'step_{step:07d}'/'policy.safetensors'
        if name=='selected': assert hashlib.sha256(path.read_bytes()).hexdigest()==training['selected_checkpoint_sha256']
        policy=ArrayPolicy(load_file(path))
        plant,inverse=(IdentityActuator(),IdentityActuator()) if name=='nominal' else (true,actuator)
        result,traces=evaluate(policy,plant,inverse,EVALUATION_SEEDS,save_traces=True)
        np.savez_compressed(folder/'traces.npz',**{f'trace_{s}':t for s,t in zip(EVALUATION_SEEDS,traces)})
        result.update({'dependencies':dependencies(),'source_policy_step':step,'policy_bytes':policy.numeric_bytes})
        p.save_json(folder/'result.json',result); results[name]=result
    p.save_json(root/'evaluation_summary.json',{'dependencies':dependencies(),'methods':results,'training':training})
    print(json.dumps({'phase':'physical_evaluation','condition':condition,'kind':kind,'seed':seed,
                      'mean_returns':{k:v['mean_return'] for k,v in results.items()}}),flush=True)


if __name__=='__main__':
    parser=argparse.ArgumentParser(); parser.add_argument('phase',choices=('train','evaluate'))
    parser.add_argument('--condition',choices=CONDITIONS,required=True); parser.add_argument('--kind',choices=KINDS,required=True)
    parser.add_argument('--seed',type=int,choices=TRAIN_SEEDS,required=True); parser.add_argument('--device',choices=('mps','cpu'),default='mps')
    args=parser.parse_args()
    if args.phase=='train': train(args.condition,args.kind,args.seed,args.device)

Verbatim archive excerpt from locomotion_policy_learning_gate333.py (companion source E38). Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.

The boundary that matters

These tests involve exposed law families and held-out resets in simulation. They do not show transfer to unknown real actuator physics. The negative gate remains negative even if individual trajectories or intermediate model metrics look encouraging.

Keep building

Other posts of interest