The architecture in context
What this comparison asks
The experiment moves beyond measuring a model in isolation. Policies are trained through different actuator representations and evaluated under the declared physical model. The question is whether the representation changes downstream control capability. The saved report includes 32 runs across four families and rejects the primary gate in every family.
Who does what in the stack
- Frozen NumPy policies
- Support independent evaluation from saved tensors.
- Simulation harness
- Pairs actuator representations with physical transitions.
- Protocol/report files
- Retain all required comparisons and the failed gate.
The custom protocol combines training receipts, checkpoint selection and a separate final evaluator. The code excerpt comes from that evaluator; the primary evidence is the report. This is an evaluation case study linked to the SAC architecture, not a new controller invented from the results.
From module map to executable structure
Inside SAC policy with added sensor channels
This results or evaluation article shares the implementation in E37. The architecture below describes that companion, not a newly trained model.
Observation grows 17→23; action dimension 6; actor and twin critics have two 256-unit ReLU layers.
| Layer or branch | Output shape | Implementation detail |
|---|---|---|
| Observation + sensors | B × 23 | Original 17 coordinates plus 6 added sensor values. |
| Actor trunk | B × 256 | Linear 23→256 ReLU → Linear 256→256 ReLU. |
| Action distribution | B × 6 | SAC policy distribution; gSDE is enabled in the configured library policy. |
| Twin critic trunks | Each B × 256 | Each receives 23 observations +6 actions =29; two 256-unit ReLU layers. |
| Two Q outputs | Each B × 1 | Critic losses and entropy-regularized actor objective; target critics are separate. |
Zero initialization preserves the first-layer output at the instant of expansion. For critics, action columns must move past the new observation columns; simply appending zeros would mix sensor and action meaning. Actor and critic branches are parallel consumers of the observation, not one sequential network.
The equation and the update
SAC configured learning rate 1e-4, replay 100,000, batch 256, learning_starts 1,000, gamma .99, tau .005, train_freq 1 / gradient_steps 1, automatic entropy initialized .01; gSDE frequency 4 and warm-up enabled.
Implementation card / no invented benchmarks
Capacity, budget and execution evidence
- Parameters / retained state
- Expansion adds 1,536 input weights to actor and each critic first layer; policy-distribution parameters, both critics and target critics must also be counted. No unsupported whole-policy total is supplied.
- Duration and hardware evidence
- The gate/report, not generic M4 hardware specifications, defines the recorded training budgets. No physical-robot timing is claimed.
- Source coordinates
- E37 lines 83–111; E38 execution gate; E39 report
- Current reproduction context
- Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.
Counts above are calculated from the stated layer shapes unless identified as saved measurements. They exclude optimizer state and nontrainable buffers. No archived training was rerun for this revision.
What these design choices change
A function-preserving initialization is useful when loading a pretrained controller, but it is not a guarantee after optimizer updates. New sensor columns start with zero influence, yet can acquire gradients from the first update. Optimizer state and target-network mapping must be reconciled independently of parameter shape.
Reproduction and measurement protocol
Compare old and expanded actor outputs with added sensors set to zero. Repeat for critics using arbitrary actions, then nonzero added sensors: initial equality should still hold because those columns are zero. The simulation study’s failed acceptance gates remain failures even if this checkpoint transformation is algebraically correct.
For a new run, save the resolved Python/framework versions, backend, dtype, seed, input shapes, batch size and exact source revision. Start with one batch and one update. Log training steps separately from epochs or environment steps. Do not equate the configured maximum with a completed budget or convergence.
Measure initialization/compilation, data preparation, warmed forward pass, training updates and evaluation separately. Synchronize accelerator work around timed regions using the chosen framework’s supported mechanism. Report peak process memory and framework allocation separately; parameter bytes exclude activations, gradients, optimizer state and input buffers. On a shared machine, begin with a single CPU worker and a small batch rather than claiming all available resources.
A closer look at the implementation
The code that carries the idea
The evaluation path restores frozen policy arrays and evaluates declared conditions without further optimization. The report’s two SAC seeds and eleven checkpoint levels describe the experiment’s structure; checkpoint count is not an independent-seed count.
def physical_evaluation(condition,kind,seed):
root=case_path(condition,kind,seed); training=json.loads((root/'trained.json').read_text()); assert training['dependencies']==dependencies()
actuator,_=law(condition,kind); true=OracleActuator(source.sample(condition,SOURCE_CASE)); results={}
for name,step in (('selected',training['selected_step']),('initial',0),('nominal',0)):
folder=root/'evaluation'/name; folder.mkdir(parents=True,exist_ok=False)
path=root/'checkpoints'/f'step_{step:07d}'/'policy.safetensors'
if name=='selected': assert hashlib.sha256(path.read_bytes()).hexdigest()==training['selected_checkpoint_sha256']
policy=ArrayPolicy(load_file(path))
plant,inverse=(IdentityActuator(),IdentityActuator()) if name=='nominal' else (true,actuator)
result,traces=evaluate(policy,plant,inverse,EVALUATION_SEEDS,save_traces=True)
np.savez_compressed(folder/'traces.npz',**{f'trace_{s}':t for s,t in zip(EVALUATION_SEEDS,traces)})
result.update({'dependencies':dependencies(),'source_policy_step':step,'policy_bytes':policy.numeric_bytes})
p.save_json(folder/'result.json',result); results[name]=result
p.save_json(root/'evaluation_summary.json',{'dependencies':dependencies(),'methods':results,'training':training})
print(json.dumps({'phase':'physical_evaluation','condition':condition,'kind':kind,'seed':seed,
'mean_returns':{k:v['mean_return'] for k,v in results.items()}}),flush=True)
if __name__=='__main__':
parser=argparse.ArgumentParser(); parser.add_argument('phase',choices=('train','evaluate'))
parser.add_argument('--condition',choices=CONDITIONS,required=True); parser.add_argument('--kind',choices=KINDS,required=True)
parser.add_argument('--seed',type=int,choices=TRAIN_SEEDS,required=True); parser.add_argument('--device',choices=('mps','cpu'),default='mps')
args=parser.parse_args()
if args.phase=='train': train(args.condition,args.kind,args.seed,args.device)Verbatim archive excerpt from locomotion_policy_learning_gate333.py (companion source E38). Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.
The boundary that matters
These tests involve exposed law families and held-out resets in simulation. They do not show transfer to unknown real actuator physics. The negative gate remains negative even if individual trajectories or intermediate model metrics look encouraging.