Training systems · E38 · Implementation

Make each checkpoint an experiment receipt

Weights, validation traces, resource measurements and an immutable-world check turn a checkpoint into something another engineer can inspect.

Stable-Baselines3 callbackSafetensorsHashing + JSON / NPZ
A checkpoint should identify its weights, configuration and observed validation context, not merely provide a binary file.
Figure 1. Make a checkpoint inspectable. A checkpoint should identify its weights, configuration and observed validation context, not merely provide a binary file. Artifact anatomy. Original vector illustration.

Follow the information

From input to outcome

A callback snapshots the policy and checks its tensors. Validation runs with that saved policy and returns measurements to the receipt; receipt metadata is not itself a model input.

A callback snapshots the policy and checks its tensors. Validation runs with that saved policy and returns measurements to the receipt; receipt metadata is not itself a model input.
Figure 2. Information flow. Solid arrows carry observations, tensors or artifacts; other routes are explicitly labelled. Signal shapes, matrices and network icons are schematic, not measured samples or literal neuron counts. Open full-size SVG ↗ On narrow screens, scroll the diagram horizontally.

Read this alongside Figure 1: A checkpoint should identify its weights, configuration and observed validation context, not merely provide a binary file. The module map and layer-level figures below expand the operations in this route.

Make each checkpoint an experiment receipt: architectureTraining callback: Declared step schedule → Policy state: Finite tensor check → Safetensors snapshot: Weights + file hash → Frozen validation: Saved trajectories → Receipt JSON: Metrics / resources / digests. A high-level module map; comparison branches and training details are explained in the article.TRAINING SYSTEMS / E38 / MODULE MAP01 INPUTTraining callbackDeclared step schedule02 MODULEPolicy stateFinite tensor check03 MODULESafetensors snapshotWeights + file hash04 MODULEFrozen validationSaved trajectories05 OUTPUTReceipt JSONMetrics / resources / digests
Source-grounded module map. Boxes summarize operations, not individual neurons; comparison arms and training paths are detailed below. On a small screen, scroll the diagram horizontally.
Training callback — Declared step schedule

The architecture in context

The system we are building

A weights file is not enough to explain an experiment. The callback saves the policy and critic tensors, checks that they are finite, evaluates a frozen adapter and writes trajectories. A receipt then records training step, update count, elapsed time, resource measurements and content hashes.

Who does what in the stack

Stable-Baselines3 callback
Triggers checkpoint events during training.
Safetensors
Stores tensor state without Python-object pickling.
Hashing + JSON / NPZ
Links receipts and trajectories to exact saved weights.

The project extends the standard training callback rather than modifying SAC. An actuator digest is checked before and after validation, providing a targeted assertion that the declared world-model object did not change. Validation seeds and final evaluation seeds have distinct roles.

Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph.
Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph. Open full-size SVG ↗

From module map to executable structure

Inside SAC policy with added sensor channels

This results or evaluation article shares the implementation in E37. The architecture below describes that companion, not a newly trained model.

Observation grows 17→23; action dimension 6; actor and twin critics have two 256-unit ReLU layers.

Layer-level implementation. B denotes batch size; parameter and shape conventions are expanded in the table.
Layer-level implementation. B denotes batch size; parameter and shape conventions are expanded in the table. Open full-size SVG ↗
Layer / tensor / operation ledger
Layer or branchOutput shapeImplementation detail
Observation + sensorsB × 23Original 17 coordinates plus 6 added sensor values.
Actor trunkB × 256Linear 23→256 ReLU → Linear 256→256 ReLU.
Action distributionB × 6SAC policy distribution; gSDE is enabled in the configured library policy.
Twin critic trunksEach B × 256Each receives 23 observations +6 actions =29; two 256-unit ReLU layers.
Two Q outputsEach B × 1Critic losses and entropy-regularized actor objective; target critics are separate.

Zero initialization preserves the first-layer output at the instant of expansion. For critics, action columns must move past the new observation columns; simply appending zeros would mix sensor and action meaning. Actor and critic branches are parallel consumers of the observation, not one sequential network.

The equation and the update

Wactor+=[Wactor  0],Wcritic+=[Wobs  0  Waction]W_{\rm actor}^{+}=[W_{\rm actor}\;0],\qquad W_{\rm critic}^{+}=[W_{\rm obs}\;0\;W_{\rm action}]

SAC configured learning rate 1e-4, replay 100,000, batch 256, learning_starts 1,000, gamma .99, tau .005, train_freq 1 / gradient_steps 1, automatic entropy initialized .01; gSDE frequency 4 and warm-up enabled.

Learning or solution path. A parameter-update path is different from the forward inference path; see text for target-network, frozen-feature and local-loss boundaries.
Learning or solution path. A parameter-update path is different from the forward inference path; see text for target-network, frozen-feature and local-loss boundaries. Open full-size SVG ↗

Implementation card / no invented benchmarks

Capacity, budget and execution evidence

Parameters / retained state
Expansion adds 1,536 input weights to actor and each critic first layer; policy-distribution parameters, both critics and target critics must also be counted. No unsupported whole-policy total is supplied.
Duration and hardware evidence
The gate/report, not generic M4 hardware specifications, defines the recorded training budgets. No physical-robot timing is claimed.
Source coordinates
E37 lines 83–111; E38 execution gate; E39 report
Current reproduction context
Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.

Counts above are calculated from the stated layer shapes unless identified as saved measurements. They exclude optimizer state and nontrainable buffers. No archived training was rerun for this revision.

What these design choices change

A function-preserving initialization is useful when loading a pretrained controller, but it is not a guarantee after optimizer updates. New sensor columns start with zero influence, yet can acquire gradients from the first update. Optimizer state and target-network mapping must be reconciled independently of parameter shape.

Reproduction and measurement protocol

Compare old and expanded actor outputs with added sensors set to zero. Repeat for critics using arbitrary actions, then nonzero added sensors: initial equality should still hold because those columns are zero. The simulation study’s failed acceptance gates remain failures even if this checkpoint transformation is algebraically correct.

For a new run, save the resolved Python/framework versions, backend, dtype, seed, input shapes, batch size and exact source revision. Start with one batch and one update. Log training steps separately from epochs or environment steps. Do not equate the configured maximum with a completed budget or convergence.

Measure initialization/compilation, data preparation, warmed forward pass, training updates and evaluation separately. Synchronize accelerator work around timed regions using the chosen framework’s supported mechanism. Report peak process memory and framework allocation separately; parameter bytes exclude activations, gradients, optimizer state and input buffers. On a shared machine, begin with a single CPU worker and a small batch rather than claiming all available resources.

A closer look at the implementation

The code that carries the idea

The snippet uses exist_ok=False for a new checkpoint folder, writes contiguous CPU tensors through safetensors and hashes the serialized file. It also records the numeric size of the inference adapter. That size excludes Python runtime, training replay and device allocator overhead.

Python · file · lines 80–96
    def checkpoint(self):
        step=int(self.model.num_timesteps); folder=self.destination/'checkpoints'/f'step_{step:07d}'
        folder.mkdir(parents=True,exist_ok=False)
        tensors={k:v.detach().cpu().contiguous() for k,v in self.model.policy.state_dict().items()}
        if not all(torch.isfinite(v).all() for v in tensors.values()): raise RuntimeError('Nonfinite policy/critic tensor')
        save_file(tensors,str(folder/'policy.safetensors'))
        policy=ArrayPolicy({k:v.numpy() for k,v in tensors.items()})
        result,traces=evaluate(policy,self.actuator,self.actuator,VALIDATION_SEEDS,save_traces=True)
        np.savez_compressed(folder/'validation.npz',**{f'trace_{s}':t for s,t in zip(VALIDATION_SEEDS,traces)})
        assert self.before==actuator_digest(self.actuator)
        result.update({'step':step,'updates':self.model._n_updates,'seconds':time.perf_counter()-self.began,
                       'policy_array_bytes':policy.numeric_bytes,'policy_digest':policy.digest(),
                       'checkpoint_sha256':hashlib.sha256((folder/'policy.safetensors').read_bytes()).hexdigest(),
                       'world_model_immutable':True,**self.resources()})
        p.save_json(folder/'receipt.json',result); self.history.append(result)
        print(json.dumps({'phase':'checkpoint','step':step,'virtual_return':result['mean_return'],
                          'seconds':result['seconds'],'mps_driver_bytes':result['mps_driver_bytes']}),flush=True)

Verbatim archive excerpt from locomotion_policy_learning_gate333.py. Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.

The boundary that matters

A content hash detects changes, not scientific validity or authorship. Safetensors avoids pickle-style executable object loading but does not certify that a policy is safe or correctly trained. Repeated checkpoint validation consumes feedback and must not be renamed final evaluation.

Keep building

Other posts of interest