Evaluation practice · E36 · Negative result

A nine-second RL run can validate plumbing, not competence

A Pong smoke test completes on MPS but does not learn a useful policy. Its action distribution is more informative than its runtime headline.

Gymnasium / ALEPyTorch MPSGreedy evaluator
The archived nine-second execution and −21 greedy return describe a plumbing test, not a successfully trained Pong agent.
Figure 1. A smoke test is not a trained agent. The archived nine-second execution and −21 greedy return describe a plumbing test, not a successfully trained Pong agent. Archived smoke-test record. Original vector illustration.

Follow the information

From input to outcome

This short run checks that collection, replay and updates can execute. The evaluation outcome is not evidence of game-playing competence, regardless of the successful runtime.

This short run checks that collection, replay and updates can execute. The evaluation outcome is not evidence of game-playing competence, regardless of the successful runtime.
Figure 2. Information flow. Solid arrows carry observations, tensors or artifacts; other routes are explicitly labelled. Signal shapes, matrices and network icons are schematic, not measured samples or literal neuron counts. Open full-size SVG ↗ On narrow screens, scroll the diagram horizontally.

Read this alongside Figure 1: The archived nine-second execution and −21 greedy return describe a plumbing test, not a successfully trained Pong agent. The module map and layer-level figures below expand the operations in this route.

A nine-second RL run can validate plumbing, not competence: system and evaluation mapPong observations: Preprocessed stacked frames → CNN Q network: Blockwise reverse updates → 3,000 training steps: MPS smoke run → Greedy evaluation: No exploration → Diagnostics: Return + action distribution. A high-level module map; comparison branches and training details are explained in the article.EVALUATION PRACTICE / E36 / MODULE MAP01 INPUTPong observationsPreprocessed stacked frames02 MODULECNN Q networkBlockwise reverse updates03 MODULE3,000 training stepsMPS smoke run04 MODULEGreedy evaluationNo exploration05 OUTPUTDiagnosticsReturn + action distribution
Source-grounded module map. Boxes summarize operations, not individual neurons; comparison arms and training paths are detailed below. On a small screen, scroll the diagram horizontally.
Pong observations — Preprocessed stacked frames

The architecture in context

What this comparison asks

Short reinforcement-learning runs are valuable for catching shape errors, device problems and broken replay. They are poor evidence of capability. The saved Pong run records 3,000 steps, roughly nine seconds of runtime and a final greedy return of −21. The recorded random reference is −20.7.

Who does what in the stack

Gymnasium / ALE
Supplies the Pong environment and episode semantics.
PyTorch MPS
Executed the archived smoke training loop.
Greedy evaluator
Measures policy behavior without epsilon exploration.

The companion evaluator switches to greedy actions and aggregates episode returns. The primary artifact is a run receipt rather than another model implementation. It supports the narrow conclusion that the loop executed and produced evaluable parameters.

Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph.
Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph. Open full-size SVG ↗

From module map to executable structure

Inside Atari Q-network with explicit reverse VJPs

This results or evaluation article shares the implementation in E35. The architecture below describes that companion, not a newly trained model.

Stack of four 84×84 grayscale frames; A environment actions; online and target networks.

Layer-level implementation. B denotes batch size; parameter and shape conventions are expanded in the table.
Layer-level implementation. B denotes batch size; parameter and shape conventions are expanded in the table. Open full-size SVG ↗
Layer / tensor / operation ledger
Layer or branchOutput shapeImplementation detail
Conv 4→32, 8×8 / stride 4B × 32 × 20 × 20Valid convolution → ReLU.
Conv 32→64, 4×4 / stride 2B × 64 × 9 × 9Valid convolution → ReLU.
Conv 64→64, 3×3 / stride 1B × 64 × 7 × 7Valid convolution → ReLU.
Flatten + Linear 3136→512B × 512ReLU feature layer.
Action-value headB × ALinear 512→A; gather only sampled action for TD error.

The equation shows the Double-DQN branch; the vanilla branch takes the target network’s own maximum. Target construction is under no_grad. Each online block builds a local graph. The reverse loop calls autograd.grad with the next block’s input cotangent, assigns parameter gradients and finally steps Adam. This still implements a reverse chain rule using PyTorch autograd; it is not derivative-free or independent local learning.

The equation and the update

y=Rn+γn(1−d)Qθˉ(s+,arg⁡max⁡aQθ(s+,a));ga=2Bclip⁡(Qθ(s,a)−y,−1,1)y=R_n+\gamma^n(1-d)Q_{\bar\theta}(s^+,\arg\max_a Q_\theta(s^+,a));\quad g_a=\frac{2}{B}\operatorname{clip}(Q_\theta(s,a)-y,-1,1)

Defaults 400,000 environment steps, replay 50,000, learning starts 10,000, batch 32, Adam 1e-4, gamma .99, target sync 2,000, global gradient-norm cap 10. Smoke overrides:3,000 steps, replay 2,000, start 500, sync 500. Double-DQN and n-step return are selectable flags, not assumptions about every run.

Learning or solution path. A parameter-update path is different from the forward inference path; see text for target-network, frozen-feature and local-loss boundaries.
Learning or solution path. A parameter-update path is different from the forward inference path; see text for target-network, frozen-feature and local-loss boundaries. Open full-size SVG ↗

Implementation card / no invented benchmarks

Capacity, budget and execution evidence

Parameters / retained state
1,684,128 + 513A scalars per Q-network; target doubles parameter storage. At A=6:1,687,206 per network.
Duration and hardware evidence
E36 records a 3,000-step MPS smoke run at 9.0 seconds and final return−21. It does not establish convergence or practical gameplay.
Source coordinates
E35 lines 65–108 and 140–160; E36 smoke metrics
Current reproduction context
Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.

Counts above are calculated from the stated layer shapes unless identified as saved measurements. They exclude optimizer state and nontrainable buffers. No archived training was rerun for this revision.

What these design choices change

Strided convolutions reduce the 84×84 image to 7×7 before the large dense layer. That dense layer alone has 1,606,144 parameters, so shrinking it is a materially different capacity choice. Clipping TD residuals and clipping global gradient norm address different instabilities; the factor 2 in the manual seed differs from the usual unit-threshold Huber derivative.

Reproduction and measurement protocol

First compare the manual sweep’s parameter gradients to a monolithic loss with the same scaling, on a fixed tiny replay batch and no optimizer step. Then verify that terminal transitions cannot bootstrap. Count environment steps, replay updates and evaluated episodes separately. A nine-second plumbing test is not a nine-second trained Atari agent.

For a new run, save the resolved Python/framework versions, backend, dtype, seed, input shapes, batch size and exact source revision. Start with one batch and one update. Log training steps separately from epochs or environment steps. Do not equate the configured maximum with a completed budget or convergence.

Measure initialization/compilation, data preparation, warmed forward pass, training updates and evaluation separately. Synchronize accelerator work around timed regions using the chosen framework’s supported mechanism. Report peak process memory and framework allocation separately; parameter bytes exclude activations, gradients, optimizer state and input buffers. On a shared machine, begin with a single CPU worker and a small batch rather than claiming all available resources.

A closer look at the implementation

The code that carries the idea

The snippet runs the current Q network under no_grad and chooses argmax actions. The saved intermediate action distribution is concentrated entirely on one action. That diagnostic helps explain why merely completing optimizer steps does not imply that the policy has discovered useful behavior.

Python · file · lines 127–137
def eval_greedy(env, units, n_act, n_ep, max_steps=10000):
    scores = []; acts = np.zeros(n_act)
    for ep in range(n_ep):
        obs, _ = env.reset(seed=9000 + ep); tot = 0.0
        for _ in range(max_steps):
            x = torch.tensor(np.asarray(obs)[None], dtype=torch.float32, device=DEV) / 255.0
            a = int(q_forward(units, x).argmax(1)); acts[a] += 1
            obs, r, term, trunc, _ = env.step(a); tot += r
            if term or trunc: break
        scores.append(tot)
    return float(np.mean(scores)), (acts / acts.sum()).round(2)

Verbatim archive excerpt from closed_form_neat_dqn_atari.py (companion source E35). Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.

The saved comparison

Archived results, not new training. The article states the comparison’s scope and limitations.

A nine-second RL run can validate plumbing, not competence — selected recorded values
Recorded itemValue
Training steps3,000
Final greedy return−21.0
Random reference−20.7
Saved runtime9.0 seconds

The boundary that matters

The runtime is not a benchmark against standard DQN, and the short horizon is not a serious sample-efficiency comparison. A smoke failure is not an impossibility result; it is a reason to avoid publishing a capability claim from this run.

Keep building

Other posts of interest