The architecture in context
What this comparison asks
Short reinforcement-learning runs are valuable for catching shape errors, device problems and broken replay. They are poor evidence of capability. The saved Pong run records 3,000 steps, roughly nine seconds of runtime and a final greedy return of −21. The recorded random reference is −20.7.
Who does what in the stack
- Gymnasium / ALE
- Supplies the Pong environment and episode semantics.
- PyTorch MPS
- Executed the archived smoke training loop.
- Greedy evaluator
- Measures policy behavior without epsilon exploration.
The companion evaluator switches to greedy actions and aggregates episode returns. The primary artifact is a run receipt rather than another model implementation. It supports the narrow conclusion that the loop executed and produced evaluable parameters.
From module map to executable structure
Inside Atari Q-network with explicit reverse VJPs
This results or evaluation article shares the implementation in E35. The architecture below describes that companion, not a newly trained model.
Stack of four 84×84 grayscale frames; A environment actions; online and target networks.
| Layer or branch | Output shape | Implementation detail |
|---|---|---|
| Conv 4→32, 8×8 / stride 4 | B × 32 × 20 × 20 | Valid convolution → ReLU. |
| Conv 32→64, 4×4 / stride 2 | B × 64 × 9 × 9 | Valid convolution → ReLU. |
| Conv 64→64, 3×3 / stride 1 | B × 64 × 7 × 7 | Valid convolution → ReLU. |
| Flatten + Linear 3136→512 | B × 512 | ReLU feature layer. |
| Action-value head | B × A | Linear 512→A; gather only sampled action for TD error. |
The equation shows the Double-DQN branch; the vanilla branch takes the target network’s own maximum. Target construction is under no_grad. Each online block builds a local graph. The reverse loop calls autograd.grad with the next block’s input cotangent, assigns parameter gradients and finally steps Adam. This still implements a reverse chain rule using PyTorch autograd; it is not derivative-free or independent local learning.
The equation and the update
Defaults 400,000 environment steps, replay 50,000, learning starts 10,000, batch 32, Adam 1e-4, gamma .99, target sync 2,000, global gradient-norm cap 10. Smoke overrides:3,000 steps, replay 2,000, start 500, sync 500. Double-DQN and n-step return are selectable flags, not assumptions about every run.
Implementation card / no invented benchmarks
Capacity, budget and execution evidence
- Parameters / retained state
- 1,684,128 + 513A scalars per Q-network; target doubles parameter storage. At A=6:1,687,206 per network.
- Duration and hardware evidence
- E36 records a 3,000-step MPS smoke run at 9.0 seconds and final return−21. It does not establish convergence or practical gameplay.
- Source coordinates
- E35 lines 65–108 and 140–160; E36 smoke metrics
- Current reproduction context
- Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.
Counts above are calculated from the stated layer shapes unless identified as saved measurements. They exclude optimizer state and nontrainable buffers. No archived training was rerun for this revision.
What these design choices change
Strided convolutions reduce the 84×84 image to 7×7 before the large dense layer. That dense layer alone has 1,606,144 parameters, so shrinking it is a materially different capacity choice. Clipping TD residuals and clipping global gradient norm address different instabilities; the factor 2 in the manual seed differs from the usual unit-threshold Huber derivative.
Reproduction and measurement protocol
First compare the manual sweep’s parameter gradients to a monolithic loss with the same scaling, on a fixed tiny replay batch and no optimizer step. Then verify that terminal transitions cannot bootstrap. Count environment steps, replay updates and evaluated episodes separately. A nine-second plumbing test is not a nine-second trained Atari agent.
For a new run, save the resolved Python/framework versions, backend, dtype, seed, input shapes, batch size and exact source revision. Start with one batch and one update. Log training steps separately from epochs or environment steps. Do not equate the configured maximum with a completed budget or convergence.
Measure initialization/compilation, data preparation, warmed forward pass, training updates and evaluation separately. Synchronize accelerator work around timed regions using the chosen framework’s supported mechanism. Report peak process memory and framework allocation separately; parameter bytes exclude activations, gradients, optimizer state and input buffers. On a shared machine, begin with a single CPU worker and a small batch rather than claiming all available resources.
A closer look at the implementation
The code that carries the idea
The snippet runs the current Q network under no_grad and chooses argmax actions. The saved intermediate action distribution is concentrated entirely on one action. That diagnostic helps explain why merely completing optimizer steps does not imply that the policy has discovered useful behavior.
def eval_greedy(env, units, n_act, n_ep, max_steps=10000):
scores = []; acts = np.zeros(n_act)
for ep in range(n_ep):
obs, _ = env.reset(seed=9000 + ep); tot = 0.0
for _ in range(max_steps):
x = torch.tensor(np.asarray(obs)[None], dtype=torch.float32, device=DEV) / 255.0
a = int(q_forward(units, x).argmax(1)); acts[a] += 1
obs, r, term, trunc, _ = env.step(a); tot += r
if term or trunc: break
scores.append(tot)
return float(np.mean(scores)), (acts / acts.sum()).round(2)Verbatim archive excerpt from closed_form_neat_dqn_atari.py (companion source E35). Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.
The saved comparison
Archived results, not new training. The article states the comparison’s scope and limitations.
| Recorded item | Value |
|---|---|
| Training steps | 3,000 |
| Final greedy return | −21.0 |
| Random reference | −20.7 |
| Saved runtime | 9.0 seconds |
The boundary that matters
The runtime is not a benchmark against standard DQN, and the short horizon is not a serious sample-efficiency comparison. A smoke failure is not an impossibility result; it is a reason to avoid publishing a capability claim from this run.