Evaluation practice · E30 · Negative result

A local learning rule still has to beat random features

A saved CIFAR-10 comparison turns a disappointing result into a useful engineering decision: isolate the representation before expanding the training recipe.

PyTorchRidge solveJSON result record
The saved comparison includes raw inputs and random features, not only trained models. Here the local recipe did not establish an advantage.
Figure 1. A local loss still needs a control. The saved comparison includes raw inputs and random features, not only trained models. Here the local recipe did not establish an advantage. Archived CIFAR-10 accuracies. Original vector illustration.

Follow the information

From input to outcome

This diagram is a comparison, not a four-stage representation model. Raw inputs, random features and locally learned features are separate arms; the supervised control uses its declared trained classifier.

This diagram is a comparison, not a four-stage representation model. Raw inputs, random features and locally learned features are separate arms; the supervised control uses its declared trained classifier.
Figure 2. Information flow. Solid arrows carry observations, tensors or artifacts; other routes are explicitly labelled. Signal shapes, matrices and network icons are schematic, not measured samples or literal neuron counts. Open full-size SVG ↗ On narrow screens, scroll the diagram horizontally.

Read this alongside Figure 1: The saved comparison includes raw inputs and random features, not only trained models. Here the local recipe did not establish an advantage. The module map and layer-level figures below expand the operations in this route.

A local learning rule still has to beat random features: system and evaluation mapCIFAR-10 images: Same task across arms → Representation arms: Raw / random / local / supervised → Frozen feature probes: Ridge where applicable → Accuracy comparison: Common evaluation labels → Decision: Local recipe not established. A high-level module map; comparison branches and training details are explained in the article.EVALUATION PRACTICE / E30 / MODULE MAP01 INPUTCIFAR-10 imagesSame task across arms02 MODULERepresentation armsRaw / random / local / supervised03 MODULEFrozen feature probesRidge where applicable04 MODULEAccuracy comparisonCommon evaluation labels05 OUTPUTDecisionLocal recipe not established
Source-grounded module map. Boxes summarize operations, not individual neurons; comparison arms and training paths are detailed below. On a small screen, scroll the diagram horizontally.
CIFAR-10 images — Same task across arms

The architecture in context

What this comparison asks

The blockwise contrastive experiment produced a representation, but the relevant question is whether it improved the downstream task. The saved comparison includes raw features, a random convolutional representation and supervised training. That makes it possible to distinguish the value of the architecture from the value of the local learning rule.

Who does what in the stack

PyTorch
Produces the compared representations.
Ridge solve
Tests feature usefulness with a low-capacity head.
JSON result record
Retains the negative comparison and run context.

The evaluation code fits a ridge readout after training the representation, using training-derived feature normalization. The result file records the arm outputs and run context. This article is a case study in interpreting a failed learning recipe, not a second implementation of the blockwise network.

Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph.
Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph. Open full-size SVG ↗

From module map to executable structure

Inside Greedy local contrastive CNN

This results or evaluation article shares the implementation in E29. The architecture below describes that companion, not a newly trained model.

CIFAR-10, three blocks with channels 32,64,128; two 3×3 convolutions per block.

Layer-level implementation. B denotes batch size; parameter and shape conventions are expanded in the table.
Layer-level implementation. B denotes batch size; parameter and shape conventions are expanded in the table. Open full-size SVG ↗
Layer / tensor / operation ledger
Layer or branchOutput shapeImplementation detail
Block 1: 3→32→32B × 32 × 16 × 16Each convolution → BatchNorm → ReLU; finish with average pool 2.
Block 2: 32→64→64B × 64 × 8 × 8Train after block 1 is frozen and in eval mode.
Block 3: 64→128→128B × 128 × 4 × 4Train after both preceding blocks are frozen.
Temporary local headB × 128At the current block: pool → Linear(c,256) → ReLU → Linear(256,128).
Frozen feature probeClass scoresLocal projection head is not the final ridge classifier.

This is a different learning rule from a blockwise reverse sweep. A local loss updates only the current block and its projection head. Previously trained blocks receive neither an input gradient nor an optimizer step. The NT-Xent matrix contains 2B views, excludes self-similarity and identifies the matching view by a half-batch offset.

The equation and the update

θb←θb−η∇θbLb(gb(fb(stopgrad⁡(hb−1))))\theta_b\leftarrow\theta_b-\eta\nabla_{\theta_b}\mathcal L_b(g_b(f_b(\operatorname{stopgrad}(h_{b-1}))))

Configured 80 epochs per block, batch 512, Adam 1e-3, temperature .2; subset 20,000. The saved comparison is negative versus random convolutional features. Do not reinterpret the training budget as evidence of a benefit.

Learning or solution path. A parameter-update path is different from the forward inference path; see text for target-network, frozen-feature and local-loss boundaries.
Learning or solution path. A parameter-update path is different from the forward inference path; see text for target-network, frozen-feature and local-loss boundaries. Open full-size SVG ↗

Implementation card / no invented benchmarks

Capacity, budget and execution evidence

Parameters / retained state
287,904 backbone scalars. Each temporary projection head adds 256(c+1)+128×257; heads are not all simultaneously trained.
Duration and hardware evidence
E30 records 326.8 seconds for its saved experiment on MPS; this is not a per-block or per-epoch timing and does not identify the current workstation.
Source coordinates
E29 lines 24–25 and 50–84
Current reproduction context
Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.

Counts above are calculated from the stated layer shapes unless identified as saved measurements. They exclude optimizer state and nontrainable buffers. No archived training was rerun for this revision.

What these design choices change

Greedy freezing can reduce simultaneously trainable state but may discard information that later blocks need. A useful local contrastive objective is not automatically a useful end-to-end representation. The supervised and random-feature controls are necessary to interpret the negative result.

Reproduction and measurement protocol

After one optimizer step, compare every earlier block’s parameters and BatchNorm buffers byte-for-byte. Check that only the current block and projection head move. A zero parameter gradient is insufficient if running statistics continue changing.

For a new run, save the resolved Python/framework versions, backend, dtype, seed, input shapes, batch size and exact source revision. Start with one batch and one update. Log training steps separately from epochs or environment steps. Do not equate the configured maximum with a completed budget or convergence.

Measure initialization/compilation, data preparation, warmed forward pass, training updates and evaluation separately. Synchronize accelerator work around timed regions using the chosen framework’s supported mechanism. Report peak process memory and framework allocation separately; parameter bytes exclude activations, gradients, optimizer state and input buffers. On a shared machine, begin with a single CPU worker and a small batch rather than claiming all available resources.

A closer look at the implementation

The code that carries the idea

The companion snippet fits the classifier in feature space. The archived local-learning accuracy is 0.4907, slightly below random convolutional features at 0.5010 and well below the supervised arm at 0.7376. Raw features reach 0.3896. These values describe the saved setting, not a full sweep of the method family.

Python · file · lines 98–104
def probe(Ftr, ytr, Fte, yte, lam=1e2):
    mu, sd = Ftr.mean(0), Ftr.std(0) + 1e-6; Ftr = (Ftr - mu) / sd; Fte = (Fte - mu) / sd
    Ftr = torch.cat([Ftr, torch.ones(len(Ftr), 1, device=DEV)], 1); Fte = torch.cat([Fte, torch.ones(len(Fte), 1, device=DEV)], 1)
    yoh = torch.eye(10, device=DEV)[ytr]
    W = torch.linalg.solve(Ftr.T @ Ftr + lam * torch.eye(Ftr.shape[1], device=DEV), Ftr.T @ yoh)
    return float((Fte @ W).argmax(1).eq(yte).float().mean())

Verbatim archive excerpt from closed_form_neat_biolocal.py (companion source E29). Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.

The saved comparison

Archived results, not new training. The article states the comparison’s scope and limitations.

A local learning rule still has to beat random features — selected recorded values
ArmSaved accuracy
Raw features0.3896
Random CNN0.5010
Local contrastive0.4907
Supervised0.7376

The boundary that matters

The result does not establish that local learning is impossible, but it does remove the basis for claiming that this particular recipe improved representation quality. More epochs or another projection head would be a new experiment, not a correction to the saved evidence.

Keep building

Other posts of interest