Evaluation practice · E34 · Evaluation

Match the joint readout without claiming universal zero forgetting

The saved DINOv2-feature experiment supports a precise result: additive statistics recover the same fixed-feature fit. It does not freeze every earlier prediction.

Saved JSON metricsPyTorch linear headsJoint ridge control
Across three saved runs, task-wise and joint Gram readouts agree at the final objective; this is not a guarantee about every earlier prediction.
Figure 1. Recover the joint readout. Across three saved runs, task-wise and joint Gram readouts agree at the final objective; this is not a guarantee about every earlier prediction. Archived accuracies. Original vector illustration.

Follow the information

From input to outcome

The task-wise and joint Gram paths are independent arms fitting the same final fixed-feature objective. Their agreement does not imply invariant intermediate predictions.

The task-wise and joint Gram paths are independent arms fitting the same final fixed-feature objective. Their agreement does not imply invariant intermediate predictions.
Figure 2. Information flow. Solid arrows carry observations, tensors or artifacts; other routes are explicitly labelled. Signal shapes, matrices and network icons are schematic, not measured samples or literal neuron counts. Open full-size SVG ↗ On narrow screens, scroll the diagram horizontally.

Read this alongside Figure 1: Across three saved runs, task-wise and joint Gram readouts agree at the final objective; this is not a guarantee about every earlier prediction. The module map and layer-level figures below expand the operations in this route.

Match the joint readout without claiming universal zero forgetting: system and evaluation mapSame cached features: Shared preprocessing → Task-wise Gram arm: Accumulate class batches → Joint Gram control: Fit all classes together → Other head controls: Nearest mean / backprop → Final accuracy: Three saved seeds. A high-level module map; comparison branches and training details are explained in the article.EVALUATION PRACTICE / E34 / MODULE MAP01 INPUTSame cached featuresShared preprocessing02 MODULETask-wise Gram armAccumulate class batches03 MODULEJoint Gram controlFit all classes together04 MODULEOther head controlsNearest mean / backprop05 OUTPUTFinal accuracyThree saved seeds
Source-grounded module map. Boxes summarize operations, not individual neurons; comparison arms and training paths are detailed below. On a small screen, scroll the diagram horizontally.
Same cached features — Shared preprocessing

The architecture in context

What this comparison asks

The strongest comparison in this record is between the task-wise Gram learner and the joint fixed-feature fit. Both optimize the same ridge objective over the same representation, so agreement is a useful implementation check. The saved accuracies match at 0.8457, 0.8458 and 0.8458 across three seeds.

Who does what in the stack

Saved JSON metrics
Preserves seed-level results for all recorded arms.
PyTorch linear heads
Implements the sequential optimization comparator.
Joint ridge control
Tests the exact fixed-feature objective being reconstructed.

The experiment adds nearest-class-mean and sequential backpropagation heads to put the result in context. Those controls use different head structures and optimization procedures, so their gap should not be described as an isolated advantage of update order alone. The article links to the implementation instead of inventing a separate model.

Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph.
Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph. Open full-size SVG ↗

From module map to executable structure

Inside Frozen encoder with additive Gram readout

This results or evaluation article shares the implementation in E33. The architecture below describes that companion, not a newly trained model.

Cached DINOv 2 features → R random Fourier features plus bias; default R=10,000, ten class-incremental tasks.

Layer-level implementation. B denotes batch size; parameter and shape conventions are expanded in the table.
Layer-level implementation. B denotes batch size; parameter and shape conventions are expanded in the table. Open full-size SVG ↗
Layer / tensor / operation ledger
Layer or branchOutput shapeImplementation detail
Frozen visual encoderN × dDINOv 2 ViT-S/14 feature cache; encoder remains necessary to process a new image.
Standardize + RFFN × (R+1)[1, sqrt(2/R) cos(FΩ+b)]; transformation fixed across tasks.
Accumulate G and BG: (R+1)²; B: (R+1)×CAdd ΦᵀΦ and ΦᵀY for each task.
Regularized solveW: (R+1) × CSolve (G+λI)W=B, rather than invert explicitly.
InferenceN × CCompute the same fixed features and multiply by W.

The additive objects are sufficient statistics for one fixed quadratic objective. Associativity removes task-order dependence of that objective, up to floating-point effects. It does not freeze predictions after adding a new task: W is solved again and can change everywhere. A growing or retrained encoder invalidates the old statistics unless their cross-products are reconstructed.

The equation and the update

G=∑tΦtTΦt,B=∑tΦtTYt,(G+λI)W=BG=\sum_t\Phi_t^T\Phi_t,\quad B=\sum_t\Phi_t^TY_t,\quad (G+\lambda I)W=B

No SGD epochs for the readout solve. Default ridge λ=10; additive statistics are order-invariant in exact arithmetic for a fixed feature map. Feature standardization uses the full training pool, so this is not a strictly online unknown-future preprocessing protocol.

Learning or solution path. A parameter-update path is different from the forward inference path; see text for target-network, frozen-feature and local-loss boundaries.
Learning or solution path. A parameter-update path is different from the forward inference path; see text for target-network, frozen-feature and local-loss boundaries. Open full-size SVG ↗

Implementation card / no invented benchmarks

Capacity, budget and execution evidence

Parameters / retained state
Readout has(R+1)C coefficients; G has(R+1)² entries. For R=10,000, float 32 G alone is 400,080,004 bytes. Frozen encoder weights are additional.
Duration and hardware evidence
Saved accuracies support the particular comparison, not a measured edge-device latency or memory-free training claim.
Source coordinates
E33 feature transformation, accumulator and ridge solve; E34 saved metrics
Current reproduction context
Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.

Counts above are calculated from the stated layer shapes unless identified as saved measurements. They exclude optimizer state and nontrainable buffers. No archived training was rerun for this revision.

What these design choices change

R controls approximation capacity but also quadratic Gram storage and solve cost. Increasing R from 4,000 to 10,000 multiplies G storage by roughly 6.25, even though feature width increases only 2.5 times. A small trainable readout is therefore not the same as a small total-memory training system.

Reproduction and measurement protocol

Compare accumulated G/B with a single concatenated-batch computation in float 64 on a tiny example. Then reverse task order and report coefficient and prediction tolerances. Keep encoder, normalization, RFF seed and ridge convention in the checkpoint; weights alone are not a complete inference artifact.

For a new run, save the resolved Python/framework versions, backend, dtype, seed, input shapes, batch size and exact source revision. Start with one batch and one update. Log training steps separately from epochs or environment steps. Do not equate the configured maximum with a completed budget or convergence.

Measure initialization/compilation, data preparation, warmed forward pass, training updates and evaluation separately. Synchronize accelerator work around timed regions using the chosen framework’s supported mechanism. Report peak process memory and framework allocation separately; parameter bytes exclude activations, gradients, optimizer state and input buffers. On a shared machine, begin with a single CPU worker and a small batch rather than claiming all available resources.

A closer look at the implementation

The code that carries the idea

The excerpt shows the sequential backpropagation control and its final evaluation. The metrics also retain a zero observed task-order standard deviation in this run. That numerical observation is narrower than a proof of bitwise invariance for every accumulation order.

Python · file · lines 66–82
            m = torch.isin(yo, torch.tensor(cls)); Xc, yc = ftr[m], yo[m]
            for _ in range(a.epochs):
                bi = torch.randint(0, len(Xc), (256,)); opt.zero_grad(); loss = lf(head(Xc[bi]), yc[bi]); loss.backward(); opt.step()
        with torch.no_grad():
            res["backprop"].append(float((head(fte).argmax(1) == yeo).float().mean()))
        # data-efficiency (gram vs backprop) at few shots/class
        for n in [5, 10, 25, 100]:
            keep = torch.cat([torch.where(yo == c)[0][torch.randperm(len(torch.where(yo == c)[0]), generator=g)[:n]] for c in range(nclass)])
            deff["gram"][n].append(gram_cil_acc(Ztr[keep], yo[keep], Zte, yeo, tasks, nclass, a.lam))
            torch.manual_seed(7 + s); h2 = nn.Linear(ftr.shape[1], nclass); o2 = torch.optim.Adam(h2.parameters(), 1e-3)
            for cls in tasks:
                m = torch.isin(yo[keep], torch.tensor(cls)); Xc, yc = ftr[keep][m], yo[keep][m]
                for _ in range(a.epochs):
                    bi = torch.randint(0, len(Xc), (min(256, len(Xc)),)); o2.zero_grad(); loss = lf(h2(Xc[bi]), yc[bi]); loss.backward(); o2.step()
            with torch.no_grad():
                deff["backprop"][n].append(float((h2(fte).argmax(1) == yeo).float().mean()))
        # order-invariance of Gram (4 random task orders)

Verbatim archive excerpt from closed_form_distill_frontier.py (companion source E33). Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.

The saved comparison

Archived results, not new training. The article states the comparison’s scope and limitations.

Match the joint readout without claiming universal zero forgetting — selected recorded values
ArmRun 1Run 2Run 3
Task-wise / joint Gram0.84570.84580.8458
Sequential backprop head0.53090.55260.5369

The boundary that matters

The encoder is pretrained and preprocessing sees the full training feature set. Final class-incremental accuracy is not a guarantee that performance on each old class never worsened during updates. The correction in E46 explains the distinction with a scalar example.

Keep building

Other posts of interest