The architecture in context
What this comparison asks
The strongest comparison in this record is between the task-wise Gram learner and the joint fixed-feature fit. Both optimize the same ridge objective over the same representation, so agreement is a useful implementation check. The saved accuracies match at 0.8457, 0.8458 and 0.8458 across three seeds.
Who does what in the stack
- Saved JSON metrics
- Preserves seed-level results for all recorded arms.
- PyTorch linear heads
- Implements the sequential optimization comparator.
- Joint ridge control
- Tests the exact fixed-feature objective being reconstructed.
The experiment adds nearest-class-mean and sequential backpropagation heads to put the result in context. Those controls use different head structures and optimization procedures, so their gap should not be described as an isolated advantage of update order alone. The article links to the implementation instead of inventing a separate model.
From module map to executable structure
Inside Frozen encoder with additive Gram readout
This results or evaluation article shares the implementation in E33. The architecture below describes that companion, not a newly trained model.
Cached DINOv 2 features → R random Fourier features plus bias; default R=10,000, ten class-incremental tasks.
| Layer or branch | Output shape | Implementation detail |
|---|---|---|
| Frozen visual encoder | N × d | DINOv 2 ViT-S/14 feature cache; encoder remains necessary to process a new image. |
| Standardize + RFF | N × (R+1) | [1, sqrt(2/R) cos(FΩ+b)]; transformation fixed across tasks. |
| Accumulate G and B | G: (R+1)²; B: (R+1)×C | Add ΦᵀΦ and ΦᵀY for each task. |
| Regularized solve | W: (R+1) × C | Solve (G+λI)W=B, rather than invert explicitly. |
| Inference | N × C | Compute the same fixed features and multiply by W. |
The additive objects are sufficient statistics for one fixed quadratic objective. Associativity removes task-order dependence of that objective, up to floating-point effects. It does not freeze predictions after adding a new task: W is solved again and can change everywhere. A growing or retrained encoder invalidates the old statistics unless their cross-products are reconstructed.
The equation and the update
No SGD epochs for the readout solve. Default ridge λ=10; additive statistics are order-invariant in exact arithmetic for a fixed feature map. Feature standardization uses the full training pool, so this is not a strictly online unknown-future preprocessing protocol.
Implementation card / no invented benchmarks
Capacity, budget and execution evidence
- Parameters / retained state
- Readout has(R+1)C coefficients; G has(R+1)² entries. For R=10,000, float 32 G alone is 400,080,004 bytes. Frozen encoder weights are additional.
- Duration and hardware evidence
- Saved accuracies support the particular comparison, not a measured edge-device latency or memory-free training claim.
- Source coordinates
- E33 feature transformation, accumulator and ridge solve; E34 saved metrics
- Current reproduction context
- Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.
Counts above are calculated from the stated layer shapes unless identified as saved measurements. They exclude optimizer state and nontrainable buffers. No archived training was rerun for this revision.
What these design choices change
R controls approximation capacity but also quadratic Gram storage and solve cost. Increasing R from 4,000 to 10,000 multiplies G storage by roughly 6.25, even though feature width increases only 2.5 times. A small trainable readout is therefore not the same as a small total-memory training system.
Reproduction and measurement protocol
Compare accumulated G/B with a single concatenated-batch computation in float 64 on a tiny example. Then reverse task order and report coefficient and prediction tolerances. Keep encoder, normalization, RFF seed and ridge convention in the checkpoint; weights alone are not a complete inference artifact.
For a new run, save the resolved Python/framework versions, backend, dtype, seed, input shapes, batch size and exact source revision. Start with one batch and one update. Log training steps separately from epochs or environment steps. Do not equate the configured maximum with a completed budget or convergence.
Measure initialization/compilation, data preparation, warmed forward pass, training updates and evaluation separately. Synchronize accelerator work around timed regions using the chosen framework’s supported mechanism. Report peak process memory and framework allocation separately; parameter bytes exclude activations, gradients, optimizer state and input buffers. On a shared machine, begin with a single CPU worker and a small batch rather than claiming all available resources.
A closer look at the implementation
The code that carries the idea
The excerpt shows the sequential backpropagation control and its final evaluation. The metrics also retain a zero observed task-order standard deviation in this run. That numerical observation is narrower than a proof of bitwise invariance for every accumulation order.
m = torch.isin(yo, torch.tensor(cls)); Xc, yc = ftr[m], yo[m]
for _ in range(a.epochs):
bi = torch.randint(0, len(Xc), (256,)); opt.zero_grad(); loss = lf(head(Xc[bi]), yc[bi]); loss.backward(); opt.step()
with torch.no_grad():
res["backprop"].append(float((head(fte).argmax(1) == yeo).float().mean()))
# data-efficiency (gram vs backprop) at few shots/class
for n in [5, 10, 25, 100]:
keep = torch.cat([torch.where(yo == c)[0][torch.randperm(len(torch.where(yo == c)[0]), generator=g)[:n]] for c in range(nclass)])
deff["gram"][n].append(gram_cil_acc(Ztr[keep], yo[keep], Zte, yeo, tasks, nclass, a.lam))
torch.manual_seed(7 + s); h2 = nn.Linear(ftr.shape[1], nclass); o2 = torch.optim.Adam(h2.parameters(), 1e-3)
for cls in tasks:
m = torch.isin(yo[keep], torch.tensor(cls)); Xc, yc = ftr[keep][m], yo[keep][m]
for _ in range(a.epochs):
bi = torch.randint(0, len(Xc), (min(256, len(Xc)),)); o2.zero_grad(); loss = lf(h2(Xc[bi]), yc[bi]); loss.backward(); o2.step()
with torch.no_grad():
deff["backprop"][n].append(float((h2(fte).argmax(1) == yeo).float().mean()))
# order-invariance of Gram (4 random task orders)Verbatim archive excerpt from closed_form_distill_frontier.py (companion source E33). Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.
The saved comparison
Archived results, not new training. The article states the comparison’s scope and limitations.
| Arm | Run 1 | Run 2 | Run 3 |
|---|---|---|---|
| Task-wise / joint Gram | 0.8457 | 0.8458 | 0.8458 |
| Sequential backprop head | 0.5309 | 0.5526 | 0.5369 |
The boundary that matters
The encoder is pretrained and preprocessing sees the full training feature set. Final class-incremental accuracy is not a guarantee that performance on each old class never worsened during updates. The correction in E46 explains the distinction with a scalar example.