Representation learning · E33 · Implementation

Add classes by accumulating statistics over frozen features

A pretrained representation, random Fourier features and a ridge solve form a simple continual readout. Its guarantees stop at the fixed feature map.

Pretrained DINOv2 featuresPyTorch random featurestorch.linalg.solve
With fixed features, a quadratic objective can be accumulated through additive statistics and solved for the readout.
Figure 1. Memory as an additive quadratic. With fixed features, a quadratic objective can be accumulated through additive statistics and solved for the readout. Exact fixed-feature objective. Original vector illustration.

Follow the information

From input to outcome

Class-batch labels contribute to the cross-statistic B; G describes fixed features. Later queries use the same fixed map and the current W. Addition is order-independent, but the optimum can change old predictions.

Class-batch labels contribute to the cross-statistic B; G describes fixed features. Later queries use the same fixed map and the current W. Addition is order-independent, but the optimum can change old predictions.
Figure 2. Information flow. Solid arrows carry observations, tensors or artifacts; other routes are explicitly labelled. Signal shapes, matrices and network icons are schematic, not measured samples or literal neuron counts. Open full-size SVG ↗ On narrow screens, scroll the diagram horizontally.

Read this alongside Figure 1: With fixed features, a quadratic objective can be accumulated through additive statistics and solved for the readout. The module map and layer-level figures below expand the operations in this route.

Add classes by accumulating statistics over frozen features: architectureCached DINOv2 features: Training-standardized → Fixed Fourier map: Bias + cosine features → Class batches: One-hot targets → Sufficient statistics: G += ZᵀZ · B += ZᵀY → Ridge solve: (G + λI) W = B → Prediction: argmax(ZW). A high-level module map; comparison branches and training details are explained in the article.REPRESENTATION LEARNING / E33 / MODULE MAP01 INPUTCached DINOv2 featuresTraining-standardized02 MODULEFixed Fourier mapBias + cosine features03 MODULEClass batchesOne-hot targets04 MODULESufficient statisticsG += ZᵀZ · B += ZᵀY05 MODULERidge solve(G + λI) W = B06 OUTPUTPredictionargmax(ZW)
Source-grounded module map. Boxes summarize operations, not individual neurons; comparison arms and training paths are detailed below. On a small screen, scroll the diagram horizontally.
Cached DINOv2 features — Training-standardized

The architecture in context

The system we are building

Once the feature map is fixed, ridge regression can be reconstructed from additive statistics rather than all individual examples. The experiment expands cached visual features with a random Fourier map, accumulates a Gram matrix and label cross-products across class batches, then solves a regularized linear system. The accumulation is equivalent to the pooled fixed-feature objective in exact arithmetic.

Who does what in the stack

Pretrained DINOv2 features
Supply a fixed visual representation.
PyTorch random features
Build a fixed cosine expansion.
torch.linalg.solve
Fits the pooled regularized readout.

The project combines a pretrained visual encoder with a task-wise readout update and explicit joint-fit controls. It is not teacher-weight distillation: the encoder remains a required feature generator for new images. The custom work is the representation adapter and the experiment around additive learning.

Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph.
Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph. Open full-size SVG ↗

From module map to executable structure

Inside Frozen encoder with additive Gram readout

Cached DINOv 2 features → R random Fourier features plus bias; default R=10,000, ten class-incremental tasks.

Layer-level implementation. B denotes batch size; parameter and shape conventions are expanded in the table.
Layer-level implementation. B denotes batch size; parameter and shape conventions are expanded in the table. Open full-size SVG ↗
Layer / tensor / operation ledger
Layer or branchOutput shapeImplementation detail
Frozen visual encoderN × dDINOv 2 ViT-S/14 feature cache; encoder remains necessary to process a new image.
Standardize + RFFN × (R+1)[1, sqrt(2/R) cos(FΩ+b)]; transformation fixed across tasks.
Accumulate G and BG: (R+1)²; B: (R+1)×CAdd ΦᵀΦ and ΦᵀY for each task.
Regularized solveW: (R+1) × CSolve (G+λI)W=B, rather than invert explicitly.
InferenceN × CCompute the same fixed features and multiply by W.

The additive objects are sufficient statistics for one fixed quadratic objective. Associativity removes task-order dependence of that objective, up to floating-point effects. It does not freeze predictions after adding a new task: W is solved again and can change everywhere. A growing or retrained encoder invalidates the old statistics unless their cross-products are reconstructed.

The equation and the update

G=∑tΦtTΦt,B=∑tΦtTYt,(G+λI)W=BG=\sum_t\Phi_t^T\Phi_t,\quad B=\sum_t\Phi_t^TY_t,\quad (G+\lambda I)W=B

No SGD epochs for the readout solve. Default ridge λ=10; additive statistics are order-invariant in exact arithmetic for a fixed feature map. Feature standardization uses the full training pool, so this is not a strictly online unknown-future preprocessing protocol.

Learning or solution path. A parameter-update path is different from the forward inference path; see text for target-network, frozen-feature and local-loss boundaries.
Learning or solution path. A parameter-update path is different from the forward inference path; see text for target-network, frozen-feature and local-loss boundaries. Open full-size SVG ↗

Implementation card / no invented benchmarks

Capacity, budget and execution evidence

Parameters / retained state
Readout has(R+1)C coefficients; G has(R+1)² entries. For R=10,000, float 32 G alone is 400,080,004 bytes. Frozen encoder weights are additional.
Duration and hardware evidence
Saved accuracies support the particular comparison, not a measured edge-device latency or memory-free training claim.
Source coordinates
E33 feature transformation, accumulator and ridge solve; E34 saved metrics
Current reproduction context
Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.

Counts above are calculated from the stated layer shapes unless identified as saved measurements. They exclude optimizer state and nontrainable buffers. No archived training was rerun for this revision.

What these design choices change

R controls approximation capacity but also quadratic Gram storage and solve cost. Increasing R from 4,000 to 10,000 multiplies G storage by roughly 6.25, even though feature width increases only 2.5 times. A small trainable readout is therefore not the same as a small total-memory training system.

Reproduction and measurement protocol

Compare accumulated G/B with a single concatenated-batch computation in float 64 on a tiny example. Then reverse task order and report coefficient and prediction tolerances. Keep encoder, normalization, RFF seed and ridge convention in the checkpoint; weights alone are not a complete inference artifact.

For a new run, save the resolved Python/framework versions, backend, dtype, seed, input shapes, batch size and exact source revision. Start with one batch and one update. Log training steps separately from epochs or environment steps. Do not equate the configured maximum with a completed budget or convergence.

Measure initialization/compilation, data preparation, warmed forward pass, training updates and evaluation separately. Synchronize accelerator work around timed regions using the chosen framework’s supported mechanism. Report peak process memory and framework allocation separately; parameter bytes exclude activations, gradients, optimizer state and input buffers. On a shared machine, begin with a single CPU worker and a small batch rather than claiming all available resources.

A closer look at the implementation

The code that carries the idea

The excerpt updates G and B inside the task loop and solves only after accumulation. That preserves the information needed for the declared ridge objective, but not arbitrary future features or old-task predictions. With 10,000 Fourier coordinates plus a bias, a dense float32 Gram alone is about 400 MB.

Python · file · lines 29–35
def gram_cil_acc(Ztr, yo, Zte, yeo, tasks, nclass, lam):
    Fdim = Ztr.shape[1]; G = torch.zeros(Fdim, Fdim); B = torch.zeros(Fdim, nclass)
    for cls in tasks:
        m = torch.isin(yo, torch.tensor(cls)); Y = torch.zeros(int(m.sum()), nclass); Y[torch.arange(int(m.sum())), yo[m]] = 1
        G = G + Ztr[m].T @ Ztr[m]; B = B + Ztr[m].T @ Y
    W = torch.linalg.solve(G + lam * torch.eye(Fdim), B)
    return float(((Zte @ W).argmax(1) == yeo).float().mean())

Verbatim archive excerpt from closed_form_distill_frontier.py. Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.

The boundary that matters

Training-wide feature standardization is computed before the simulated task stream. This is a declared full-training preprocessing choice, not a strictly online normalization scheme. Floating-point sums are also not bitwise order-independent.

Keep building

Other posts of interest