Continual learning · Research & Algorithms

Can a frozen encoder keep learning new classes?

A good visual representation can make continual learning look surprisingly simple. The important question is what the small updating component actually learns—and what the frozen encoder already knows.

EXPLORE THE IDEA

Expensive perception. Small adaptation.

Count what the frozen representation already contributes.

FROZEN REPRESENTATIONPerception does not start from zeroDINOv2 ViT-B featuresThe encoder stays fixed; only the readout changesARCHIVED CIFAR-100 PANELWhat does the readout add?88.47%ordinary linear ridge89.10%cosine-expanded ridge+0.63 percentage points in this panel
50%
Actual archived final CIFAR-100 accuracies: linear ridge 88.47%, cosine-expanded ridge 89.10%, on the named frozen DINOv2 features. The graphic reveals the incremental difference; it does not fabricate source images or a confusion matrix.

Follow the information

From input to outcome

Labels enter the sufficient statistics, not the frozen visual backbone. At inference, an image still passes through the encoder and fixed features before the fitted head scores its class; no statistics update is required for that query.

Scroll the diagram horizontally to follow the route. Keyboard: focus the diagram, then use the arrow keys.

Incoming image → Frozen DINOv2 → Fixed feature expansion → Accumulate G and b → Ridge class readout. Labels enter the sufficient statistics, not the frozen visual backbone. At inference, an image still passes through the encoder and fixed features before the fitted head scores its class; no statistics update is required for that query.
Information-flow map. Training-set-wide normalization is not strictly causal online preprocessing. Original vector schematic based on the method and evidence discussed in this article; signal shapes and icons are illustrative, not additional measurements. Open full-size diagram ↗

Read the main route from left to right; labelled side branches show additional inputs, checks or feedback. The sections below explain the operations and their experimental limits.

Let perception be expensive once

Imagine adding new object categories without retraining a visual backbone each time. A pretrained encoder turns images into features; a small decision layer associates those features with labels. This division is attractive when a good encoder is already available. It also changes how we should describe the system: inexpensive adaptation does not mean perception was acquired cheaply.

Follow an image all the way to its class

The expensive visual encoder remains in the prediction path. Its frozen feature vector is standardized, optionally expanded through random features, and multiplied by a fitted readout. Incoming supervision updates cross-products, not the backbone. The architecture is therefore a large fixed representation with a replaceable analytic head, not a small standalone vision model.

The ablations tell us where the accuracy resides. Raw-feature ridge already reaches 88.47%, compared with 89.10% for the cosine expansion. The much larger gap to sequential cross-entropy mixes representation, loss, and evidence retention. It cannot be credited solely to one special nonlinear activation.

Frozen perception, changing evidence
Frozen perception, changing evidence. Original scientific diagram; the stated component and information flow, not an additional experiment. Open full-size figure ↗

Keep the objective, not a replay buffer

For fixed features and a squared-error readout, every labeled batch contributes two matrices. Their sums define exactly the same ridge objective as all the feature rows together. A later solve can use the combined evidence without replaying individual examples. This is a classical analytic-learning mechanism, closely related to RanPAC, rather than a new consequence of calling the features a biological substrate.

G←G+ZTZ,B←B+ZTY,(G+λI)W=B\begin{gathered}G\leftarrow G+Z^TZ,\quad B\leftarrow B+Z^TY,\quad (G+\lambda I)W=B\end{gathered}
These statistics preserve the fixed-feature squared-error objective. They do not freeze every old prediction or permit arbitrary changes to the encoder.

The strong number has a revealing control

On the named frozen DINOv2 ViT-B feature panel, the cosine-expanded readout reaches 89.10% final CIFAR-100 accuracy. Linear ridge on those same frozen features already reaches 88.47%. The nonlinear expansion adds about 0.63 percentage points. A ReLU-plus-raw expansion nearly ties cosine. Most of the capability is already present in the representation and ordinary pooled fitting.

Complete three-entry readout ablation from the named archived panel. The inherited representation plus linear ridge accounts for most of the final accuracy; no new vision training was run.
Complete three-entry readout ablation from the named archived panel. The inherited representation plus linear ridge accounts for most of the final accuracy; no new vision training was run.

Equality to a joint fit is a precise benefit

The accumulated and joint ridge implementations have identical saved final accuracies in every seed. A sequential cross-entropy head reaches a much lower 64.70%, but uses different features, a different loss, and different retention. That comparison does not isolate the optimizer. Nor does pooled-objective preservation guarantee that a conflicting new class cannot change an old prediction.

There is more behind the small head

The runner normalizes features using the complete training set before task arrival, so this is not a strictly causal online preprocessing experiment. Its default 10,001-coordinate Gram occupies about 400 MB in binary32 before other arrays. Inference still needs the encoder. These costs and information assumptions are part of the method, not footnotes to omit from an edge-learning story.

Update supervision without rebuilding perception

This can be practical when labels change more often than the visual world and a stable encoder is already available. It is less convincing as an edge-memory claim when a dense 10,000-feature Gram and the encoder are omitted from the budget. The manuscript now makes both the computation and the retained state visible.

A useful division of labor

The opportunity is modularity: retain a high-quality representation, update a precisely defined decision objective, and measure whether the resulting system serves the application. That can be valuable without inventing a new perception model or promising universal zero forgetting. The paper separates the observed accuracy, the exact estimator identity, and the remaining deployment costs.

Evidence & further reading

The links below distinguish the project record from foundational literature. This revised story does not add a new application-validation experiment.

  1. Consolidated research results, including constitutive edges and continual memory. Daniel Schmitter (2026). Local archive snapshot.
  2. Full experimental record. Daniel Schmitter (2026). Local archive snapshot.
  3. Negative-result appendix. Daniel Schmitter (2026). Local archive snapshot.
  4. Correction: what exact Gram memory does and does not establish. Spline research archive (2026). Local archive snapshot.