Continual learning · Research & Algorithms

Continual learning without a replay buffer: what can we actually guarantee?

A model can retain every contribution to its training objective and still get worse on an old example. That distinction changes what “memory without replay” can honestly promise.

EXPLORE THE IDEA

Keep every observation. Still change your mind.

Objective preservation is not behavioral preservation.

EVIDENCE LEDGERBoth targets remain in memoryEarlier observationtarget +1 · weight 1Later observationtarget −1 · weight 0.50MODEL STATEA changed optimumc* = 0.200Old squared error0.640No observation was erased
50%
Exact scalar ridge example with λ=1: target +1 receives weight one and target −1 receives the slider weight w. The optimum is (1−w)/(2+w). Old squared error changes even though no evidence is discarded.

Follow the information

From input to outcome

The summary preserves a fixed quadratic fitting question. A conflicting new label can move its exact optimum and worsen an old prediction even though no fitting evidence was forgotten.

Scroll the diagram horizontally to follow the route. Keyboard: focus the diagram, then use the arrow keys.

New fixed features Φ → Add sufficient statistics → Retained objective → Pooled ridge solution → Updated predictions. The summary preserves a fixed quadratic fitting question. A conflicting new label can move its exact optimum and worsen an old prediction even though no fitting evidence was forgotten.
Information-flow map. Objective retention is not behavioral retention. Original vector schematic based on the method and evidence discussed in this article; signal shapes and icons are illustrative, not additional measurements. Open full-size diagram ↗

Read the main route from left to right; labelled side branches show additional inputs, checks or feedback. The sections below explain the operations and their experimental limits.

A perfect notebook can contain a conflict

Imagine learning one constant prediction. The first observation asks for +1; the next asks for −1. A learner that remembers both cannot satisfy them simultaneously with that constant. This tiny example exposes an ambiguity: remembering evidence is not the same as preserving the behavior that evidence once produced. Something has to give, even with a perfect notebook.

A sufficient statistic is a memory of a question

The retained Gram and right-hand side summarize a fixed least-squares question: how would these same features fit all accumulated labels? They allow a later solve without replaying the rows. The feature map, weighting, normalization, and target meaning must remain compatible. Changing any of those can change which historical products are needed.

This object is different from a frozen function. When a new label conflicts with an old one, the correct pooled optimum moves. The one-coefficient counterexample is deliberately small because it makes the information distinction impossible to hide behind optimization noise: both observations are retained exactly, yet the old prediction worsens.

A proposed update is not a committed memory
A proposed update is not a committed memory. Original scientific diagram; the stated component and information flow, not an additional experiment. Open full-size figure ↗

What the compact summary preserves

For fixed features and squared loss, a Gram matrix and a target cross-product summarize every observation’s contribution. A ridge solve using them matches the corresponding batch calculation, provided features, weighting, and regularization stay fixed. Raw replay is unnecessary for that objective. The summary does not promise that its evolving optimum leaves each earlier prediction unchanged.

G=∑nznznT,b=∑nznyn,c=(G+λI)−1b\begin{gathered}G=\sum_nz_nz_n^T,\quad b=\sum_nz_ny_n,\quad c=(G+\lambda I)^{-1}b\end{gathered}
These statistics preserve a fixed-feature ridge objective, not the optimizer’s earlier answer.

One coefficient disproves the stronger claim

Set the feature and ridge coefficient to one. After target +1, the fitted coefficient is 1/2 and old squared error is 1/4. Add target −1: the optimum becomes zero and old error rises to one. No observation has been lost. The worse prediction is the compromise prescribed by the objective that was retained exactly.

Local features make the notebook smaller

A derivative observation of a cubic spline touches at most four neighboring coefficients. Its Gram contribution is local, producing a narrow band. The band and right-hand side use roughly five doubles per coefficient. That is useful accounting, but excludes coefficients, masks, validation evidence, and audit histories. It is not the full memory of an autonomous learner.

The experiments found the same boundary

Pooled-statistics adaptation accurately accumulated its normal equations but showed old-region regression. Protecting sampled predictions was also insufficient: the model could change between samples. Those failures motivated continuous-region protection. They were not inconvenient exceptions to a zero-forgetting claim; they established which claim was justified and which required a genuinely different mechanism.

Three different promises a memory can make

This matters for lifelong-learning systems that advertise compact memory. One can preserve past evidence, a past model, or a protected behavior; those are different interfaces with different costs. The rest of the memory sequence builds a continuous behavior guarantee on top of this distinction rather than using statistical accumulation as a substitute for it.

Name the memory object

The summary works for a fixed feature space and objective. New features may require historical cross-products never retained, while a new task may need information absent from that summary. A practical system must name its memory object: fitting evidence, sampled outputs, a continuous function, or immutable versions. Each is useful, and each has different preservation conditions.

Evidence & further reading

The links below distinguish the project record from foundational literature. This revised story does not add a new application-validation experiment.

  1. Correction: what exact Gram memory does and does not establish. Spline research archive (2026). Local archive snapshot.
  2. Consolidated research results, including constitutive edges and continual memory. Daniel Schmitter (2026). Local archive snapshot.
  3. Operator-spline theory: consolidated research manuscript. Daniel Schmitter (2026). Local archive snapshot.