Evaluation practice · E46 · Engineering practice

Turn an overbroad memory claim into a regression test

Retaining sufficient statistics is not the same as preserving old predictions. An eleven-line counterexample keeps that distinction executable.

NumPy assertionsPriorMomentStateRegression tests
Adding a conflicting observation changes a pooled ridge optimum from 1/2 to zero, even though no sufficient statistics were discarded.
Figure 1. Order invariance ≠ zero forgetting. Adding a conflicting observation changes a pooled ridge optimum from 1/2 to zero, even though no sufficient statistics were discarded. Exact scalar ridge counterexample. Original vector illustration.

Follow the information

From input to outcome

The second sample contributes to the pooled sufficient statistics. The updated solve keeps the full pooled objective but changes the prediction at the old point. No loss of statistics is needed for this counterexample.

The second sample contributes to the pooled sufficient statistics. The updated solve keeps the full pooled objective but changes the prediction at the old point. No loss of statistics is needed for this counterexample.
Figure 2. Information flow. Solid arrows carry observations, tensors or artifacts; other routes are explicitly labelled. Signal shapes, matrices and network icons are schematic, not measured samples or literal neuron counts. Open full-size SVG ↗ On narrow screens, scroll the diagram horizontally.

Read this alongside Figure 1: Adding a conflicting observation changes a pooled ridge optimum from 1/2 to zero, even though no sufficient statistics were discarded. The module map and layer-level figures below expand the operations in this route.

Turn an overbroad memory claim into a regression test: system and evaluation mapOld observation: x=1, y=1 → Ridge state: λ=1 → w=0.5 → Conflicting new label: x=1, y=−1 → Pooled statistics: No records forgotten algebraically → Updated optimum: w=0 · old error increases. A high-level module map; comparison branches and training details are explained in the article.EVALUATION PRACTICE / E46 / MODULE MAP01 INPUTOld observationx=1, y=102 MODULERidge stateλ=1 → w=0.503 MODULEConflicting new labelx=1, y=−104 MODULEPooled statisticsNo records forgotten algebraically05 OUTPUTUpdated optimumw=0 · old error increases
Source-grounded module map. Boxes summarize operations, not individual neurons; comparison arms and training paths are detailed below. On a small screen, scroll the diagram horizontally.
Old observation — x=1, y=1

The architecture in context

The system we are building

The archive corrected an overly broad interpretation of Gram memory. Additive statistics preserve the pooled fixed-feature ridge objective. They do not guarantee that every old prediction, old-task loss or old decision remains unchanged. The smallest counterexample uses one scalar coefficient and two conflicting labels.

Who does what in the stack

NumPy assertions
Make the counterexample numerically executable.
PriorMomentState
Implements the pooled ridge summary.
Regression tests
Prevent corrected claims from reappearing after refactoring.

The correction is backed by executable tests rather than prose alone. The companion test constructs a PriorMomentState, fits the first observation, adds the conflicting observation and checks both the new optimum and retained statistics. It distinguishes a data-summary guarantee from a behavioral guarantee.

Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph.
Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph. Open full-size SVG ↗

Open up the implementation

Order invariance does not freeze old predictions

A concrete operation-level view of this implementation; no unobserved neural architecture is implied.
A concrete operation-level view of this implementation; no unobserved neural architecture is implied. Open full-size SVG ↗

The simplest counterexample uses one coefficient. With one old sample x=1,y=1 and ridge λ=1, the optimum is 1/2. Adding x=1,y=−1 changes the optimum to 0. Both task orders produce the same final solution, but the prediction on the old sample has changed. Additivity therefore cannot by itself justify zero forgetting.

The mathematical contract

(G0+G1+λI)W=B0+B1(G_0+G_1+\lambda I)W=B_0+B_1

Protecting outputs requires an additional constraint, fixed subspace or other mechanism that defines exactly what is protected. Keeping all Gram information solves a joint least-squares objective; it need not minimize the old-task loss after a distribution change.

Implementation and resource card

Capacity / budget
Stored sufficient statistics, not a separate neural architecture. Exact-arithmetic order invariance and retention are different properties.
Execution evidence
This revision inspects and explains the archived implementation. It does not rerun the original workload. No unrecorded convergence time, throughput or accelerator result is supplied.
Current reproduction context
Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.

From explanation to a reproducible check

Calculate the scalar example by hand, then run both task orders. Report final equality across orders and old-prediction drift as separate assertions. Neither should be used as a substitute for the other.

Preserve input identities, configuration and failure records with the result. A successful numerical check only establishes the operation it exercises: it does not certify an entire dataset, model or deployed system. Reproduce the interface on a small deterministic input before optimizing throughput or increasing workload size.

A closer look at the implementation

The code that carries the idea

With unit regularization, the first solution is 0.5. After observing the opposite label at the same input, the solution becomes zero. Error on the old target increases from 0.25 to 1, while the Gram and cross-product values are exactly the expected pooled values. Nothing has been discarded from the declared sufficient statistics.

Python · file · lines 6–16
def test_pooled_optimum_can_worsen_an_old_task_without_losing_statistics():
    state = PriorMomentState(np.zeros((1, 1)), 1.)
    state.observe([[1.]], [[1.]])
    old = state.solve().copy(); old_error = float((old[0, 0]-1.)**2)
    state.observe([[1.]], [[-1.]])
    new = state.solve()
    np.testing.assert_allclose(old, [[.5]])
    np.testing.assert_allclose(new, [[0.]])
    assert float((new[0, 0]-1.)**2) > old_error
    np.testing.assert_array_equal(state.gram, [[2.]])
    np.testing.assert_array_equal(state.prior, [[0.]])

Verbatim archive excerpt from test_gram_claim_boundaries.py (companion source X01). Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.

The boundary that matters

Other tests in the same file cover finite-precision order dependence, disclosure from statistics and missing cross-statistics after feature growth. “Raw-data-free” does not imply private, and a fixed-feature guarantee does not automatically survive representation changes.

Keep building

Other posts of interest