The architecture in context
The system we are building
The archive corrected an overly broad interpretation of Gram memory. Additive statistics preserve the pooled fixed-feature ridge objective. They do not guarantee that every old prediction, old-task loss or old decision remains unchanged. The smallest counterexample uses one scalar coefficient and two conflicting labels.
Who does what in the stack
- NumPy assertions
- Make the counterexample numerically executable.
- PriorMomentState
- Implements the pooled ridge summary.
- Regression tests
- Prevent corrected claims from reappearing after refactoring.
The correction is backed by executable tests rather than prose alone. The companion test constructs a PriorMomentState, fits the first observation, adds the conflicting observation and checks both the new optimum and retained statistics. It distinguishes a data-summary guarantee from a behavioral guarantee.
Open up the implementation
Order invariance does not freeze old predictions
The simplest counterexample uses one coefficient. With one old sample x=1,y=1 and ridge λ=1, the optimum is 1/2. Adding x=1,y=−1 changes the optimum to 0. Both task orders produce the same final solution, but the prediction on the old sample has changed. Additivity therefore cannot by itself justify zero forgetting.
The mathematical contract
Protecting outputs requires an additional constraint, fixed subspace or other mechanism that defines exactly what is protected. Keeping all Gram information solves a joint least-squares objective; it need not minimize the old-task loss after a distribution change.
Implementation and resource card
- Capacity / budget
- Stored sufficient statistics, not a separate neural architecture. Exact-arithmetic order invariance and retention are different properties.
- Execution evidence
- This revision inspects and explains the archived implementation. It does not rerun the original workload. No unrecorded convergence time, throughput or accelerator result is supplied.
- Current reproduction context
- Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.
From explanation to a reproducible check
Calculate the scalar example by hand, then run both task orders. Report final equality across orders and old-prediction drift as separate assertions. Neither should be used as a substitute for the other.
Preserve input identities, configuration and failure records with the result. A successful numerical check only establishes the operation it exercises: it does not certify an entire dataset, model or deployed system. Reproduce the interface on a small deterministic input before optimizing throughput or increasing workload size.
A closer look at the implementation
The code that carries the idea
With unit regularization, the first solution is 0.5. After observing the opposite label at the same input, the solution becomes zero. Error on the old target increases from 0.25 to 1, while the Gram and cross-product values are exactly the expected pooled values. Nothing has been discarded from the declared sufficient statistics.
def test_pooled_optimum_can_worsen_an_old_task_without_losing_statistics():
state = PriorMomentState(np.zeros((1, 1)), 1.)
state.observe([[1.]], [[1.]])
old = state.solve().copy(); old_error = float((old[0, 0]-1.)**2)
state.observe([[1.]], [[-1.]])
new = state.solve()
np.testing.assert_allclose(old, [[.5]])
np.testing.assert_allclose(new, [[0.]])
assert float((new[0, 0]-1.)**2) > old_error
np.testing.assert_array_equal(state.gram, [[2.]])
np.testing.assert_array_equal(state.prior, [[0.]])Verbatim archive excerpt from test_gram_claim_boundaries.py (companion source X01). Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.
The boundary that matters
Other tests in the same file cover finite-precision order dependence, disclosure from statistics and missing cross-statistics after feature growth. “Raw-data-free” does not imply private, and a fixed-feature guarantee does not automatically survive representation changes.