Historical source. Some claims in older records were subsequently corrected. The associated article states the adopted interpretation. This record preserves the original source alongside its rendered reading view.
Rendered archival Markdown
This reading view preserves headings, tables, lists, code fragments and mathematical notation from the local research record.
Correction: what exact Gram memory does and does not establish
The new real-data and matched-memory results prompted an additional manuscript consistency audit. Older text in the main theory and building blocks incorrectly inferred that an exact pooled ridge optimum prevents later tasks from degrading earlier predictions. This is a mathematical error, not merely a missing empirical confidence interval. The corrected manuscript supersedes that inference; historical versions and measurements remain in Git.
With scalar feature x=1 and ridge strength 1, the first target y=1 gives w=1/2 and old-example squared error 1/4. Adding an equally weighted target y=-1 gives the exact pooled optimum w=0 and old error 1. Both observations remain fully represented in G=2, b=0. No statistic was forgotten, yet old prediction quality worsened. Tests exercise the actual moment implementation.
The valid guarantee is equality to batch ridge for a fixed feature map, fixed output schema, fixed sample-sum regularizer and correctly accumulated statistics, up to numerical solve/accumulation error. It removes sequential optimizer dependence relative to that batch estimator. It does not remove conflicting targets, representation mismatch, class competition or distribution shift. Finite forgetting deliberately changes the objective.
Four additional boundaries are made executable:
- An immutable old coefficient program remains bitwise unchanged even when a
newly pooled program predicts differently. This is the stronger versioned preservation mechanism used elsewhere, not a property of every pooled fit.
- Floating-point positive Gram sums can differ with addition order. With
feature magnitudes 1e8, 1, 1, the two accumulation orders differ by two at scale 1e16. Mathematical order independence is not bitwise reproducibility.
- A one-record scalar Gram and cross-statistic can reconstruct the original
positive feature and target exactly. Keeping no raw exemplar is therefore not itself a privacy guarantee; previous attack-specific empirical findings cannot establish universal privacy or absence of membership leakage.
- Two datasets can have the same old Gram but different cross-statistics for
a new nonlinear feature. Exact feature growth needs extra retained evidence, sufficient cross-statistics, or a restricted transformation of the old map.
Exact unlearning by subtraction also requires the deleted contribution in the same fixed feature/schema/regularization regime. The aggregate does not magically identify an arbitrary individual's contribution after it is lost, and subtraction does not erase knowledge from a previously trained backbone.
These are elementary counterexamples and qualifications, not new theorems or a new benchmark win. They narrow the headline claim while preserving valid observations of low forgetting, protected-support identity, archived-program immutability and empirical performance. The subsequent SAC experiments also use ordinary gradient training: gradient-free coefficient compilation must not be presented as no backpropagation anywhere in the enlarged project.
Regression evidence: tests/test_gram_claim_boundaries.py.
The later clipping-aware extension adds a distinct boundary: identical full cubic empirical Grams need not encode the same clipped fitting objective. Its two new tests leave the five original counterexamples and all historical experiment code unchanged.
Five new counterexample tests and two existing moment tests pass. The main proof, abstract, vision, memory block, related-work interpretation and limitations are corrected; the theory abstract states the same scope. Two misleading schematic images are replaced in the main manuscript by editable vector diagrams with the corrected contracts. Original images and historical experiments are preserved. Changed main pages 1, 4, 6–7, 11, 17–18, 24, 92–93 and theory pages 1–2 were rendered and visually inspected. This is not an exhaustive revalidation of all historical benchmark/SOTA statements or all pre-existing layout issues in the long archival manuscripts.
Follow-through in historical result paragraphs
An additional consistency pass corrects the old continual-learning subsection, lifelong-memory summary table and RL introduction. Reversing a particular stream to 1e-6 agreement is empirical numerical stability, not proof that any gradient method must forget. The table's attack AUC is now attack-specific, not labelled near-private; retained-task statistic subtraction is not labelled general O(1) unlearning. Existing measured numbers are unchanged, and a SOTA label unsupported by a current matched comparison is removed.
The two-game RL example jointly exposed its trunk to both games before freezing it. Old-policy preservation by frozen modules can also be used with gradient training, and this example is not unseen-game representation transfer. The gated-delta update is already identified as a gradient step; the introduction now agrees with that explanation. Visual inspection also shows the historical EMA self-improvement curve plateauing and declining after its peak, so its caption no longer calls the trajectory monotonic. These are interpretation corrections, not new runs or changes to archived outcomes.
The ten Gram/SSP boundary tests pass again. The rebuilt main manuscript has 145 pages; changed result pages 27, 31, 32 and 38 are rendered and inspected. Older graphical shorthand is explicitly qualified in the caption rather than altering the archived raster experiment figure. The rest of the historical ledger is not thereby certified correct.
View raw MD source
# Correction: what exact Gram memory does and does not establish
The new real-data and matched-memory results prompted an additional
manuscript consistency audit. Older text in the main theory and building
blocks incorrectly inferred that an exact pooled ridge optimum prevents
later tasks from degrading earlier predictions. This is a mathematical error,
not merely a missing empirical confidence interval. The corrected manuscript
supersedes that inference; historical versions and measurements remain in Git.
With scalar feature x=1 and ridge strength 1, the first target y=1 gives
w=1/2 and old-example squared error 1/4. Adding an equally weighted target
y=-1 gives the exact pooled optimum w=0 and old error 1. Both observations
remain fully represented in G=2, b=0. No statistic was forgotten, yet old
prediction quality worsened. Tests exercise the actual moment implementation.
The valid guarantee is equality to batch ridge for a fixed feature map,
fixed output schema, fixed sample-sum regularizer and correctly accumulated
statistics, up to numerical solve/accumulation error. It removes sequential
optimizer dependence relative to that batch estimator. It does not remove
conflicting targets, representation mismatch, class competition or distribution
shift. Finite forgetting deliberately changes the objective.
Four additional boundaries are made executable:
- An immutable old coefficient program remains bitwise unchanged even when a
newly pooled program predicts differently. This is the stronger versioned
preservation mechanism used elsewhere, not a property of every pooled fit.
- Floating-point positive Gram sums can differ with addition order. With
feature magnitudes 1e8, 1, 1, the two accumulation orders differ by two at
scale 1e16. Mathematical order independence is not bitwise reproducibility.
- A one-record scalar Gram and cross-statistic can reconstruct the original
positive feature and target exactly. Keeping no raw exemplar is therefore
not itself a privacy guarantee; previous attack-specific empirical findings
cannot establish universal privacy or absence of membership leakage.
- Two datasets can have the same old Gram but different cross-statistics for
a new nonlinear feature. Exact feature growth needs extra retained evidence,
sufficient cross-statistics, or a restricted transformation of the old map.
Exact unlearning by subtraction also requires the deleted contribution in
the same fixed feature/schema/regularization regime. The aggregate does not
magically identify an arbitrary individual's contribution after it is lost,
and subtraction does not erase knowledge from a previously trained backbone.
These are elementary counterexamples and qualifications, not new theorems or
a new benchmark win. They narrow the headline claim while preserving valid
observations of low forgetting, protected-support identity, archived-program
immutability and empirical performance. The subsequent SAC experiments also
use ordinary gradient training: gradient-free coefficient compilation must
not be presented as no backpropagation anywhere in the enlarged project.
Regression evidence: `tests/test_gram_claim_boundaries.py`.
The later [clipping-aware extension](CLIPPED_GRAM_MEMORY_BOUNDARY_2026-09-13.md)
adds a distinct boundary: identical full cubic empirical Grams need not encode
the same clipped fitting objective. Its two new tests leave the five original
counterexamples and all historical experiment code unchanged.
Five new counterexample tests and two existing moment tests pass. The main
proof, abstract, vision, memory block, related-work interpretation and
limitations are corrected; the theory abstract states the same scope.
Two misleading schematic images are replaced in the main manuscript by
editable vector diagrams with the corrected contracts. Original images and
historical experiments are preserved. Changed main pages 1, 4, 6–7, 11, 17–18,
24, 92–93 and theory pages 1–2 were rendered and visually inspected. This is
not an exhaustive revalidation of all historical benchmark/SOTA statements or
all pre-existing layout issues in the long archival manuscripts.
## Follow-through in historical result paragraphs
An additional consistency pass corrects the old continual-learning subsection,
lifelong-memory summary table and RL introduction. Reversing a particular
stream to 1e-6 agreement is empirical numerical stability, not proof that any
gradient method must forget. The table's attack AUC is now attack-specific,
not labelled near-private; retained-task statistic subtraction is not labelled
general O(1) unlearning. Existing measured numbers are unchanged, and a SOTA
label unsupported by a current matched comparison is removed.
The two-game RL example jointly exposed its trunk to both games before freezing
it. Old-policy preservation by frozen modules can also be used with gradient
training, and this example is not unseen-game representation transfer. The
gated-delta update is already identified as a gradient step; the introduction
now agrees with that explanation. Visual inspection also shows the historical
EMA self-improvement curve plateauing and declining after its peak, so its
caption no longer calls the trajectory monotonic. These are interpretation
corrections, not new runs or changes to archived outcomes.
The ten Gram/SSP boundary tests pass again. The rebuilt main manuscript has
145 pages; changed result pages 27, 31, 32 and 38 are rendered and inspected.
Older graphical shorthand is explicitly qualified in the caption rather than
altering the archived raster experiment figure. The rest of the historical
ledger is not thereby certified correct.