The attractive story has two claims
An agent solves a problem, keeps what worked, and uses it to solve the next. It is tempting to draw an ever-rising capability curve. But the mechanism makes two empirically different claims: experience helps relative to a control, and accumulating more experience keeps improving the system. Our small language-model study supports an exploratory instance of the first, not the second.
A feedback loop needs a capability measure
The store grows when the generator produces an accepted answer, but the prompt sees only a recency window. More stored examples need not mean more information used on the next problem. A rising memory count is therefore not a learning curve. The saved verified arm scores 30 of the first sixty problems and 29 of the last sixty.
Three different properties should remain separate: accepted-record correctness, aggregate benefit from filtered context, and improvement in future competence as experience accumulates. The first is an invariant of the verifier. The second is an exploratory observed comparison. The third is not supported by this record.
The record does not contain an upward curve
Verified memory answers 30 of the first 60 problems correctly and 29 of the last 60. Its six consecutive twenty-problem blocks contain 11, 8, 11, 12, 9, and 8 correct answers. We report the full sequence of block counts rather than selecting a favorable window. This is not evidence of compounding accuracy, even though the full-run total beats both tested controls.

A growing store is not a growing learner
Only the most recent accepted examples enter the prompt. Older records remain in the list but fall out of view; the neural weights never change. More stored bytes therefore need not mean more accessible knowledge. Nor does correct-answer filtering guarantee that the accompanying explanation teaches a transferable procedure. Those are separate mechanisms the experiment did not implement or isolate.
Three properties should not share one headline
Record retention asks whether accepted information is still stored. Behavioral retention asks whether an earlier capability still works. Recursive improvement asks whether changes to the system produce further improved learning or problem solving. A system may have one without the others. Our continuous-function preservation experiments concern yet another precisely declared object—not a universal guarantee about agent intelligence.
The useful research question becomes sharper
To establish progressive improvement, a study would have to compare changing memory against strong fixed-context controls, retain complete trajectories, and measure transfer beyond the examples that supplied feedback. That is not a recipe already validated here. It identifies what the current result leaves open, without erasing the observed benefit of filtered context.
Measure new capability, not a growing archive
A serious self-improvement experiment would need independently measured new capability and controlled alternatives such as a fixed clean memory. The current result is useful because it identifies that missing step. Ambition belongs in the proposed system; demonstrated progress belongs in the measured capability, not the size of its archive.
Ambition survives a more exact description
A reliable small agent may combine a frozen model, external checks, retrieved evidence, and local adaptation. No single component earns the label of recursive self-improvement by itself. The scientific opportunity is to connect them in a system whose new capabilities can be measured—and to recognize when the evidence shows a plateau rather than draw the growth curve we hoped to see.
Evidence & further reading
The links below distinguish the project record from foundational literature. This revised story does not add a new application-validation experiment.
- Consolidated research results, including constitutive edges and continual memory. Daniel Schmitter (2026). Local archive snapshot.
- Negative-result appendix. Daniel Schmitter (2026). Local archive snapshot.
- Flagship stories, source corrections and visual direction. Publication audit (2026). Local archive snapshot.
- Archived 120-problem multiplication outcomes. Spline research archive (2026). Local archive snapshot.