Learning algorithms · Research & Algorithms

Why self-improvement stalls—and what changes the outcome

A system can perform better with memory without becoming progressively better as it accumulates memory. The distinction is the difference between a useful mechanism and an unsupported growth story.

EXPLORE THE IDEA

A useful memory is not a rising curve

Every chronological block stays in the picture.

ALL CHRONOLOGICAL BLOCKSCorrect answers out of twenty111283114125·6·Blocks reveal in order; the counts do not changeTHE TWO HALVESBenefit is not compounding improvement30 / 60first half29 / 60second halfOne adaptive stream, not independent trials
50%
Actual verified-memory block counts: 11, 8, 11, 12, 9, 8 correct out of 20; totals 30/60 and 29/60 for the two halves. Revealing blocks does not fit a trend or imply independent replications.

Follow the information

From input to outcome

This diagram follows the evaluation branch rather than equating storage growth with learning. A growing store can feed a fixed-size context while aggregate accuracy remains flat. The full block sequence is retained, not a favorable selected window.

Scroll the diagram horizontally to follow the route. Keyboard: focus the diagram, then use the arrow keys.

Accepted experience store → Bounded recency window → Frozen generator → Chronological accuracy → Capability over time. This diagram follows the evaluation branch rather than equating storage growth with learning. A growing store can feed a fixed-size context while aggregate accuracy remains flat. The full block sequence is retained, not a favorable selected window.
Information-flow map. More stored examples did not establish compounding improvement. Original vector schematic based on the method and evidence discussed in this article; signal shapes and icons are illustrative, not additional measurements. Open full-size diagram ↗

Read the main route from left to right; labelled side branches show additional inputs, checks or feedback. The sections below explain the operations and their experimental limits.

The attractive story has two claims

An agent solves a problem, keeps what worked, and uses it to solve the next. It is tempting to draw an ever-rising capability curve. But the mechanism makes two empirically different claims: experience helps relative to a control, and accumulating more experience keeps improving the system. Our small language-model study supports an exploratory instance of the first, not the second.

A feedback loop needs a capability measure

The store grows when the generator produces an accepted answer, but the prompt sees only a recency window. More stored examples need not mean more information used on the next problem. A rising memory count is therefore not a learning curve. The saved verified arm scores 30 of the first sixty problems and 29 of the last sixty.

Three different properties should remain separate: accepted-record correctness, aggregate benefit from filtered context, and improvement in future competence as experience accumulates. The first is an invariant of the verifier. The second is an exploratory observed comparison. The third is not supported by this record.

Verification controls the memory boundary
Verification controls the memory boundary. Original scientific diagram; the stated component and information flow, not an additional experiment. Open full-size figure ↗

The record does not contain an upward curve

Verified memory answers 30 of the first 60 problems correctly and 29 of the last 60. Its six consecutive twenty-problem blocks contain 11, 8, 11, 12, 9, and 8 correct answers. We report the full sequence of block counts rather than selecting a favorable window. This is not evidence of compounding accuracy, even though the full-run total beats both tested controls.

Same experiment, different question: the overall context advantage is real in the saved counts, but the chronological halves do not support increasing competence. There is no fitted trend or independent replication here.
Same experiment, different question: the overall context advantage is real in the saved counts, but the chronological halves do not support increasing competence. There is no fitted trend or independent replication here.

A growing store is not a growing learner

Only the most recent accepted examples enter the prompt. Older records remain in the list but fall out of view; the neural weights never change. More stored bytes therefore need not mean more accessible knowledge. Nor does correct-answer filtering guarantee that the accompanying explanation teaches a transferable procedure. Those are separate mechanisms the experiment did not implement or isolate.

Three properties should not share one headline

Record retention asks whether accepted information is still stored. Behavioral retention asks whether an earlier capability still works. Recursive improvement asks whether changes to the system produce further improved learning or problem solving. A system may have one without the others. Our continuous-function preservation experiments concern yet another precisely declared object—not a universal guarantee about agent intelligence.

stored information ≠ accessible information ≠ improved capability\begin{gathered}\text{stored information}\ \neq\ \text{accessible information}\ \neq\ \text{improved capability}\end{gathered}
These are different properties of the complete system, each requiring its own definition and evidence.

The useful research question becomes sharper

To establish progressive improvement, a study would have to compare changing memory against strong fixed-context controls, retain complete trajectories, and measure transfer beyond the examples that supplied feedback. That is not a recipe already validated here. It identifies what the current result leaves open, without erasing the observed benefit of filtered context.

Measure new capability, not a growing archive

A serious self-improvement experiment would need independently measured new capability and controlled alternatives such as a fixed clean memory. The current result is useful because it identifies that missing step. Ambition belongs in the proposed system; demonstrated progress belongs in the measured capability, not the size of its archive.

Ambition survives a more exact description

A reliable small agent may combine a frozen model, external checks, retrieved evidence, and local adaptation. No single component earns the label of recursive self-improvement by itself. The scientific opportunity is to connect them in a system whose new capabilities can be measured—and to recognize when the evidence shows a plateau rather than draw the growth curve we hoped to see.

Evidence & further reading

The links below distinguish the project record from foundational literature. This revised story does not add a new application-validation experiment.

  1. Consolidated research results, including constitutive edges and continual memory. Daniel Schmitter (2026). Local archive snapshot.
  2. Negative-result appendix. Daniel Schmitter (2026). Local archive snapshot.
  3. Flagship stories, source corrections and visual direction. Publication audit (2026). Local archive snapshot.
  4. Archived 120-problem multiplication outcomes. Spline research archive (2026). Local archive snapshot.