Learning algorithms · Research & Algorithms

Can a frozen language model improve without changing its weights?

A model’s weights can stay fixed while its working context changes. In one arithmetic stream, keeping externally checked examples helped; keeping unfiltered generated examples hurt.

EXPLORE THE IDEA

Put verification at the memory boundary

A frozen generator can inherit better context.

ILLUSTRATIVE PROPOSED EXAMPLECheck the final integer234 × 567proposed: 132677exact: 132678MEMORY BOUNDARYOnly checked final answers enterREJECT EXAMPLEThe explanation is not verified.Not a retained historical model transcript
50%
Illustrative arithmetic examples evaluated by exact integer multiplication. They demonstrate the stored-final-answer rule, not historical model transcripts (which were not retained). Actual complete-stream counts appear in the article.

Follow the information

From input to outcome

The verifier checks the parsed final integer, not every reasoning step in the response. Accepted examples may enter future prompts through recency selection. Rejected responses do not update the verified store.

Scroll the diagram horizontally to follow the route. Keyboard: focus the diagram, then use the arrow keys.

Next arithmetic task → Assemble prompt → Frozen language model → External final-answer check → Accept into memory. The verifier checks the parsed final integer, not every reasoning step in the response. Accepted examples may enter future prompts through recency selection. Rejected responses do not update the verified store.
Information-flow map. Exploratory benefit on one saved stream, not general recursive self-improvement. Original vector schematic based on the method and evidence discussed in this article; signal shapes and icons are illustrative, not additional measurements. Open full-size diagram ↗

Read the main route from left to right; labelled side branches show additional inputs, checks or feedback. The sections below explain the operations and their experimental limits.

Experience can live outside the weights

An assistant can reuse an earlier solution by placing it in its next prompt. That is a different mechanism from gradient-based learning: the generator stays fixed while its available examples change. The interesting design question is which experiences deserve to enter the store. A confident wrong answer is still wrong—and may become an unhelpful demonstration.

What exactly is verified before memory changes?

The external routine verifies the parsed final product. A generated explanation with a correct final integer can still contain an invalid intermediate step. If accepted, the complete generated example enters a store, and a bounded recency slice contributes to later prompts. The model’s weights remain unchanged.

The paired accounting prevents aggregate accuracy from hiding regressions: verified memory solves 36 positions missed by the no-memory arm but loses 14 positions that arm solves. The saved response texts are absent, so the paper does not reconstruct persuasive reasoning traces as if they were observed.

Verification controls the memory boundary
Verification controls the memory boundary. Original scientific diagram; the stated component and information flow, not an additional experiment. Open full-size figure ↗

Put an independent check at the memory boundary

The retained experiment asks a frozen Qwen2.5-VL 7B model to multiply three-digit integers. An external arithmetic routine checks the parsed final answer. One arm keeps only successful examples, another has no memory, and a third keeps every response with a parseable integer. The implementation uses recent examples, not semantic nearest-neighbor retrieval.

Mt+1={Mt∪{(xt,ot)},V(xt,ot)=1,Mt,otherwise,\begin{gathered}M_{t+1}=\begin{cases}M_t\cup\{(x_t,o_t)\},&V(x_t,o_t)=1,\\M_t,&\text{otherwise},\end{cases}\end{gathered}
The invariant concerns accepted records: every stored final integer passes the checker. It does not guarantee the model’s next answer.

The observed difference is substantial, but specific

Across the same 120 problem positions, verified memory gets 59 correct, the no-memory arm 37, and unfiltered memory 16. Verified memory therefore improves aggregate accuracy on this saved stream by 18.33 percentage points over no memory. It is not a per-question guarantee: there are fourteen positions where the baseline succeeds and verified memory fails.

Every saved attempt is included. Aggregate context benefit and progressive improvement are different questions; the verified arm does not improve between the two chronological halves.
Every saved attempt is included. Aggregate context benefit and progressive improvement are different questions; the verified arm does not improve between the two chronological halves.

A correct final integer is not a verified explanation

The entire generated response is stored, but only its parsed final number is checked. A flawed derivation can pass if it ends with the right integer. Parsing also matters: the available code selects the last number after the final marker. Our deterministic tests make that acceptance rule explicit; the historical responses were not retained, so we cannot retrospectively inspect their reasoning.

What this experiment cannot separate

The histories adapt as the run proceeds, the three arms execute in a fixed order, and complete prompts, errors, and model digests are missing. A clean fixed few-shot control was not tested. We therefore report all descriptive outcomes without presenting an iid significance test or attributing the difference to a single isolated cause. An exact calculator would already solve this task, so this is not an efficient arithmetic product.

Verification belongs at the memory boundary

The useful agent-design principle is an independent acceptance predicate at the memory boundary. It can prevent certain bad records from entering context. It does not guarantee good retrieval, correct reasoning, or a monotone increase in capability. The missing fixed-correct-example control is especially important before attributing the gain to continual accumulation.

A useful boundary for agent design

The example illustrates how an external acceptance rule can improve the contents of an agent’s context. It does not demonstrate general recursive self-improvement, permanent neural learning, or a spline-based memory mechanism. The compelling principle is narrower: when a trustworthy check exists, use it to distinguish evidence from repetition before experience is fed back into the system.

Evidence & further reading

The links below distinguish the project record from foundational literature. This revised story does not add a new application-validation experiment.

  1. Verification-guided in-context learning runner. Spline research archive (2026). Local archive snapshot.
  2. Archived 120-problem multiplication outcomes. Spline research archive (2026). Local archive snapshot.
  3. Flagship stories, source corrections and visual direction. Publication audit (2026). Local archive snapshot.
  4. Consolidated research results, including constitutive edges and continual memory. Daniel Schmitter (2026). Local archive snapshot.