Experience can live outside the weights
An assistant can reuse an earlier solution by placing it in its next prompt. That is a different mechanism from gradient-based learning: the generator stays fixed while its available examples change. The interesting design question is which experiences deserve to enter the store. A confident wrong answer is still wrong—and may become an unhelpful demonstration.
What exactly is verified before memory changes?
The external routine verifies the parsed final product. A generated explanation with a correct final integer can still contain an invalid intermediate step. If accepted, the complete generated example enters a store, and a bounded recency slice contributes to later prompts. The model’s weights remain unchanged.
The paired accounting prevents aggregate accuracy from hiding regressions: verified memory solves 36 positions missed by the no-memory arm but loses 14 positions that arm solves. The saved response texts are absent, so the paper does not reconstruct persuasive reasoning traces as if they were observed.
Put an independent check at the memory boundary
The retained experiment asks a frozen Qwen2.5-VL 7B model to multiply three-digit integers. An external arithmetic routine checks the parsed final answer. One arm keeps only successful examples, another has no memory, and a third keeps every response with a parseable integer. The implementation uses recent examples, not semantic nearest-neighbor retrieval.
The observed difference is substantial, but specific
Across the same 120 problem positions, verified memory gets 59 correct, the no-memory arm 37, and unfiltered memory 16. Verified memory therefore improves aggregate accuracy on this saved stream by 18.33 percentage points over no memory. It is not a per-question guarantee: there are fourteen positions where the baseline succeeds and verified memory fails.

A correct final integer is not a verified explanation
The entire generated response is stored, but only its parsed final number is checked. A flawed derivation can pass if it ends with the right integer. Parsing also matters: the available code selects the last number after the final marker. Our deterministic tests make that acceptance rule explicit; the historical responses were not retained, so we cannot retrospectively inspect their reasoning.
What this experiment cannot separate
The histories adapt as the run proceeds, the three arms execute in a fixed order, and complete prompts, errors, and model digests are missing. A clean fixed few-shot control was not tested. We therefore report all descriptive outcomes without presenting an iid significance test or attributing the difference to a single isolated cause. An exact calculator would already solve this task, so this is not an efficient arithmetic product.
Verification belongs at the memory boundary
The useful agent-design principle is an independent acceptance predicate at the memory boundary. It can prevent certain bad records from entering context. It does not guarantee good retrieval, correct reasoning, or a monotone increase in capability. The missing fixed-correct-example control is especially important before attributing the gain to continual accumulation.
A useful boundary for agent design
The example illustrates how an external acceptance rule can improve the contents of an agent’s context. It does not demonstrate general recursive self-improvement, permanent neural learning, or a spline-based memory mechanism. The compelling principle is narrower: when a trustworthy check exists, use it to distinguish evidence from repetition before experience is fed back into the system.
Evidence & further reading
The links below distinguish the project record from foundational literature. This revised story does not add a new application-validation experiment.
- Verification-guided in-context learning runner. Spline research archive (2026). Local archive snapshot.
- Archived 120-problem multiplication outcomes. Spline research archive (2026). Local archive snapshot.
- Flagship stories, source corrections and visual direction. Publication audit (2026). Local archive snapshot.
- Consolidated research results, including constitutive edges and continual memory. Daniel Schmitter (2026). Local archive snapshot.