ML systems · Research & Algorithms

Can dynamical models shrink a transformer’s KV cache?

A transformer does not remember its past in order to redraw a tensor. It remembers so that its next query can use that past. That distinction changes what a useful compression algorithm must preserve.

EXPLORE THE IDEA

Tiny key error. A different answer.

The future query determines what compression must preserve.

KEY RECONSTRUCTIONError can become almost invisibleε = 3.16e-4k₁ = (1, +ε)k₂ = (1, −ε)q = (0, 1/ε)ATTENTION OUTPUTThe query amplifies the lost directionOriginal keys0.7616Both keys collapsed0Output error stays tanh(1); query norm grows
50%
Computed two-key counterexample: keys (1,±ε), values ±1, query (0,1/ε). Replacing both keys with (1,0) drives key error to zero while attention-output error remains tanh(1). Query norm grows; typical-model behavior is not implied.

Follow the information

From input to outcome

The diagram shows the required codec-to-consumer interface, not a validated implementation from the archive. The audited prototype omitted causal masking in its diagnostic and normalization in its auxiliary query extraction, and some byte counts did not match decoded precision.

Scroll the diagram horizontally to follow the route. Keyboard: focus the diagram, then use the arrow keys.

Transformer K and V → Encode and serialize → Decode stored bytes → Causal attention → Consumer output. The diagram shows the required codec-to-consumer interface, not a validated implementation from the archive. The audited prototype omitted causal masking in its diagnostic and normalization in its auxiliary query extraction, and some byte counts did not match decoded precision.
Information-flow map. Required interface; the historical proxy did not establish a usable KV-cache codec. Original vector schematic based on the method and evidence discussed in this article; signal shapes and icons are illustrative, not additional measurements. Open full-size diagram ↗

Read the main route from left to right; labelled side branches show additional inputs, checks or feedback. The sections below explain the operations and their experimental limits.

The next token is the customer

An autoregressive model keeps keys and values from earlier tokens so it can use them when generating the next one. This growing cache is an attractive compression target. Our operator-based prototype asks whether temporal modes and low-dimensional factors can represent that history compactly. But a beautiful reconstruction of a cache is not the application: useful future attention is.

Put serialization inside the model diagram

An encoder produces a representation; a decoder reconstructs keys and values from its actual stored bytes; attention applies the model’s real query and permitted-key mask. A reduced-precision factor must be rounded before its quality is scored. A data-dependent graph basis must be stored or regenerated without borrowing the original activation tensor.

The prototype crosses those boundaries inconsistently. Its auxiliary queries omit a normalization stage, its attention diagnostic omits causality, and some payloads are priced at a precision not used in reconstruction. A large sweep cannot establish cache quality when the measured consumer is not the real transformer computation.

A compressed cache must answer the real query
A compressed cache must answer the real query. Original scientific diagram; the stated component and information flow, not an additional experiment. Open full-size figure ↗

A small error can face a large question

Consider two almost identical keys. Replace them with the same reconstructed key and the relative tensor error can be made arbitrarily small. Now ask a query that magnifies their small difference. The original attention favors one value; the compressed attention averages the two. The output error stays fixed even as the reconstruction score approaches perfection. The animation implements this deliberately adverse example.

The missing quantity is sensitivity

This does not mean cache compression is hopeless, or that real queries grow without bound. It means a guarantee needs the query, its norm, the allowed attention positions, and the values. For a fixed query, the range of the change in attention logits gives a useful bound. Value reconstruction and key-induced changes in weights enter separately. The right representation is tied to the calculation it must survive.

∥o^−o∥≤ϵV+2Vmax⁡tanh⁡(Ω/4)\begin{gathered}\|\widehat o-o\|\le\epsilon_V+2V_{\max}\tanh(\Omega/4)\end{gathered}
Value error contributes directly. Omega is the range of the key-induced logit errors for the same permitted attention positions.

What the prototype actually measured

The archive contains 259,200 diagnostic rows across ten files, exploring polynomial and exponential modes, low-rank factors, graphs, and quantization. Source inspection found that the output diagnostic omitted a causal mask and the extracted GPT-2 queries omitted pre-attention normalization. Several byte estimates also charged reduced precision without reconstructing rounded factors. Those records cannot establish a usable cache codec, regardless of an attractive row in the table.

A smaller, sound result survives

We preserve the original experiments and add deterministic tests of the attention counterexample and masking discrepancy. No language model is rerun and no failed benchmark is rescued by searching another configuration. The resulting technical note identifies exactly what the representation-to-query interface must preserve. It is a diagnostic contribution, not a claim of beating current cache systems.

Analytic two-token counterexample from the stated equations. Relative key error shrinks, but attention-output discrepancy stays at tanh(1/2). Query magnitude changes with epsilon; this is not measured transformer performance.
Analytic two-token counterexample from the stated equations. Relative key error shrinks, but attention-output discrepancy stays at tanh(1/2). Query magnitude changes with epsilon; this is not measured transformer performance. Open full-size figure ↗

Compress for causal generation

The useful contribution is the query-sensitive error analysis and the interface it demands. A key tensor can look close in relative norm while a particular attention output changes substantially. A future codec should therefore be judged on decoded causal generation, actual memory, and latency—not on a promising reconstruction leaderboard alone.

Compress for the operation, not the picture

The same idea extends beyond language models: a compact scientific memory should protect the questions that will be asked of it. An operator description could still help if it preserves query-sensitive directions more efficiently than ordinary low-rank projection or quantization. Here that remains a hypothesis. Demonstrating it requires actual encoded bytes and causal model outputs, not only tensor agreement.

Evidence & further reading

The links below distinguish the project record from foundational literature. This revised story does not add a new application-validation experiment.

  1. Operator-native activation compression: prototype documentation. Local project archive (2026). Local archive snapshot.
  2. Rethinking Attention with Performers. Krzysztof Choromanski et al. (2020). Primary literature.
  3. Compression that preserves future computation. Spline research archive (2026). Local archive snapshot.