The next token is the customer
An autoregressive model keeps keys and values from earlier tokens so it can use them when generating the next one. This growing cache is an attractive compression target. Our operator-based prototype asks whether temporal modes and low-dimensional factors can represent that history compactly. But a beautiful reconstruction of a cache is not the application: useful future attention is.
Put serialization inside the model diagram
An encoder produces a representation; a decoder reconstructs keys and values from its actual stored bytes; attention applies the model’s real query and permitted-key mask. A reduced-precision factor must be rounded before its quality is scored. A data-dependent graph basis must be stored or regenerated without borrowing the original activation tensor.
The prototype crosses those boundaries inconsistently. Its auxiliary queries omit a normalization stage, its attention diagnostic omits causality, and some payloads are priced at a precision not used in reconstruction. A large sweep cannot establish cache quality when the measured consumer is not the real transformer computation.
A small error can face a large question
Consider two almost identical keys. Replace them with the same reconstructed key and the relative tensor error can be made arbitrarily small. Now ask a query that magnifies their small difference. The original attention favors one value; the compressed attention averages the two. The output error stays fixed even as the reconstruction score approaches perfection. The animation implements this deliberately adverse example.
The missing quantity is sensitivity
This does not mean cache compression is hopeless, or that real queries grow without bound. It means a guarantee needs the query, its norm, the allowed attention positions, and the values. For a fixed query, the range of the change in attention logits gives a useful bound. Value reconstruction and key-induced changes in weights enter separately. The right representation is tied to the calculation it must survive.
What the prototype actually measured
The archive contains 259,200 diagnostic rows across ten files, exploring polynomial and exponential modes, low-rank factors, graphs, and quantization. Source inspection found that the output diagnostic omitted a causal mask and the extracted GPT-2 queries omitted pre-attention normalization. Several byte estimates also charged reduced precision without reconstructing rounded factors. Those records cannot establish a usable cache codec, regardless of an attractive row in the table.
A smaller, sound result survives
We preserve the original experiments and add deterministic tests of the attention counterexample and masking discrepancy. No language model is rerun and no failed benchmark is rescued by searching another configuration. The resulting technical note identifies exactly what the representation-to-query interface must preserve. It is a diagnostic contribution, not a claim of beating current cache systems.
Compress for causal generation
The useful contribution is the query-sensitive error analysis and the interface it demands. A key tensor can look close in relative norm while a particular attention output changes substantially. A future codec should therefore be judged on decoded causal generation, actual memory, and latency—not on a promising reconstruction leaderboard alone.
Compress for the operation, not the picture
The same idea extends beyond language models: a compact scientific memory should protect the questions that will be asked of it. An operator description could still help if it preserves query-sensitive directions more efficiently than ordinary low-rank projection or quantization. Here that remains a hypothesis. Demonstrating it requires actual encoded bytes and causal model outputs, not only tensor agreement.
Evidence & further reading
The links below distinguish the project record from foundational literature. This revised story does not add a new application-validation experiment.
- Operator-native activation compression: prototype documentation. Local project archive (2026). Local archive snapshot.
- Rethinking Attention with Performers. Krzysztof Choromanski et al. (2020). Primary literature.
- Compression that preserves future computation. Spline research archive (2026). Local archive snapshot.