Compression · E48 · Implementation audit

Collect a transformer’s internal state at the correct boundary

A KV-cache experiment needs more than hidden-state tensors. Queries, keys and values must come from the same actual attention computation.

Hugging Face TransformersPyTorchCustom GPT-2 adapter
A hook before normalization observes a different tensor from the attention projection’s actual input, changing a reconstructed query.
Figure 1. Where exactly did you hook the model?. A hook before normalization observes a different tensor from the attention projection’s actual input, changing a reconstructed query. Q/K/V boundary schematic. Original vector illustration.

Follow the information

From input to outcome

Q reconstruction must use the same normalized attention input as the model. K/V and Q are sibling observations of the attention computation, not a sequence in which the cache generates the query.

Q reconstruction must use the same normalized attention input as the model. K/V and Q are sibling observations of the attention computation, not a sequence in which the cache generates the query.
Figure 2. Information flow. Solid arrows carry observations, tensors or artifacts; other routes are explicitly labelled. Signal shapes, matrices and network icons are schematic, not measured samples or literal neuron counts. Open full-size SVG ↗ On narrow screens, scroll the diagram horizontally.

Read this alongside Figure 1: A hook before normalization observes a different tensor from the attention projection’s actual input, changing a reconstructed query. The module map and layer-level figures below expand the operations in this route.

Collect a transformer’s internal state at the correct boundary: architecturePrompt tokens: Bounded input length → Pretrained GPT-2: Evaluation / no_grad → KV cache extraction: Layers × heads × tokens × d → Query reconstruction: Attention input boundary → Offline record: Compression diagnostics. A high-level module map; comparison branches and training details are explained in the article.COMPRESSION / E48 / MODULE MAP01 INPUTPrompt tokensBounded input length02 MODULEPretrained GPT-2Evaluation / no_grad03 MODULEKV cache extractionLayers × heads × tokens × d04 MODULEQuery reconstructionAttention input boundary05 OUTPUTOffline recordCompression diagnostics
Source-grounded module map. Boxes summarize operations, not individual neurons; comparison arms and training paths are detailed below. On a small screen, scroll the diagram horizontally.
Prompt tokens — Bounded input length

The architecture in context

The system we are building

The collector runs a pretrained causal language model with cache and hidden-state outputs enabled, detaches the returned tensors and records them for offline compression experiments. Keeping model execution separate from codec evaluation makes experiments repeatable without rerunning the transformer for every candidate representation.

Who does what in the stack

Hugging Face Transformers
Loads the model and exposes cache/hidden-state outputs.
PyTorch
Reshapes, detaches and moves tensors for offline analysis.
Custom GPT-2 adapter
Reconstructs queries and must match the normalization boundary.

The custom collector adds a GPT-2-specific query reconstruction path because a KV cache does not include Q. It reshapes projected queries into per-head tensors and includes a fallback attention reconstruction with a causal mask. Hugging Face supplies model loading and forward outputs; PyTorch supplies tensor operations.

Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph.
Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph. Open full-size SVG ↗

Open up the implementation

Trace the exact boundary of a transformer hook

A concrete operation-level view of this implementation; no unobserved neural architecture is implied.
A concrete operation-level view of this implementation; no unobserved neural architecture is implied. Open full-size SVG ↗

A hook observes a specific tensor in a model’s computation graph. The collector’s reconstructed Q path omits the block’s ln_1 normalization before c_attn. That means its reconstructed Q is not the same tensor as the model’s true pre-attention query. Using it to score K/V distortion can produce a neatly computed but mismatched attention proxy.

The mathematical contract

[Q,K,V]=LN⁡(H)WQKV+bQKV[Q,K,V]=\operatorname{LN}(H)W_{QKV}+b_{QKV}

Layer, head, prompt, sequence length and dtype define an activation sample. Averaging many rows from the same prompt does not create many independent prompts. Frozen model collection still incurs the entire encoder/decoder weight and activation cost; only the codec’s incremental state can be called small.

Implementation and resource card

Capacity / budget
GPT-2 is a frozen upstream model; collection adds no trained parameters. max_tokens 256 is a cap, not proof that every prompt has 256 tokens.
Execution evidence
This revision inspects and explains the archived implementation. It does not rerun the original workload. No unrecorded convergence time, throughput or accelerator result is supplied.
Current reproduction context
Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.

From explanation to a reproducible check

Compare the captured Q with a hook at the actual normalized projection input on a fixed short prompt. Assert equality before benchmarking codecs. Preserve the existing flawed proxy as historical evidence and label a corrected collector as a new version.

Preserve input identities, configuration and failure records with the result. A successful numerical check only establishes the operation it exercises: it does not certify an entire dataset, model or deployed system. Reproduce the interface on a small deterministic input before optimizing throughput or increasing workload size.

A closer look at the implementation

The code that carries the idea

The excerpt applies block.attn.c_attn directly to the recorded hidden state. GPT-2’s attention input normally passes through the block’s first layer normalization. Omitting that step means the reconstructed queries need not match the model’s actual queries, even if every tensor has the expected shape.

Python · file · lines 38–46
def _extract_gpt2_queries(model, hidden_states: List[torch.Tensor]) -> List[torch.Tensor]:
    out = []
    for l, block in enumerate(model.transformer.h):
        h_in = hidden_states[l]  # [B, T, d_model]
        qkv = block.attn.c_attn(h_in)
        q, _, _ = qkv.split(block.attn.split_size, dim=2)
        bsz, t, _ = q.shape
        q = q.view(bsz, t, block.attn.num_heads, block.attn.head_dim).permute(0, 2, 1, 3).contiguous()
        out.append(q[0].detach())

Verbatim archive excerpt from collect_activations.py. Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.

The boundary that matters

Stored Q/K/V consistency must be verified before interpreting an attention-error metric. The later metric utility also has its own masking boundary, discussed in E51. This article does not download weights, execute prompts or claim an improved language model.

Keep building

Other posts of interest