The architecture in context
The system we are building
The collector runs a pretrained causal language model with cache and hidden-state outputs enabled, detaches the returned tensors and records them for offline compression experiments. Keeping model execution separate from codec evaluation makes experiments repeatable without rerunning the transformer for every candidate representation.
Who does what in the stack
- Hugging Face Transformers
- Loads the model and exposes cache/hidden-state outputs.
- PyTorch
- Reshapes, detaches and moves tensors for offline analysis.
- Custom GPT-2 adapter
- Reconstructs queries and must match the normalization boundary.
The custom collector adds a GPT-2-specific query reconstruction path because a KV cache does not include Q. It reshapes projected queries into per-head tensors and includes a fallback attention reconstruction with a causal mask. Hugging Face supplies model loading and forward outputs; PyTorch supplies tensor operations.
Open up the implementation
Trace the exact boundary of a transformer hook
A hook observes a specific tensor in a model’s computation graph. The collector’s reconstructed Q path omits the block’s ln_1 normalization before c_attn. That means its reconstructed Q is not the same tensor as the model’s true pre-attention query. Using it to score K/V distortion can produce a neatly computed but mismatched attention proxy.
The mathematical contract
Layer, head, prompt, sequence length and dtype define an activation sample. Averaging many rows from the same prompt does not create many independent prompts. Frozen model collection still incurs the entire encoder/decoder weight and activation cost; only the codec’s incremental state can be called small.
Implementation and resource card
- Capacity / budget
- GPT-2 is a frozen upstream model; collection adds no trained parameters. max_tokens 256 is a cap, not proof that every prompt has 256 tokens.
- Execution evidence
- This revision inspects and explains the archived implementation. It does not rerun the original workload. No unrecorded convergence time, throughput or accelerator result is supplied.
- Current reproduction context
- Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.
From explanation to a reproducible check
Compare the captured Q with a hook at the actual normalized projection input on a fixed short prompt. Assert equality before benchmarking codecs. Preserve the existing flawed proxy as historical evidence and label a corrected collector as a new version.
Preserve input identities, configuration and failure records with the result. A successful numerical check only establishes the operation it exercises: it does not certify an entire dataset, model or deployed system. Reproduce the interface on a small deterministic input before optimizing throughput or increasing workload size.
A closer look at the implementation
The code that carries the idea
The excerpt applies block.attn.c_attn directly to the recorded hidden state. GPT-2’s attention input normally passes through the block’s first layer normalization. Omitting that step means the reconstructed queries need not match the model’s actual queries, even if every tensor has the expected shape.
def _extract_gpt2_queries(model, hidden_states: List[torch.Tensor]) -> List[torch.Tensor]:
out = []
for l, block in enumerate(model.transformer.h):
h_in = hidden_states[l] # [B, T, d_model]
qkv = block.attn.c_attn(h_in)
q, _, _ = qkv.split(block.attn.split_size, dim=2)
bsz, t, _ = q.shape
q = q.view(bsz, t, block.attn.num_heads, block.attn.head_dim).permute(0, 2, 1, 3).contiguous()
out.append(q[0].detach())Verbatim archive excerpt from collect_activations.py. Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.
The boundary that matters
Stored Q/K/V consistency must be verified before interpreting an attention-error metric. The later metric utility also has its own masking boundary, discussed in E51. This article does not download weights, execute prompts or claim an improved language model.