The architecture in context
The system we are building
A cache exists to serve attention, so its distortion should be measured at that operation as well as at the tensor level. The metric utility computes outputs from original and reconstructed keys/values with fixed queries, then compares relative error and cosine similarity. This can reveal a damaging direction that a global reconstruction norm underweights.
Who does what in the stack
- NumPy matrix products
- Implement the attention proxy.
- Stable softmax arithmetic
- Avoids unnecessary exponential overflow.
- Relative norm / cosine
- Report complementary output distortions.
The custom NumPy proxy isolates one head’s computation from the full transformer. Subtracting each row’s largest logit stabilizes exponentiation without changing the softmax probabilities. The utility is small enough to test against an independently implemented reference.
Open up the implementation
Open the attention-output error metric
Subtracting the row maximum before exponentiation stabilizes softmax without changing its mathematical output. Causal masking, however, changes which keys are visible. The selected helper does not apply a causal mask, so its output should not be presented as the exact GPT-2 causal attention result.
The mathematical contract
A compressed V can change output directly; a compressed K changes the weights through softmax. These are different error pathways. Relative error also becomes unstable when the reference norm is near zero, so the denominator convention needs to be explicit.
Implementation and resource card
- Capacity / budget
- No optimization or extra model parameters. This metric allocates attention scores and must declare mask semantics.
- Execution evidence
- This revision inspects and explains the archived implementation. It does not rerun the original workload. No unrecorded convergence time, throughput or accelerator result is supplied.
- Current reproduction context
- Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.
From explanation to a reproducible check
Build a three-token example where the final token has a huge value. Verify earlier causal outputs cannot depend on it in a corrected causal metric. Report the existing unmasked proxy and the proposed causal version as different metrics.
Preserve input identities, configuration and failure records with the result. A successful numerical check only establishes the operation it exercises: it does not certify an entire dataset, model or deployed system. Reproduce the interface on a small deterministic input before optimizing throughput or increasing workload size.
A closer look at the implementation
The code that carries the idea
The excerpt has no causal mask. Every query can attend to every key in the supplied matrix, including later token positions. That may be a legitimate noncausal proxy, but it is not the full causal GPT-2 attention operation. Query consistency must also be fixed at collection time.
def attention_output(q: np.ndarray, k: np.ndarray, v: np.ndarray) -> np.ndarray:
d = q.shape[1]
logits = (q @ k.T) / math.sqrt(max(d, 1))
logits = logits - np.max(logits, axis=1, keepdims=True)
exps = np.exp(logits)
probs = exps / np.sum(exps, axis=1, keepdims=True)
return probs @ v
def attention_output_error(q, k, v, k_hat, v_hat):
base = attention_output(q, k, v)
comp = attention_output(q, k_hat, v_hat)
rel = relative_fro_error(base, comp)
cos = float(np.dot(base.ravel(), comp.ravel()) / ((np.linalg.norm(base.ravel()) * np.linalg.norm(comp.ravel())) + 1e-12))
return rel, cosVerbatim archive excerpt from metrics.py. Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.
The boundary that matters
Output cosine can remain high while magnitudes drift; a relative norm can become unstable near a zero reference. Neither metric establishes language-model quality or generation speed. The archived CSV must therefore be interpreted as proxy evidence.