Compression · E51 · Implementation audit

Evaluate compression at the operation that consumes it

Attention-output error is more task-connected than tensor error—but only if the proxy matches the attention semantics you intend to preserve.

NumPy matrix productsStable softmax arithmeticRelative norm / cosine
The cover illustrates causal consumer semantics. The archived proxy below is an unmasked row-softmax helper, so the causal mask is a requirement to consider—not a feature already implemented.
Figure 1. The consumer defines the error. The cover illustrates causal consumer semantics. The archived proxy below is an unmasked row-softmax helper, so the causal mask is a requirement to consider—not a feature already implemented. Illustrative causal mask. Original vector illustration.

Follow the information

From input to outcome

The same attention operation is repeated with reconstructed K/V while Q stays fixed. This archived helper does not apply a causal mask; the cover illustrates why consumer semantics matter, not a mask implemented in this proxy.

The same attention operation is repeated with reconstructed K/V while Q stays fixed. This archived helper does not apply a causal mask; the cover illustrates why consumer semantics matter, not a mask implemented in this proxy.
Figure 2. Information flow. Solid arrows carry observations, tensors or artifacts; other routes are explicitly labelled. Signal shapes, matrices and network icons are schematic, not measured samples or literal neuron counts. Open full-size SVG ↗ On narrow screens, scroll the diagram horizontally.

Read this alongside Figure 1: The cover illustrates causal consumer semantics. The archived proxy below is an unmasked row-softmax helper, so the causal mask is a requirement to consider—not a feature already implemented. The module map and layer-level figures below expand the operations in this route.

Evaluate compression at the operation that consumes it: architectureQ, K, V tensors: One head → Scaled dot products: QKᵀ / √d → Row softmax: Numerically stabilized → Output: softmax(scores) V → Reconstructed K / V: Repeat same operation → Error + cosine: Compare outputs. A high-level module map; comparison branches and training details are explained in the article.COMPRESSION / E51 / MODULE MAP01 INPUTQ, K, V tensorsOne head02 MODULEScaled dot productsQKᵀ / √d03 MODULERow softmaxNumerically stabilized04 MODULEOutputsoftmax(scores) V05 MODULEReconstructed K / VRepeat same operation06 OUTPUTError + cosineCompare outputs
Source-grounded module map. Boxes summarize operations, not individual neurons; comparison arms and training paths are detailed below. On a small screen, scroll the diagram horizontally.
Q, K, V tensors — One head

The architecture in context

The system we are building

A cache exists to serve attention, so its distortion should be measured at that operation as well as at the tensor level. The metric utility computes outputs from original and reconstructed keys/values with fixed queries, then compares relative error and cosine similarity. This can reveal a damaging direction that a global reconstruction norm underweights.

Who does what in the stack

NumPy matrix products
Implement the attention proxy.
Stable softmax arithmetic
Avoids unnecessary exponential overflow.
Relative norm / cosine
Report complementary output distortions.

The custom NumPy proxy isolates one head’s computation from the full transformer. Subtracting each row’s largest logit stabilizes exponentiation without changing the softmax probabilities. The utility is small enough to test against an independently implemented reference.

Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph.
Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph. Open full-size SVG ↗

Open up the implementation

Open the attention-output error metric

A concrete operation-level view of this implementation; no unobserved neural architecture is implied.
A concrete operation-level view of this implementation; no unobserved neural architecture is implied. Open full-size SVG ↗

Subtracting the row maximum before exponentiation stabilizes softmax without changing its mathematical output. Causal masking, however, changes which keys are visible. The selected helper does not apply a causal mask, so its output should not be presented as the exact GPT-2 causal attention result.

The mathematical contract

O=softmax⁡(QKT/d+M)V,eO=∥O−O^∥F∥O∥FO=\operatorname{softmax}(QK^T/\sqrt d+M)V,\qquad e_O=\frac{\|O-\widehat O\|_F}{\|O\|_F}

A compressed V can change output directly; a compressed K changes the weights through softmax. These are different error pathways. Relative error also becomes unstable when the reference norm is near zero, so the denominator convention needs to be explicit.

Implementation and resource card

Capacity / budget
No optimization or extra model parameters. This metric allocates attention scores and must declare mask semantics.
Execution evidence
This revision inspects and explains the archived implementation. It does not rerun the original workload. No unrecorded convergence time, throughput or accelerator result is supplied.
Current reproduction context
Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.

From explanation to a reproducible check

Build a three-token example where the final token has a huge value. Verify earlier causal outputs cannot depend on it in a corrected causal metric. Report the existing unmasked proxy and the proposed causal version as different metrics.

Preserve input identities, configuration and failure records with the result. A successful numerical check only establishes the operation it exercises: it does not certify an entire dataset, model or deployed system. Reproduce the interface on a small deterministic input before optimizing throughput or increasing workload size.

A closer look at the implementation

The code that carries the idea

The excerpt has no causal mask. Every query can attend to every key in the supplied matrix, including later token positions. That may be a legitimate noncausal proxy, but it is not the full causal GPT-2 attention operation. Query consistency must also be fixed at collection time.

Python · file · lines 12–26
def attention_output(q: np.ndarray, k: np.ndarray, v: np.ndarray) -> np.ndarray:
    d = q.shape[1]
    logits = (q @ k.T) / math.sqrt(max(d, 1))
    logits = logits - np.max(logits, axis=1, keepdims=True)
    exps = np.exp(logits)
    probs = exps / np.sum(exps, axis=1, keepdims=True)
    return probs @ v


def attention_output_error(q, k, v, k_hat, v_hat):
    base = attention_output(q, k, v)
    comp = attention_output(q, k_hat, v_hat)
    rel = relative_fro_error(base, comp)
    cos = float(np.dot(base.ravel(), comp.ravel()) / ((np.linalg.norm(base.ravel()) * np.linalg.norm(comp.ravel())) + 1e-12))
    return rel, cos

Verbatim archive excerpt from metrics.py. Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.

The boundary that matters

Output cosine can remain high while magnitudes drift; a relative norm can become unstable near a zero reference. Neither metric establishes language-model quality or generation speed. The archived CSV must therefore be interpreted as proxy evidence.

Keep building

Other posts of interest