Evaluation practice · E52 · Evaluation

Nineteen thousand metric rows are not nineteen thousand experiments

A head-by-head compression table needs hierarchy-aware aggregation before it can support a model-level claim.

CSV schemaNumPy metric implementationGrouped analysis
Metric rows share prompt, layer and head ancestry; a large table is not the same as a large collection of independent trials.
Figure 1. 19,440 rows are not 19,440 trials. Metric rows share prompt, layer and head ancestry; a large table is not the same as a large collection of independent trials. Source row count; schematic hierarchy. Original vector illustration.

Follow the information

From input to outcome

Rows inherit shared prompt, layer and head identities. Aggregation must preserve that dependence; treating every method row as a separate independent experiment inflates the evidence.

Rows inherit shared prompt, layer and head identities. Aggregation must preserve that dependence; treating every method row as a separate independent experiment inflates the evidence.
Figure 2. Information flow. Solid arrows carry observations, tensors or artifacts; other routes are explicitly labelled. Signal shapes, matrices and network icons are schematic, not measured samples or literal neuron counts. Open full-size SVG ↗ On narrow screens, scroll the diagram horizontally.

Read this alongside Figure 1: Metric rows share prompt, layer and head ancestry; a large table is not the same as a large collection of independent trials. The module map and layer-level figures below expand the operations in this route.

Nineteen thousand metric rows are not nineteen thousand experiments: system and evaluation mapSaved metric table: 19,440 rows → Group identities: Prompt / layer / head / method → Matched comparisons: Same source tensor → Prompt-level summaries: Preserve within-prompt dependence → Interpretation: Proxy distortion, not deployment. A high-level module map; comparison branches and training details are explained in the article.EVALUATION PRACTICE / E52 / MODULE MAP01 INPUTSaved metric table19,440 rows02 MODULEGroup identitiesPrompt / layer / head / method03 MODULEMatched comparisonsSame source tensor04 MODULEPrompt-level summariesPreserve within-prompt dependence05 OUTPUTInterpretationProxy distortion, not deployment
Source-grounded module map. Boxes summarize operations, not individual neurons; comparison arms and training paths are detailed below. On a small screen, scroll the diagram horizontally.
Saved metric table — 19,440 rows

The architecture in context

What this comparison asks

The selected CSV contains 19,440 rows spanning methods, heads, layers and prompts. These rows are useful for locating where a representation fails, but they are not independent language-model trials. A fair comparison first pairs methods on the same source tensor and then preserves the prompt-level grouping when summarizing variation.

Who does what in the stack

CSV schema
Retains configuration and metric identities.
NumPy metric implementation
Defines the reported tensor error.
Grouped analysis
Must respect nested prompts, layers and heads.

The saved schema records token count, rank, poles, window size, reconstruction error, attention error, cosine and byte estimates. That makes it possible to inspect accounting alongside distortion rather than reporting whichever column looks most favorable. The companion snippet defines the relative reconstruction metric.

Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph.
Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph. Open full-size SVG ↗

Open up the implementation

Rows, prompts and tokens are different sample counts

A concrete operation-level view of this implementation; no unobserved neural architecture is implied.
A concrete operation-level view of this implementation; no unobserved neural architecture is implied. Open full-size SVG ↗

Multiple rows may share the same activation tensor and differ only in method or budget. Others share a prompt but use different layers or heads. Treating all rows as independent would make uncertainty look artificially small. Pair methods on identical inputs before aggregating.

The mathematical contract

Δp=mean⁡l,h(ep,l,h(A)−ep,l,h(B))\Delta_p=\operatorname{mean}_{l,h}(e_{p,l,h}^{(A)}-e_{p,l,h}^{(B)})

A global average can hide a small number of severely damaged heads or short prompts where metadata dominates. Report distributions and coverage rather than selecting only favorable cases. Estimated compression ratio and attention-proxy fidelity do not establish generation quality or decode throughput.

Implementation and resource card

Capacity / budget
19,440 result rows, not 19,440 independent prompts. Sequence lengths vary despite the 256-token filename.
Execution evidence
This revision inspects and explains the archived implementation. It does not rerun the original workload. No unrecorded convergence time, throughput or accelerator result is supplied.
Current reproduction context
Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.

From explanation to a reproducible check

Group by the full tensor identity before comparing methods. Confirm that no configuration is missing selectively, retain failed cases and compute uncertainty at a defensible independent unit such as prompt—not raw CSV row.

Preserve input identities, configuration and failure records with the result. A successful numerical check only establishes the operation it exercises: it does not certify an entire dataset, model or deployed system. Reproduce the interface on a small deterministic input before optimizing throughput or increasing workload size.

A closer look at the implementation

The code that carries the idea

The first recorded prompt/head example illustrates a tradeoff: int8 has relative reconstruction error about 0.012 and an estimated ratio about 1.99; int4 has error about 0.213 and an estimated ratio about 3.97. These are individual rows, not aggregate winners. Actual token counts vary despite the filename’s 256-token limit.

Python · file · lines 7–9
def relative_fro_error(x: np.ndarray, x_hat: np.ndarray) -> float:
    denom = np.linalg.norm(x, ord="fro") + 1e-12
    return float(np.linalg.norm(x - x_hat, ord="fro") / denom)

Verbatim archive excerpt from metrics.py (companion source E51). Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.

The boundary that matters

Estimated bytes inherit the accounting limitations in E50. Attention results inherit the proxy boundaries in E48 and E51. This article does not turn the table into a new significance test or select a favorable subset of heads.

Keep building

Other posts of interest