Evaluation practice · E32 · Negative result

A smaller student is not automatically a compressed teacher

Parameter accounting and capability retention are separate tests. A two-run distillation record makes both visible.

PyTorch model definitionSaved metricsGram / classifier evaluation
The two saved runs show a substantial teacher–student accuracy gap under the feature-plus-Gram readout; fewer parameters alone do not establish retention.
Figure 1. Smaller is not the same as preserved. The two saved runs show a substantial teacher–student accuracy gap under the feature-plus-Gram readout; fewer parameters alone do not establish retention. Archived feature + Gram accuracies. Original vector illustration.

Follow the information

From input to outcome

A teacher reference and compact students are evaluated on the same task. Comparisons must preserve readout choice: feature-plus-Gram scores are not interchangeable with learned classifier-head scores.

A teacher reference and compact students are evaluated on the same task. Comparisons must preserve readout choice: feature-plus-Gram scores are not interchangeable with learned classifier-head scores.
Figure 2. Information flow. Solid arrows carry observations, tensors or artifacts; other routes are explicitly labelled. Signal shapes, matrices and network icons are schematic, not measured samples or literal neuron counts. Open full-size SVG ↗ On narrow screens, scroll the diagram horizontally.

Read this alongside Figure 1: The two saved runs show a substantial teacher–student accuracy gap under the feature-plus-Gram readout; fewer parameters alone do not establish retention. The module map and layer-level figures below expand the operations in this route.

A smaller student is not automatically a compressed teacher: system and evaluation mapTeacher features: Reference representation → Compact student: 210,596 parameters → Training arms: Distillation / scratch → Readout choices: Gram or learned classifier → Paired evaluation: Capability retained?. A high-level module map; comparison branches and training details are explained in the article.EVALUATION PRACTICE / E32 / MODULE MAP01 INPUTTeacher featuresReference representation02 MODULECompact student210,596 parameters03 MODULETraining armsDistillation / scratch04 MODULEReadout choicesGram or learned classifier05 OUTPUTPaired evaluationCapability retained?
Source-grounded module map. Boxes summarize operations, not individual neurons; comparison arms and training paths are detailed below. On a small screen, scroll the diagram horizontally.
Teacher features — Reference representation

The architecture in context

What this comparison asks

The saved compact-student experiment asks whether a lightweight CNN can inherit a stronger representation. It reports parameter count alongside downstream scores and feature MSE. That is the right direction for compression evaluation: a byte or parameter reduction only becomes useful when the desired capability survives.

Who does what in the stack

PyTorch model definition
Provides the student parameter structure.
Saved metrics
Retains both runs rather than selecting the better one.
Gram / classifier evaluation
Separates feature quality from head training.

The primary artifact is a metrics record; the displayed companion code defines the student architecture. Keeping those roles distinct avoids presenting a network diagram as evidence that the network worked. The count includes the archived model’s parameters, not a measured packed deployment footprint.

Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph.
Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph. Open full-size SVG ↗

From module map to executable structure

Inside Compact distillation student

This results or evaluation article shares the implementation in E31. The architecture below describes that companion, not a newly trained model.

CIFAR-100 images 3×32×32; teacher feature width 512; student classifier 100 classes.

Layer-level implementation. B denotes batch size; parameter and shape conventions are expanded in the table.
Layer-level implementation. B denotes batch size; parameter and shape conventions are expanded in the table. Open full-size SVG ↗
Layer / tensor / operation ledger
Layer or branchOutput shapeImplementation detail
Conv 3→32 + ReLU + pool 2B × 32 × 16 × 163×3, stride 1, padding 1.
Conv 32→64 + ReLU + pool 2B × 64 × 8 × 83×3, stride 1, padding 1.
Conv 64→128 + ReLUB × 128 × 8 × 8Adaptive average pool 1×1 → flatten 128.
Feature projectionB × 512Linear 128→512; raw output is the feature-MSE target space.
Optional classifierB × 100ReLU(features) → Linear 512→100. Unused by feature-only distillation.

The teacher features are cached targets, so the teacher is not in the student’s backward graph. The feature projection is linear; applying its later classifier ReLU before computing MSE would change the experiment. Downstream Gram evaluation uses frozen features and must be reported separately from the student’s direct classifier accuracy.

The equation and the update

Ldistill=1512B∑i∥fθ(xi)−ti∥22\mathcal L_{\rm distill}=\frac{1}{512B}\sum_i\|f_\theta(x_i)-t_i\|_2^2

Adam 1e-3, batch 256; default 8 epochs, two seeds. Feature-only loss is MSE; hybrid adds supervised cross-entropy; scratch uses cross-entropy alone. These objectives update different parts of the same allocated model.

Learning or solution path. A parameter-update path is different from the forward inference path; see text for target-network, frozen-feature and local-loss boundaries.
Learning or solution path. A parameter-update path is different from the forward inference path; see text for target-network, frozen-feature and local-loss boundaries. Open full-size SVG ↗

Implementation card / no invented benchmarks

Capacity, budget and execution evidence

Parameters / retained state
210,596 allocated scalars; 51,300 classifier scalars receive no gradient in feature-only mode. The feature path has 159,296.
Duration and hardware evidence
The training function returns elapsed time, but E32’s JSON does not retain it; no duration is supplied as if measured.
Source coordinates
E31 lines 34–42, 71–91; E32 saved metrics
Current reproduction context
Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.

Counts above are calculated from the stated layer shapes unless identified as saved measurements. They exclude optimizer state and nontrainable buffers. No archived training was rerun for this revision.

What these design choices change

Global pooling makes the feature projection much smaller than flattening the last spatial map. A512-wide representation is chosen to match the teacher target, not because every student needs 512 features. Reducing it requires a new teacher projection or a different distillation objective.

Reproduction and measurement protocol

Assert teacher/student row identity before comparing features; cached targets in a different shuffle order create a valid-shaped but meaningless loss. Check classifier gradients are absent in feature-only mode. The saved teacher-to-student accuracy gap shows that parameter reduction did not preserve the teacher’s usefulness.

For a new run, save the resolved Python/framework versions, backend, dtype, seed, input shapes, batch size and exact source revision. Start with one batch and one update. Log training steps separately from epochs or environment steps. Do not equate the configured maximum with a completed budget or convergence.

Measure initialization/compilation, data preparation, warmed forward pass, training updates and evaluation separately. Synchronize accelerator work around timed regions using the chosen framework’s supported mechanism. Report peak process memory and framework allocation separately; parameter bytes exclude activations, gradients, optimizer state and input buffers. On a shared machine, begin with a single CPU worker and a small batch rather than claiming all available resources.

A closer look at the implementation

The code that carries the idea

The student uses convolutional channels 32, 64 and 128, then a 512-dimensional projection and a 100-class head. Saved teacher-feature Gram accuracies are 0.5862 and 0.5866; the distilled student’s Gram accuracies are 0.3257 and 0.3312. Scratch classification reaches 0.3545 and 0.3675.

Python · file · lines 34–42
class StudentCNN(nn.Module):
    def __init__(self, out=512, nclass=100):
        super().__init__()
        self.enc = nn.Sequential(nn.Conv2d(3, 32, 3, 1, 1), nn.ReLU(), nn.MaxPool2d(2),
                                 nn.Conv2d(32, 64, 3, 1, 1), nn.ReLU(), nn.MaxPool2d(2),
                                 nn.Conv2d(64, 128, 3, 1, 1), nn.ReLU(), nn.AdaptiveAvgPool2d(1), nn.Flatten())
        self.feat = nn.Linear(128, out); self.cls = nn.Linear(out, nclass)
    def features(self, x): return self.feat(self.enc(x))
    def forward(self, x): return self.cls(torch.relu(self.features(x)))

Verbatim archive excerpt from closed_form_distill_v3_compact.py (companion source E31). Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.

The saved comparison

Archived results, not new training. The article states the comparison’s scope and limitations.

A smaller student is not automatically a compressed teacher — selected recorded values
ArmRun 1Run 2
Teacher features + Gram0.58620.5866
Distilled student + Gram0.32570.3312
Scratch classifier0.35450.3675

The boundary that matters

The student did not preserve teacher capability in this setting, and the reported distillation arm does not establish an advantage over scratch classification. Two saved runs are not enough to characterize a broad distillation frontier. Feature MSE also needs its normalization convention to be interpretable.

Keep building

Other posts of interest