Evaluation practice · E61 · Evaluation

Before improving the network, test the representation

A saved text-as-image comparison shows close visual training results—but the straightforward raw-text baseline is stronger.

Saved JSON metricsTfidfVectorizerLogisticRegression
The saved comparison places visual models beside a raw-text control; improvements inside the visual pipeline must be judged against that representation choice.
Figure 1. Before the network, test the input. The saved comparison places visual models beside a raw-text control; improvements inside the visual pipeline must be judged against that representation choice. Archived accuracies; two runs where available. Original vector illustration.

Follow the information

From input to outcome

Raw text and rendered pixels follow different model paths. Local reverse updates and standard backprop share the visual representation; TF-IDF tests whether that representation was useful in the first place.

Raw text and rendered pixels follow different model paths. Local reverse updates and standard backprop share the visual representation; TF-IDF tests whether that representation was useful in the first place.
Figure 2. Information flow. Solid arrows carry observations, tensors or artifacts; other routes are explicitly labelled. Signal shapes, matrices and network icons are schematic, not measured samples or literal neuron counts. Open full-size SVG ↗ On narrow screens, scroll the diagram horizontally.

Read this alongside Figure 1: The saved comparison places visual models beside a raw-text control; improvements inside the visual pipeline must be judged against that representation choice. The module map and layer-level figures below expand the operations in this route.

Before improving the network, test the representation: system and evaluation mapSame document task: Fourteen categories → Visual branch: Rendered text + CNN → Training comparison: Local reverse / backprop → Random-feature control: Gram readout → Raw-text branch: TF-IDF + logistic → Saved accuracy: Interpret representation cost. A high-level module map; comparison branches and training details are explained in the article.EVALUATION PRACTICE / E61 / MODULE MAP01 INPUTSame document taskFourteen categories02 MODULEVisual branchRendered text + CNN03 MODULETraining comparisonLocal reverse / backprop04 MODULERandom-feature controlGram readout05 MODULERaw-text branchTF-IDF + logistic06 OUTPUTSaved accuracyInterpret representation cost
Source-grounded module map. Boxes summarize operations, not individual neurons; comparison arms and training paths are detailed below. On a small screen, scroll the diagram horizontally.
Same document task — Fourteen categories

The architecture in context

What this comparison asks

The saved experiment separates two questions: does the custom reverse schedule train the visual network, and is rendering text into pixels a good representation for this task? The local-reverse and backpropagation arms are close in the recorded runs, while the text-native control reaches the highest recorded accuracy.

Who does what in the stack

Saved JSON metrics
Retains visual training runs and controls.
TfidfVectorizer
Learns a training-only word n-gram representation.
LogisticRegression
Fits the text-native classifier.

The result file preserves both visual runs and a random-feature Gram control. The code excerpt comes from the raw-text comparator. It uses a fitted TF-IDF vocabulary and logistic regression, despite a stale comment that describes ridge and character-plus-word features. The actual constructor uses word unigrams and bigrams.

Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph.
Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph. Open full-size SVG ↗

From module map to executable structure

Inside Rendered-text convolutional classifier

This results or evaluation article shares the implementation in E60. The architecture below describes that companion, not a newly trained model.

Rendered grayscale 48×192 canvas;14 classes; four blocks of widths 32,64,128,128.

Layer-level implementation. B denotes batch size; parameter and shape conventions are expanded in the table.
Layer-level implementation. B denotes batch size; parameter and shape conventions are expanded in the table. Open full-size SVG ↗
Layer / tensor / operation ledger
Layer or branchOutput shapeImplementation detail
Render textB × 1 × 48 × 192Fixed font, wrapping and clipping are part of preprocessing.
Two-convolution block 32B × 32 × 24 × 96Each 3×3 SAME conv → BatchNorm → ReLU, then average pool 2.
Blocks 64 and 128B × 128 × 6 × 24Same two-convolution pattern; spatial resolution halves after each block.
Fourth block 128B × 128 × 3 × 12Two more convolutions and pool 2.
Pool + class headB × 14 logitsAdaptive average 1×1 → flatten 128 → Linear 128→14.

The explicit sweep changes how reverse-mode differentiation is scheduled, not the mathematical supervised objective. It still propagates a global loss gradient backward through every block. This is unlike greedy local contrastive training, which uses independent block losses and freezes preceding modules.

The equation and the update

dlogits=(softmax⁡(z)−onehot⁡(y))/B,dl−1=JlTdld_{\rm logits}=(\operatorname{softmax}(z)-\operatorname{onehot}(y))/B,\qquad d_{l-1}=J_l^Td_l

Adam 2e-3, decay 5e-4, batch 128, cosine learning-rate schedule; configured 30 epochs,2,000 examples per class and two seeds. Compare ordinary backward with an explicit reverse VJP sweep using the same cross-entropy gradient.

Learning or solution path. A parameter-update path is different from the forward inference path; see text for target-network, frozen-feature and local-loss boundaries.
Learning or solution path. A parameter-update path is different from the forward inference path; see text for target-network, frozen-feature and local-loss boundaries. Open full-size SVG ↗

Implementation card / no invented benchmarks

Capacity, budget and execution evidence

Parameters / retained state
584,814 scalars, excluding BatchNorm running buffers.
Duration and hardware evidence
No elapsed runtime is retained in the selected saved metrics; do not infer training hours from workstation specifications.
Source coordinates
E60 lines 20, 58–80 and 91–105; E61 saved metrics
Current reproduction context
Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.

Counts above are calculated from the stated layer shapes unless identified as saved measurements. They exclude optimizer state and nontrainable buffers. No archived training was rerun for this revision.

What these design choices change

Rendering makes familiar CNN tools applicable but truncates and spatially rearranges text. The TF-IDF control operates on words and word bigrams without paying for a pixel representation. Its stronger saved result is informative: a more elaborate neural input representation is not automatically the better engineering choice.

Reproduction and measurement protocol

Keep a long-text clipping test and a font-availability check next to model tests. Compare gradients before optimizer updates, with identical model state and BatchNorm behavior. Report the word-based logistic-regression baseline as its actual algorithm, not the stale character/ridge description in a comment.

For a new run, save the resolved Python/framework versions, backend, dtype, seed, input shapes, batch size and exact source revision. Start with one batch and one update. Log training steps separately from epochs or environment steps. Do not equate the configured maximum with a completed budget or convergence.

Measure initialization/compilation, data preparation, warmed forward pass, training updates and evaluation separately. Synchronize accelerator work around timed regions using the chosen framework’s supported mechanism. Report peak process memory and framework allocation separately; parameter bytes exclude activations, gradients, optimizer state and input buffers. On a shared machine, begin with a single CPU worker and a small batch rather than claiming all available resources.

A closer look at the implementation

The code that carries the idea

fit_transform learns the vocabulary on training documents; transform applies that vocabulary to test documents. The saved local-reverse accuracies are 0.9342 and 0.9390, backpropagation 0.9342 and 0.9354, random Gram 0.4434 and 0.4379, and TF-IDF 0.9719.

Python · file · lines 121–131
def tfidf_text(txt_tr, ytr, txt_te, yte):
    """Honesty ceiling: TF-IDF char+word n-grams + ridge on the RAW TEXT (what text-native NLP gets)."""
    try:
        from sklearn.feature_extraction.text import TfidfVectorizer
        from sklearn.linear_model import LogisticRegression
    except Exception as e:
        return None
    v = TfidfVectorizer(max_features=50000, ngram_range=(1, 2), sublinear_tf=True)
    Xtr = v.fit_transform(txt_tr); Xte = v.transform(txt_te)
    clf = LogisticRegression(max_iter=300, C=10.0); clf.fit(Xtr, ytr.numpy())
    return round(float((clf.predict(Xte) == yte.numpy()).mean()), 4)

Verbatim archive excerpt from closed_form_neat_textvision.py (companion source E60). Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.

The saved comparison

Archived results, not new training. The article states the comparison’s scope and limitations.

Before improving the network, test the representation — selected recorded values
ArmRun 1Run 2
Visual local reverse0.93420.9390
Visual backprop0.93420.9354
Random visual Gram0.44340.4379
Text TF-IDF0.9719Not separately recorded

The boundary that matters

The visual methods’ closeness is not evidence that either improves on the text-native representation, and two saved runs are not an equivalence test. The result is useful because it prevents optimizer novelty from distracting from a stronger, simpler input representation.

Keep building

Other posts of interest