Representation learning · E60 · Implementation

What happens when a text classifier reads pixels instead of tokens?

A language-as-image experiment turns rendering into a lossy front end and compares two ways of training the resulting CNN.

PillowPyTorch CNNscikit-learn
Rendering text makes layout and clipping part of the input representation before the first convolution is applied.
Figure 1. What if a classifier reads pixels?. Rendering text makes layout and clipping part of the input representation before the first convolution is applied. Illustrative text raster; source canvas ratio. Original vector illustration.

Follow the information

From input to outcome

The label supervises a visual classifier over rendered text. A separate TF-IDF/logistic-regression control consumes the original text directly; it is not a downstream layer of the CNN.

The label supervises a visual classifier over rendered text. A separate TF-IDF/logistic-regression control consumes the original text directly; it is not a downstream layer of the CNN.
Figure 2. Information flow. Solid arrows carry observations, tensors or artifacts; other routes are explicitly labelled. Signal shapes, matrices and network icons are schematic, not measured samples or literal neuron counts. Open full-size SVG ↗ On narrow screens, scroll the diagram horizontally.

Read this alongside Figure 1: Rendering text makes layout and clipping part of the input representation before the first convolution is applied. The module map and layer-level figures below expand the operations in this route.

What happens when a text classifier reads pixels instead of tokens?: architectureRaw text: Document label → Pillow renderer: Wrapped grayscale page → CNN blocks: Visual representation → Training schedule: Backprop or blockwise VJP → Class logits: Fourteen categories → Text-native control: TF-IDF + logistic regression. A high-level module map; comparison branches and training details are explained in the article.REPRESENTATION LEARNING / E60 / MODULE MAP01 INPUTRaw textDocument label02 MODULEPillow rendererWrapped grayscale page03 MODULECNN blocksVisual representation04 MODULETraining scheduleBackprop or blockwise VJP05 MODULEClass logitsFourteen categories06 OUTPUTText-native controlTF-IDF + logistic regression
Source-grounded module map. Boxes summarize operations, not individual neurons; comparison arms and training paths are detailed below. On a small screen, scroll the diagram horizontally.
Raw text — Document label

The architecture in context

The system we are building

The experiment renders text into fixed-size grayscale images and asks a CNN to predict document categories. This removes the tokenizer from the visual branch but introduces a new front end: font choice, wrapping, raster resolution and clipping. The model sees only the information that survives that renderer.

Who does what in the stack

Pillow
Renders wrapped text to grayscale pixels.
PyTorch CNN
Learns visual document features.
scikit-learn
Supplies the text-native TF-IDF/logistic control.

The project adds the text-to-image adapter and a blockwise reverse schedule alongside ordinary backpropagation. The latter still uses PyTorch’s local derivatives; it is not the greedy self-supervised block training in E29. A separate raw-text baseline tests whether the visual representation is useful at all.

Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph.
Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph. Open full-size SVG ↗

From module map to executable structure

Inside Rendered-text convolutional classifier

Rendered grayscale 48×192 canvas;14 classes; four blocks of widths 32,64,128,128.

Layer-level implementation. B denotes batch size; parameter and shape conventions are expanded in the table.
Layer-level implementation. B denotes batch size; parameter and shape conventions are expanded in the table. Open full-size SVG ↗
Layer / tensor / operation ledger
Layer or branchOutput shapeImplementation detail
Render textB × 1 × 48 × 192Fixed font, wrapping and clipping are part of preprocessing.
Two-convolution block 32B × 32 × 24 × 96Each 3×3 SAME conv → BatchNorm → ReLU, then average pool 2.
Blocks 64 and 128B × 128 × 6 × 24Same two-convolution pattern; spatial resolution halves after each block.
Fourth block 128B × 128 × 3 × 12Two more convolutions and pool 2.
Pool + class headB × 14 logitsAdaptive average 1×1 → flatten 128 → Linear 128→14.

The explicit sweep changes how reverse-mode differentiation is scheduled, not the mathematical supervised objective. It still propagates a global loss gradient backward through every block. This is unlike greedy local contrastive training, which uses independent block losses and freezes preceding modules.

The equation and the update

dlogits=(softmax⁡(z)−onehot⁡(y))/B,dl−1=JlTdld_{\rm logits}=(\operatorname{softmax}(z)-\operatorname{onehot}(y))/B,\qquad d_{l-1}=J_l^Td_l

Adam 2e-3, decay 5e-4, batch 128, cosine learning-rate schedule; configured 30 epochs,2,000 examples per class and two seeds. Compare ordinary backward with an explicit reverse VJP sweep using the same cross-entropy gradient.

Learning or solution path. A parameter-update path is different from the forward inference path; see text for target-network, frozen-feature and local-loss boundaries.
Learning or solution path. A parameter-update path is different from the forward inference path; see text for target-network, frozen-feature and local-loss boundaries. Open full-size SVG ↗

Implementation card / no invented benchmarks

Capacity, budget and execution evidence

Parameters / retained state
584,814 scalars, excluding BatchNorm running buffers.
Duration and hardware evidence
No elapsed runtime is retained in the selected saved metrics; do not infer training hours from workstation specifications.
Source coordinates
E60 lines 20, 58–80 and 91–105; E61 saved metrics
Current reproduction context
Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.

Counts above are calculated from the stated layer shapes unless identified as saved measurements. They exclude optimizer state and nontrainable buffers. No archived training was rerun for this revision.

What these design choices change

Rendering makes familiar CNN tools applicable but truncates and spatially rearranges text. The TF-IDF control operates on words and word bigrams without paying for a pixel representation. Its stronger saved result is informative: a more elaborate neural input representation is not automatically the better engineering choice.

Reproduction and measurement protocol

Keep a long-text clipping test and a font-availability check next to model tests. Compare gradients before optimizer updates, with identical model state and BatchNorm behavior. Report the word-based logistic-regression baseline as its actual algorithm, not the stale character/ridge description in a comment.

For a new run, save the resolved Python/framework versions, backend, dtype, seed, input shapes, batch size and exact source revision. Start with one batch and one update. Log training steps separately from epochs or environment steps. Do not equate the configured maximum with a completed budget or convergence.

Measure initialization/compilation, data preparation, warmed forward pass, training updates and evaluation separately. Synchronize accelerator work around timed regions using the chosen framework’s supported mechanism. Report peak process memory and framework allocation separately; parameter bytes exclude activations, gradients, optimizer state and input buffers. On a shared machine, begin with a single CPU worker and a small batch rather than claiming all available resources.

A closer look at the implementation

The code that carries the idea

The excerpt wraps text and draws only the lines that fit the canvas. This is an explicit information bottleneck. Two long documents with the same visible prefix can become identical images even if their later content distinguishes their labels.

Python · file · lines 24–32
def render_text(text):
    from PIL import Image, ImageDraw, ImageFont
    font = ImageFont.truetype(FONT_PATH, 10)
    img = Image.new("L", (W, H), 255); d = ImageDraw.Draw(img)
    lines = textwrap.wrap(text, CHARS)[:LINES]
    d.multiline_text((1, 0), "\n".join(lines), font=font, fill=0, spacing=2)
    return np.asarray(img, dtype=np.float32) / 255.0

Verbatim archive excerpt from closed_form_neat_textvision.py. Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.

The boundary that matters

Removing tokenization does not make the model language-agnostic or more efficient end to end. Rendering and image tensors have their own cost. The raw-text comparator receives more directly accessible lexical information; that difference is part of the representation question.

Keep building

Other posts of interest