Compression · E50 · Implementation audit

Four-bit arithmetic is not yet a four-bit representation

A tiny quantizer makes the distinction between simulated distortion, resident array storage and a real packed payload impossible to miss.

NumPyCustom accountingRequired deployment addition
Quantization changes representable values; packing changes their physical storage. An int32 array does not become four-bit storage by containing small integers.
Figure 1. Four-bit values need four-bit storage. Quantization changes representable values; packing changes their physical storage. An int32 array does not become four-bit storage by containing small integers. Illustrative uniform quantizer. Original vector illustration.

Follow the information

From input to outcome

Rounding changes values; packing changes storage. The archived integer workspace retains int32 elements, so nominal four-bit arithmetic does not demonstrate a four-bit serialized buffer.

Rounding changes values; packing changes storage. The archived integer workspace retains int32 elements, so nominal four-bit arithmetic does not demonstrate a four-bit serialized buffer.
Figure 2. Information flow. Solid arrows carry observations, tensors or artifacts; other routes are explicitly labelled. Signal shapes, matrices and network icons are schematic, not measured samples or literal neuron counts. Open full-size SVG ↗ On narrow screens, scroll the diagram horizontally.

Read this alongside Figure 1: Quantization changes representable values; packing changes their physical storage. An int32 array does not become four-bit storage by containing small integers. The module map and layer-level figures below expand the operations in this route.

Four-bit arithmetic is not yet a four-bit representation: architectureFloat tensor: One calibration range → Symmetric scale: max |x| / qmax → Round + clip: Integer levels → int32 workspace: Actual temporary representation → Float reconstruction: Distortion measurement → Estimated bytes: Nominal bits + metadata. A high-level module map; comparison branches and training details are explained in the article.COMPRESSION / E50 / MODULE MAP01 INPUTFloat tensorOne calibration range02 MODULESymmetric scalemax |x| / qmax03 MODULERound + clipInteger levels04 MODULEint32 workspaceActual temporary representation05 MODULEFloat reconstructionDistortion measurement06 OUTPUTEstimated bytesNominal bits + metadata
Source-grounded module map. Boxes summarize operations, not individual neurons; comparison arms and training paths are detailed below. On a small screen, scroll the diagram horizontally.
Float tensor — One calibration range

The architecture in context

The system we are building

Uniform quantization maps a tensor onto a finite grid using a scale derived from its largest absolute value. The archived function is concise enough to inspect completely. It is a useful baseline for asking how much distortion a nominal four- or eight-bit grid introduces before investing in a more complicated codec.

Who does what in the stack

NumPy
Calibrates, rounds, clips and reconstructs tensors.
Custom accounting
Estimates a hypothetical low-bit payload.
Required deployment addition
A real bitstream format and decoder, not supplied by this function.

The function owns per-tensor calibration, clipping and a metadata estimate. NumPy performs rounding and reconstruction. There is no packed encoder/decoder, low-bit matrix kernel or persistent compressed object in this excerpt.

Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph.
Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph. Open full-size SVG ↗

Open up the implementation

An integer code is not yet a packed buffer

A concrete operation-level view of this implementation; no unobserved neural architecture is implied.
A concrete operation-level view of this implementation; no unobserved neural architecture is implied. Open full-size SVG ↗

The scale maps the largest magnitude to the chosen code range. One extreme outlier can enlarge the step for every element. Rounding reduces precision even though the reconstructed array may still use float 32 storage. Quantization error and physical compression are therefore distinct results in this implementation.

The mathematical contract

q=clip⁡(round⁡(x/s),−qmax⁡−1,qmax⁡),x^=sqq=\operatorname{clip}(\operatorname{round}(x/s),-q_{\max}-1,q_{\max}),\qquad \widehat x=sq

Per-channel or grouped scaling can reduce distortion at the price of additional scale metadata and a different kernel interface. Those are possible extensions, not measurements established here. A4-bit claim needs a packing layout, odd-length handling and a decoder that actually consumes that layout.

Implementation and resource card

Capacity / budget
Estimated payload:nbits/8+8 bytes for n values. The implementation uses int 32 intermediates and float 32 reconstruction; it does not store packed 4-bit codes.
Execution evidence
This revision inspects and explains the archived implementation. It does not rerun the original workload. No unrecorded convergence time, throughput or accelerator result is supplied.
Current reproduction context
Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.

From explanation to a reproducible check

Test zeros, a single outlier, positive and negative endpoints, and odd element counts. For nine 4-bit codes a packed payload needs five bytes before metadata. Compare the estimated payload, serialized payload and runtime array nbytes separately.

Preserve input identities, configuration and failure records with the result. A successful numerical check only establishes the operation it exercises: it does not certify an entire dataset, model or deployed system. Reproduce the interface on a small deterministic input before optimizing throughput or increasing workload size.

A closer look at the implementation

The code that carries the idea

The quantized values are stored as np.int32 and immediately reconstructed as float32. Meanwhile bytes_used is calculated as element_count × nominal_bits / 8 plus eight metadata bytes. That estimate describes an assumed representation; it is not q.nbytes and does not measure peak memory.

Python · file · lines 6–14
def uniform_quantize(x: np.ndarray, bits: int = 8):
    qmax = (1 << (bits - 1)) - 1
    max_abs = np.max(np.abs(x))
    scale = 1.0 if max_abs == 0 else max_abs / qmax
    zero_point = 0
    q = np.clip(np.round(x / scale), -qmax - 1, qmax).astype(np.int32)
    x_hat = (q.astype(np.float32) * scale).astype(np.float32)
    bytes_used = q.size * bits / 8 + 8  # scale + zero-point metadata
    return x_hat, {"bits": bits, "scale": float(scale), "zero_point": int(zero_point), "bytes": float(bytes_used)}

Verbatim archive excerpt from quant.py. Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.

The boundary that matters

Actual four-bit packing must round odd element counts up to a whole byte and specify signed coding, scale precision and format overhead. Small bit depths, empty tensors and nonfinite values also need validation. A zero tensor is handled by selecting scale one.

Keep building

Other posts of interest