Compression · E47 · Evaluation

A compressed memory needs an answer budget as well as a byte budget

A bounded streaming capsule keeps a small state, but its numerical assurance can fail before memory runs out.

Native accumulatorPython / NumPy serializationInterval query machinery
The source byte counts compare retained representations, not whole-process memory or arbitrary access to the original observations.
Figure 1. Retained bytes are only one budget. The source byte counts compare retained representations, not whole-process memory or arbitrary access to the original observations. Source byte counts; schematic size comparison. Original vector illustration.

Follow the information

From input to outcome

A later query uses only the retained capsule. Its answer must carry the declared uncertainty bound or be rejected; the capsule is not a lossless copy of every observation.

A later query uses only the retained capsule. Its answer must carry the declared uncertainty bound or be rejected; the capsule is not a lossless copy of every observation.
Figure 2. Information flow. Solid arrows carry observations, tensors or artifacts; other routes are explicitly labelled. Signal shapes, matrices and network icons are schematic, not measured samples or literal neuron counts. Open full-size SVG ↗ On narrow screens, scroll the diagram horizontally.

Read this alongside Figure 1: The source byte counts compare retained representations, not whole-process memory or arbitrary access to the original observations. The module map and layer-level figures below expand the operations in this route.

A compressed memory needs an answer budget as well as a byte budget: system and evaluation mapObservation blocks: At most 4,096 values → Native accumulator: 144,016 allocated bytes → Retained capsule: About 12.7 KB with metadata → Later model query: Likelihood intervals → Answer gate: Posterior bound or rejection. A high-level module map; comparison branches and training details are explained in the article.COMPRESSION / E47 / MODULE MAP01 INPUTObservation blocksAt most 4,096 values02 MODULENative accumulator144,016 allocated bytes03 MODULERetained capsuleAbout 12.7 KB with metadata04 MODULELater model queryLikelihood intervals05 OUTPUTAnswer gatePosterior bound or rejection
Source-grounded module map. Boxes summarize operations, not individual neurons; comparison arms and training paths are detailed below. On a small screen, scroll the diagram horizontally.
Observation blocks — At most 4,096 values

The architecture in context

What this comparison asks

The physical-memory study retains selected statistics for a declared family of later queries instead of preserving arbitrary raw history. Its final streaming panel separates acquisition memory, retained bytes, process memory and answer accuracy. This is a practical compression pattern: specify what future questions the representation must still support.

Who does what in the stack

Native accumulator
Maintains bounded live statistics.
Python / NumPy serialization
Stores arrays with schema and uncertainty metadata.
Interval query machinery
Checks the declared answer budget after capture.

The project combines a native accumulator with Python query and verification machinery. The excerpt comes from the companion capsule serializer, where numeric payload, schema metadata, uncertainty radii and an integrity hash travel together. A retained capsule is a query snapshot, not a restart image of the full live accumulator.

Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph.
Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph. Open full-size SVG ↗

Open up the implementation

Budget retained state and query cost separately

A concrete operation-level view of this implementation; no unobserved neural architecture is implied.
A concrete operation-level view of this implementation; no unobserved neural architecture is implied. Open full-size SVG ↗

The compact state preserves a specified family of queries with an associated error bound. It does not retain arbitrary historical access. Serialization includes metadata, radii and integrity information; those bytes belong in the comparison. A query that needs preparation can cost much more than its final dot product.

The mathematical contract

∥f−f^∥≤εonly under the declared query norm and domain\|f-\widehat f\|\leq\varepsilon\quad\text{only under the declared query norm and domain}

Tighter tolerance can increase retained state or reject a stream. The saved panel has failed as well as passed cases. Repeated prefixes from the same streams are correlated evidence, not independent datasets. Compact representation and fast end-to-end service are two separately measured objectives.

Implementation and resource card

Capacity / budget
Selected artifacts:144,016 native bytes versus about 12,700 retained bytes. Peak process RSS is over 100 MiB; retained bytes are not whole-process memory.
Execution evidence
This revision inspects and explains the archived implementation. It does not rerun the original workload. No unrecorded convergence time, throughput or accelerator result is supplied.
Current reproduction context
Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.

From explanation to a reproducible check

Round-trip the serialized state and verify query bounds against a reference for the declared query family. Report capture, preparation and steady query time separately, alongside retained bytes and peak RSS.

Preserve input identities, configuration and failure records with the result. A successful numerical check only establishes the operation it exercises: it does not certify an entire dataset, model or deployed system. Reproduce the interface on a small deterministic input before optimizing throughput or increasing workload size.

A closer look at the implementation

The code that carries the idea

The final report records capsules of 12,700–12,703 bytes and fixed native allocation of 144,016 bytes. All eight million-sample cases pass the declared posterior bound; only two of eight four-million-sample cases do. The complete scale-up gate therefore fails despite small measured discrepancies against the floating reference.

Python · file · lines 60–68
        E+=self.n*point(offset)**2
        return outward(E)[1]
    def dumps(self):
        payload=np.r_[self.acf,self.head,self.tail].astype('<f8').tobytes()
        meta={'schema':'pairwise_enclosed_lag_v1','n':self.n,'lag':self.lag,'dt':self.step,
            'total':self.total,'energy_upper':self.energy_upper,'acf_radius':self.acf_radius,
            'total_radius':self.total_radius,'dtype':'<f8','noise_variance':.04,
            'input_nonzero_range':[2.**-200,2.**200],'arithmetic':'binary64_nearest_pairwise_no_reassociation'}
        header=json.dumps(meta,sort_keys=True,separators=(',',':')).encode()

Verbatim archive excerpt from certified_capture.py (companion source X02). Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.

The boundary that matters

Allocated native bytes are not whole-process RSS or an embedded-device measurement. The final paired prefixes represent four underlying streams, not sixteen independent physical recordings. The checksum protects integrity, not the numerical truth of the enclosed answer.

Keep building

Other posts of interest