Compression & memory · Research & Algorithms

Merge learning histories without repeatedly recompressing them

Distributed learning histories should combine without paying a new approximation penalty at every branch. Fixed parameter nodes give a surprisingly simple way to preserve that property.

EXPLORE THE IDEA

Merge without another approximation

Fixed parameter nodes survive an order-preserving tree.

CHRONOLOGICAL BLOCKSRetain the same parameter nodes01121023120324104123Block values add component by component.MERGE RESULTParenthesization does not resample(A + B) + (C + D)A + (B + (C + D))Both give [7, 9, 7]Time order remains part of the state contract
50%
Illustrative three-node vectors from four chronological blocks. Their component-wise sum is identical under the two displayed parenthesizations. This does not permit rearranging time-dependent state transitions.

Follow the information

From input to outcome

At common nodes, ordinary affine/quadratic composition combines blocks directly. Parenthesization may change; temporal order may not. Hermite derivatives follow the product rule, while finite-precision rounding remains a separate error source.

Scroll the diagram horizontally to follow the route. Keyboard: focus the diagram, then use the arrow keys.

Chronological blocks → Block state + loss maps → Compose in time order → Merged nodal record → Later objective query. At common nodes, ordinary affine/quadratic composition combines blocks directly. Parenthesization may change; temporal order may not. Hermite derivatives follow the product rule, while finite-precision rounding remains a separate error source.
Information-flow map. Associative chronological algebra, not commutative shuffling of history. Original vector schematic based on the method and evidence discussed in this article; signal shapes and icons are illustrative, not additional measurements. Open full-size diagram ↗

Read the main route from left to right; labelled side branches show additional inputs, checks or feedback. The sections below explain the operations and their experimental limits.

A memory becomes more useful when it can travel

Suppose a sensor produces one compact learning record per hour. A recipient may want to join hours, devices may upload in batches, or a long record may be processed as a tree. If every merge reconstructs an approximate function and compresses it again, approximation errors can become a property of the merge schedule rather than the data.

Merge time blocks without inventing a new projection

A block maps its incoming state to an outgoing state and adds a quadratic loss. To append another block, substitute the first state map into the second loss. This yields an associative chronological algebra. Evaluating that algebra at common parameter nodes gives the same nodal values whether blocks are merged sequentially or in a balanced tree in exact arithmetic.

Hermite data obey the same rule when derivatives are propagated through sums and products. The guarantee does not apply to every compressed representation: projecting intermediate products into an arbitrary low-order space can lose terms needed by the final product. Rounding at checkpoints is another, separate source of error.

Store a family of future fitting problems
Store a family of future fitting problems. Original scientific diagram; the stated component and information flow, not an additional experiment. Open full-size figure ↗

Preserve the coordinates where composition is exact

Each block in our scalar-filter example describes a state transition and a quadratic loss. At a fixed time-constant node, combining two blocks is ordinary affine/quadratic algebra. If every block retains values at the same nodes, that algebra acts directly on the retained values. We interpolate only to answer queries between them.

A small algebraic fact does the heavy lifting

Evaluation at a point preserves sums and products. Consequently, any parenthesization that preserves chronology gives the same exact node values as processing the entire record directly. Hermite value-and-derivative pairs also work because derivatives obey the product rule. The representation’s interpolation error remains, but repeated merging does not introduce a fresh real-arithmetic interpolation error at every internal node.

E(fg)=E(f)⊙E(g),E(f)=(f(τ1),…,f(τm))\begin{gathered}E(fg)=E(f)\odot E(g),\quad E(f)=(f(\tau_1),\ldots,f(\tau_m))\end{gathered}
Fixed-node evaluation commutes with products. That is the reason node-based block algebra survives merging; chronology still matters.

Not every projection has this property

A generic least-squares projection can discard a component that later multiplication would make important. Projecting intermediate products can therefore differ from projecting their final product. The practical lesson is to preserve defining data that respect the operation you need. Merely calling a representation a spline or a low-dimensional summary does not guarantee safe composition.

What the implementation verified

The experiment compares direct capture, sequential merging, and balanced merging of 32 chronological blocks. Their node-based summaries agree near floating-point roundoff. A separate checkpoint panel repeatedly rounds and restores the stored object through a million samples and reports all four precision/representation arms. This is evidence for a particular stable scalar mechanism, not arbitrary distributed neural training.

A composable record of computation

This is the foundation for a mergeable computational record. Sensors or processing chunks can retain compatible objects and combine them later, while the order of time remains meaningful. The archive verifies the algebra and finite-precision behavior for the scalar family; it does not establish unrestricted distributed neural training from deleted data.

Three kinds of error stay separate

Exact composition algebra, between-node approximation, and machine roundoff are distinct. The first can be proved, the second bounded under analytic assumptions, and the third tested or enclosed. Keeping those distinctions makes a mergeable learning record intelligible. It also explains why a duration-independent mean-loss bound does not automatically protect a posterior built from a growing total likelihood.

Evidence & further reading

The links below distinguish the project record from foundational literature. This revised story does not add a new application-validation experiment.

  1. Compression that preserves future computation. Spline research archive (2026). Local archive snapshot.
  2. Retunable objective memory: study 02 findings. Spline research archive (2026). Local archive snapshot.
  3. PASS-GLM: polynomial approximate sufficient statistics for scalable Bayesian GLM inference. Jonathan H. Huggins, Ryan P. Adams and Tamara Broderick (2017). Primary literature.