Continual learning · Research & Algorithms

Grow a spline network without disrupting its predictions

What happens inside a model when its resolution increases? We follow a six-edge spline network through fitting, exact expansion, and continued learning—and separate a stable transition from a better final prediction.

Six inputs pass through six spline edges and a summation. Each edge expands from 24 to 96 coefficients.Open full-size figure ↗
Actual six-edge additive architecture. Refinement changes edge resolution, not connectivity. Inset dots schematically indicate grid density.

Follow the information

From input to outcome

This route follows a model checkpoint through growth, rather than depicting refinement as an inference layer. Refinement changes the edge coordinates, not the network wiring. It preserves the function at the transition; subsequent training can change it. Independent testing decides whether extra capacity was useful.

Scroll the diagram horizontally to follow the route. Keyboard: focus the diagram, then use the arrow keys.

Coarse training data → Coarse checkpoint → Exact grid transfer → Expanded network → Scalar prediction. This route follows a model checkpoint through growth, rather than depicting refinement as an inference layer. Refinement changes the edge coordinates, not the network wiring. It preserves the function at the transition; subsequent training can change it. Independent testing decides whether extra capacity was useful.
Information-flow map. 144 → 576 coefficients; no claim that Adam moments were transferred. Original vector schematic based on the method and evidence discussed in this article; signal shapes and icons are illustrative, not additional measurements. Open full-size diagram ↗

Read the main route from left to right; labelled side branches show additional inputs, checks or feedback. The sections below explain the operations and their experimental limits.

A model that can make room for detail

A compact learned response often captures broad trends before it can represent narrow features. Increasing capacity should let the model add detail. But an enlarged model is less useful if enlargement itself unexpectedly changes what was already learned.

For spline networks, capacity lives partly inside the edges. Each edge is a learned function, represented by basis functions and coefficients. A finer grid provides more local freedom without another input or output. Refinement is therefore an attractive ingredient for adaptive models—and raises a precise question: what should the new coefficients be at the instant of growth?

Our experiment separates three effects: transferring the old function, adding capacity, and learning how to use that capacity.

Here is the network

Six scalar inputs feed six spline edges, and a summation node produces one prediction. Each training example supplies a six-dimensional input and a noisy scalar target. The model is not given six separately labeled edge responses.

Each coarse edge stores 24 coefficients: 144 parameters in total. Refinement increases that to 96 per edge, or 576 coefficients, without changing the wiring. Within a cubic edge, one input activates only four neighbors. More coefficients increase resolution across the domain, not the number active at a point.

This is a one-layer additive network, not a deep KAN. It is nonlinear in its inputs but linear in its coefficients. That simplicity isolates grid transfer. Individual edge offsets are not uniquely identified: moving a constant between edges leaves their sum unchanged.

Follow one checkpoint through three operations

First, train the coarse model for 500 full-batch Adam updates on 4,096 examples. The synthetic target contains broad oscillations, a narrow localized feature, and a high-frequency component. Training labels have 1% noise; 4,096 separate test inputs have clean labels.

Second, expand every edge from the same coarse checkpoint. The arms use continuous projection, coefficient-array interpolation, values at new knots treated as coefficients, or zero restart.

Third, train each arm for 250 further updates with a fresh optimizer. This does not test transfer of Adam moments. We report five paired seeds. The diagram separates an instantaneous representation change from the subsequent learning.

The archived experiment separates fitting, transfer, and continued learning. Every fine-grid arm has a fresh optimizer.Open full-size figure ↗
The archived experiment separates fitting, transfer, and continued learning. Every fine-grid arm has a fresh optimizer.

Transfer the function, not the array

Spline coefficients generally are not samples of the represented function. Interpolating their array can change the output even if the new grid looks like a denser version of the old one.

Continuous projection instead asks which function in the new space best approximates the old function over the domain. The cross-Gram matrix measures overlap between the two bases; the new-space Gram matrix converts those overlaps into coefficients. For aligned nested grids, the old function already belongs to the fine space, giving zero approximation error.

For dyadic cubic refinement, a short classical mask performs the same embedding more cheaply than a dense solve. Projection also handles nonnested changes and coarsening. Our independent checks agree with classical refinement and full-rank sampled refitting. Exact preservation is not something existing KAN methods are intrinsically unable to achieve.

cfine=Gfine−1Cfine,coarseccoarse\begin{gathered}c_{\mathrm{fine}}=G_{\mathrm{fine}}^{-1}C_{\mathrm{fine,coarse}}c_{\mathrm{coarse}}\end{gathered}
Implemented as a linear solve. Nested spaces preserve the function; otherwise this is continuous least-squares projection.

What changes at the instant of growth?

Exact transfer changes the trained edge functions by a median relative amount of 1.97 × 10⁻¹⁵: roundoff. Coefficient interpolation changes them by 2.63% and raises immediate test loss by 5.67%. Zero restart discards the inherited function, increasing loss by roughly 71 times.

These are transition results, not final accuracy. They matter when another component immediately consumes a prediction and when a learning comparison should start from the same function. They do not establish a safety guarantee. The values-as-coefficients control even slightly improves median immediate loss while changing the function.

Below, reveal the actual archived test-error checkpoints after expansion. The historical record retained learning curves, not every coefficient state; the animation does not invent those missing states.

ARCHIVED TRAINING CHECKPOINTS

Growth is not the end of learning.

Reveal actual saved test-error curves. All arms have the same fine-grid capacity; their initial functions differ.

Median over five archived seeds. Curves reveal saved ten-update checkpoints, not reconstructed model states.

New capacity, not a magical initialization

After 250 updates, the three nonzero warm starts reach almost identical test errors. Exact transfer cleans up the transition but does not win the endpoint comparison by itself. That is more useful than selecting a dramatic transient and calling it better learning.

The large improvement comes from giving a deliberately unresolved target enough resolution. Coarse median test MSE is about 0.0201; refinement reaches roughly 2.09 × 10⁻⁵. That change primarily reflects capacity, not the transfer operator.

Continuous curvature regularization gives a smaller independent benefit: median final error falls from 2.26 × 10⁻⁵ to 2.09 × 10⁻⁵, improving all five paired seeds. Integrating the squared curvature measures the represented function, rather than the size of its coefficient array. Its meaning survives an exact refinement.

R(c)=∫01∣fc′′(x)∣2 dx=cTQc\begin{gathered}R(c)=\int_0^1|f_c^{\prime\prime}(x)|^2\,dx=c^TQc\end{gathered}
Exact refinement preserves this energy because it preserves the function.
Archived learning curves and paired regularization resultsOpen full-size figure ↗
Five archived seeds. Shading is the observed minimum-to-maximum range, not a confidence interval. The paired panel isolates regularization from the larger capacity effect.

The experiment where growth was wrong

An earlier smooth target was already fitted well. Exact refinement followed by further training worsened median test error from 5.02 × 10⁻⁶ to 2.24 × 10⁻⁵. The transition itself still preserved the function.

Exact transfer prevents a numerical disturbance; it does not decide whether extra freedom will fit signal or noise. A practical adaptive system needs an independent growth decision and perhaps a way to coarsen. Those decisions remain unvalidated here.

Because this additive model is linear in its coefficients, regularized least squares is also a natural fitting alternative. Adam studies a controlled checkpoint transition; it is not evidence that iterative optimization is necessary or optimal.

A continuous meaning for changing a model

The useful abstraction is a representation change with a specified effect. Nested refinement preserves a function; nonnested transfer gives its best approximation under a declared continuous metric; coarsening trades capacity for quantified error. Continuous regularization measures the same physical quantity before and after refinement.

For this additive architecture, integral-preserving edge projection has another consequence: under independent uniform inputs, integrated squared output disturbance equals the sum of the edge disturbances. The paper derives and numerically checks this relation. Correlated inputs and changing hidden coordinates require different analysis.

The vision is an adaptive model whose resolution can evolve without confusing coordinate changes with learning. We have tested the transition machinery and its limits. A system that decides where and when to grow is the next layer to build.

Evidence & further reading

The companion paper includes the complete protocol, derivation, archived measurements and an independent refinement check.

  1. Consolidated research results, including constitutive edges and continual memory. Daniel Schmitter (2026). Local archive snapshot.
  2. An Inner-Product Calculus for Periodic Functions and Curves. Anaïs Badoual, Daniel Schmitter and Michael Unser (2016). Primary literature.
  3. Operator-spline theory: consolidated research manuscript. Daniel Schmitter (2026). Local archive snapshot.