The complete collection

One foundation.
Many perspectives.

42 articles and 22 focused technical drafts. Stars identify the 12 flagship stories.

IDArticlePaperAnimated example
B01 ★When physics learning needs a solve, not a training loopM22An explicit architecture separates supplied physics, noisy observations, fitted coordinates, and prediction; interactive studies reveal the role of sensor placement and diffusion.
B02 ★4.2× faster spline-layer evaluation through cardinal structureM01The coefficient-strip interaction is schematic. The 512-coefficient forward timings, 5.435 ms dense and 1.286 ms streamed, are archived MPS measurements for batch 1024, 32 inputs and 64 outputs. The smaller-grid route can lose; the full result plot below retains that crossover.
B03Exact smoothness losses for spline networks—without sampling the integralM01Computed periodic cubic curvature Gram for eight basis functions, integrated with four-point Gauss quadrature per cell (exact for these products up to rounding). The curve and quadratic energy use the displayed Gram matrix. Unit cell spacing; not a neural training-speed benchmark.
B04 ★Grow a spline network without disrupting its predictionsM02The actual six-edge additive network expands its spline resolution; archived learning checkpoints distinguish exact transfer, added capacity, and regularization.
B05Build boundary conditions into a learned model instead of penalizing themM02Two computed cubic Hermite cells share an interior value and derivative. The slider changes that value; both outer values and outer derivatives remain fixed. This illustrates a parameterization, not an optimized physical trajectory.
B06 ★Beyond a direct KAN: learning the physical law that extrapolatesM03The colored x–t field is an analytic schematic, not a PDE rollout. The flux plot evaluates F(u)=u²/2+0.08 sin(4u). The learned-versus-direct comparisons below are archived experiments; the flux dictionary contains this generating component.
B07Learning physical laws from noisy data without differentiating the noiseM03Deterministic synthetic noisy samples and a compact cosine test window. The readout is their normalized weighted average, not an estimated physical law or a denoising-accuracy result.
B08Learning aircraft vibrations with a compact dynamical modelM04AI-generated editorial ground-test scene. The animated sensor channel is a computed damped sinusoid, not a recorded aircraft trace or deformation. The benchmark results discussed below remain separate.
B09Better coordinates or a bigger network?M04Computed free-motion example: x₁(t)=t and x₂(t)=−t. The position view merges the two at t=0; the phase portrait keeps their velocities distinct. No measured robot data is implied.
B10 ★How small can a learned drone model become?M04AI-generated editorial nano-drone. Sizes are archived retained-program counts: 24,288 bytes compact versus approximately 24.79 MB for source-plus-adapted neural programs. They are not total process memory. The adapted ensemble was more accurate. The reveal is not a flight replay.
B11Continual learning without a replay buffer: what can we actually guarantee?M05Exact scalar ridge example with λ=1: target +1 receives weight one and target −1 receives the slider weight w. The optimum is (1−w)/(2+w). Old squared error changes even though no evidence is discarded.
B12 ★Zero forgetting where it can be guaranteedM05Computed compact cubic B-spline update on a fixed coordinate system. Its support is entirely to the right of the protected region; the left-hand field and curve stay unchanged. Not a general neural-network zero-forgetting guarantee.
B13Version control for learned models: commit, roll back and forget explicitlyM05Interactive design sketch of propose, validate, commit and reject states. Curves are computed illustrations; displayed decisions are not logs or a shipped version-control product.
B14Can a frozen encoder keep learning new classes?M06Actual archived final CIFAR-100 accuracies: linear ridge 88.47%, cosine-expanded ridge 89.10%, on the named frozen DINOv2 features. The graphic reveals the incremental difference; it does not fabricate source images or a confusion matrix.
B15Evolve the network; solve the readoutM07Four illustrative graph candidates, not saved evolutionary champions. The parity construction counts in the article are separate archived results. The animation changes topology without inventing a fitness improvement.
B16Is your “local learning” algorithm actually backpropagation?M08Illustrative adjoint propagation δ at one block at a time. The direction and Jacobian-transpose operation describe reverse-mode differentiation, not biological activity or a gradient-free learner.
B17Learning between events instead of stepping through timeM09Computed linear exponential responses to three specified impulses. The cursor reveals the exact state between arrivals; event positions are synthetic and do not represent a recorded spiking network.
B18The same memory dynamics, different results at low precisionM09Computed rounding of the fixed point (0.62,0.34) on a lattice of step 0.25, after rotating coordinates. Both the lattice and decoded point are calculated. This geometric example is not an archived low-precision memory trajectory.
B19Update a robot’s dynamics model without rewriting its entire memoryM10An editorial quadruped rendering and interactive synthetic response curve explain a local model update. The actual control experiment used HalfCheetah-v5 simulation; the pictured robot is not experimental evidence.
B20Why a better dynamics model can still produce a worse controllerM10Actual compound-lag restoration summaries: shared prior 0.8966; supplied true inverse 0.8977; predeclared target 0.90. These are aggregate archived results with the same frozen actor, not a simulated rollout video or an optimal-control bound.
B21Your physics loss went down. Did your predictions improve?M11Computed counterexample: two constant velocity fields (1,0) and (1,0.2) both have zero divergence, but their particle trajectories separate. This is not a replay of the measured turbulence experiment.
B22Discard the measurements, change the prior, reconstruct againM12Actual archived cardinal-32 walnut reconstructions, before and after the separately frozen TV-prior diagnostic, displayed with a common grayscale window. The control blends two stored images; it does not solve a continuum of priors. Diagnostic feasibility limitations remain in the paper. Dataset: Hämäläinen et al., CC BY 4.0.
B23A real CT scan tests our compression ideaM12Interactive comparison of actual archived walnut CT reconstructions, with a common grayscale window. The source is a measured 2D slice, not a 3D or medical-imaging experiment. Dataset: Hämäläinen et al., CC BY 4.0.
B24Why sharing models is not enough for collective intelligenceM13Illustrative grid mission with visit, avoid and ordered inspection states. The missing second-inspection state is explicitly marked; no physical swarm footage or successful collective capability is fabricated.
B25 ★Keep the loss function, discard the training streamM14Computed toy objective for a damped exponential with a variable time constant. The visible data and loss curve are generated from the same deterministic samples. It explains objective retention; measured compression results remain below.
B26Merge learning histories without repeatedly recompressing themM14Illustrative three-node vectors from four chronological blocks. Their component-wise sum is identical under the two displayed parenthesizations. This does not permit rearranging time-dependent state transitions.
B27Recalibrate a physical model after the raw data is goneM15AI-generated editorial oscillator bench with a computed candidate sinusoid. The slider changes frequency, not experimental calibration evidence. The retained-products mechanism and archived physical-memory results appear below.
B28How long can compressed memory remain trustworthy?M15Synthetic query: estimated response 0.5 with a fixed numerical allowance 0.03. The control changes the requested tolerance and therefore acceptance. This is an exact toy decision, not a fitted trajectory through archived checkpoints.
B29Can dynamical models shrink a transformer’s KV cache?M16Computed two-key counterexample: keys (1,±ε), values ±1, query (0,1/ε). Replacing both keys with (1,0) drives key error to zero while attention-output error remains tanh(1). Query norm grows; typical-model behavior is not implied.
B30A physics solver that computes at interfaces instead of everywhereM17Computed source-free solution of −u″+qu=0 in a single cell, with q=4 and a variable right endpoint. Neighboring cells are a schematic interface chain; they do not form a recorded pricing solve.
B31Why exponential models become numerically unstable—and how to fix itM17Computed exponential divided difference versus its derivative limit at t=1, α=−1. The stable value uses expm1; the expanded value uses subtraction in JavaScript binary64. This is a numerical illustration, not a timing benchmark.
B32 ★A useful numerical service on one CPU threadM17Interface replay of the documented native-service contract, not a live request or screen recording. Reported 4.793 ms warm median and 2.219 MiB peak native memory are archived Apple M4 Max measurements, not edge-device or hard-real-time claims.
B33Return an error bound with the predictionM18Computed signed Gaussian readout example: two unit-height bumps approach and cancel. The displayed signed L² norm is obtained by numerical integration of the illustrated functions; it is not the paper’s certified Green-kernel bound.
B34Don’t trust the model’s warning—verify it independentlyM18Protocol illustration: validate a request-bound pair of model fields, recompute observation compatibility, and check response separation. Status transitions illustrate the protocol; they are not live parser logs or market certification.
B35Two models fit the data. Can they still disagree about risk?M19Illustration of a shrinking coefficient layer. Readouts state the analytical limiting relation for a₁=2a₀: the price difference tends to zero while point gamma tends to half the background value. The drawing does not fabricate a finite-width price solution.
B36When a fast, accurate model fails on real dataM19Actual archive counts: 12 homogeneous, 60 clock-study and 20 direct-feasibility proposals; none admitted. Twenty overlapping views are not independent trials. The downstream 120 risk directions remained untested.
B37 ★How much speech can a tiny dynamical model preserve?M20Computed synthetic source–filter spectral illustration, not recorded speech, a listening test or a measured codec rate. Harmonic excitation and two resonance envelopes are explicitly generated; no audio plays.
B38 ★Can a laptop hear you breathe?Report / perspectiveOriginal acoustic-path diagram and computed synthetic modulation. This is an audio motion prototype, not Wi-Fi or validated respiratory measurement. The archive used a 20 kHz carrier and 48 kHz audio sampling. No sound plays.
B39What quantum computing teaches us about representation costReport / perspectiveComputed Gaussian control envelope beside a geometric Bloch-sphere illustration. The pulse is not propagated through a Hamiltonian; the state marker is conceptual. No quantum advantage or experiment is implied.
B40A time-series model that looked promising until we tested it chronologicallyReport / perspectiveSynthetic five-second-bar example with an earlier chart pivot, later confirmation trigger, and a shaded region beyond the current information cutoff. No historical market data, trading outcome or profitability is represented.
B41 ★Can a frozen language model improve without changing its weights?M21Illustrative arithmetic examples evaluated by exact integer multiplication. They demonstrate the stored-final-answer rule, not historical model transcripts (which were not retained). Actual complete-stream counts appear in the article.
B42 ★Why self-improvement stalls—and what changes the outcomeM21Actual verified-memory block counts: 11, 8, 11, 12, 9, 8 correct out of 20; totals 30/60 and 29/60 for the two halves. Revealing blocks does not fit a trend or imply independent replications.