# Flagship stories and visual direction ## Positioning Lead with consequential ML capabilities and the strongest qualified results: physical extrapolation beyond the tested direct KAN, efficient spline-layer execution, structure-exact physics representation, scoped zero-forgetting updates, and verification-guided improvement. The mathematical toolbox explains the result; it should not displace the result from the headline. The [publication map](PUBLICATION_MAP.md) remains the evidence inventory. This brief changes its editorial hierarchy: zero-forgetting and self-improvement belong among the principal stories, not only in a reserve list. Boldness belongs in the question and demonstrated capability, with the comparator and conditions specified in the opening paragraph and figure captions. ## Flagship lineup These titles are proposed editorial framing, not newly validated experiments. Existing blog/manuscript IDs are retained where possible. | Flagship | Proposed title | Evidence and claim boundary | Animated example | Manuscript connection | | --- | --- | --- | --- | --- | | F01 | **Beyond a direct KAN: learning the physical law that extrapolates** | In the controlled oscillatory-law experiment, median amplitude-OOD rollout nRMSE is 9.29e-7 for the selected reproduction atlas versus 0.330 for the tested direct KAN. This uses known outer physics and a favorable candidate library; it is not architecture-wide KAN superiority. | Train inside a shaded amplitude range, then drive beyond it; display reference, direct KAN, polynomial control and operator-law rollouts with the recovered flux below. | B06 / M03 | | F02 | **4.2× faster spline-layer evaluation through cardinal structure** | Archived MPS batch-1024, 32-input, 64-output, 512-knot comparison: 4.23× forward and 2.95× forward/backward versus dense explicit-cardinal evaluation. Smaller grids can favor dense execution. This is not a 4.2× end-to-end training claim. | Reveal four active coefficients versus full basis materialization, then switch grid size to reveal the measured crossover. | B02 / M01 | | F03 | **When physics learning needs a solve, not a training loop** | Matched synthetic Helmholtz/null-space reproduction reaches numerical precision under its stated sampling, operator and boundary assumptions. Exact satisfaction alone does not establish recovery of an arbitrary PDE solution. Classical direct/spectral methods are essential context. | Fit a wave from sufficient observations; reveal that every candidate already satisfies the homogeneous equation, then show what changes when the operator is misspecified. | Expanded B01 / M01, with M03 distinguishing unknown-law learning | | F04 | **Zero forgetting where it can be guaranteed** | Protected-support coefficient updates preserve a declared continuous input region; immutable versions preserve their own programs. Pooled Gram accumulation alone does not preserve old-task predictions. | Learn a new local feature while the protected curve stays fixed; contrast this with a sample-only constraint that changes the curve between samples. | B12–B13 / M05 | | F05 | **Can a frozen language model improve without changing its weights?** | The archived 120-problem multiplication stream gives 59/120 correct with execution-filtered exemplars, 37/120 without memory and 16/120 with unfiltered exemplars. This is an in-context memory/control result, not open-ended RSI or a spline-specific gain. | A problem enters three lanes: no memory, all prior answers, and verified answers only; show what enters the next prompt and final aggregate accuracy. | Promote R03 into the core series; candidate M21 below | | F06 | **Why self-improvement stalls—and what changes the outcome** | Fixed-feature self-labeling, simulated verifier quality, trainable representations and teaching produce different outcomes in the archived panels. Do not convert these bounded experiments into a universal impossibility theorem. | Show separate curves for confidence filtering, an oracle-style checker and externally supplied teaching; label where each source of information comes from. | Companion to F05 / candidate M21 | | F07 | **Keep the loss function, discard the training stream** | Compact summaries retain a fixed family of stable-filter loss and terminal-state queries; both Hermite and Chebyshev pass the recorded panel. Not arbitrary neural-network retraining from a capsule. | Discard the waveform, change a permitted time constant and compare the retained-query answer with the full-record reference. | B25–B26 / M14 | | F08 | **How small can a learned drone model become?** | Strong measured storage/fitting advantages coexist with worse accuracy than the matched adapted neural comparator. No unmeasured drone flight, MCU or energy claim. | Recorded-flight trajectory replay next to a three-axis accuracy/storage/fitting-cost comparison, not staged footage implying new autonomous control. | B10 / M04 | | F09 | **Grow a spline network without disrupting its predictions** | Exact nested transport preserves the represented function in the tested additive periodic model. Growth still requires validation; the smooth-target case worsens when capacity is increased unnecessarily. | Insert knots while the function remains stationary; compare with heuristic transfer and show the eventual validation gate. | B04 / M02 | | F10 | **Can a laptop hear you breathe? A contactless acoustic-sensing prototype** | Consumer-audio code and a local report describe a short recording and a respiration-band peak. Independent reference agreement and robustness are not established; the existing medical-grade and uncertainty-limit claims are not adopted. | Show the physical speaker–reflector–microphone path beside a clearly labeled replay of the recorded signal; a future reference trace is added only after actual synchronized acquisition. | Promote B38 to a practical flagship; technical case study first | | F11 | **How much speech can a tiny dynamical model preserve?** | Existing source–filter and codebook prototypes support an engineering walkthrough, not a verified competitive bitrate/quality result. | Switch between original and reconstructed audio while the excitation, poles and bit-accounting display explain what is retained. | B37 / M20 | | F12 | **A useful numerical service on one CPU thread** | The recorded native calculator and verifier run on this workstation with measured request costs and process memory. This is a working engineering artifact, not proof of profitable trading or low-power-device performance. | A screen recording shows a request, native result, validation and real process measurements; visually distinguish the calculator from the verifier. | B32–B34 / M17–M18 | F01, F02 and F03 are different claims. F01 is a learning/representation result; F02 is an implementation result; F03 is a structure-exact matched problem. Their numbers must not migrate into one another's headlines. In particular, the CPU Cox comparison is not a neural inner-product speed measurement. The existing continuous-curvature penalty also belongs in the ML story: the additive-KAN refinement study already applies an analytic curvature Gram. Its recorded five-seed benefit is approximately a 1.10 geometric factor versus unregularized fine-grid training in that case. This supports an exact functional-regularization article; it does not establish a tenfold training speedup. See the [main results](../../paper/v2_sections/04_results.tex), constitutive-edge subsection starting at line 1447. ## Self-improvement: a concrete paper candidate and necessary corrections Add **M21 — Verification-Guided In-Context Improvement: Memory, Feedback and the Limits of Self-Training** as a consolidation candidate, not a ready claim of recursive learning-rule improvement. It deserves a distinct question from M05's retention guarantees and M06's continual classification. The existing reserve post R03 becomes a core flagship rather than expanding the list merely to meet a quota. The multiplication evidence was checked against the saved boolean outcomes and the current runner, not only the manuscript: | Arm | Correct / 120 | Accuracy | | --- | ---: | ---: | | No memory | 37 | 30.83% | | Execution-filtered exemplars | 59 | 49.17% | | Unfiltered exemplars | 16 | 13.33% | The runner uses the latest four stored exemplars, not nearest-neighbor retrieval despite the older prose. Verification checks the parsed final answer against integer multiplication, not every intermediate reasoning step. The exemplar store is an ordinary growing list, not a demonstration of spline/Gram memory. It evaluates a frozen `qwen2.5vl:7b` model with temperature zero and an external execution signal; that signal is information, so “without any supervision” would be misleading. The verified arm is 50.0% in the first half and 48.33% in the second half. Thus the aggregate lift does not demonstrate progressive compounding across this stream. A faithful animation must not draw a steadily rising curve that the record does not contain. There is also no learned change to the improvement algorithm itself. “Toward recursive self-improvement” is a motivation; “verification-guided in-context improvement” describes this experiment. The runner catches request exceptions as empty responses, and the metric file does not retain full prompts/responses or explicit failure provenance. Before publishing a quantitative flagship, audit these details, problem duplication, ordering and control comparability, and locate the original run configuration. Do not silently rerun or rescue a revealed experiment. Any fresh confirmation needs a frozen protocol and should be distinguishable from this archive. The self-training manuscripts also contain overly general claims such as “only teaching” can improve representations. Their observations should be stated as results of the tested protocols, not universal theorems about self-supervised learning. Likewise, an unchanged exemplar list does not establish invariant future answers from a language model. Sources: [runner](../../bio_growth/self_improve_llm_verifier.py), [saved outcomes](../../bio_growth/closed_form_neat_outputs/metrics_self_improve_llm_verifier.json), [scoreboard](../../RESULTS_SCOREBOARD.md), [main self-improvement section](../../paper/v2_sections/04_results.tex), and [negative-result appendix](../../paper/v2_sections/A3_negatives.tex). ## Reference articles and visual lessons The article structures were read; selected published assets were inspected visually. Full rendered-page layout and responsive behavior remain unreviewed because browser discovery returned no available browser. This is an asset and editorial review, not a claim to have watched every embedded video or inspected the live site's typography and spacing. | Google Research reference | Observed material | Design lesson for this series | | --- | --- | --- | | [Learning Better Simulation Methods for Partial Differential Equations](https://research.google/blog/learning-better-simulation-methods-for-partial-differential-equations/), July 23, 2019 | Article plus three frames of its 101-frame Burgers GIF: baseline and learned method share axes, colors and a visible simulation time. | A synchronized comparison can carry the result more directly than an architecture diagram. Borrow the explanatory principle, not the asset. | | [Titans + MIRAS: Helping AI have long-term memory](https://research.google/blog/titans-miras-helping-ai-have-long-term-memory/), December 4, 2025 | Conceptual hero artwork and three-layer architecture diagram visually inspected; article exposes paper links and a looping-video section. That video was not played. | Give the hero, mechanism diagram and empirical figure different jobs. The hero attracts; the diagram explains; the results substantiate. | | [Image Compression with Neural Networks](https://research.google/blog/image-compression-with-neural-networks/), September 29, 2016 | Article and magnified three-way reconstruction comparison inspected. | Use a meaningful close-up to make a technical tradeoff visible; explain both improved artifacts and lost detail. | | [An All-Neural On-Device Speech Recognizer](https://research.google/blog/an-all-neural-on-device-speech-recognizer/), March 12, 2019 | Article and video/diagram captions read; its video was not played. | Lead with a user-observable capability, then connect it to model architecture and deployment constraints. | These observations concern communication, not independent endorsement of every technical assertion in those posts. In particular, adopt the clarity and layering without copying overgeneralized biological or learning claims. ## Proposed visual language This is an original design recommendation, not a measured specification of Google's site. The editorial mix has three equal purposes: research results that change how a model learns; engineering results that change its resource cost; and practical demonstrations that make the capability tangible. A breathing-sensing prototype or audible vocoder can be more immediately understandable than a benchmark plot. Practical work should not be relegated to a miscellaneous section merely because it is signal processing rather than a large neural network. - Light neutral background, dark readable text, generous whitespace and strong typographic hierarchy. Keep body lines around 65–75 characters as an initial design target; let figures extend wider than the prose. - One distinctive hero composition per story: waves becoming an operator, a model growing around a protected region, or verified examples entering memory. Avoid a generic glowing brain on every post. - A small consistent palette: blue/teal for the proposed method, charcoal for reference, amber for a competing method/error. Supplement color with line patterns and labels; do not encode success only as green versus red. - Flat, precise diagrams with a restrained use of depth. Do not reproduce Google's logos, exact artwork or branding. The research should have its own recognizable visual identity. - A concise title, one-sentence explanatory deck, author/date, and visible Paper / Code / Results links. The reader should understand the result and its scope before encountering detailed mathematics. - Prefer a single purposeful visual per section over dense dashboard layouts. Each caption should state the comparison, observation and relevant condition. ## Media package for each flagship 1. **Hero image:** an original conceptual graphic that still works as a social preview or static poster, with no unsupported performance number embedded. 2. **Short silent loop:** roughly 8–15 seconds explaining one mechanism or comparison; provide play/pause and a useful static fallback. 3. **Result figure:** real archived data with units, comparator and conditions; include unfavorable cases or the crossover where relevant. 4. **Optional 30–60-second video:** for a trajectory, multi-stage workflow or narrated capability demonstration. Recorded-data visualization is labeled as such, not presented as a new physical deployment. 5. **Optional interaction:** one meaningful control, such as grid resolution, protected interval, time constant or query amplitude. Do not add interaction merely to decorate the page. SVG/Canvas is appropriate for explanatory diagrams and small interactive models; prerecorded video can keep heavy result replays cheap for readers. Use MP4/WebM with a poster and reduced-motion fallback for longer loops rather than making large GIF downloads the production default. These are proposed delivery choices; no page or media implementation has been created yet. Scientific animations must obey the same accounting as figures. Synchronize simulation time when comparing predictions; label separately if comparing wall-clock runtime. Do not make one method look slow by animating it slowly. Keep axes and error scales comparable, expose train/test ranges, distinguish schematic motion from measured results, and do not turn selected examples into a claim about the complete cohort. ## First storyboards ### F01: physical extrapolation beyond the tested direct KAN - Opening: two models agree inside the observed state-amplitude band. - Transition: an input crosses outside the band; show the archived competing rollouts and a reference on common scales. - Mechanism: reveal the known PDE structure and the learned flux curve. - Result: show all registered seeds, the polynomial control and the unmatched law/noise limitations, not only the most dramatic oscillatory example. - Closing: paper link and an explicit sentence on supplied physics and library. The comparison should not insinuate that a KAN could not receive the same physics. The key scientific question is where learning is placed, not whether one architecture has universal inferiority. ### F04: zero forgetting within a protected region - Opening: a learned function and a clearly shaded protected interval. - Update: new observations arrive outside it; only permitted coefficients move. - Contrast: an ordinary update changes old predictions; a sample-nullspace update hides drift between samples; support protection preserves the region. - Boundary: conflicting demands inside the protected region require rejection, an explicit version or a changed retention policy. Do not imply unlimited plasticity, automatic context routing or retention of all downstream behavior. The restriction is the mechanism, not fine print. ### F05: verified memory for a frozen LLM - Opening: a multiplication prompt enters a frozen model. - Branch: a final-answer checker admits a correct outcome to exemplar memory or rejects an incorrect one; label the checker as external computation. - Contrast: unfiltered memory also retains wrong answers and can contaminate later prompts. Show stored content versus the last four examples actually used. - Result: show 37/120, 59/120 and 16/120, plus an honest chronological plot. - Boundary: weights and learning algorithm are unchanged; the evidence is one bounded stream and not proof of open-ended recursive self-improvement. ## Production decision Make F01, F03, F04 and F05 the research-facing flagship candidates; F02 supplies the immediately concrete ML-systems story. Add F10 as the leading hands-on prototype and F12 as the working small-resource application. The launch order should follow claim/artifact readiness rather than the largest headline number. The first production deliverable should be one complete article–animation–paper-outline package and a checked evidence manifest, not forty partially designed pages. No new models were trained, no old experiments altered, no site published and no third-party artwork copied into the proposed publication assets. The visual references were inspected as temporary local files only. ## Acoustic breathing: practical story and validation boundary The [local report](../../../spline_signal_sensing/report.md) and [implementation](../../../spline_signal_sensing/ghost_v2.py) describe a 20 kHz carrier, 48 kHz audio sampling, a 50 ms rolling analysis context and 10 ms processing increments. The report's 44-second example has a dominant peak near 0.134 Hz, or about 8.1 cycles per minute. Its interpretation as an accurate breathing-rate measurement is not independently established by finding that peak. The reported cardiac-band peak is even less suitable as a physiological headline without a synchronized reference and artifact controls. The recurrence estimator is exact for an ideal single sinusoid under its assumptions; a real multipath, filtered recording is not that model. Its current implementation estimates a recurrence coefficient using dot products and arccos. That is a signal-model/annihilation connection to the operator toolbox, not by itself a new canonical exponential B-spline construction or a trained neural network. Overlapping windows also do not create independent 10 ms measurements or establish 10 ms end-to-end latency. Audio callbacks use 50 ms blocks, another distinction an animation should not hide. For the first article, use a hardware photograph or original schematic, a 30–45-second recorded-data visualization, a compact explanation of the estimator, and an explicit account of what was actually observed. Do not fabricate synchronized chest video or a reference-sensor trace; none was established in this inspection. Avoid transmitting the high-frequency carrier through the webpage—illustrate it visually rather than auto-playing it. A manuscript initially belongs in the prototype/methods category. A stronger measurement paper would need reference agreement, empty-room/still-object and motion-confound controls, multiple recording conditions, uncertainty and an ordinary phase-demodulation or spectral comparator under the same acquisition contract. These are publication prerequisites, not experiments authorized or run in this assessment. Sensor and participant data also need appropriate consent/provenance before publication.