Historical source. Some claims in older records were subsequently corrected. The associated article states the adopted interpretation. This record preserves the original source alongside its rendered reading view.
Rendered archival TeX
This is an HTML reading rendition of the local TeX record. Mathematical notation is rendered with KaTeX; archived figures are included when their source assets are part of this collection.
The new research direction is a colony of inexpensive machines that develops useful skills locally, exchanges what it learns, and acquires new collective capabilities without repeatedly retraining a central model. This vision is independent of whether a skill is acquired from scratch or distilled from a teacher. It does not imply unlimited capability, biological emulation, or learner-rule recursive self-improvement. Collective lifelong edge learning, reward machines, and compositional navigation have substantial prior art; the dated selection document records primary-source links and reading scope. Development observations and historical independent interims follow. The completed thirty-pool result and mechanism decision at the end of this section supersede their provisional execution status; their evidence is retained.
Representation and learned temporal state.
An agent queries a synthetic binary acceptance evaluator on coordinate sequences. Prefixes are distinguished by their responses to a finite set of continuations. The learner does not receive the evaluator's state labels or a temporal-template identifier. It then fits coordinate-dependent transitions between the identified states. For a two-dimensional cardinal cubic basis, A(x,y)=i,j∑Bi(x)Bj(y)Cij,(Cij)rs≥0,s∑(Cij)rs=1. Partition of unity makes A row-stochastic in exact arithmetic on the declared domain. A temporal acceptance score is α⊤A(x1,y1)⋯A(xT,yT)ω. The initial P0 planner instead uses maximum-probability transitions; that discretization is not equivalent to the soft acceptance model. Independently learned memories are combined in a classical product-state dynamic program, not by averaging donor actions. The learned objects specify temporal task acceptance, not motor policies or unknown physical dynamics. Object-coordinate interfaces, a motion graph and an action-synthesizing planner are supplied; successful exchange would not by itself establish autonomous acquisition of physical motor abilities.
Where the spline calculus enters.
Fitting uses continuous mass and derivative Grams. Four-point Gaussian quadrature per cell exactly integrates the relevant polynomial products in exact arithmetic. The finite-domain Grams are banded, not circulant. Cardinal coefficients can also be converted by fixed extraction matrices to cellwise bicubic Bernstein coefficients. Evaluation then becomes local matrix–vector contraction, without Cox–de Boor recursion. This exact representation interchange does not guarantee that the learned temporal semantics are correct, nor that a product-state planner remains small as the number of skills increases.
P0 development acquisition.
At source checkpoint ced30cc1e, seed 36200, three synthetic teachers produced learned models with two, two and three states: visiting a region, avoiding an exclusion region, and inspecting two sides in order. Their compressed exported messages total 17,319 bytes. Recipient Bernstein coefficients require 999,424 numeric bytes, before planning tables and runtime memory. All 73,699 evaluator calls, including 1,536 development validation calls, are retained as query tapes. These are distillation queries to cheap programs, not physical robot trials or free human supervision. In-process acquisition, export and validation took 0.185 s; peak RSS was 110.6 MB on the shared host. Hard-transition validation accuracy was 98.05, 95.70 and 99.80%, respectively, but the inspection validation had only eight positive sequences among 512. Its aggregate accuracy is therefore weak evidence of inspection competence. Reloaded float32 cards and their float64 Bernstein evaluation agree within 7.8×10−16 on the checked points; serialization leaves probability-row sum errors up to 2.3×10−8.
P0 navigation exposes a substantive failure.
At navigation source checkpoint eadfef868, three development scenes were tested at each of two, three and four simultaneous requirements. The learned product planner succeeded on 3/3, 0/3 and 2/3, while its hand-coded teacher-state ceiling succeeded on 3/3, 2/3 and 2/3. Success here requires reaching the goal, physical collision avoidance, and semantic acceptance on a trajectory sampled eight times more finely than planning waypoints. All three failures in the three-skill learned panel violated exclusion semantics. The hand-coded ceiling also exposed between-waypoint exclusion violations. Median learned planning times were 28.3, 61.3 and 91.7 ms, including transition-table preparation; the entire development comparison used 72.4 MB peak RSS. No-transfer solved none of these missions, but the best single donor already solved two of the three two-requirement scenes. These observations do not establish a competitive breakthrough: high standalone predictive accuracy is insufficient under planning-induced distribution shift, and waypoint correctness is not pathwise correctness.
Historical development decision.
The second and final substantive revision in this mechanism family will retain sets of plausible temporal states and inspect transitions along motion segments. Such sets must not be advertised as calibrated safety guarantees without supporting assumptions and evidence. Strong alternative learned representations and modular planning controls, independent confirmation, retention under later acquisition, and end-to-end resource accounting remain required. The detailed chronology and raw-evidence pointers are maintained in COLONY\_DEVELOPMENT\_LOG\_2026-09-14.md.
P1 support-state development result.
At source 11ebd16f9, unchanged donor messages were compiled with a fixed probability-support cutoff of 0.02 and eight semantic subdivisions per motion segment. Possible temporal states are retained as subsets. Exact minimization over the finite emitted scene alphabet reduces the three subset memories from 3/3/7 to 2/2/3 states. On the same development scenes, learned success becomes 3/3, 3/3 and 2/3 for two, three and four requirements. The remaining failure returns no feasible plan under the conservative model and fixed horizon; all eight returned plans pass the dense semantic check. The segment-aware hand-coded ceiling succeeds on all nine scenes. The visit-only donor solves all two-requirement scenes, but no individual donor solves a three- or four-requirement scene. Median learned planning latency increases to 825/881/929 ms, primarily through segment-table evaluation. Peak RSS is 81.8 MB. This is evidence that temporal uncertainty matters in this prototype, not a safety certificate, independent replication, or a spline-specific competitive result. The next permitted steps are stronger controls and exact output-preserving compilation, not another semantic revision or threshold sweep.
Matched-query representation controls.
At source e85ac39ad, the same inferred temporal states and recovered transition examples train a bounded decision tree, a 64-tree ensemble and a 2–64–64–S2 ReLU MLP. No additional teacher queries are supplied. On the same development panel, success at two/three/four requirements is 3/3,3/3,2/3 for cardinal fields, 3/3,2/3,2/3 for the tree, zero for the ensemble under the common 0.02 support rule, and 3/3 at every size for the MLP. Thus this panel does not establish a spline-exclusive capability advantage. Three-message sizes are respectively 17,319, 9,834, 560,042 and 59,083 bytes. With an exact Bernstein support specialization, cardinal median full-plan time is 336/386/430 ms versus MLP 486/534/592 ms; this modest runtime difference does not erase the MLP's additional successful mission. The full comparison takes 52.86 s and peaks at 498.2 MB RSS. Fit timers include first library imports and are not warmed training-speed ratios. Ordinary maximum-probability decisions for every representation, all-order sequential reuse, and a memoryless additive cost-map control are the next fixed comparisons. The ensemble result particularly cautions that a common probability threshold is not a common uncertainty calibration.
Classical reuse limits the novelty claim.
At source 80f9f326d, all-order sequential composition of the same learned modules matches joint-product development success: cardinal 3/3,3/3,2/3, and MLP 3/3 at every size. A memoryless additive learned cost-map control succeeds only on the two-requirement scenes. With ordinary maximum-probability transitions, all four learned representations lose two or three of the three-requirement scenes. These 72 additional plans take 76.57 s and peak at 237.1 MB RSS. The support rule is therefore useful in this development panel, but neither temporal composition itself nor a capability beyond classical sequential reuse has been established. The mechanism is retained for a bounded acquisition/interchange/disconnection/ retention confirmation, not promoted to a practical or SOTA breakthrough.
Historical first-pool confirmation interim.
The 30-pool protocol is frozen and pushed at 50cf21aa1 before any independent seed is opened. In its first completed pool, inspection state identification collapses to one nonaccepting state. All learned representations then fail all 16 three-/four-requirement missions; the hand-coded ceiling succeeds on all 16. Inspection balanced accuracy is 0.5 on 1,024 independent validation sequences, including 139 positives. A one-state stochastic operator with a nonaccepting output cannot accept any word, regardless of its static function approximator. Disconnection, exact experience replay and structural retention audits nevertheless pass. The 7,751-byte payload is partly smaller because it encodes an incorrect collapsed model, not because it attains the same capability more efficiently. This one-pool interim result is retained while the fixed panel continues without changing grids, thresholds or query budgets. The dated confirmation report records the complete eventual panel.
Historical six-pool interim and completed secondary panel.
All 36 prespecified heading-lattice missions now complete: cardinal support planning succeeds on 22/36 (61.1%), below the fixed 80% capability target. MLP support succeeds on 25/36 and the hand-coded ceiling on 34/36. This secondary capability target fails; the test concerns discrete motion graphs, not deployment on a physical drone or car. In the still-running primary panel, cardinal succeeds on 41/96 three-/four-requirement missions, MLP on 60/96 and the ceiling on 93/96. Sequential reuse matches both learned models' respective counts. Two of six inspection acquisitions collapse to one state. Saved-query coverage explains these two failures without supplying any new evidence to learners. The full 30-pool protocol remains unchanged.
Historical ten-pool decision: capability cannot be recovered.
At ten completed pools, cardinal support succeeds on 68/160 composed missions, MLP on 91/160 and the ceiling on 154/160; sequential reuse again matches each learned model. Even if all 320 remaining primary composed missions succeeded, cardinal would reach only 388/480 (80.83%), below the fixed 85% target. The capability component has therefore irreversibly failed without treating any pending case as an observed failure. This version is rejected as a breakthrough candidate. The full frozen panel continues for complete estimates and controls, not another grid, anchor or probability-threshold refinement.
Qualification of the supplied-model reference.
The method named oracle, called a ceiling in the development chronology, is not a mathematical upper bound on task success. It also plans with eight semantic substeps per edge. In independent pool 13, mission r4\_m6, its returned path fails the 64-substep avoidance check. An oracle nonreturn does not establish continuous infeasibility. The separately frozen interval audit reported below inspects every returned path, including oracle paths, without replacing the primary score or claiming continuous certification of every temporal requirement.
Completed independent confirmation and mechanism decision
Complete panel, unchanged method.
All thirty donor pools finish between September 13, 21:56:57 UTC and September 14, 01:03:58 UTC. Every planned pool is retained: 720 primary missions, including 480 with three or four requirements, and 36 secondary heading-lattice missions. Frozen source hashes verify after execution. No learner setting, seed, grid, support threshold or horizon is changed after confirmation begins. Table [tab:colony-final-030] gives the complete primary capability frontier.
Cardinal support attains 161/480=33.5%, MLP support 242/480=50.4%, and the supplied-model reference 464/480=96.7%. The cardinal and MLP sequential controls match their respective success counts. Descriptive 95% donor-cluster bootstrap intervals, using the frozen 10,000 resamples of thirty pools, are 21.5–46.7% for cardinal and 34.0–67.3% for MLP. Their paired difference is −16.875 percentage points, interval [−25.2,−9.4]. The cardinal versus cardinal-sequential difference and its interval are zero. These intervals are not scene-independent or multiplicity-adjusted novelty tests. The bounded prototype, portability, beyond-classical composition and matched-capability MLP comparison all fail. The separate resource screen passes. The secondary result is unchanged: cardinal 22/36, MLP 25/36, reference 34/36. Cardinal heading planning has median/p95 1.255/1.438 s.
Coverage, not just query count.
Fourteen inspection acquisitions collapse to a single nonaccepting state. The saved-query audit predicts all thirty inferred state counts from whether both inspection regions occur in the finite 64-anchor discovery alphabet. When one region is missed, discovery responses are all negative; the collapsed model then selects only the empty continuation, so later single-point fitting queries cannot expose the two-step order either. Many queries therefore do not imply informative temporal coverage. This is a diagnosis of this procedure, not a universal lower bound against distillation. Cardinal balanced accuracies are 97.77% (visit), 97.53% (avoid) and 72.57% (inspection). Inspection ordinary accuracy is 90.71% on 30,720 sequences with 4,915 positives, illustrating the importance of class counts.
Of the 480 composed cardinal cases, 161 return valid paths, 224 do not return with collapsed inspection models, 85 do not return despite a valid reference route, and ten do not return without such a route. MLP shares the 224 collapsed-state failures, but has only six and eight other nonreturns in those two reference strata. Missing states thus do not explain the entire cardinal deficit. These are post-hoc strata, not causal interventions or redefined denominators. Across all 720 primary cases, cardinal returns 401 valid paths and no invalid ones; its remaining 319 cases still fail.
Separate continuous-avoidance verification.
After the main timing campaign, all 9,158 returned method–mission paths are inspected with outward-rounded interval evaluation at 40 decimal digits. The count includes duplicate geometrical paths from different methods, not 9,158 independent safety trials. There are 7,926 avoidance-clear paths, 1,232 violation witnesses and zero unresolved checks. No avoidance violation is found that the frozen 64-substep check missed. A constructed unit-test crossing nevertheless shows why this observed agreement is not a theorem about finite sampling. All 401 cardinal-support primary and 22 heading returns are clear, as are 482/25 MLP-support returns. The supplied reference has three primary avoidance violations, already rejected by the frozen score.
This audit uses 483,448 interval evaluations, 55.193 s of summed active stage time, maximum RSS 80,265,216 bytes and zero new teacher queries. It assumes the exact supplied stored geometry and mpmath's experimental interval runtime. It checks robot-center avoidance, not every continuous temporal condition, unknown physics or real-world safety. Unreturned paths are not clearance successes. Neither the original score nor any criterion is replaced.
Payload, memory and complete measured costs.
Median three-message sizes are 17,096 bytes for cardinal, 9,318 for the tree, 531,570 for ExtraTrees and 59,050.5 for MLP. The cardinal maximum is 17,690 bytes. Its lower latency and payload than MLP do not pass the comparison because capability is more than five percentage points worse. The common support threshold is not matched uncertainty calibration. Primary-only maximum RSS is 111,640,576 bytes (106.47 MiB), and primary planning p95 is 0.415 s. Mission metadata adds median/max 1,121,003/2,217,324 bytes per inbox; runtime and planning working memory are additional to the small messages. The acquisition/control/evaluation parent reaches 549,928,960 bytes RSS. These are single-threaded CPU measurements on a shared M4 Max, not MPS or microcontroller deployment measurements.
Independent acquisition consumes 1,813,994 non-validation teacher queries; shared held-out sequence validation adds 92,160 labels, counted once per skill rather than once per representation. Development/rehearsals add 218,793 queries including validation. Supervision comes from cheap synthetic programs, not real robot trials or a hidden trained foundation model. Exact original transcript replay reproduces every message bitwise, with median 0.100 s replay time and 447,643 transcript bytes. It gives no accuracy advantage to transfer over access to the same experience.
Nine measured development stages sum to 316.210 s; independent whole-pool timers sum to 11,205.877 s, versus campaign elapsed time 11,221.971 s. Including the separate interval audit yields 11,577.280 s of measured numerical stages. Nested role timers are not added again. Literature, engineering/assistant effort, unit tests, un-timed QA, PDF/Git work and older campaigns are excluded; these are not energy measurements or the entire project cost. All thirty disconnection, primary-isolation, exact replay and unchanged-old-module checks pass. Twenty-five selected regression tests pass, with one sklearn column-vector warning; the full repository suite is not rerun. Structural isolation of separate immutable modules is not learned zero-forgetting or RSI.
Retained algorithmic contribution and next investment.
On the reused development panel, exact Bernstein support specialization preserves the audited transition tables, paths and memories while reducing full-plan time by 2.10–2.46×. Retained numeric program storage falls from 999,424 to 142,325 bytes. The audit checks 100,000 points per card and four alternating-order timing pairs per scene. Cardinal local contractions and exact continuous derivative-product Grams are genuinely exercised, without assuming that the finite-boundary fitting operators are circulant. This is a scoped compiler improvement, not rescued task competence.
The tested task-description pipeline is retired as a breakthrough candidate; no third semantic refinement or opportunistic second family is launched. A future capability must make independently learned physical behavior useful to a recipient under strong modular, classical, whole-policy and shared-data controls. Distillation remains an allowed acquisition route. Unknown joint interactions require a stated physical interface or additional distinguishing evidence; compact messages cannot manufacture missing information. The dated confirmation report, next-capability decision and research index record all raw evidence, sources, commands, cost exclusions and this prospective boundary.
Original: paper/colony_flagship_20260914.tex · Raw source file
View raw TEX source
\section{Colony flagship: portable temporal skill operators}
\label{sec:colony-flagship-20260914}
\paragraph{Fixed vision and current claim.}
The new research direction is a colony of inexpensive machines that develops
useful skills locally, exchanges what it learns, and acquires new collective
capabilities without repeatedly retraining a central model. This vision is
independent of whether a skill is acquired from scratch or distilled from a
teacher. It does not imply unlimited capability, biological emulation, or
learner-rule recursive self-improvement. Collective lifelong edge learning,
reward machines, and compositional navigation have substantial prior art;
the dated selection document records primary-source links and reading scope.
Development observations and historical independent interims follow. The
completed thirty-pool result and mechanism decision at the end of this section
supersede their provisional execution status; their evidence is retained.
\paragraph{Representation and learned temporal state.}
An agent queries a synthetic binary acceptance evaluator on coordinate
sequences. Prefixes are distinguished by their responses to a finite set of
continuations. The learner does not receive the evaluator's state labels or a
temporal-template identifier. It then fits coordinate-dependent transitions
between the identified states. For a two-dimensional cardinal cubic basis,
\[
A(x,y)=\sum_{i,j} B_i(x)B_j(y)C_{ij},\qquad
(C_{ij})_{rs}\geq 0,\quad \sum_s(C_{ij})_{rs}=1.
\]
Partition of unity makes $A$ row-stochastic in exact arithmetic on the declared
domain. A temporal acceptance score is
$\alpha^\top A(x_1,y_1)\cdots A(x_T,y_T)\omega$.
The initial P0 planner instead uses maximum-probability transitions;
that discretization is not equivalent to the soft acceptance model.
Independently learned memories are combined in a classical product-state
dynamic program, not by averaging donor actions.
The learned objects specify temporal task acceptance, not motor policies or
unknown physical dynamics. Object-coordinate interfaces, a motion graph and
an action-synthesizing planner are supplied; successful exchange would not by
itself establish autonomous acquisition of physical motor abilities.
\paragraph{Where the spline calculus enters.}
Fitting uses continuous mass and derivative Grams. Four-point Gaussian
quadrature per cell exactly integrates the relevant polynomial products in
exact arithmetic. The finite-domain Grams are banded, not circulant.
Cardinal coefficients can also be converted by fixed extraction matrices to
cellwise bicubic Bernstein coefficients. Evaluation then becomes local
matrix--vector contraction, without Cox--de Boor recursion. This exact
representation interchange does not guarantee that the learned temporal
semantics are correct, nor that a product-state planner remains small as the
number of skills increases.
\paragraph{P0 development acquisition.}
At source checkpoint \texttt{ced30cc1e}, seed 36200, three synthetic teachers
produced learned models with two, two and three states: visiting a region,
avoiding an exclusion region, and inspecting two sides in order. Their
compressed exported messages total 17,319 bytes. Recipient Bernstein
coefficients require 999,424 numeric bytes, before planning tables and runtime
memory. All 73,699 evaluator calls, including 1,536 development validation
calls, are retained as query tapes. These are distillation queries to cheap
programs, not physical robot trials or free human supervision. In-process
acquisition, export and validation took 0.185 s; peak RSS was 110.6 MB on the
shared host. Hard-transition validation accuracy was 98.05, 95.70 and
99.80\%, respectively, but the inspection validation had only eight positive
sequences among 512. Its aggregate accuracy is therefore weak evidence of
inspection competence. Reloaded float32 cards and their float64 Bernstein
evaluation agree within $7.8\times10^{-16}$ on the checked points; serialization
leaves probability-row sum errors up to $2.3\times10^{-8}$.
\paragraph{P0 navigation exposes a substantive failure.}
At navigation source checkpoint \texttt{eadfef868}, three development scenes
were tested at each of two, three and four simultaneous requirements. The
learned product planner succeeded on 3/3, 0/3 and 2/3, while its hand-coded
teacher-state ceiling succeeded on 3/3, 2/3 and 2/3. Success here requires
reaching the goal, physical collision avoidance, and semantic acceptance on
a trajectory sampled eight times more finely than planning waypoints.
All three failures in the three-skill learned panel violated exclusion
semantics. The hand-coded ceiling also exposed between-waypoint exclusion
violations. Median learned planning times were 28.3, 61.3 and 91.7 ms,
including transition-table preparation; the entire development comparison used
72.4 MB peak RSS. No-transfer solved none of these missions, but the best
single donor already solved two of the three two-requirement scenes.
These observations do not establish a competitive breakthrough: high
standalone predictive accuracy is insufficient under planning-induced
distribution shift, and waypoint correctness is not pathwise correctness.
\paragraph{Historical development decision.}
The second and final substantive revision in this mechanism family will
retain sets of plausible temporal states and inspect transitions along motion
segments. Such sets must not be advertised as calibrated safety guarantees
without supporting assumptions and evidence. Strong alternative learned
representations and modular planning controls, independent confirmation,
retention under later acquisition, and end-to-end resource accounting remain
required. The detailed chronology and raw-evidence pointers are maintained in
\texttt{COLONY\_DEVELOPMENT\_LOG\_2026-09-14.md}.
\paragraph{P1 support-state development result.}
At source \texttt{11ebd16f9}, unchanged donor messages were compiled with a
fixed probability-support cutoff of 0.02 and eight semantic subdivisions per
motion segment. Possible temporal states are retained as subsets. Exact
minimization over the finite emitted scene alphabet reduces the three subset
memories from 3/3/7 to 2/2/3 states. On the same development scenes, learned
success becomes 3/3, 3/3 and 2/3 for two, three and four requirements. The
remaining failure returns no feasible plan under the conservative model and
fixed horizon; all eight returned plans pass the dense semantic check. The
segment-aware hand-coded ceiling succeeds on all nine scenes. The visit-only
donor solves all two-requirement scenes, but no individual donor solves a
three- or four-requirement scene. Median learned planning latency increases
to 825/881/929 ms, primarily through segment-table evaluation. Peak RSS is
81.8 MB. This is evidence that temporal uncertainty matters in this prototype,
not a safety certificate, independent replication, or a spline-specific
competitive result. The next permitted steps are stronger controls and exact
output-preserving compilation, not another semantic revision or threshold sweep.
\paragraph{Matched-query representation controls.}
At source \texttt{e85ac39ad}, the same inferred temporal states and recovered
transition examples train a bounded decision tree, a 64-tree ensemble and a
2--64--64--$S^2$ ReLU MLP. No additional teacher queries are supplied. On the
same development panel, success at two/three/four requirements is
$3/3,3/3,2/3$ for cardinal fields, $3/3,2/3,2/3$ for the tree, zero for the
ensemble under the common 0.02 support rule, and $3/3$ at every size for the
MLP. Thus this panel does not establish a spline-exclusive capability
advantage. Three-message sizes are respectively 17,319, 9,834, 560,042 and
59,083 bytes. With an exact Bernstein support specialization, cardinal median
full-plan time is 336/386/430\,ms versus MLP 486/534/592\,ms; this modest
runtime difference does not erase the MLP's additional successful mission.
The full comparison takes 52.86\,s and peaks at 498.2\,MB RSS. Fit timers
include first library imports and are not warmed training-speed ratios.
Ordinary maximum-probability decisions for every representation, all-order
sequential reuse, and a memoryless additive cost-map control are the next
fixed comparisons. The ensemble result particularly cautions that a common
probability threshold is not a common uncertainty calibration.
\paragraph{Classical reuse limits the novelty claim.}
At source \texttt{80f9f326d}, all-order sequential composition of the same
learned modules matches joint-product development success: cardinal
$3/3,3/3,2/3$, and MLP $3/3$ at every size. A memoryless additive learned
cost-map control succeeds only on the two-requirement scenes. With ordinary
maximum-probability transitions, all four learned representations lose two or
three of the three-requirement scenes. These 72 additional plans take
76.57\,s and peak at 237.1\,MB RSS. The support rule is therefore useful in
this development panel, but neither temporal composition itself nor a
capability beyond classical sequential reuse has been established. The
mechanism is retained for a bounded acquisition/interchange/disconnection/
retention confirmation, not promoted to a practical or SOTA breakthrough.
\paragraph{Historical first-pool confirmation interim.}
The 30-pool protocol is frozen and pushed at \texttt{50cf21aa1} before any
independent seed is opened. In its first completed pool, inspection state
identification collapses to one nonaccepting state. All learned representations
then fail all 16 three-/four-requirement missions; the hand-coded ceiling
succeeds on all 16. Inspection balanced accuracy is 0.5 on 1,024 independent
validation sequences, including 139 positives. A one-state stochastic operator
with a nonaccepting output cannot accept any word, regardless of its static
function approximator. Disconnection, exact experience replay and structural
retention audits nevertheless pass. The 7,751-byte payload is partly smaller
because it encodes an incorrect collapsed model, not because it attains the
same capability more efficiently. This one-pool interim result is retained
while the fixed panel continues without changing grids, thresholds or query
budgets. The dated confirmation report records the complete eventual panel.
\paragraph{Historical six-pool interim and completed secondary panel.}
All 36 prespecified heading-lattice missions now complete: cardinal support
planning succeeds on 22/36 (61.1\%), below the fixed 80\% capability target.
MLP support succeeds on 25/36 and the hand-coded ceiling on 34/36. This
secondary capability target fails; the test concerns discrete motion graphs,
not deployment on a physical drone or car. In the still-running primary
panel, cardinal succeeds on 41/96 three-/four-requirement missions, MLP on
60/96 and the ceiling on 93/96. Sequential reuse matches both learned
models' respective counts. Two of six inspection acquisitions collapse to
one state. Saved-query coverage explains these two failures without supplying
any new evidence to learners. The full 30-pool protocol remains unchanged.
\paragraph{Historical ten-pool decision: capability cannot be recovered.}
At ten completed pools, cardinal support succeeds on 68/160 composed missions,
MLP on 91/160 and the ceiling on 154/160; sequential reuse again matches each
learned model. Even if all 320 remaining primary composed missions succeeded,
cardinal would reach only 388/480 (80.83\%), below the fixed 85\% target.
The capability component has therefore irreversibly failed without treating
any pending case as an observed failure. This version is rejected as a
breakthrough candidate. The full frozen panel continues for complete estimates
and controls, not another grid, anchor or probability-threshold refinement.
\paragraph{Qualification of the supplied-model reference.}
The method named \texttt{oracle}, called a ceiling in the development
chronology, is not a mathematical upper bound on task success. It also plans
with eight semantic substeps per edge. In independent pool 13, mission
\texttt{r4\_m6}, its returned path fails the 64-substep avoidance check.
An oracle nonreturn does not establish continuous infeasibility. The separately
frozen interval audit reported below inspects every returned path, including
oracle paths, without replacing the primary score or claiming continuous
certification of every temporal requirement.
\subsection{Completed independent confirmation and mechanism decision}
\paragraph{Complete panel, unchanged method.}
All thirty donor pools finish between September 13, 21:56:57 UTC and
September 14, 01:03:58 UTC. Every planned pool is retained: 720 primary missions,
including 480 with three or four requirements, and 36 secondary heading-lattice
missions. Frozen source hashes verify after execution. No learner setting,
seed, grid, support threshold or horizon is changed after confirmation begins.
Table~\ref{tab:colony-final-030} gives the complete primary capability frontier.
\begingroup
% Keep the new table local to its result and independent of legacy float hooks.
% A section-local printed number and explicit reference give a unique target.
\renewcommand{\thetable}{\thesection.1}
\renewcommand{\theHtable}{colony.20260914.1}
\par\medskip\noindent
\begin{minipage}{\linewidth}
\refstepcounter{table}
\centering
\small
\begin{tabular}{lrrrrr}
\toprule
Method & 2 req. & 3 req. & 4 req. & Composed & Median ms \\
\midrule
Cardinal support & 240 & 85 & 76 & 161 & 349.7 \\
Cardinal maximum & 213 & 77 & 70 & 147 & 878.2 \\
Tree support & 195 & 66 & 46 & 112 & 439.2 \\
Tree maximum & 170 & 47 & 23 & 70 & 478.7 \\
ExtraTrees support & 15 & 0 & 0 & 0 & 3887.7 \\
ExtraTrees maximum & 192 & 78 & 66 & 144 & 3936.9 \\
MLP support & 240 & 121 & 121 & 242 & 498.3 \\
MLP maximum & 216 & 85 & 77 & 162 & 536.7 \\
Cardinal sequential & 240 & 85 & 76 & 161 & 470.3 \\
MLP sequential & 240 & 121 & 121 & 242 & 615.7 \\
Cardinal cost map & 240 & 0 & 0 & 0 & 219.9 \\
MLP cost map & 239 & 0 & 0 & 0 & 74.1 \\
No transfer & 0 & 0 & 0 & 0 & 140.1 \\
Best individual donor & 215 & 0 & 0 & 0 & --- \\
Supplied-model reference & 240 & 233 & 231 & 464 & 428.6 \\
\bottomrule
\end{tabular}
\par\smallskip\raggedright\noindent
Table~\thetable: Final successful missions, out of 240 in each requirement column and
480 in the combined three-/four-requirement column. Full-plan median latency
includes transition-table preparation and all 720 primary cases, measured in
the common control harness. The best individual donor is selected after
scoring, not an implemented oracle-free selector. The supplied-model planner
is an implementation reference, not a mathematical ceiling.
\label{tab:colony-final-030}
\end{minipage}\par\medskip
\endgroup
\paragraph{Capability rejection with cluster uncertainty.}
Cardinal support attains $161/480=33.5\%$, MLP support $242/480=50.4\%$,
and the supplied-model reference $464/480=96.7\%$. The cardinal and MLP
sequential controls match their respective success counts. Descriptive
95\% donor-cluster bootstrap intervals, using the frozen 10,000 resamples of
thirty pools, are 21.5--46.7\% for cardinal and 34.0--67.3\% for MLP. Their
paired difference is $-16.875$ percentage points, interval $[-25.2,-9.4]$.
The cardinal versus cardinal-sequential difference and its interval are zero.
These intervals are not scene-independent or multiplicity-adjusted novelty
tests. The bounded prototype, portability, beyond-classical composition and
matched-capability MLP comparison all fail. The separate resource screen
passes. The secondary result is unchanged: cardinal $22/36$, MLP $25/36$,
reference $34/36$. Cardinal heading planning has median/p95 1.255/1.438 s.
\paragraph{Coverage, not just query count.}
Fourteen inspection acquisitions collapse to a single nonaccepting state.
The saved-query audit predicts all thirty inferred state counts from whether
both inspection regions occur in the finite 64-anchor discovery alphabet.
When one region is missed, discovery responses are all negative; the
collapsed model then selects only the empty continuation, so later
single-point fitting queries cannot expose the two-step order either.
Many queries therefore do not imply informative temporal coverage. This is
a diagnosis of this procedure, not a universal lower bound against distillation.
Cardinal balanced accuracies are 97.77\% (visit), 97.53\% (avoid) and
72.57\% (inspection). Inspection ordinary accuracy is 90.71\% on 30,720
sequences with 4,915 positives, illustrating the importance of class counts.
Of the 480 composed cardinal cases, 161 return valid paths, 224 do not return
with collapsed inspection models, 85 do not return despite a valid reference
route, and ten do not return without such a route. MLP shares the 224
collapsed-state failures, but has only six and eight other nonreturns in
those two reference strata. Missing states thus do not explain the entire
cardinal deficit. These are post-hoc strata, not causal interventions or
redefined denominators. Across all 720 primary cases, cardinal returns 401
valid paths and no invalid ones; its remaining 319 cases still fail.
\paragraph{Separate continuous-avoidance verification.}
After the main timing campaign, all 9,158 returned method--mission paths
are inspected with outward-rounded interval evaluation at 40 decimal digits.
The count includes duplicate geometrical paths from different methods, not
9,158 independent safety trials. There are 7,926 avoidance-clear paths,
1,232 violation witnesses and zero unresolved checks. No avoidance violation
is found that the frozen 64-substep check missed. A constructed unit-test
crossing nevertheless shows why this observed agreement is not a theorem
about finite sampling. All 401 cardinal-support primary and 22 heading
returns are clear, as are 482/25 MLP-support returns. The supplied reference
has three primary avoidance violations, already rejected by the frozen score.
This audit uses 483,448 interval evaluations, 55.193 s of summed active stage
time, maximum RSS 80,265,216 bytes and zero new teacher queries. It assumes
the exact supplied stored geometry and mpmath's experimental interval runtime.
It checks robot-center avoidance, not every continuous temporal condition,
unknown physics or real-world safety. Unreturned paths are not clearance
successes. Neither the original score nor any criterion is replaced.
\paragraph{Payload, memory and complete measured costs.}
Median three-message sizes are 17,096 bytes for cardinal, 9,318 for the tree,
531,570 for ExtraTrees and 59,050.5 for MLP. The cardinal maximum is 17,690
bytes. Its lower latency and payload than MLP do not pass the comparison
because capability is more than five percentage points worse. The common
support threshold is not matched uncertainty calibration. Primary-only
maximum RSS is 111,640,576 bytes (106.47 MiB), and primary planning p95 is
0.415 s. Mission metadata adds median/max 1,121,003/2,217,324 bytes per inbox;
runtime and planning working memory are additional to the small messages.
The acquisition/control/evaluation parent reaches 549,928,960 bytes RSS.
These are single-threaded CPU measurements on a shared M4 Max, not MPS or
microcontroller deployment measurements.
Independent acquisition consumes 1,813,994 non-validation teacher queries;
shared held-out sequence validation adds 92,160 labels, counted once per
skill rather than once per representation. Development/rehearsals add 218,793
queries including validation. Supervision comes from cheap synthetic programs,
not real robot trials or a hidden trained foundation model. Exact original
transcript replay reproduces every message bitwise, with median 0.100 s replay
time and 447,643 transcript bytes. It gives no accuracy advantage to transfer
over access to the same experience.
Nine measured development stages sum to 316.210 s; independent whole-pool
timers sum to 11,205.877 s, versus campaign elapsed time 11,221.971 s. Including
the separate interval audit yields 11,577.280 s of measured numerical stages.
Nested role timers are not added again. Literature, engineering/assistant
effort, unit tests, un-timed QA, PDF/Git work and older campaigns are excluded;
these are not energy measurements or the entire project cost. All thirty
disconnection, primary-isolation, exact replay and unchanged-old-module
checks pass. Twenty-five selected regression tests pass, with one sklearn
column-vector warning; the full repository suite is not rerun. Structural
isolation of separate immutable modules is not learned zero-forgetting or RSI.
\paragraph{Retained algorithmic contribution and next investment.}
On the reused development panel, exact Bernstein support specialization
preserves the audited transition tables, paths and memories while reducing
full-plan time by $2.10$--$2.46\times$. Retained numeric program storage falls
from 999,424 to 142,325 bytes. The audit checks 100,000 points per card and
four alternating-order timing pairs per scene. Cardinal local contractions
and exact continuous derivative-product Grams are genuinely exercised, without
assuming that the finite-boundary fitting operators are circulant. This is a
scoped compiler improvement, not rescued task competence.
The tested task-description pipeline is retired as a breakthrough candidate;
no third semantic refinement or opportunistic second family is launched.
A future capability must make independently learned physical behavior useful
to a recipient under strong modular, classical, whole-policy and shared-data
controls. Distillation remains an allowed acquisition route. Unknown joint
interactions require a stated physical interface or additional distinguishing
evidence; compact messages cannot manufacture missing information. The dated
confirmation report, next-capability decision and research index record all
raw evidence, sources, commands, cost exclusions and this prospective boundary.