Historical source. Some claims in older records were subsequently corrected. The associated article states the adopted interpretation. This record preserves the original source alongside its rendered reading view.
Rendered archival Markdown
This reading view preserves headings, tables, lists, code fragments and mathematical notation from the local research record.
Single-thread deployment succeeds on the shared workstation
September17,2026. This is a positive, bounded engineering result: two native applications meet their predeclared correctness, latency and memory requirements using one observed computation thread on this Apple M4 Max. It is not a fresh market test, an edge-hardware measurement or a newly discovered pricing model.
What now runs
bin/v2/evaluate exact accepts a maturity and15 spatial log-volatility parameters and returns23 price intervals plus two signed spread finite-response intervals under the archived synthetic model. Configuration is compiled in. Native parsing, calculation, serialization and pipe transport are measured. The polynomial comparator builds its focused mesh in C++, removing the prior Python-side mesh dependency from this deployment comparison.
bin/v2/verify receives an already prepared request-bound warning packet and independently checks its model parameters, quote-compatibility conditions, numerical enclosures and material risk separation against the fixed synthetic contract. It does not perform calibration or find the packet. A verified packet does not establish that this contract describes any actual market.
Neither executable uses Python, NumPy, BLAS, a GPU, a thread pool or network access at runtime. Linked dependencies are the system C++ and system libraries. The exact numerical code, interval primitives, contour and packet contract are reused unchanged from the completed September16 campaign. Numerical computation is not new here; native mesh generation, a bounded line interface, thread/CPU instrumentation and the complete deployment comparison are the new additions.
Primary measurements
All values below include the declared request path. Warm medians/p95 include request/response transport. Cold measurements include new-process launch, configuration, response and clean exit; they do not flush the OS page cache.
| Native application | Warm median | Warm p95 | Cold p95 | Peak process RSS | Executable |
|---|
| Certified exact-cell evaluation | 4.793ms | 5.045ms | 8.718ms | 2,326,528bytes /2.219MiB | 139,064bytes |
| Positive risk-packet verification | 42.978ms | 44.956ms | 48.673ms | 2,506,752bytes /2.391MiB | 123,096bytes |
The calculator's36 warm requests cover12 exposed cases repeated three times; its12 cold requests cover each case once. Recipient warm timing shown above is for24 positive requests (eight positive cases, three repetitions), not the much cheaper negative cases. All42 positive/negative cases were exercised:126 warm, 42 warm-up and42 cold classifications. The recipient cold timing shown is the eight positives. Raw results for every phase and classification remain available.
Calculator native CPU median is4.751ms and native wall median4.751ms; recipient positive medians are42.795ms CPU and42.799ms native wall. Every before/after thread observation is1. These are single-threaded processes, not physical-core affinity measurements. macOS may migrate a thread and other applications remain free to run on the machine. Source inspection finds no thread creation or parallel numerical library. No claim of zero concurrent host activity is made.
The standard-library benchmark driver has peak RSS38,010,880bytes (~36.25MiB). That is disclosed separately: native process memory is not the total memory of the test harness plus worker. The executable can be used directly from a shell with the supplied request files, without that harness. End-to-end timings are not the earlier unverified point-evaluation kernel timings; comparing4.79ms here with0.6-1.4ms from a different contract would conflate workloads.
Matched numerical comparator, including failures
Both calculators return the same25 outputs under the same bounded-domain model, source/readouts,24-node time inversion and binary64 interval arithmetic. The polynomial comparator uses consistent-mass FEM, cubic primal-dual correction, verified residual bounds, and reused tridiagonal factorizations. Its source-focused sinh grid preserves all interfaces and readouts and includes midpoint subdivision. Native mesh construction is checked against the archived Python algorithm on all 48 case/budget combinations; maximum node deviation is below2e-14.
Qualification requires full containment of the saved independent reference intervals and widths <=2e-8 for all prices and <=4e-8 for both responses.
| Method/base mesh budget | Cases qualifying in all three warm repetitions | Warm median across all cases | Maximum price width | Peak native RSS |
|---|
| Exact cells | 12/12 | 4.793ms | 6.177e-10 | 2.219MiB |
| Polynomial32 | 0/12 | 29.112ms | 4.430e2 | 2.344MiB |
| Polynomial128 | 0/12 | 69.931ms | 3.361 | 2.516MiB |
| Polynomial512 | 1/12 | 244.119ms | 8.829e-3 | 3.844MiB |
| Polynomial2048 | 6/12 | 926.699ms | 3.541e-5 | 8.406MiB |
Every polynomial output contains the reference; failures are insufficiently tight certificates, not incorrect reference values. Do not divide all-case latency medians and label that an accuracy-matched speedup. On the six jointly qualifying cases, the smallest tested qualifying polynomial budget costs52.75-206.04x as much as the exact method. This is an observed frontier on exposed cases; it is not a preselected adaptive deployment rule. The other six cases have no qualified polynomial comparator within the tested budgets, so no speed ratio is assigned.
All six qualifying polynomial cases are the longer-maturity cases. This repeats the known short-maturity certification bottleneck; it is not broad evidence against conventional numerical PDE methods. Better adaptive/high-order schemes, different certificate constructions, optimized libraries and commercial solvers were not tested. The earlier Green-norm variant did not improve the qualification frontier, but that does not exhaust possible comparator improvements. No claim against SOTA or against ordinary uncertified Black-Scholes formulas follows.
Verification and preservation
Nine prerequisite unit tests passed, covering source/binary hashes, all12 exact references, all42 packet classifications, native mesh construction, parser and budget rejection, overlong/unterminated lines, stale identity and valid recovery, one-worker enforcement and fail-closed result checks.
The frozen benchmark completed in83.609seconds. It retains510 responses: 300 numerical (five methods x12 cases xfive phase/repetition instances) and210 recipient responses, plus eight stale-identity/recovery pairs. All exact outputs meet accuracy requirements; every recipient classification and reference check passes. All ten predeclared application gates pass (correctness, one thread, <8MiB RSS, warm median threshold, cold p95 threshold for each application).
The separate read-only audit verifies eight frozen input identities and all518 result-file hashes, independently checks strict output schemas and exact rational interval-width comparisons, verifies request/repetition coverage and recomputes the warm medians and resource gates. It passes. It also tests that a missing status cannot count as a correct negative packet. See AUDIT.json and audit.py. Both documented native shell examples were executed successfully after the benchmark; those smoke checks are not added to the frozen timing sample.
The first wrapper build's JSON line-framing issue was corrected before timing; its build record and binaries are retained. A sandbox-denied hardware metadata query stopped the first harness invocation before any numerical worker started. The unchanged harness then ran with permission for that query. These setup events are recorded in IMPLEMENTATION_LOG.md. No timing failures or outcomes were removed.
The protocol, implementation and passing prerequisites were committed as 53dd78052 before timing. Results, all negative comparator qualifications, examples, this report and reproduction instructions are committed in a closing checkpoint. Earlier experiments and market reports are unchanged. No data downloads, registrations, production changes or packages were needed.
Interpretation
This supports a concrete claim: the existing operator-based calculation and fixed-contract verification can be packaged as small, responsive, single-threaded native applications on this machine. The delta is a deployment result, not proof of commercial usefulness or a replacement for failed market admission.
It strengthens the plausibility of resource-constrained deployment, but a genuine edge claim still requires actual target hardware, portability work, representative input coverage and measured resource/energy budgets there. A financial deployment also needs a valid observation/model contract and a demonstrated decision use. The previous market-fit failure is neither erased nor solved by these timings.
There is no need to optimize these exposed cases further to answer the present question. The runnable examples are the handoff; any next experiment should test a new application requirement or actual target device rather than repeat this benchmark to chase a better timing headline.
View raw MD source
# Single-thread deployment succeeds on the shared workstation
September17,2026. This is a positive, bounded engineering result: two native
applications meet their predeclared correctness, latency and memory requirements
using one observed computation thread on this Apple M4 Max. It is not a fresh
market test, an edge-hardware measurement or a newly discovered pricing model.
## What now runs
`bin/v2/evaluate exact` accepts a maturity and15 spatial log-volatility parameters
and returns23 price intervals plus two signed spread finite-response intervals
under the archived synthetic model. Configuration is compiled in. Native parsing,
calculation, serialization and pipe transport are measured. The polynomial
comparator builds its focused mesh in C++, removing the prior Python-side mesh
dependency from this deployment comparison.
`bin/v2/verify` receives an already prepared request-bound warning packet and
independently checks its model parameters, quote-compatibility conditions,
numerical enclosures and material risk separation against the fixed synthetic
contract. It does not perform calibration or find the packet. A verified packet
does not establish that this contract describes any actual market.
Neither executable uses Python, NumPy, BLAS, a GPU, a thread pool or network
access at runtime. Linked dependencies are the system C++ and system libraries.
The exact numerical code, interval primitives, contour and packet contract are
reused unchanged from the completed September16 campaign. Numerical computation
is not new here; native mesh generation, a bounded line interface, thread/CPU
instrumentation and the complete deployment comparison are the new additions.
## Primary measurements
All values below include the declared request path. Warm medians/p95 include
request/response transport. Cold measurements include new-process launch,
configuration, response and clean exit; they do not flush the OS page cache.
| Native application | Warm median | Warm p95 | Cold p95 | Peak process RSS | Executable |
| --- | ---: | ---: | ---: | ---: | ---: |
| Certified exact-cell evaluation | 4.793ms | 5.045ms | 8.718ms | 2,326,528bytes /2.219MiB | 139,064bytes |
| Positive risk-packet verification | 42.978ms | 44.956ms | 48.673ms | 2,506,752bytes /2.391MiB | 123,096bytes |
The calculator's36 warm requests cover12 exposed cases repeated three times;
its12 cold requests cover each case once. Recipient warm timing shown above is
for24 positive requests (eight positive cases, three repetitions), not the much
cheaper negative cases. All42 positive/negative cases were exercised:126 warm,
42 warm-up and42 cold classifications. The recipient cold timing shown is the
eight positives. Raw results for every phase and classification remain available.
Calculator native CPU median is4.751ms and native wall median4.751ms; recipient
positive medians are42.795ms CPU and42.799ms native wall. Every before/after
thread observation is1. These are single-threaded processes, not physical-core
affinity measurements. macOS may migrate a thread and other applications remain
free to run on the machine. Source inspection finds no thread creation or
parallel numerical library. No claim of zero concurrent host activity is made.
The standard-library benchmark driver has peak RSS38,010,880bytes (~36.25MiB).
That is disclosed separately: native process memory is not the total memory of
the test harness plus worker. The executable can be used directly from a shell
with the supplied request files, without that harness. End-to-end timings are
not the earlier unverified point-evaluation kernel timings; comparing4.79ms here
with0.6-1.4ms from a different contract would conflate workloads.
## Matched numerical comparator, including failures
Both calculators return the same25 outputs under the same bounded-domain model,
source/readouts,24-node time inversion and binary64 interval arithmetic. The
polynomial comparator uses consistent-mass FEM, cubic primal-dual correction,
verified residual bounds, and reused tridiagonal factorizations. Its source-focused
sinh grid preserves all interfaces and readouts and includes midpoint subdivision.
Native mesh construction is checked against the archived Python algorithm on all
48 case/budget combinations; maximum node deviation is below2e-14.
Qualification requires full containment of the saved independent reference
intervals and widths <=2e-8 for all prices and <=4e-8 for both responses.
| Method/base mesh budget | Cases qualifying in all three warm repetitions | Warm median across all cases | Maximum price width | Peak native RSS |
| --- | ---: | ---: | ---: | ---: |
| Exact cells | 12/12 | 4.793ms | 6.177e-10 | 2.219MiB |
| Polynomial32 | 0/12 | 29.112ms | 4.430e2 | 2.344MiB |
| Polynomial128 | 0/12 | 69.931ms | 3.361 | 2.516MiB |
| Polynomial512 | 1/12 | 244.119ms | 8.829e-3 | 3.844MiB |
| Polynomial2048 | 6/12 | 926.699ms | 3.541e-5 | 8.406MiB |
Every polynomial output contains the reference; failures are insufficiently tight
certificates, not incorrect reference values. Do not divide all-case latency
medians and label that an accuracy-matched speedup. On the six jointly qualifying
cases, the smallest tested qualifying polynomial budget costs52.75-206.04x as
much as the exact method. This is an observed frontier on exposed cases; it is
not a preselected adaptive deployment rule. The other six cases have no qualified
polynomial comparator within the tested budgets, so no speed ratio is assigned.
All six qualifying polynomial cases are the longer-maturity cases. This repeats
the known short-maturity certification bottleneck; it is not broad evidence
against conventional numerical PDE methods. Better adaptive/high-order schemes,
different certificate constructions, optimized libraries and commercial solvers
were not tested. The earlier Green-norm variant did not improve the qualification
frontier, but that does not exhaust possible comparator improvements. No claim
against SOTA or against ordinary uncertified Black-Scholes formulas follows.
## Verification and preservation
Nine prerequisite unit tests passed, covering source/binary hashes, all12 exact
references, all42 packet classifications, native mesh construction, parser and
budget rejection, overlong/unterminated lines, stale identity and valid recovery,
one-worker enforcement and fail-closed result checks.
The frozen benchmark completed in83.609seconds. It retains510 responses:
300 numerical (five methods x12 cases xfive phase/repetition instances) and210
recipient responses, plus eight stale-identity/recovery pairs. All exact outputs
meet accuracy requirements; every recipient classification and reference check
passes. All ten predeclared application gates pass (correctness, one thread,
<8MiB RSS, warm median threshold, cold p95 threshold for each application).
The separate read-only audit verifies eight frozen input identities and all518
result-file hashes, independently checks strict output schemas and exact rational
interval-width comparisons, verifies request/repetition coverage and recomputes
the warm medians and resource gates. It passes. It also tests that a missing
status cannot count as a correct negative packet. See AUDIT.json and audit.py.
Both documented native shell examples were executed successfully after the
benchmark; those smoke checks are not added to the frozen timing sample.
The first wrapper build's JSON line-framing issue was corrected before timing;
its build record and binaries are retained. A sandbox-denied hardware metadata
query stopped the first harness invocation before any numerical worker started.
The unchanged harness then ran with permission for that query. These setup events
are recorded in IMPLEMENTATION_LOG.md. No timing failures or outcomes were removed.
The protocol, implementation and passing prerequisites were committed as
53dd78052 before timing. Results, all negative comparator qualifications, examples,
this report and reproduction instructions are committed in a closing checkpoint.
Earlier experiments and market reports are unchanged. No data downloads,
registrations, production changes or packages were needed.
## Interpretation
This supports a concrete claim: **the existing operator-based calculation and
fixed-contract verification can be packaged as small, responsive, single-threaded
native applications on this machine.** The delta is a deployment result, not
proof of commercial usefulness or a replacement for failed market admission.
It strengthens the plausibility of resource-constrained deployment, but a genuine
edge claim still requires actual target hardware, portability work, representative
input coverage and measured resource/energy budgets there. A financial deployment
also needs a valid observation/model contract and a demonstrated decision use.
The previous market-fit failure is neither erased nor solved by these timings.
There is no need to optimize these exposed cases further to answer the present
question. The runnable examples are the handoff; any next experiment should test
a new application requirement or actual target device rather than repeat this
benchmark to chase a better timing headline.