ML systems · Research & Algorithms

A useful numerical service on one CPU thread

A useful numerical component does not always need a GPU—or even a large runtime. The important question is what the complete service returns, at what accuracy, and with which costs included.

EXPLORE THE IDEA

A whole service on one thread

Request, calculation, enclosure, response.

NATIVE SERVICE / REPLAYA complete request has a complete cost01 parse supplied field02 assemble exact cells03 enclose 23 prices + 2 responses04 serialize outputARCHIVED ONE-THREAD MEASUREMENTApple M4 Max · shared host4.793 mswarm median · complete service2.219 MiBpeak native-process memoryNot a live monitor or low-power-device test
50%
Interface replay of the documented native-service contract, not a live request or screen recording. Reported 4.793 ms warm median and 2.219 MiB peak native memory are archived Apple M4 Max measurements, not edge-device or hard-real-time claims.

Follow the information

From input to outcome

The timed path includes parsing, construction, computation and serialization. There is no model training in the native service. Its one-thread workstation measurements do not establish a hard real-time or low-power-device guarantee.

Scroll the diagram horizontally to follow the route. Keyboard: focus the diagram, then use the arrow keys.

Native request → Validate and parse → Construct and solve → Enclose output batch → Serialized reply. The timed path includes parsing, construction, computation and serialization. There is no model training in the native service. Its one-thread workstation measurements do not establish a hard real-time or low-power-device guarantee.
Information-flow map. One native computation thread; driver memory is a separate process. Original vector schematic based on the method and evidence discussed in this article; signal shapes and icons are illustrative, not additional measurements. Open full-size diagram ↗

Read the main route from left to right; labelled side branches show additional inputs, checks or feedback. The sections below explain the operations and their experimental limits.

From a notebook to a bounded service

A fast kernel is not yet a usable component. A caller supplies a coefficient field and maturity, the program validates and parses the request, constructs its numerical state, computes answers, and serializes them. We measured that whole route for 23 European price intervals and two signed finite spot responses. No model training happens in the native service.

A real request includes more than the kernel

The native program parses a supplied model request, constructs its cells, evaluates resolvents and time inversion, encloses a batch of prices and signed finite responses, and serializes the result. The benchmark times that path through a pipe, not just an isolated basis evaluation. One computation thread is observed and there is no GPU, Python interpreter, or BLAS thread pool in the application.

The comparator must meet the same output-width requirement before a speed ratio is meaningful. At its largest tested budget the polynomial method qualifies six of twelve cases; the other cases remain unqualified rather than being assigned an infinite speedup. Cold startup and warm requests are reported separately.

Compile the operator before serving the request
Compile the operator before serving the request. Original scientific diagram; the stated component and information flow, not an additional experiment. Open full-size figure ↗

What one thread achieved

On the shared Apple M4 Max, the exact-cell calculator has a 4.793 millisecond warm median and a 5.045 millisecond warm p95. Launch-inclusive cold p95 is 8.718 milliseconds. Peak native process memory is 2.219 MiB; the executable is 139,064 bytes. The observed calculation uses one thread and no Python, BLAS, GPU, or thread pool at runtime. The benchmark driver is a separate, larger process.

Accuracy is part of the comparison

Both the exact and polynomial paths must contain independent reference intervals and meet fixed width limits. The exact path qualifies all twelve exposed cases. The largest tested polynomial budget qualifies six. On those six—and only those six—the smallest qualifying tested polynomial budget is 52.75–206.04 times more expensive. The other cases receive no speed ratio. Wide but containing intervals are not falsely described as inaccurate point predictions.

pref∈[p‾,p‾],p‾−p‾≤ε\begin{gathered}p_{\mathrm{ref}}\in[\underline p,\overline p],\\ \overline p-\underline p\le\varepsilon\end{gathered}
Qualification requires reference containment and a useful interval width. Timing a method that fails this criterion does not establish an accuracy-matched speedup.
All methods and all twelve cases remain in the chart. Latencies are not directly comparable as qualified speedups where the width requirement fails.
All methods and all twelve cases remain in the chart. Latencies are not directly comparable as qualified speedups where the width requirement fails.

Why this is not an edge-device benchmark

A single thread on a powerful workstation is not a measurement on a microcontroller or low-power laptop. The runs use a shared host, no fixed core affinity, and a small exposed synthetic panel. They do not establish energy use, hard deadlines, or superiority over every optimized numerical library. Those limitations constrain the headline, not the usefulness of having a genuinely small service.

Transfer the numerical contract to a target device

This is a concrete small-runtime application of the theory. It ran on a powerful workstation, so it should not yet be advertised as a measured microcontroller or low-energy deployment. The next engineering transfer would preserve its numerical request contract while measuring the actual target device.

The application opportunity

An independently usable calculation engine can sit beside a learned model: validate candidate physical parameters, supply precise targets, or answer a bounded downstream query. In this project, the next component checks model-ambiguity warnings without trusting the search that produced them. Financial utility still depends on valid observations and models; the real-data study did not cross that gate. Small computation is a capability to build with, not a promise of profitable decisions.

Evidence & further reading

The links below distinguish the project record from foundational literature. This revised story does not add a new application-validation experiment.

  1. Single-thread deployment succeeds on the shared workstation. Spline research archive (2026). Local archive snapshot.
  2. Compact operator representations for option pricing and risk. Spline research archive (2026). Local archive snapshot.
  3. Compact, verifiable operator-based options risk. Spline research archive (2026). Local archive snapshot.