Compact computation and model ambiguity in option risk
Historical source. Some claims in older records were subsequently corrected. The associated article states the adopted interpretation. This record preserves the original source alongside its rendered reading view.
Rendered archival Markdown
This reading view preserves headings, tables, lists, code fragments and mathematical notation from the local research record.
Compact computation and model ambiguity in option risk
Executive assessment
The compact operator representation makes the declared nine-model risk calculation substantially less expensive, but the tested stress catalog does not meet the application-usefulness criterion. Both numerical methods complete all 360 requests and converge on all 3,240 candidate fits. Their eligibility decisions agree exactly. Median complete Python request time is 164.94 milliseconds for the compact engine versus 3.704 seconds for the nested finite-element comparator, a paired median ratio of 22.225.[^12]
The application limitation is decisive. On the combined out-of-catalog families, false-reassurance events for point and finite-bump gamma fall from 311 to 190, a 38.91% reduction rather than the required 50%. Finite-bump and move-output actionability improves from 24.54% to 34.77%, but remains below the required 50%. Across the whole cohort, 133 of 360 panels have no compatible catalog member. The result supports efficient model-sensitivity computation for a declared restricted family, not a generally reliable model-ambiguity risk service or a demonstrated trading breakthrough.
Research question and scope
An inexpensive option-pricing engine is useful only if its outputs answer the intended risk question. Numerical accuracy, sensitivity to quote noise and sensitivity to the assumed volatility model are different properties. A solver can compute a derivative accurately while the available observations leave that derivative poorly determined. This investigation asks whether a compact operator-derived representation can make explicit model-sensitivity diagnostics affordable, and whether those diagnostics remain useful on models outside the chosen stress catalog.
The comparison is deliberately synthetic. Ground-truth volatility fields, independent price references and controlled observation noise make it possible to separate numerical error from model error. No market dataset, execution record, transaction-cost assumption or physical-return forecast enters the study. Consequently, even successful truth coverage would establish a model-based diagnostic capability, not profitable trading or reliable live hedging.
The central comparison is between one fitted model and nine independently fitted model hypotheses. A compact exact-cell engine and an ordinary nested finite-element implementation perform the same calculation. The stress catalog, optimizer, observation noise, portfolios, output tolerances and continuation criterion were fixed before the application cohort was evaluated. The application has 15 independent volatility-field draws, each observed at three maturities with eight independent quote-noise realizations: 45 reference states and 360 noisy panels per numerical representation.[^1]
Prior evidence and scientific motivation
The preceding application investigation established compact computation of prices, sensitivities and signed-portfolio risk for a restricted one-dimensional diffusion family. A native implementation matched 144 tested Python portfolio comparison requests and completed the six-portfolio request with median compute time of 1.765 milliseconds and median round trip of 2.904 milliseconds on the local Apple M4 Max. Its measured process memory was approximately 2.05 MiB. Those are measurements of that native program and workload, not of the larger nine-model Python calculation examined here.[^2]
That earlier investigation also exposed a substantive limitation. Conditional quote-noise propagation could be well calibrated about a fitted model's own center while missing the true risk of a different generating model. A thin-layer construction sharpened the issue: a narrow change in local volatility could become almost invisible to the tested option-price panel while changing point gamma substantially. At the narrowest tested layer, the point-gamma ratios approached four and one quarter, whereas the fixed 0.5% bump-gamma changes were below 0.384%.[^2]
The associated bounded-domain argument holds two positive diffusion coefficients fixed while the width of a central layer tends to zero. For positive maturity, prices converge to the background-model prices, but gamma at the layer center retains the inverse local-coefficient factor. Fixed, nonzero spot-bump responses converge to the background response. The result distinguishes the limit of shrinking coefficient support from the limit of shrinking the risk-measurement bump. It does not prove that finite-bump risk is identifiable from every finite quote panel, nor is external peer review or priority established for the argument.[^3]
These findings motivate output-specific assessment. The intended application is a risk service that reports which quantities are stable under its declared model alternatives, which are sensitive and when the calculation is incomplete. Reporting a wider range is not automatically useful: an interval covering almost everything may prevent false confidence while offering little decision support. The present study therefore scores coverage and narrowness together, with failures and abstentions retained in the denominators.
Operator representation and ordinary comparator
The financial model is a diffusion in the drift-normalized state Xt=Stexp(−(r−q)t)/Sref, with dXt=σ(logXt)XtdWt, initial state one, rate r=0.03 and dividend yield q=0.01. Quotes are normalized by the fixed reference spot; their log strikes are relative to its maturity-forward value. Price and finite-move changes are dimensionless normalized prices, while normalized gamma is Sref times ordinary spot gamma. Each fitted request concerns one maturity. Independently fitting three maturities does not establish a single coherent time-dependent volatility model.
For the transformed one-dimensional diffusion problem, the numerical core solves a family of resolvent equations at complex inversion frequencies. Within a cell of constant coefficient, the homogeneous solutions are exponentials. Their boundary values and fluxes determine an exact local dynamic-stiffness relation. Assembly joins these local relations into a complex symmetric tridiagonal system; selected Green-function rows and their parameter derivatives provide prices and risk outputs.[^4]
For cell width h and transformed coefficient Q, the exact entries are D=pcoth(ph) and E=−pcsch(ph), with p=Q. The matched linear finite-element entries are D=1/h+Qh/3 and E=−1/h+Qh/6. Both methods use the same tridiagonal machinery and inversion rule. The comparator uses nested coefficient-aligned meshes and a fixed Richardson combination, rather than deliberately coarse finite elements. Construction, derivatives and all requested outputs are charged to both implementations.
The compact method removes interior spatial degrees of freedom by solving the local differential equation analytically. This is the relevant connection to the operator-spline toolbox: representation follows the operator and supports exact local calculus and economical assembly. The calculation does not use a cardinal KAN activation, Cox-de Boor recursion or a tensor-core matrix-multiplication reformulation. Its heterogeneous bounded-domain system is not circulant, so FFT inversion of a circulant matrix is not the measured source of acceleration.
Five log-volatility parameters describe the base field. Each stress hypothesis multiplies volatility by a fixed factor over a fixed central interval, without adding fitted parameters. The nine hypotheses are the unstressed field and the eight combinations of half-widths 0.04, 0.004, 0.0004 and 0.00004 with multipliers 0.5 and 2. The same base parameters are fitted afresh for every hypothesis and for each numerical method. No fitted state is transferred from one representation to the other.
The risk outputs are normalized price, delta, normalized point gamma, finite-bump gamma and the two finite-move price changes. The bump is 0.5% of the initial normalized spot. Bumped prices are evaluated in the same fixed volatility field; the field is not recentered when spot moves. Three signed portfolios are used: an at-the-money call, a vertical spread and a butterfly. Portfolio sensitivities are aggregated with their signs before noise standard deviations are calculated.
Numerical qualification
Study01 tested three fixed base-volatility vectors, three maturities and all nine stresses, giving 81 admission states. Independent 60-decimal contour inversions, inversion-order checks and domain checks tested prices and the six risk outputs. Value tolerances were 10−8 for normalized price, 10−7 for delta, 10−5 for point gamma, 10−4 for finite-bump gamma and 10−8 for each move-price change. Reference refinement had to be tighter by a factor of ten.[^5]
All value gates passed for both representations. The original derivative self-check, however, admitted 81 of 81 compact cases and only 79 of 81 finite-element cases. The two finite-element point-gamma Jacobian differences were approximately 1.53×10−5 and 1.01×10−5 against a 10−5 scaled-error threshold. These failures remain part of the record; they were not removed by changing the original tolerance or rerunning an altered solver.
Study02 independently audited the derivatives on the same 81 states. Sixty-decimal reference calculations used two parameter-difference steps and Richardson extrapolation. Both unchanged methods then passed 81 of 81 independently qualified derivative checks. The maximum scaled point-gamma Jacobian error was approximately 5.89×10−11 for the compact representation and 3.91×10−7 for finite elements. This supports cancellation in the original double-precision self-difference diagnostic, rather than a detected analytic-Jacobian error. It is additional numerical evidence on the same states, not 81 new independent application examples.[^6]
The new application references were separately qualified before fitting. Layered truths used independent 48- and 64-node high-precision inversions. Smooth truths used 2,048-, 4,096- and 8,192-cell coefficient approximations, successive Richardson estimates, inversion-order checks and domain checks. All 45 states qualified. The largest smooth-reference finite-bump-gamma refinement difference was approximately 4.43×10−7, below its 10−5 reference threshold.[^7]
Observation cohort and uncertainty definitions
There are three independent field draws in each of five families: an unstressed base field, a stress included in the catalog, an unseen central stress, a shifted stress and a smooth volatility field. Each field is evaluated at maturities 0.03, 0.30 and 1.20. The unseen, shifted and smooth families deliberately test model-class limitations rather than only interpolation within a fitted family.
Each panel contains 21 equispaced log-strike quotes from -0.30 to 0.30. Independent Gaussian noise has normalized-price standard deviation 5×10−5(1+∣y∣/0.30). Twenty midpoint strikes are held out from fitting. Negative noisy observations are retained. This convenient synthetic noise model is not asserted to describe market bid-ask spreads, asynchronous quotes or arbitrage violations.
Every fit starts from volatility 0.20 and has volatility bounds 0.08 to 0.60. The objective combines whitened quote residuals, second-difference regularization and a weak penalty toward the starting log-volatility. The same bounded trust-region optimizer, tolerances and limit of 80 function evaluations apply to every fit. There are no restarts, regularization sweeps or post-outcome stress additions.[^1]
Conditional observation-noise propagation differentiates the penalized calibration map. If J is the whitened price Jacobian, H is the full objective Hessian and B is a portfolio-output Jacobian, the local influence matrix is BH−1J⊤. The row norms give conditional standard deviations for independent unit-variance whitened quote noise. The Hessian includes residual curvature and penalty curvature, with a two-step refinement check and a positive-definiteness requirement. This is not the same as substituting a posterior covariance BH−1B⊤.
A converged, numerically qualified fitted representative is quote-compatible when its unpenalized squared whitened residual is below the fixed 99th percentile of a chi-square distribution with 21 degrees of freedom. This is a declared compatibility screen, not a fitted-model confidence theorem. Each eligible representative contributes its fitted output plus or minus 1.96 conditional standard deviations. Their lower and upper extrema form a finite stress envelope; the unexpanded fitted-output range is also retained.
The base-only control uses exactly the same numerical and compatibility requirements. An empty envelope is an abstention, not a zero-width interval. An all-nine request is computationally complete only if every fit converges with finite outputs and every compatible representative has a qualified conditional diagnostic. A converged representative rejected by the quote screen does not need a conditional interval, but its rejection cannot rule out every other parameter setting in that model family.
These envelopes omit untested hypotheses and the within-family set of all quote-compatible parameters. Neither the factor 1.96 nor taking a union converts them into certified bounds, posterior credible regions or guaranteed 95% coverage intervals. Their practical value must be established empirically against the declared synthetic truth distribution.
Scoring and uncertainty accounting
Unconditional truth coverage counts unavailable intervals as not covered; conditional-on-availability coverage is reported separately. Actionability means a half-width below a fixed output-specific tolerance. False reassurance means an actionable interval that excludes the true output. A reduction in false reassurance can arise because a diagnostic becomes more accurate, wider or unavailable, so the full combination of coverage, availability, width and actionability matters.
Per unit of portfolio L1 exposure, the actionability tolerances are 10−4 for price, 0.01 for delta, 0.25 for both gamma definitions and 5×10−5 for each finite-move price change. These are declared engineering tolerances, not a trading desk's validated economic utility function. Both raw results and resolution-qualified results are retained; the latter use the same base-model standard-deviation mask when comparing base and stress estimates within one representation.
The continuation criterion requires at least 90% computationally complete requests. On the combined unseen-central, shifted and smooth families, it also requires at least a 50% reduction in false-reassurance frequency for point and finite-bump gamma, at least 20 baseline false-reassurance events, and finite-bump/move-output actionability of at least 50% and at least 80% of the base-only level. This is a project continuation rule, not a statistical discovery threshold.
Uncertainty summaries resample the 15 independent fields in 2,000 seeded bootstrap draws, retaining all maturities, noise panels and portfolios of a field together. The resulting percentile intervals are descriptive: there are only three independent fields per family. Neither 360 noisy panels nor 3,240 candidate fits per representation should be reported as independent model cases. All family and output breakdowns are retained regardless of whether the pooled continuation criterion passes.
Application results and transfer limits
All 3,240 candidate fits converge under each representation. There are 959 quote-compatible contributors per method and 2,894 conditionally qualified fits in total. The remaining 346 fits are incompatible with the quote screen, so their unavailable conditional diagnostics do not violate the declared definition of computational completeness. This distinction is operational: 360 completed requests do not imply 360 available risk ranges.[^12]
The base-only control provides a range on 169 of 360 panels; the stress catalog provides one on 227. The 133 stress abstentions comprise two unseen-central panels, all 72 shifted panels and 59 smooth panels. The following table averages the three fixed portfolios equally. Each family has 72 noisy panels but only three independent generating fields, and unavailable intervals count as not covered.
Truth family
Stress availability
Point-gamma coverage: base / stress
Finite-gamma coverage: base / stress
Base
100%
96.76% / 100%
96.76% / 100%
Included central
100%
0% / 97.22%
18.06% / 97.69%
Unseen central
97.22%
0% / 70.83%
3.24% / 50.93%
Shifted
0%
0% / 0%
0% / 0%
Smooth
18.06%
0% / 2.78%
0% / 4.17%
The included-family result verifies a limited mechanism: when the stress family contains a relevant alternative, accounting for that alternative can greatly improve truth coverage relative to a misspecified base model. The unseen-central result demonstrates partial transfer to unlisted widths and amplitudes, but its finite-gamma coverage remains only 50.93%. Neither subgroup establishes the full application claim.
There are two distinct out-of-family failures. The shifted fields are rejected entirely, making their missing risk estimates explicit. The smooth fields are mostly rejected, but acceptance is not sufficient either: among their 13 accepted panels, pooled conditional point-gamma coverage is only 15.38%, and finite-gamma coverage only 23.08%. The diagnostic can therefore return a narrow wrong interval even after quote compatibility has been checked. Expanding the catalog increases false-reassurance frequency on this smooth family rather than reliably correcting it.
Across all five families, pooled unconditional coverage increases from 36.30% to 54.17% for price, 37.41% to 53.70% for delta, 19.35% to 54.17% for point gamma and 23.61% to 50.56% for finite-bump gamma. The downward and upward move-price channels increase from 30.19% and 29.81% to 52.69% and 52.22%, respectively. These gains coexist with substantial abstention and cannot be interpreted as nominal 95% uncertainty calibration.
Width also matters where the true field belongs to the base family. There, stress coverage reaches 100% for both gamma channels, but point-gamma actionability falls from 100% to 40.74%, and finite-gamma actionability to 55.56%. More conservative intervals can improve coverage at the cost of precision. A service intended to support a decision must make this tradeoff explicit rather than optimize coverage alone.
The field-cluster bootstrap yields an overall point-gamma coverage gain of 34.81 percentage points, with a descriptive 95% percentile interval from 14.63 to 58.52 points. The point-gamma false-reassurance difference is -19.07 points, but its interval spans -41.57 to +2.31 points. For finite gamma the corresponding differences are +26.94 points, interval +11.11 to +47.41, and -13.33 points, interval -30.74 to +3.80. The false-reassurance intervals crossing zero reinforce the limited independent evidence; they do not overturn the prespecified failed continuation gate.
Computational cost and numerical agreement
The comparison charges construction, all nine independent calibrations, Jacobian calculations, full-Hessian noise diagnostics, signed-portfolio aggregation and rejected candidates. There is no best-repeat selection. The compact and finite-element methods run sequentially with rotated order on the same shared Apple M4 Max host. These are Python computation timings, not native-service round trips, network latency or energy measurements.
Measured workload
Compact engine
Nested finite elements
Nine-fit request, median
164.94 ms
3,704.30 ms
Nine-fit request, empirical p99
196.23 ms
4,337.90 ms
Base-only fit and diagnostic, median
18.00 ms
405.73 ms
Converged candidate fits
3,240 / 3,240
3,240 / 3,240
Computationally complete requests
360 / 360
360 / 360
The paired request-time ratio has median 22.225 and ranges from 16.006 to 24.485. This is a scoped comparison against the qualified implementation used here. It does not establish superiority over optimized production libraries, all finite-element methods or existing fast calibration schemes. The larger calculation should not inherit the preceding native service's approximately two-MiB process-memory claim.
The application driver has peak resident memory of approximately 157.3 MiB while executing both implementations sequentially. Numerical admission measured retained model arrays of roughly 65 kB for the compact representation versus 9.74 MB for nested finite elements. Retained arrays and process peak memory are different quantities; neither measurement establishes a low-power-device or energy advantage without a target-device test.
All 9,720 paired portfolio values per output channel are finite. Maximum differences between the independently fitted methods are approximately 6.71×10−9 for price, 1.49×10−7 for delta, 1.59×10−5 for point gamma, 5.79×10−6 for finite gamma and less than 7.78×10−10 for the move-price channels. Because the two optimizers fit independently, these differences combine parameter differences with numerical approximation error; they are not same-parameter error estimates. The two methods nevertheless make identical eligibility decisions and reach the same failed application gate.
A separately frozen fitted-state audit checks every model contributing to a reported stress range. All 959 contributors per representation pass independent high-precision value, reference-refinement, contour and domain checks at their own fitted parameters. The maximum point-gamma value error per unit portfolio L1 exposure is approximately 2.00×10−10 for the compact engine and 6.38×10−8 for finite elements; finite-bump gamma errors are 2.15×10−9 and 1.95×10−8. This additional evidence qualifies the reported risk values, not the statistical meaning of their uncertainty envelopes. Its 606.58-second offline reference cost is separate from the online request timings.[^13]
Prior art and interpretation boundaries
Model uncertainty in calibration and robust derivative pricing is established prior art. Gupta and Reisinger discuss Bayesian calibration, model-class assumptions, prior influence and robust pricing. The present finite envelope does not implement their posterior construction, and does not claim novelty for considering multiple quote-compatible models. Its narrower question is whether the compact numerical representation makes a specified diagnostic less costly and whether that diagnostic is useful.[^8]
Cont's earlier framework measures model uncertainty in derivative portfolios using coherent and convex risk measures compatible with benchmark prices. This reinforces the distinction between a numerical engine and a model-risk methodology: fast evaluation alone does not establish a new risk measure or an exhaustive ambiguity set.[^10]
Goal-oriented inverse-problem methods also provide an important comparison. Spantini and colleagues derive low-rank approximations aimed at specified quantities of interest in linear-Gaussian problems, rather than requiring accurate recovery of every parameter direction. Their guarantees do not transfer automatically to nonlinear option calibration, unknown volatility families or finite stress catalogs. Output-specific inference is therefore a relevant established framework, not a novelty claim arising merely from using portfolio outputs here.[^9]
Another relevant practical baseline is pre-calibration smoothing. Yang and colleagues propose automatic local regression of implied-volatility quotes before local-volatility calibration, with Greek stability as an application objective. Such a method would be a meaningful comparator for future end-to-end calibration work. Smoothness and numerical stability, however, must not be equated with identification of the true local dynamics; the thin-layer example illustrates why this distinction matters.[^11]
Application decision and next evidence
This catalog comparison is closed as a failed application continuation gate with a substantial, narrowly demonstrated computational advantage. Adding shifted stresses after inspecting these outcomes would construct a different model search, not validate the frozen claim. Loosening quote compatibility would likewise change the acceptance contract without solving the smooth-field undercoverage. Neither intervention is justified as a rescue of this result.
The most useful surviving application is narrower: an inspectable component that rapidly recomputes specified portfolios under explicitly declared volatility scenarios, reports numerical qualification and separates conditional quote noise from scenario spread. In that role, the engine need not pretend that a chosen catalog exhausts market uncertainty. The current implementation and synthetic tests support further evaluation of that calculation capability, but do not establish demand, a validated production workflow or reliable hedging decisions.
The higher-impact scientific target is output-specific identification under a defensible model class. That requires explicit restrictions on coefficient regularity, location, scale and temporal evolution, then an assessment of which portfolio responses can be bounded or estimated from the available observations. A collection of fitted representatives is not a substitute for that argument. Finite-move outputs may offer a more useful contract than point gamma in some settings, but the present failures rule out assuming that changing the output alone solves model ambiguity.
The next practical evidence should be tied to one such contract. A genuine modest-CPU benchmark can test whether the computational advantage survives deployment constraints. A separately frozen, licensed or already authorized quote cohort can test observation handling and model adequacy, with established calibration baselines. Neither should be launched merely to avoid the present negative finding. Market data are unnecessary to establish the numerical result or this catalog's synthetic failure, but become necessary before claiming live-market usefulness.
The broader operator-spline vision remains a hypothesis about economical continuous representations and usable calculus. The evidence now identifies both a concrete strength and a boundary: compact exact-cell computations can make a larger sensitivity calculation affordable, while the information content of the observations and the model-class assumptions remain separate bottlenecks. Solving that boundary, rather than accumulating alternative catalog members, is the substantive research challenge.
Sources
[^1]: Project research record, Study03: independent synthetic model-ambiguity application, pricing_model_ambiguity_20260916/PROTOCOL_03.md, with frozen cohort.py, ambiguity_fit.py, run_application.py and summarize_application.py. Repository-local primary evidence, September 2026. [^2]: Project research record, Compact quote-to-portfolio risk with operator-derived representations, pricing_application_20260916/REPORT.md, native portfolio and thin-layer findings. Completed repository-local investigation, September 2026; not an external benchmark. [^3]: Project research note, pricing_application_20260916/THIN_LAYER_LIMIT_THEOREM.md, bounded-domain thin-layer limit and proof. Repository-local mathematical argument, September 2026. [^4]: Project numerical implementation, pricing_model_ambiguity_20260916/stress_models.py and the hash-identified selected-chain, parameter-chain and diffusion implementations imported from the preceding investigations. Repository-local source evidence. [^5]: Project research record, pricing_model_ambiguity_20260916/PROTOCOL_01.md, FINDINGS_01.md and RESULTS_01.json. Original admission results, including both failed self-difference checks. [^6]: Project research record, pricing_model_ambiguity_20260916/PROTOCOL_02.md, FINDINGS_02.md and RESULTS_02.json. Independent high-precision derivative audit of unchanged methods. [^7]: Project research record, pricing_model_ambiguity_20260916/COHORT_03_FINDINGS.md, COHORT_03.json and cohort03/. Independent synthetic truth generation, refinement evidence and saved observations. [^8]: A. Gupta and C. Reisinger, Robust Calibration of Financial Models Using Bayesian Estimators. Author manuscript dated February 7, 2012; journal publication 2014, 10.21314/JCF.2014.285. Author manuscript, opening discussion and Sections 2-3.1, beginning of 3.2. [^9]: A. Spantini, T. Cui, K. Willcox, L. Tenorio and Y. Marzouk, Goal-oriented optimal approximations of Bayesian linear inverse problems, arXiv version 2, March 14, 2017. Primary manuscript, introduction and selected theory passages in Sections 2.1-2.3. [^10]: R. Cont, Model Uncertainty and Its Impact on the Pricing of Derivative Instruments, Mathematical Finance 16(3), 519-547, 2006. Publisher abstract and metadata. [^11]: R. Yang, H. Qin, C. Che and L. Feng, Volatility Calibration via Automatic Local Regression, arXiv version 1, September 19, 2025. Primary manuscript, introduction, existing-method discussion and opening methodology. [^12]: Project research record, pricing_model_ambiguity_20260916/RESULTS_03.json, ANALYSIS_03.json, FINDINGS_03.md and the full per-case study03/ records. Frozen 360-panel paired application, September 2026; synthetic data, 15 independent field clusters. [^13]: Project research record, pricing_model_ambiguity_20260916/PROTOCOL_04.md, RESULTS_04.json, FINDINGS_04.md and study04/. Independent value audit of all 1,918 contributing fitted states, with no changes to application fits or scoring.
Original: research/pricing_model_ambiguity_20260916/REPORT.md · Raw source file
View raw MD source
# Compact computation and model ambiguity in option risk
## Executive assessment
The compact operator representation makes the declared nine-model risk
calculation substantially less expensive, but the tested stress catalog does
not meet the application-usefulness criterion. Both numerical methods
complete all 360 requests and converge on all 3,240 candidate fits. Their
eligibility decisions agree exactly. Median complete Python request time is
164.94 milliseconds for the compact engine versus 3.704 seconds for the
nested finite-element comparator, a paired median ratio of 22.225.[^12]
The application limitation is decisive. On the combined out-of-catalog
families, false-reassurance events for point and finite-bump gamma fall from
311 to 190, a 38.91% reduction rather than the required 50%. Finite-bump and
move-output actionability improves from 24.54% to 34.77%, but remains below
the required 50%. Across the whole cohort, 133 of 360 panels have no
compatible catalog member. The result supports efficient model-sensitivity
computation for a declared restricted family, not a generally reliable
model-ambiguity risk service or a demonstrated trading breakthrough.
## Research question and scope
An inexpensive option-pricing engine is useful only if its outputs answer the
intended risk question. Numerical accuracy, sensitivity to quote noise and
sensitivity to the assumed volatility model are different properties. A solver
can compute a derivative accurately while the available observations leave
that derivative poorly determined. This investigation asks whether a compact
operator-derived representation can make explicit model-sensitivity diagnostics
affordable, and whether those diagnostics remain useful on models outside the
chosen stress catalog.
The comparison is deliberately synthetic. Ground-truth volatility fields,
independent price references and controlled observation noise make it possible
to separate numerical error from model error. No market dataset, execution
record, transaction-cost assumption or physical-return forecast enters the
study. Consequently, even successful truth coverage would establish a
model-based diagnostic capability, not profitable trading or reliable live
hedging.
The central comparison is between one fitted model and nine independently
fitted model hypotheses. A compact exact-cell engine and an ordinary nested
finite-element implementation perform the same calculation. The stress catalog,
optimizer, observation noise, portfolios, output tolerances and continuation
criterion were fixed before the application cohort was evaluated. The
application has 15 independent volatility-field draws, each observed at three
maturities with eight independent quote-noise realizations: 45 reference states
and 360 noisy panels per numerical representation.[^1]
## Prior evidence and scientific motivation
The preceding application investigation established compact computation of
prices, sensitivities and signed-portfolio risk for a restricted one-dimensional
diffusion family. A native implementation matched 144 tested Python portfolio
comparison requests and completed the six-portfolio request with median compute time of
1.765 milliseconds and median round trip of 2.904 milliseconds on the local
Apple M4 Max. Its measured process memory was approximately 2.05 MiB. Those
are measurements of that native program and workload, not of the larger
nine-model Python calculation examined here.[^2]
That earlier investigation also exposed a substantive limitation. Conditional
quote-noise propagation could be well calibrated about a fitted model's own
center while missing the true risk of a different generating model. A
thin-layer construction sharpened the issue: a narrow change in local
volatility could become almost invisible to the tested option-price panel
while changing point gamma substantially. At the narrowest tested layer,
the point-gamma ratios approached four and one quarter, whereas the fixed
0.5% bump-gamma changes were below 0.384%.[^2]
The associated bounded-domain argument holds two positive diffusion
coefficients fixed while the width of a central layer tends to zero. For
positive maturity, prices converge to the background-model prices, but gamma
at the layer center retains the inverse local-coefficient factor. Fixed,
nonzero spot-bump responses converge to the background response. The result
distinguishes the limit of shrinking coefficient support from the limit of
shrinking the risk-measurement bump. It does not prove that finite-bump risk
is identifiable from every finite quote panel, nor is external peer review
or priority established for the argument.[^3]
These findings motivate output-specific assessment. The intended application
is a risk service that reports which quantities are stable under its declared
model alternatives, which are sensitive and when the calculation is
incomplete. Reporting a wider range is not automatically useful: an interval
covering almost everything may prevent false confidence while offering little
decision support. The present study therefore scores coverage and narrowness
together, with failures and abstentions retained in the denominators.
## Operator representation and ordinary comparator
The financial model is a diffusion in the drift-normalized state
$X_t=S_t\exp(-(r-q)t)/S_{ref}$, with
$dX_t=\sigma(\log X_t)X_t\,dW_t$, initial state one, rate $r=0.03$ and
dividend yield $q=0.01$. Quotes are normalized by the fixed reference spot;
their log strikes are relative to its maturity-forward value. Price and
finite-move changes are dimensionless normalized prices, while normalized
gamma is $S_{ref}$ times ordinary spot gamma. Each fitted request concerns
one maturity. Independently fitting three maturities does not establish a
single coherent time-dependent volatility model.
For the transformed one-dimensional diffusion problem, the numerical core
solves a family of resolvent equations at complex inversion frequencies.
Within a cell of constant coefficient, the homogeneous solutions are
exponentials. Their boundary values and fluxes determine an exact local
dynamic-stiffness relation. Assembly joins these local relations into a
complex symmetric tridiagonal system; selected Green-function rows and their
parameter derivatives provide prices and risk outputs.[^4]
For cell width $h$ and transformed coefficient $Q$, the exact entries are
$D=p\coth(ph)$ and $E=-p\operatorname{csch}(ph)$, with $p=\sqrt Q$.
The matched linear finite-element entries are $D=1/h+Qh/3$ and
$E=-1/h+Qh/6$. Both methods use the same tridiagonal machinery and inversion
rule. The comparator uses nested coefficient-aligned meshes and a fixed
Richardson combination, rather than deliberately coarse finite elements.
Construction, derivatives and all requested outputs are charged to both
implementations.
The compact method removes interior spatial degrees of freedom by solving
the local differential equation analytically. This is the relevant connection
to the operator-spline toolbox: representation follows the operator and
supports exact local calculus and economical assembly. The calculation does
not use a cardinal KAN activation, Cox-de Boor recursion or a tensor-core
matrix-multiplication reformulation. Its heterogeneous bounded-domain system
is not circulant, so FFT inversion of a circulant matrix is not the measured
source of acceleration.
Five log-volatility parameters describe the base field. Each stress hypothesis
multiplies volatility by a fixed factor over a fixed central interval, without
adding fitted parameters. The nine hypotheses are the unstressed field and
the eight combinations of half-widths 0.04, 0.004, 0.0004 and 0.00004 with
multipliers 0.5 and 2. The same base parameters are fitted afresh for every
hypothesis and for each numerical method. No fitted state is transferred
from one representation to the other.
The risk outputs are normalized price, delta, normalized point gamma,
finite-bump gamma and the two finite-move price changes. The bump is 0.5% of
the initial normalized spot. Bumped prices are evaluated in the same fixed
volatility field; the field is not recentered when spot moves. Three signed
portfolios are used: an at-the-money call, a vertical spread and a butterfly.
Portfolio sensitivities are aggregated with their signs before noise
standard deviations are calculated.
## Numerical qualification
Study01 tested three fixed base-volatility vectors, three maturities and all
nine stresses, giving 81 admission states. Independent 60-decimal contour
inversions, inversion-order checks and domain checks tested prices and the
six risk outputs. Value tolerances were $10^{-8}$ for normalized price,
$10^{-7}$ for delta, $10^{-5}$ for point gamma, $10^{-4}$ for finite-bump
gamma and $10^{-8}$ for each move-price change. Reference refinement had to
be tighter by a factor of ten.[^5]
All value gates passed for both representations. The original derivative
self-check, however, admitted 81 of 81 compact cases and only 79 of 81
finite-element cases. The two finite-element point-gamma Jacobian differences
were approximately $1.53\times10^{-5}$ and $1.01\times10^{-5}$ against a
$10^{-5}$ scaled-error threshold. These failures remain part of the record;
they were not removed by changing the original tolerance or rerunning an
altered solver.
Study02 independently audited the derivatives on the same 81 states.
Sixty-decimal reference calculations used two parameter-difference steps
and Richardson extrapolation. Both unchanged methods then passed 81 of 81
independently qualified derivative checks. The maximum scaled point-gamma
Jacobian error was approximately $5.89\times10^{-11}$ for the compact
representation and $3.91\times10^{-7}$ for finite elements. This supports
cancellation in the original double-precision self-difference diagnostic,
rather than a detected analytic-Jacobian error. It is additional numerical
evidence on the same states, not 81 new independent application examples.[^6]
The new application references were separately qualified before fitting.
Layered truths used independent 48- and 64-node high-precision inversions.
Smooth truths used 2,048-, 4,096- and 8,192-cell coefficient approximations,
successive Richardson estimates, inversion-order checks and domain checks.
All 45 states qualified. The largest smooth-reference finite-bump-gamma
refinement difference was approximately $4.43\times10^{-7}$, below its
$10^{-5}$ reference threshold.[^7]
## Observation cohort and uncertainty definitions
There are three independent field draws in each of five families: an
unstressed base field, a stress included in the catalog, an unseen central
stress, a shifted stress and a smooth volatility field. Each field is
evaluated at maturities 0.03, 0.30 and 1.20. The unseen, shifted and smooth
families deliberately test model-class limitations rather than only
interpolation within a fitted family.
Each panel contains 21 equispaced log-strike quotes from -0.30 to 0.30.
Independent Gaussian noise has normalized-price standard deviation
$5\times10^{-5}(1+|y|/0.30)$. Twenty midpoint strikes are held out from
fitting. Negative noisy observations are retained. This convenient synthetic
noise model is not asserted to describe market bid-ask spreads, asynchronous
quotes or arbitrage violations.
Every fit starts from volatility 0.20 and has volatility bounds 0.08 to
0.60. The objective combines whitened quote residuals, second-difference
regularization and a weak penalty toward the starting log-volatility.
The same bounded trust-region optimizer, tolerances and limit of 80 function
evaluations apply to every fit. There are no restarts, regularization sweeps
or post-outcome stress additions.[^1]
Conditional observation-noise propagation differentiates the penalized
calibration map. If $J$ is the whitened price Jacobian, $H$ is the full
objective Hessian and $B$ is a portfolio-output Jacobian, the local
influence matrix is $BH^{-1}J^\top$. The row norms give conditional standard
deviations for independent unit-variance whitened quote noise. The Hessian
includes residual curvature and penalty curvature, with a two-step
refinement check and a positive-definiteness requirement. This is not the
same as substituting a posterior covariance $BH^{-1}B^\top$.
A converged, numerically qualified fitted representative is quote-compatible
when its unpenalized squared whitened residual is below the fixed 99th
percentile of a chi-square distribution with 21 degrees of freedom.
This is a declared compatibility screen, not a fitted-model confidence
theorem. Each eligible representative contributes its fitted output plus or
minus 1.96 conditional standard deviations. Their lower and upper extrema
form a finite stress envelope; the unexpanded fitted-output range is also
retained.
The base-only control uses exactly the same numerical and compatibility
requirements. An empty envelope is an abstention, not a zero-width interval.
An all-nine request is computationally complete only if every fit converges
with finite outputs and every compatible representative has a qualified
conditional diagnostic. A converged representative rejected by the quote
screen does not need a conditional interval, but its rejection cannot rule
out every other parameter setting in that model family.
These envelopes omit untested hypotheses and the within-family set of all
quote-compatible parameters. Neither the factor 1.96 nor taking a union
converts them into certified bounds, posterior credible regions or
guaranteed 95% coverage intervals. Their practical value must be established
empirically against the declared synthetic truth distribution.
## Scoring and uncertainty accounting
Unconditional truth coverage counts unavailable intervals as not covered;
conditional-on-availability coverage is reported separately. Actionability
means a half-width below a fixed output-specific tolerance. False reassurance
means an actionable interval that excludes the true output. A reduction in
false reassurance can arise because a diagnostic becomes more accurate,
wider or unavailable, so the full combination of coverage, availability,
width and actionability matters.
Per unit of portfolio L1 exposure, the actionability tolerances are
$10^{-4}$ for price, 0.01 for delta, 0.25 for both gamma definitions and
$5\times10^{-5}$ for each finite-move price change. These are declared
engineering tolerances, not a trading desk's validated economic utility
function. Both raw results and resolution-qualified results are retained;
the latter use the same base-model standard-deviation mask when comparing
base and stress estimates within one representation.
The continuation criterion requires at least 90% computationally complete
requests. On the combined unseen-central, shifted and smooth families, it
also requires at least a 50% reduction in false-reassurance frequency for
point and finite-bump gamma, at least 20 baseline false-reassurance events,
and finite-bump/move-output actionability of at least 50% and at least 80%
of the base-only level. This is a project continuation rule, not a
statistical discovery threshold.
Uncertainty summaries resample the 15 independent fields in 2,000 seeded
bootstrap draws, retaining all maturities, noise panels and portfolios of a
field together. The resulting percentile intervals are descriptive: there
are only three independent fields per family. Neither 360 noisy panels nor
3,240 candidate fits per representation should be reported as independent
model cases. All family and output breakdowns are retained regardless of
whether the pooled continuation criterion passes.
## Application results and transfer limits
All 3,240 candidate fits converge under each representation. There are 959
quote-compatible contributors per method and 2,894 conditionally qualified
fits in total. The remaining 346 fits are incompatible with the quote screen,
so their unavailable conditional diagnostics do not violate the declared
definition of computational completeness. This distinction is operational:
360 completed requests do not imply 360 available risk ranges.[^12]
The base-only control provides a range on 169 of 360 panels; the stress
catalog provides one on 227. The 133 stress abstentions comprise two
unseen-central panels, all 72 shifted panels and 59 smooth panels. The
following table averages the three fixed portfolios equally. Each family
has 72 noisy panels but only three independent generating fields, and
unavailable intervals count as not covered.
| Truth family | Stress availability | Point-gamma coverage: base / stress | Finite-gamma coverage: base / stress |
|---|---|---|---|
| Base |100%|96.76% / 100%|96.76% / 100%|
| Included central |100%|0% / 97.22%|18.06% / 97.69%|
| Unseen central |97.22%|0% / 70.83%|3.24% / 50.93%|
| Shifted |0%|0% / 0%|0% / 0%|
| Smooth |18.06%|0% / 2.78%|0% / 4.17%|
The included-family result verifies a limited mechanism: when the stress
family contains a relevant alternative, accounting for that alternative
can greatly improve truth coverage relative to a misspecified base model.
The unseen-central result demonstrates partial transfer to unlisted widths
and amplitudes, but its finite-gamma coverage remains only 50.93%.
Neither subgroup establishes the full application claim.
There are two distinct out-of-family failures. The shifted fields are
rejected entirely, making their missing risk estimates explicit. The smooth
fields are mostly rejected, but acceptance is not sufficient either: among
their 13 accepted panels, pooled conditional point-gamma coverage is only
15.38%, and finite-gamma coverage only 23.08%. The diagnostic can therefore
return a narrow wrong interval even after quote compatibility has been
checked. Expanding the catalog increases false-reassurance frequency on
this smooth family rather than reliably correcting it.
Across all five families, pooled unconditional coverage increases from
36.30% to 54.17% for price, 37.41% to 53.70% for delta, 19.35% to 54.17%
for point gamma and 23.61% to 50.56% for finite-bump gamma. The downward
and upward move-price channels increase from 30.19% and 29.81% to 52.69%
and 52.22%, respectively. These gains coexist with substantial abstention
and cannot be interpreted as nominal 95% uncertainty calibration.
Width also matters where the true field belongs to the base family. There,
stress coverage reaches 100% for both gamma channels, but point-gamma
actionability falls from 100% to 40.74%, and finite-gamma actionability to
55.56%. More conservative intervals can improve coverage at the cost of
precision. A service intended to support a decision must make this tradeoff
explicit rather than optimize coverage alone.
The field-cluster bootstrap yields an overall point-gamma coverage gain of
34.81 percentage points, with a descriptive 95% percentile interval from
14.63 to 58.52 points. The point-gamma false-reassurance difference is
-19.07 points, but its interval spans -41.57 to +2.31 points. For finite
gamma the corresponding differences are +26.94 points, interval +11.11 to
+47.41, and -13.33 points, interval -30.74 to +3.80. The false-reassurance
intervals crossing zero reinforce the limited independent evidence; they
do not overturn the prespecified failed continuation gate.
## Computational cost and numerical agreement
The comparison charges construction, all nine independent calibrations,
Jacobian calculations, full-Hessian noise diagnostics, signed-portfolio
aggregation and rejected candidates. There is no best-repeat selection.
The compact and finite-element methods run sequentially with rotated order
on the same shared Apple M4 Max host. These are Python computation timings,
not native-service round trips, network latency or energy measurements.
| Measured workload | Compact engine | Nested finite elements |
|---|---|---|
| Nine-fit request, median |164.94 ms|3,704.30 ms|
| Nine-fit request, empirical p99 |196.23 ms|4,337.90 ms|
| Base-only fit and diagnostic, median |18.00 ms|405.73 ms|
| Converged candidate fits |3,240 / 3,240|3,240 / 3,240|
| Computationally complete requests |360 / 360|360 / 360|
The paired request-time ratio has median 22.225 and ranges from 16.006 to
24.485. This is a scoped comparison against the qualified implementation
used here. It does not establish superiority over optimized production
libraries, all finite-element methods or existing fast calibration schemes.
The larger calculation should not inherit the preceding native service's
approximately two-MiB process-memory claim.
The application driver has peak resident memory of approximately 157.3 MiB
while executing both implementations sequentially. Numerical admission
measured retained model arrays of roughly 65 kB for the compact
representation versus 9.74 MB for nested finite elements. Retained arrays
and process peak memory are different quantities; neither measurement
establishes a low-power-device or energy advantage without a target-device
test.
All 9,720 paired portfolio values per output channel are finite. Maximum
differences between the independently fitted methods are approximately
$6.71\times10^{-9}$ for price, $1.49\times10^{-7}$ for delta,
$1.59\times10^{-5}$ for point gamma, $5.79\times10^{-6}$ for finite
gamma and less than $7.78\times10^{-10}$ for the move-price channels.
Because the two optimizers fit independently, these differences combine
parameter differences with numerical approximation error; they are not
same-parameter error estimates. The two methods nevertheless make identical
eligibility decisions and reach the same failed application gate.
A separately frozen fitted-state audit checks every model contributing to
a reported stress range. All 959 contributors per representation pass
independent high-precision value, reference-refinement, contour and domain
checks at their own fitted parameters. The maximum point-gamma value error
per unit portfolio L1 exposure is approximately $2.00\times10^{-10}$ for
the compact engine and $6.38\times10^{-8}$ for finite elements; finite-bump
gamma errors are $2.15\times10^{-9}$ and $1.95\times10^{-8}$.
This additional evidence qualifies the reported risk values, not the
statistical meaning of their uncertainty envelopes. Its 606.58-second
offline reference cost is separate from the online request timings.[^13]
## Prior art and interpretation boundaries
Model uncertainty in calibration and robust derivative pricing is established
prior art. Gupta and Reisinger discuss Bayesian calibration, model-class
assumptions, prior influence and robust pricing. The present finite envelope
does not implement their posterior construction, and does not claim novelty
for considering multiple quote-compatible models. Its narrower question is
whether the compact numerical representation makes a specified diagnostic
less costly and whether that diagnostic is useful.[^8]
Cont's earlier framework measures model uncertainty in derivative portfolios
using coherent and convex risk measures compatible with benchmark prices.
This reinforces the distinction between a numerical engine and a model-risk
methodology: fast evaluation alone does not establish a new risk measure or
an exhaustive ambiguity set.[^10]
Goal-oriented inverse-problem methods also provide an important comparison.
Spantini and colleagues derive low-rank approximations aimed at specified
quantities of interest in linear-Gaussian problems, rather than requiring
accurate recovery of every parameter direction. Their guarantees do not
transfer automatically to nonlinear option calibration, unknown volatility
families or finite stress catalogs. Output-specific inference is therefore
a relevant established framework, not a novelty claim arising merely from
using portfolio outputs here.[^9]
Another relevant practical baseline is pre-calibration smoothing. Yang and
colleagues propose automatic local regression of implied-volatility quotes
before local-volatility calibration, with Greek stability as an application
objective. Such a method would be a meaningful comparator for future
end-to-end calibration work. Smoothness and numerical stability, however,
must not be equated with identification of the true local dynamics; the
thin-layer example illustrates why this distinction matters.[^11]
## Application decision and next evidence
This catalog comparison is closed as a failed application continuation
gate with a substantial, narrowly demonstrated computational advantage.
Adding shifted stresses after inspecting these outcomes would construct a
different model search, not validate the frozen claim. Loosening quote
compatibility would likewise change the acceptance contract without solving
the smooth-field undercoverage. Neither intervention is justified as a
rescue of this result.
The most useful surviving application is narrower: an inspectable component
that rapidly recomputes specified portfolios under explicitly declared
volatility scenarios, reports numerical qualification and separates
conditional quote noise from scenario spread. In that role, the engine need
not pretend that a chosen catalog exhausts market uncertainty. The current
implementation and synthetic tests support further evaluation of that
calculation capability, but do not establish demand, a validated production
workflow or reliable hedging decisions.
The higher-impact scientific target is output-specific identification under
a defensible model class. That requires explicit restrictions on coefficient
regularity, location, scale and temporal evolution, then an assessment of
which portfolio responses can be bounded or estimated from the available
observations. A collection of fitted representatives is not a substitute for
that argument. Finite-move outputs may offer a more useful contract than
point gamma in some settings, but the present failures rule out assuming
that changing the output alone solves model ambiguity.
The next practical evidence should be tied to one such contract. A genuine
modest-CPU benchmark can test whether the computational advantage survives
deployment constraints. A separately frozen, licensed or already authorized
quote cohort can test observation handling and model adequacy, with
established calibration baselines. Neither should be launched merely to
avoid the present negative finding. Market data are unnecessary to establish
the numerical result or this catalog's synthetic failure, but become
necessary before claiming live-market usefulness.
The broader operator-spline vision remains a hypothesis about economical
continuous representations and usable calculus. The evidence now identifies
both a concrete strength and a boundary: compact exact-cell computations can
make a larger sensitivity calculation affordable, while the information
content of the observations and the model-class assumptions remain separate
bottlenecks. Solving that boundary, rather than accumulating alternative
catalog members, is the substantive research challenge.
## Sources
[^1]: Project research record, *Study03: independent synthetic model-ambiguity application*, `pricing_model_ambiguity_20260916/PROTOCOL_03.md`, with frozen `cohort.py`, `ambiguity_fit.py`, `run_application.py` and `summarize_application.py`. Repository-local primary evidence, September 2026.
[^2]: Project research record, *Compact quote-to-portfolio risk with operator-derived representations*, `pricing_application_20260916/REPORT.md`, native portfolio and thin-layer findings. Completed repository-local investigation, September 2026; not an external benchmark.
[^3]: Project research note, `pricing_application_20260916/THIN_LAYER_LIMIT_THEOREM.md`, bounded-domain thin-layer limit and proof. Repository-local mathematical argument, September 2026.
[^4]: Project numerical implementation, `pricing_model_ambiguity_20260916/stress_models.py` and the hash-identified selected-chain, parameter-chain and diffusion implementations imported from the preceding investigations. Repository-local source evidence.
[^5]: Project research record, `pricing_model_ambiguity_20260916/PROTOCOL_01.md`, `FINDINGS_01.md` and `RESULTS_01.json`. Original admission results, including both failed self-difference checks.
[^6]: Project research record, `pricing_model_ambiguity_20260916/PROTOCOL_02.md`, `FINDINGS_02.md` and `RESULTS_02.json`. Independent high-precision derivative audit of unchanged methods.
[^7]: Project research record, `pricing_model_ambiguity_20260916/COHORT_03_FINDINGS.md`, `COHORT_03.json` and `cohort03/`. Independent synthetic truth generation, refinement evidence and saved observations.
[^8]: A. Gupta and C. Reisinger, *Robust Calibration of Financial Models Using Bayesian Estimators*. Author manuscript dated February 7, 2012; journal publication 2014, [10.21314/JCF.2014.285](https://doi.org/10.21314/JCF.2014.285). [Author manuscript](https://people.maths.ox.ac.uk/~reisinge/Publications/RobustnessPaper.pdf), opening discussion and Sections 2-3.1, beginning of 3.2.
[^9]: A. Spantini, T. Cui, K. Willcox, L. Tenorio and Y. Marzouk, *Goal-oriented optimal approximations of Bayesian linear inverse problems*, arXiv version 2, March 14, 2017. [Primary manuscript](https://arxiv.org/html/1607.01881v2), introduction and selected theory passages in Sections 2.1-2.3.
[^10]: R. Cont, *Model Uncertainty and Its Impact on the Pricing of Derivative Instruments*, Mathematical Finance 16(3), 519-547, 2006. [Publisher abstract and metadata](https://doi.org/10.1111/j.1467-9965.2006.00281.x).
[^11]: R. Yang, H. Qin, C. Che and L. Feng, *Volatility Calibration via Automatic Local Regression*, arXiv version 1, September 19, 2025. [Primary manuscript](https://arxiv.org/html/2509.16334v1), introduction, existing-method discussion and opening methodology.
[^12]: Project research record, `pricing_model_ambiguity_20260916/RESULTS_03.json`, `ANALYSIS_03.json`, `FINDINGS_03.md` and the full per-case `study03/` records. Frozen 360-panel paired application, September 2026; synthetic data, 15 independent field clusters.
[^13]: Project research record, `pricing_model_ambiguity_20260916/PROTOCOL_04.md`, `RESULTS_04.json`, `FINDINGS_04.md` and `study04/`. Independent value audit of all 1,918 contributing fitted states, with no changes to application fits or scoring.