Date: 2026-09-18. Baseline: examples commit 725a1cd.
Examples release target: v1.0.1; Numerics package version: 2.2.0.
Status: implementation, compatibility execution, quiet-machine benchmark refresh, and final local validation completed. No push or publication was performed.
The exact published RMC.Numerics 2.2.0 package was restored from NuGet.org
into the initially empty project-local packages/ directory using the checked-in
NuGet configuration. Its repository metadata identifies Numerics commit
7e8e8d1c5f26e045a35ec9fc09367de95ed05b02, matching the inspected library checkout.
No Numerics source, algorithm, default, or tolerance was changed.
Package SHA-256:
602C5F6D48BD113276FAA7EA3D5DDF99C910C438B9C4AEA9B5B847EB9F3A210C.
| Component | Reference environment |
|---|---|
| OS / CPU | Windows 11 build 26200; Intel64 family 6 model 170; 22 logical CPUs |
| Python / pythonnet | 3.12.14 / 3.1.0 |
| .NET runtime / SDK and Roslyn distribution | 10.0.11 / 10.0.400 |
| Numerics assembly / target | 2.2.0.0 / net10.0 |
| NumPy / pandas / SciPy | 2.4.6 / 3.0.6 / 1.18.1 |
| PyMC / PyTensor / ArviZ | 5.28.5 / 2.38.3 / 0.23.4 |
| scikit-learn / Matplotlib | 1.9.1 / 3.11.2 |
| nbclient / nbformat / ipykernel | 0.11.0 / 5.11.1 / 7.3.0 |
PyTensor did not find its optional C++ compiler, and no external BLAS flags were configured. PyMC results therefore describe this Python-fallback setup, not an optimized compiled PyMC installation. This limitation does not affect the controlled Python-callback versus compiled-C# Numerics comparison.
An initial NUTS run overlapped an unrelated CPU-intensive job. It was retained only as functional-check evidence, not as the final performance reference. Final benchmark refreshes were run serially after that job and all other example validation jobs had stopped. Ordinary desktop background activity remains; three repetitions and their ranges describe observed timing variability.
Notebook 11 completed a fresh-kernel run in 12.6 seconds. Benchmark environment timestamp: 2026-09-18 19:19:20 UTC. Each row uses one untimed warm-up and three fresh seeded models; values below are median [minimum, maximum] milliseconds.
| Workload | Library | Training ms | Prediction ms | Score |
|---|---|---|---|---|
| Classification | Numerics | 57.6358 [52.5811, 94.1881] | 1.9287 [0.7831, 2.1075] | Accuracy 0.951049 |
| Classification | scikit-learn | 162.4121 [160.9046, 178.9559] | 30.8972 [21.1514, 32.3177] | Accuracy 0.951049 |
| Regression | Numerics | 128.0586 [117.8278, 132.1747] | 1.6915 [1.4210, 1.8087] | R² 0.479941 |
| Regression | scikit-learn | 177.4966 [150.6325, 212.3425] | 33.5953 [23.2976, 39.7056] | R² 0.484237 |
All four rows produced bitwise-equal predictions across these seeded repetitions. This is observed repeatability on this environment, not a cross-platform promise.
Shared settings: 100 trees, depth 20, minimum split 2, seed 12345; sqrt(d) classifier features and all regressor features. Numerics regression uses mean column 3. Classification is a workflow comparison: Numerics' historical count-weighted impurity and floored interpolated-median aggregation differ from scikit-learn Gini and probability averaging. The example uses binary 0/1 labels; that Numerics aggregation is not a general replacement for multiclass voting.
Input conversion and Numerics Matrix construction precede timing; Numerics prediction extraction follows timing. Numerics' prediction call also computes tree quantiles, whereas scikit-learn returns point predictions. Tree-spread quantiles are not calibrated predictive intervals. These differences preclude attributing the ratios solely to implementation language.
Notebook 05 completed its final fresh-kernel run in 782.4 seconds, including the other illustrative sampler/distribution comparisons. Controlled benchmark environment timestamp: 2026-09-18 19:23:07 UTC. Each mode uses one reduced, untimed warm-up and three fresh seeded runs. Sampling includes sampler warmup and subsequent draws, but excludes likelihood compilation and model setup.
| Mode | Sampling median [min, max], seconds | Minimum conservative ESS/second |
|---|---|---|
| Python callback / sequential | 4.5966 [4.5318, 4.6393] | 670.2366 |
| Compiled C# / sequential | 0.0491 [0.0480, 0.0630] | 62785.1419 |
| Compiled C# / parallel | 0.0698 [0.0656, 0.0699] | 44137.8381 |
| PyMC NUTS (Python fallback) | 76.5702 [67.0923, 95.7117] | 66.6940 |
Roslyn compilation plus assembly load took 0.4464 seconds once; PyMC model/step setup had a median of 0.4211 seconds. Numerics constructor setup was not separately measured. All modes used four chains, 1000 warmup transitions and 2000 retained draws per chain: 12000 total transitions and 8000 retained draws. Numerics combines the post-warmup Markov segment and Output segment while preserving chain boundaries.
Compiled C# sequential sampling was 93.68 times faster than the Python callback on this model. Parallel C# was faster than the Python callback but slower than sequential C#: scheduling overhead outweighs the benefit for this small likelihood. These measurements do not establish parallel scaling for expensive models. The PyMC compiler fallback and different gradients/transforms make its timing a workflow comparison, not an optimized library-versus-library performance verdict. The RStan analogy is a compiled model callable from a high-level language; Roslyn emits managed IL, and this Numerics example still uses numerical gradients, unlike Stan's compiled C++ automatic differentiation.
Likelihood parity passed at interior, boundary, and invalid inputs. All three Numerics modes produced bitwise-identical seeded draws. Rank-normalized R-hat was 1.0001 for mu and 1.0015 for sigma; minimum conservative ESS was 3080.8167. All diagnostic values were finite. Numerics had zero divergences and maximum-depth hits; per-chain E-BFMI ranged from 0.9101 to 0.9827. PyMC also had zero divergences, with R-hat 1.0005 for each parameter. Posterior means agreed within three pooled Monte Carlo standard errors:
| Parameter | Numerics mean | PyMC mean | Absolute difference | Three pooled MCSE |
|---|---|---|---|---|
| mu | 12680.7974 | 12665.1892 | 15.6082 | 45.5809 |
| sigma | 4853.8868 | 4836.1329 | 17.7539 | 34.8006 |
Passing these diagnostics is evidence for this run, not proof of convergence or a general guarantee for other models.
Use the commands in the README's release-validation section. Full notebook
execution uses a fresh kernel per file; standalone scripts use separate
processes. Generated execution copies and script logs are under ignored
artifacts/validation/. Versioned outputs are limited to notebooks 00/05/06/11;
stale baseline outputs were cleared from the other notebooks.
Completed checks:
- Eleven helper/CLI tests passed (
python -m pytest -q, 7.61 seconds), including chain-boundary/warmup extraction, diagnostic input rejection, and preservation of execution receipts during schema-only checks. - Actual .NET 10 package load, idempotent repeated load, wrong assembly name/version rejected before load, and wrong-framework override rejection with an explicit restart requirement.
- Actual Numerics-versus-ArviZ multi-chain ESS/R-hat check and managed
out-parameter binding in
tests/runtime_diagnostics.py. - All thirteen notebooks 00–12 executed successfully in fresh kernels. Final-source reruns of notebooks 03/04 passed in 17.8/122.4 seconds; notebook 06 passed its diagnostic and mixture examples in 559.6 seconds.
- All three scripts passed: Bayesian regression 139.5 s, flood-frequency analysis 3.0 s, reliability analysis 6.3 s. These are execution durations, not controlled performance benchmarks.
- The time-series notebook's public USGS download passed on a targeted network-enabled retry after the sandbox rejected its socket connection.
- All thirteen notebook schemas passed,
python -m pip checkreported no broken requirements, andgit diff --checkpassed. - Retained outputs and execution counts match the fresh execution artifacts; language metadata reports the actual kernel version. No notebook error outputs, stale execution timestamps, or absolute machine paths appear in retained text outputs. The nine source-only notebooks have no outputs or execution counts.
- Independent source review and its scoped fix review found no remaining substantive source issues.
No optional C++ compiler was installed, no Python or .NET global configuration was changed, and no source was pushed or published. The original checkout's seven unrelated dirty notebooks were preserved in place.