Skip to content

Commit f5ae26c

Browse files
Fix unit error: pp*3600 on a 4-vCPU runner is per-instance-hour, not per-core-hour; lead with unambiguous $/1M-tokens (~$0.45 -> $0.08) instead
1 parent 814b301 commit f5ae26c

2 files changed

Lines changed: 13 additions & 14 deletions

File tree

‎.github/workflows/benchmark.yml‎

Lines changed: 1 addition & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -175,7 +175,6 @@ jobs:
175175
print(f"| {name} | {scan.get('i8mm',0):,} | {scan.get('dotprod',0):,} | "
176176
f"{pp:.1f} | {pp*3600:,.0f} | ${cost:.3f} | {ratio:.2f}x |")
177177
if cold and "TEPID" in rows:
178-
print(f"\n**The one-flag fix: {rows['TEPID'][1]/cold:.2f}x more tokens per core-hour, "
179-
f"the same off the cloud bill.**")
178+
print(f"\n**The one-flag fix: {rows['TEPID'][1]/cold:.2f}x cheaper per token.**")
180179
print(f"\n_*at ${PRICE_PER_VCPU_HR}/vCPU-hr x {VCPUS} vCPU. Only the $ column assumes a price._")
181180
PY

‎README.md‎

Lines changed: 12 additions & 12 deletions
Original file line numberDiff line numberDiff line change
@@ -1,8 +1,8 @@
11
# coldpath
22

33
**Popular LLM runtimes ship on Arm with the chip's matrix hardware switched off. On Azure Cobalt 100
4-
(Neoverse N2, a cloud Arm CPU) that costs 5.75x on prompt-processing throughput: 5.75x fewer tokens per
5-
core-hour, the same 5.75x on the cloud bill. I built the tool that finds it in any binary, fixed the most
4+
(Neoverse N2, a cloud Arm CPU) that is 5.75x on prompt-processing throughput and on the cost per token —
5+
about $0.45 versus $0.08 per million. I built the tool that finds it in any binary, fixed the most
66
popular offender in one line upstream, and gated it so a cold build can't reach your Arm cloud fleet.**
77

88
`coldpath` disassembles any AArch64 binary and proves whether it *contains, and can dispatch,* the chip's
@@ -35,16 +35,16 @@ so a mis-built cloud container is one flag away from this.
3535
What it costs, measured on **Azure Cobalt 100 (Neoverse N2)** — a cloud Arm CPU, on the free GitHub
3636
runner, only `-march` changing:
3737

38-
| build (same model + hardware, only `-march` changes) | coldpath sees | pp512 tok/s | tokens / core-hour | vs COLD |
38+
| build (same model + hardware, only `-march` changes) | coldpath sees | pp512 tok/s | $ / 1M tokens | vs COLD |
3939
|---|---|---:|---:|---:|
40-
| **COLD** `armv8-a` — Ollama's Windows build, and the naive Arm64 default | i8mm 0, dotprod 0 | ~95 | ~0.34M | 1.0x |
41-
| **TEPID** `armv8.2-a+dotprod` — the one-line fix I filed upstream | dotprod 1,044 | ~545 | ~1.96M | **~5.75x** |
42-
| **WARM** `armv8.6-a+i8mm` | i8mm 268, dotprod 1,044 | ~640 | ~2.30M | **~6.7x** |
40+
| **COLD** `armv8-a` — Ollama's Windows build, and the naive Arm64 default | i8mm 0, dotprod 0 | ~95 | ~$0.45 | 1.0x |
41+
| **TEPID** `armv8.2-a+dotprod` — the one-line fix I filed upstream | dotprod 1,044 | ~545 | ~$0.08 | **~5.75x** |
42+
| **WARM** `armv8.6-a+i8mm` | i8mm 268, dotprod 1,044 | ~660 | ~$0.065 | **~6.9x** |
4343

44-
_5.75x more tokens per core-hour is 5.75x lower cost per token on any Arm cloud — the workflow's
45-
`compare` job computes it live as roughly **$0.45 → $0.08 per 1M tokens** (Graviton4 on-demand rate).
46-
Figures vary a few percent per run. The fix I filed (PR #17654) is the TEPID row — dot-product, safe on
47-
every shipped Arm device; i8mm adds the rest where the silicon has it._
44+
_The tok/s and the 5.75x ratio are measured on the 4-vCPU runner; the $/1M-tokens applies a sample
45+
Graviton4 on-demand rate, so only the dollar column assumes a price. The workflow's `compare` job
46+
recomputes all of it live each run (figures vary a few percent). The fix I filed (PR #17654) is the TEPID
47+
row — dot-product, safe on every shipped Arm device; i8mm adds the rest where the silicon has it._
4848

4949
The fix, the root cause, and the reproducible measurement are in
5050
[`examples/ollama-fix/`](examples/ollama-fix/). The benchmark is
@@ -180,8 +180,8 @@ not a broken detector. `pytest` (24 tests) covers each instruction family agains
180180
encodings, the resync-through-data property, the coverage gate, and the single-word corroboration floor,
181181
so correctness is provable without any binary on disk.
182182

183-
This validation is the receipt, not the headline. The headline is the 5.75x more tokens per core-hour
184-
that one build flag recovers on Arm cloud silicon.
183+
This validation is the receipt, not the headline. The headline is the 5.75x lower cost per token that one
184+
build flag recovers on Arm cloud silicon.
185185

186186
---
187187

0 commit comments

Comments
 (0)