Skip to content

cuda+server: q8_0 K/V at the full 262k window on 12 GB (tiered VMM KV cache), MTP drafting at every depth - #285

Closed
professorpalmer wants to merge 6 commits into
PrismML-Eng:prismfrom
professorpalmer:claude/q8-262k-tier
Closed

professorpalmer wants to merge 6 commits into
PrismML-Eng:prismfrom
professorpalmer:claude/q8-262k-tier

Conversation

@professorpalmer

@professorpalmer professorpalmer commented Sep 27, 2026 •

Copy link
Copy Markdown

Summary

Bonsai 2 27B at its full 262,144-token window with q8_0 K/V on a 12 GB RTX 4070, with MTP drafting at every depth. #221 has merged (5244cea); this branch is now its own 6 commits on current prism (445fa82), replayed without conflicts (git range-diff identical to the stacked version).

q8_0 at 262k is 9.1 GB of KV and does not fit next to the weights; #221 got there with q4_0 (mean KLD 0.00218 vs f16 K/V, 97.9% top-1 agreement, against 0.00017 / 99.4% for q8_0 on this model).

RTX 4070 (sm_89, PCIe 4.0 x16, GDDR6X +1500 MHz as in #221's table), llama-server, one slot, greedy, MTP head, genuine prefill:

depth prefill tok/s decode tok/s #221 recipe at that depth
4k 1,100 83
16k 1,100 106
32k 918 100 47.6 (96k/q8_0, drafting cut off at 24k)
64k 724 87 36.9
112k (last VRAM position) 539 70 q8_0 does not fit
131k 374 41 q8_0 does not fit
180k 298 27 q8_0 does not fit
258k 229 14.7 q8_0 does not fit

Commits

  1. llama/cuda: tiered KV cache (--kv-vram-cells). llama_context_params.n_kv_vram_cells / --kv-vram-cells N keeps cells [0, N) of each layer's K/V in device memory and the rest in pinned host memory mapped into the same device range with CUDA VMM (cuMemCreate with CU_MEM_LOCATION_TYPE_HOST); ggml_backend_cuda_tier_buffer_type is exported through the registry. Kernels unchanged; output bit-identical to an all-device cache. When an attention op reaches the host tail, the used rows are copied into a VRAM staging buffer by the copy engine (a second VMM alias maps the same VRAM head plus the staging buffer) and the op reads the alias: prefill reads the tail once per op instead of once per query tile (146 -> 280 tok/s at 180k), decode moves it by DMA (+28% at 180k). Plumbed through llama_cparams into the standard KV cache and the hybrid memory with a trailing defaulted constructor parameter; other memory types untouched. Works on Windows/WDDM.
  2. cuda: quantized-KV GQA decode on the in-place MMA flash-attention kernel. Up to 8 queries with q4_0/q8_0 K/V and GQA > 4 take the MMA kernel from e5cd0f3 instead of the vector kernel, which fetches each K/V row once per Q head (6x here). 4k 78.1 -> 86.9, 16k 77.7 -> 111.6 tok/s. GGML_CUDA_FA_MMA_DECODE_MIN_KV (default 256, 0 = off).
  3. cuda: BATCH_INVARIANT: occupancy-independent FA KV split, PTQ1_0 mat-vec up to 8 columns. The vector path's KV split no longer depends on the occupancy of the instance that runs; 5-8 column PTQ1_0 batches stay on the PT mat-vec (a 5-column verify on MMQ did not match single decode). GGML_CUDA_PTQ1_MMVQ_MAX overrides the crossover; the default of 4 is confirmed (pp5 172 MMQ vs 162 mat-vec).
  4. speculative: --spec-draft-window and --spec-draft-n-max-tail. The draft (MTP) context keeps only the last N rows (the server drops older ones before each batch; the context is sized for the window, so cells are reused and a draft pass costs the same at any depth). With (2) drafting pays at every depth, which retires the --spec-draft-depth-max 24576 cutoff from cuda: Bonsai 2 27B at the full 262k window with the MTP head on 12 GB Ada: #218 + #215 hybrid PTQ1_0 dispatch, in-place q4_0/q8_0 K/V flash attention (integration checkpoint) #221 (the flag stays, default 0): 32k 54.5 -> 103.6, 64k 48.0 -> 90.1 tok/s, acceptance unchanged. --spec-draft-n-max-tail sets the draft size from --kv-vram-cells on (4 vs 2: +26% code / +15% prose at 180k); the MTP draft loop now honours the per-draft cap.
  5. server: --reasoning-effort-allow/-fallback and --reasoning-max-tokens-floor. Bonsai 2's template raises on reasoning_effort: "high" (Cline, Kilo, Open WebUI: HTTP 500 on every request); an effort word outside the allow list becomes the fallback. With thinking on, a client max_tokens below the floor is raised to it (256-4096-token app caps end a thinking model mid-thought). Both off by default.
  6. cuda: GGML_CUDA_OP_TIMING=1 per-node GPU time breakdown with graphs off. Diagnostics.

Correctness

End-to-end quality (HumanEval 164, tests executed, server's own sampling)

Killy's HumanEval plates replayed against this branch's server (262k window, q8_0 KV tiered, MTP): medium 161/164 (his best cell on PQ2_0: 160-161), thinking off 138. Harness-proofing on the same server: reasoning_effort: "high" with a 256-token cap 160/164, against 0/20 (HTTP 500) with it off; a 4096-token cap 157/164, against 148/164 with it off. Samples and summaries: artifacts/eval/killy_20260927.

Killy's voxel-pagoda plate (037P) at medium, 6 samples on this server: 3 correct on the first try, 6 of 6 after a render-check-repair loop (bench/pagoda_plate.py --repair: run the page headless, send back console errors / floating voxels / out-of-view, ask for a fix; 1-2 rounds). The first-try misses were one-token slips in the model's JavaScript, not runtime faults. Details.

Summary

All charts in one image: mega.png (decode and prefill by depth, KV precision, iGPU, identity matrix, HumanEval replays, pagoda).

Measured and not adopted

idea result
K mean-centering on q4_0 / q8_0 with the default Hadamard KV rotation q4_0 KLD 0.00206 vs 0.00218, max KLD 2.87 vs 0.94; q8_0 unchanged
GGML_CUDA_GRAPH_OPT=1 tg128 66.1 vs 67.9
one MMA tile for all <= 8-query batches (invariance) -7% at 64k, and the KV-length dependence remains
draft cache fully in VRAM, paid for with a lower KV line 16.0 vs 17.2 tok/s at 180k

Notes

🤖 Generated with Claude Code

professorpalmer and others added 6 commits September 29, 2026 15:41
… tail in one VMM range

--kv-vram-cells N (llama_context_params.n_kv_vram_cells) keeps cells [0, N) of each layer's K/V in device
memory and the rest in pinned host memory, mapped into the same device virtual range with CUDA VMM
(cuMemCreate with CU_MEM_LOCATION_TYPE_HOST), so a context can be larger than VRAM. Attention kernels are
unchanged: they only touch host pages once a sequence is that deep. Output is bit-identical to an
all-device cache.

When an attention op's K/V range reaches the host tail, the used host rows are first copied into a VRAM
staging buffer with the copy engine (a second VMM alias maps the same VRAM head pages plus the staging
buffer), and the op reads the all-VRAM alias. Prefill reads the tail once per op instead of once per
query tile; decode moves it by DMA instead of SMs reading host memory from inside the kernel.
GGML_CUDA_KV_TIER_STAGING=0 disables staging.

Bonsai 2 27B on an RTX 4070 12 GB: q8_0 K/V at the full 262,144-token window (9.1 GB of KV), first ~94k
positions in VRAM. Staging: prefill in the tail 146 -> 280 tok/s at 180k, decode +28% at 180k.
Plumbed through llama_cparams to the standard KV cache and the hybrid (attention + recurrent) memory via a
trailing defaulted constructor parameter; other memory types are untouched.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
For up to 8 queries (single-token decode and speculative verify) with q4_0/q8_0 K/V and a GQA ratio above
4, take the MMA kernel that reads K/V in place instead of the vector kernel. The vector kernel runs one
block per Q head, so every K/V row is fetched gqa_ratio times (6x for Qwen3.5 / Bonsai 2: 24 Q heads over
4 KV heads); the MMA kernel packs the GQA heads of one KV head into one tile. Same rule the F16 path uses.

Bonsai 2 27B, RTX 4070, q8_0 K/V, MTP draft n-max 2: 4k 78.1 -> 86.9 tok/s, 16k 77.7 -> 111.6.
GGML_CUDA_FA_MMA_DECODE_MIN_KV sets the shortest KV length that takes the route (default 256, 0 = never).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…vec up to 8 columns

Under GGML_CUDA_BATCH_INVARIANT:
- the non-stream-k flash-attention KV split is sized from a fixed blocks-per-SM value instead of the
  occupancy of the template instance that runs; the 1-query and multi-query instances differ in registers
  and shared memory, so their splits (and combine order) differed.
- PTQ1_0 batches up to MMVQ_MAX_BATCH_SIZE stay on the PT mat-vec, whose per-column arithmetic does not
  depend on the column count; a 5-column speculative verify on MMQ did not match the same tokens decoded
  alone. Measured cost: pp5 172 -> 162 tok/s, only on 5-8 column batches.
GGML_CUDA_PTQ1_MMVQ_MAX overrides the mat-vec / MMQ crossover (default 4, confirmed: MMQ wins from 5).

Not covered: the stream-k split of the MMA kernel follows the padded KV length (and the tile instance
follows the query count), so attention in a verify batch is not bit-identical to single-token decode past
~32k, or at any depth on the MMA decode route. Measured at 4k/20k/40k: the weight path matches, a few
continuations diverge at the rounding level.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…fting at every depth)

--spec-draft-window N: the draft (MTP) context keeps only the last N rows. The server drops older rows
before feeding each batch and the draft context is sized for the window, so its cells are reused, its
cache stays small and on the device, and a draft pass costs the same at any depth. The MTP head predicts
the next few tokens from recent context: a 16k window accepts as many drafts as the full history at 131k.
Also sizes the draft context for --spec-draft-depth-max when no window is set.

--spec-draft-n-max-tail N: draft size once the sequence reaches --kv-vram-cells. Past the tiered-KV line a
step is bound by reading the host tail over PCIe, and a wider verify reads it once for all columns. The
drafter and the output limits are built for max(n_max, n_max_tail); each slot caps a draft by depth, and the
MTP draft loop now honours that per-draft cap.

Bonsai 2 27B on an RTX 4070 with the MMA decode route: drafting pays at every depth, so the
--spec-draft-depth-max cutoff is no longer needed (32k 54.5 -> 103.6 tok/s, 64k 48.0 -> 90.1, same
acceptance; window alone +17-18% at 131k; tail draft 4 vs 2: +26% code / +15% prose at 180k).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-floor

--reasoning-effort-allow LIST: reasoning_effort values passed to the chat template; any other value (from
the request's reasoning_effort or chat_template_kwargs) becomes --reasoning-effort-fallback instead of
reaching a template that raises on it. Bonsai 2's template accepts xhigh/medium/low and raises on "high",
which Cline, Kilo and Open WebUI send: every request failed with HTTP 500.

--reasoning-max-tokens-floor N: with thinking on, a client max_tokens below N is raised to N. A thinking
model spends its first thousands of tokens inside the thinking block; 256-4096 token app caps (PrismML's
quickstart, SillyTavern, AnythingLLM, Continue) end the request there with no answer. --reasoning-budget
still bounds the thinking. Both are off by default.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With CUDA graphs off (GGML_CUDA_DISABLE_GRAPHS=1), records an event before every graph node on the main
stream, charges the time to the next event to that node (a fused group to its first node), aggregates by
op/type/shape and prints the top 30 every 100 graphs. Launch gaps inflate the small ops; the large mat-vecs
read true. Diagnostics only; off by default.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@professorpalmer professorpalmer changed the title cuda+server: q8_0 K/V at the full 262k window on 12 GB (tiered VMM KV cache), MTP drafting at every depth (stacked on #221) cuda+server: q8_0 K/V at the full 262k window on 12 GB (tiered VMM KV cache), MTP drafting at every depth Sep 29, 2026
@professorpalmer

Copy link
Copy Markdown
Author

Rebased now that #221 has merged (5244cea). The branch is just its own 6 commits on prism 445fa82, replayed without conflicts; git range-diff shows every commit identical to the stacked version (c8b8993).

This head built clean on Windows (sm_89, CUDA 13, VS 2022): llama-server, test-backend-ops, test-mtp-catchup-batch (prints ok). I have not rerun test-backend-ops on the GPU for this head yet because the card is serving; the FA/MUL_MAT counts in the description are from c8b8993 and are labelled that way.

@professorpalmer

Copy link
Copy Markdown
Author

Reran the checks on the rebased head (3164035) now that the card is free:

Description updated with these.

@professorpalmer

Copy link
Copy Markdown
Author

Split into one PR per feature so each can land on its own, as asked on #317: #319 (tiered KV cache), #320 (MTP drafting window and tail), #321 (server reasoning flags), #322 (batch-invariant FA split and PTQ1_0 mat-vec crossover), #323 (op timing, diagnostics). The quantized-KV GQA-decode FA commit is not re-submitted; its ggml_cuda_fattn_mma_kv_native_supported guard goes into #307. The description above stays as the combined write-up and receipts.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant