Repository navigation
cuda+server: q8_0 K/V at the full 262k window on 12 GB (tiered VMM KV cache), MTP drafting at every depth - #285
professorpalmer wants to merge 6 commits into
Conversation
e6fa9ee to
c8b8993
Compare
… tail in one VMM range --kv-vram-cells N (llama_context_params.n_kv_vram_cells) keeps cells [0, N) of each layer's K/V in device memory and the rest in pinned host memory, mapped into the same device virtual range with CUDA VMM (cuMemCreate with CU_MEM_LOCATION_TYPE_HOST), so a context can be larger than VRAM. Attention kernels are unchanged: they only touch host pages once a sequence is that deep. Output is bit-identical to an all-device cache. When an attention op's K/V range reaches the host tail, the used host rows are first copied into a VRAM staging buffer with the copy engine (a second VMM alias maps the same VRAM head pages plus the staging buffer), and the op reads the all-VRAM alias. Prefill reads the tail once per op instead of once per query tile; decode moves it by DMA instead of SMs reading host memory from inside the kernel. GGML_CUDA_KV_TIER_STAGING=0 disables staging. Bonsai 2 27B on an RTX 4070 12 GB: q8_0 K/V at the full 262,144-token window (9.1 GB of KV), first ~94k positions in VRAM. Staging: prefill in the tail 146 -> 280 tok/s at 180k, decode +28% at 180k. Plumbed through llama_cparams to the standard KV cache and the hybrid (attention + recurrent) memory via a trailing defaulted constructor parameter; other memory types are untouched. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
For up to 8 queries (single-token decode and speculative verify) with q4_0/q8_0 K/V and a GQA ratio above 4, take the MMA kernel that reads K/V in place instead of the vector kernel. The vector kernel runs one block per Q head, so every K/V row is fetched gqa_ratio times (6x for Qwen3.5 / Bonsai 2: 24 Q heads over 4 KV heads); the MMA kernel packs the GQA heads of one KV head into one tile. Same rule the F16 path uses. Bonsai 2 27B, RTX 4070, q8_0 K/V, MTP draft n-max 2: 4k 78.1 -> 86.9 tok/s, 16k 77.7 -> 111.6. GGML_CUDA_FA_MMA_DECODE_MIN_KV sets the shortest KV length that takes the route (default 256, 0 = never). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…vec up to 8 columns Under GGML_CUDA_BATCH_INVARIANT: - the non-stream-k flash-attention KV split is sized from a fixed blocks-per-SM value instead of the occupancy of the template instance that runs; the 1-query and multi-query instances differ in registers and shared memory, so their splits (and combine order) differed. - PTQ1_0 batches up to MMVQ_MAX_BATCH_SIZE stay on the PT mat-vec, whose per-column arithmetic does not depend on the column count; a 5-column speculative verify on MMQ did not match the same tokens decoded alone. Measured cost: pp5 172 -> 162 tok/s, only on 5-8 column batches. GGML_CUDA_PTQ1_MMVQ_MAX overrides the mat-vec / MMQ crossover (default 4, confirmed: MMQ wins from 5). Not covered: the stream-k split of the MMA kernel follows the padded KV length (and the tile instance follows the query count), so attention in a verify batch is not bit-identical to single-token decode past ~32k, or at any depth on the MMA decode route. Measured at 4k/20k/40k: the weight path matches, a few continuations diverge at the rounding level. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…fting at every depth) --spec-draft-window N: the draft (MTP) context keeps only the last N rows. The server drops older rows before feeding each batch and the draft context is sized for the window, so its cells are reused, its cache stays small and on the device, and a draft pass costs the same at any depth. The MTP head predicts the next few tokens from recent context: a 16k window accepts as many drafts as the full history at 131k. Also sizes the draft context for --spec-draft-depth-max when no window is set. --spec-draft-n-max-tail N: draft size once the sequence reaches --kv-vram-cells. Past the tiered-KV line a step is bound by reading the host tail over PCIe, and a wider verify reads it once for all columns. The drafter and the output limits are built for max(n_max, n_max_tail); each slot caps a draft by depth, and the MTP draft loop now honours that per-draft cap. Bonsai 2 27B on an RTX 4070 with the MMA decode route: drafting pays at every depth, so the --spec-draft-depth-max cutoff is no longer needed (32k 54.5 -> 103.6 tok/s, 64k 48.0 -> 90.1, same acceptance; window alone +17-18% at 131k; tail draft 4 vs 2: +26% code / +15% prose at 180k). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-floor --reasoning-effort-allow LIST: reasoning_effort values passed to the chat template; any other value (from the request's reasoning_effort or chat_template_kwargs) becomes --reasoning-effort-fallback instead of reaching a template that raises on it. Bonsai 2's template accepts xhigh/medium/low and raises on "high", which Cline, Kilo and Open WebUI send: every request failed with HTTP 500. --reasoning-max-tokens-floor N: with thinking on, a client max_tokens below N is raised to N. A thinking model spends its first thousands of tokens inside the thinking block; 256-4096 token app caps (PrismML's quickstart, SillyTavern, AnythingLLM, Continue) end the request there with no answer. --reasoning-budget still bounds the thinking. Both are off by default. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With CUDA graphs off (GGML_CUDA_DISABLE_GRAPHS=1), records an event before every graph node on the main stream, charges the time to the next event to that node (a fused group to its first node), aggregates by op/type/shape and prints the top 30 every 100 graphs. Launch gaps inflate the small ops; the large mat-vecs read true. Diagnostics only; off by default. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
c8b8993 to
3164035
Compare
|
Rebased now that #221 has merged ( This head built clean on Windows (sm_89, CUDA 13, VS 2022): |
|
Reran the checks on the rebased head (
Description updated with these. |
|
Split into one PR per feature so each can land on its own, as asked on #317: #319 (tiered KV cache), #320 (MTP drafting window and tail), #321 (server reasoning flags), #322 (batch-invariant FA split and PTQ1_0 mat-vec crossover), #323 (op timing, diagnostics). The quantized-KV GQA-decode FA commit is not re-submitted; its |
Summary
Bonsai 2 27B at its full 262,144-token window with q8_0 K/V on a 12 GB RTX 4070, with MTP drafting at every depth. #221 has merged (
5244cea); this branch is now its own 6 commits on currentprism(445fa82), replayed without conflicts (git range-diffidentical to the stacked version).q8_0 at 262k is 9.1 GB of KV and does not fit next to the weights; #221 got there with q4_0 (mean KLD 0.00218 vs f16 K/V, 97.9% top-1 agreement, against 0.00017 / 99.4% for q8_0 on this model).
RTX 4070 (sm_89, PCIe 4.0 x16, GDDR6X +1500 MHz as in #221's table),
llama-server, one slot, greedy, MTP head, genuine prefill:Commits
llama/cuda: tiered KV cache (--kv-vram-cells).llama_context_params.n_kv_vram_cells/--kv-vram-cells Nkeeps cells[0, N)of each layer's K/V in device memory and the rest in pinned host memory mapped into the same device range with CUDA VMM (cuMemCreatewithCU_MEM_LOCATION_TYPE_HOST);ggml_backend_cuda_tier_buffer_typeis exported through the registry. Kernels unchanged; output bit-identical to an all-device cache. When an attention op reaches the host tail, the used rows are copied into a VRAM staging buffer by the copy engine (a second VMM alias maps the same VRAM head plus the staging buffer) and the op reads the alias: prefill reads the tail once per op instead of once per query tile (146 -> 280 tok/s at 180k), decode moves it by DMA (+28% at 180k). Plumbed throughllama_cparamsinto the standard KV cache and the hybrid memory with a trailing defaulted constructor parameter; other memory types untouched. Works on Windows/WDDM.cuda: quantized-KV GQA decode on the in-place MMA flash-attention kernel. Up to 8 queries with q4_0/q8_0 K/V and GQA > 4 take the MMA kernel frome5cd0f3instead of the vector kernel, which fetches each K/V row once per Q head (6x here). 4k 78.1 -> 86.9, 16k 77.7 -> 111.6 tok/s.GGML_CUDA_FA_MMA_DECODE_MIN_KV(default 256, 0 = off).cuda: BATCH_INVARIANT: occupancy-independent FA KV split, PTQ1_0 mat-vec up to 8 columns. The vector path's KV split no longer depends on the occupancy of the instance that runs; 5-8 column PTQ1_0 batches stay on the PT mat-vec (a 5-column verify on MMQ did not match single decode).GGML_CUDA_PTQ1_MMVQ_MAXoverrides the crossover; the default of 4 is confirmed (pp5 172 MMQ vs 162 mat-vec).speculative: --spec-draft-window and --spec-draft-n-max-tail. The draft (MTP) context keeps only the last N rows (the server drops older ones before each batch; the context is sized for the window, so cells are reused and a draft pass costs the same at any depth). With (2) drafting pays at every depth, which retires the--spec-draft-depth-max 24576cutoff from cuda: Bonsai 2 27B at the full 262k window with the MTP head on 12 GB Ada: #218 + #215 hybrid PTQ1_0 dispatch, in-place q4_0/q8_0 K/V flash attention (integration checkpoint) #221 (the flag stays, default 0): 32k 54.5 -> 103.6, 64k 48.0 -> 90.1 tok/s, acceptance unchanged.--spec-draft-n-max-tailsets the draft size from--kv-vram-cellson (4 vs 2: +26% code / +15% prose at 180k); the MTP draft loop now honours the per-draft cap.server: --reasoning-effort-allow/-fallback and --reasoning-max-tokens-floor. Bonsai 2's template raises onreasoning_effort: "high"(Cline, Kilo, Open WebUI: HTTP 500 on every request); an effort word outside the allow list becomes the fallback. With thinking on, a clientmax_tokensbelow the floor is raised to it (256-4096-token app caps end a thinking model mid-thought). Both off by default.cuda: GGML_CUDA_OP_TIMING=1per-node GPU time breakdown with graphs off. Diagnostics.Correctness
test-backend-ops -b CUDA0:FLASH_ATTN_EXT2994/2994 andMUL_MAT1516/1516, default and withGGML_CUDA_BATCH_INVARIANT=1, on the stacked headc8b8993and again on the rebased head (GATED_DELTA_NET39/39,GET_ROWS119/119 too).MUL_MAT_IDis 1034/1035 on both: the intermittentm=70, n=1case bri-prism reported on cuda: Bonsai 2 27B at the full 262k window with the MTP head on 12 GB Ada: #218 + #215 hybrid PTQ1_0 dispatch, in-place q4_0/q8_0 K/V flash attention (integration checkpoint) #221, a store-guard bug already inprism, fixed in cuda: bound the one-column MUL_MAT_ID mat-vec store by the expert row count #295.--kv-vram-cells 2048vs all-VRAM, q8_0, drafting on, greedy, 256-token code and prose continuations at 4096 / 20480 / 40960 tokens: 6/6 identical.End-to-end quality (HumanEval 164, tests executed, server's own sampling)
Killy's HumanEval plates replayed against this branch's server (262k window, q8_0 KV tiered, MTP): medium 161/164 (his best cell on PQ2_0: 160-161), thinking off 138. Harness-proofing on the same server:
reasoning_effort: "high"with a 256-token cap 160/164, against 0/20 (HTTP 500) with it off; a 4096-token cap 157/164, against 148/164 with it off. Samples and summaries:artifacts/eval/killy_20260927.Killy's voxel-pagoda plate (037P) at medium, 6 samples on this server: 3 correct on the first try, 6 of 6 after a render-check-repair loop (
bench/pagoda_plate.py --repair: run the page headless, send back console errors / floating voxels / out-of-view, ask for a fix; 1-2 rounds). The first-try misses were one-token slips in the model's JavaScript, not runtime faults. Details.All charts in one image: mega.png (decode and prefill by depth, KV precision, iGPU, identity matrix, HumanEval replays, pagoda).
Measured and not adopted
GGML_CUDA_GRAPH_OPT=1Notes
🤖 Generated with Claude Code