Skip to content

llama/cuda: tiered KV cache (--kv-vram-cells): VRAM head, pinned-host tail in one VMM range - #319

Open
professorpalmer wants to merge 1 commit into
PrismML-Eng:prismfrom
professorpalmer:claude/tiered-kv
Open

professorpalmer wants to merge 1 commit into
PrismML-Eng:prismfrom
professorpalmer:claude/tiered-kv

Conversation

@professorpalmer

Copy link
Copy Markdown

One of the five pieces of #285, split per bri-prism's request on #317 so each can land on its own. Single commit on current prism (eaecb50c7), applies without conflicts.

What it does

llama_context_params.n_kv_vram_cells / --kv-vram-cells N keeps cells [0, N) of each layer's K/V in device memory and the rest in pinned host memory mapped into the same device range with CUDA VMM (cuMemCreate with CU_MEM_LOCATION_TYPE_HOST); ggml_backend_cuda_tier_buffer_type is exported through the registry. Kernels are unchanged and output is bit-identical to an all-device cache. When an attention op reaches the host tail, the used rows are copied into a VRAM staging buffer by the copy engine (a second VMM alias maps the same VRAM head plus the staging buffer) and the op reads the alias: prefill reads the tail once per op instead of once per query tile (146 -> 280 tok/s at 180k on the 4070), decode moves it by DMA (+28% at 180k). Plumbed through llama_cparams into the standard KV cache and the hybrid memory with a trailing defaulted constructor parameter; other memory types untouched. Works on Windows/WDDM.

Why: q8_0 at 262k is 9.1 GB of K/V for Bonsai 2 27B and does not fit next to the weights on 12 GB; #221 reached 262k with q4_0 (mean KLD 0.00218 vs f16 K/V, 97.9% top-1 agreement) where q8_0 measures 0.00017 / 99.4%.

Receipts (RTX 4070 12 GB, sm_89, PCIe 4.0 x16, GDDR6X +1500 MHz as in #221; llama-server, one slot, greedy, MTP head, genuine prefill)

depth prefill tok/s decode tok/s #221 recipe at that depth
4k 1,100 83
16k 1,100 106
32k 918 100 47.6 (96k/q8_0, drafting cut off at 24k)
64k 724 87 36.9
112k (last VRAM position) 539 70 q8_0 does not fit
131k 374 41 q8_0 does not fit
180k 298 27 q8_0 does not fit
258k 229 14.7 q8_0 does not fit

The decode numbers above 32k include the MTP drafting-window commit (separate PR) and the FA GQA-decode commit that is being folded into #307; the tiered cache itself changes no arithmetic.

  • Tiered vs all-VRAM KV, same flags, drafting on: identical sha256 of 256-token code and prose continuations at 4k, 20k and 40k; rechecked on the rebased head with --kv-vram-cells 2048 vs all-VRAM, q8_0, greedy, at 4096 / 20480 / 40960 tokens: 6/6 identical.
  • test-backend-ops -b CUDA0 on the stacked head and again on the rebased head: FLASH_ATTN_EXT 2994/2994, MUL_MAT 1516/1516, GATED_DELTA_NET 39/39, GET_ROWS 119/119 (MUL_MAT_ID 1034/1035: the intermittent m=70, n=1 case already in prism, fixed in cuda: bound the one-column MUL_MAT_ID mat-vec store by the expert row count #295).
  • On WDDM the VRAM line must stay below the point where Windows demotes a background process: ~1.3 GB below free VRAM with the display on the card, ~1.0 GB with it on an iGPU (10-minute soaks). Past the line decode is PCIe-bound.

Serving recipe and all receipts: https://github.com/professorpalmer/bonsai-ada-surgery/blob/main/docs/Q8_FULL_CONTEXT.md

… tail in one VMM range

--kv-vram-cells N (llama_context_params.n_kv_vram_cells) keeps cells [0, N) of each layer's K/V in device
memory and the rest in pinned host memory, mapped into the same device virtual range with CUDA VMM
(cuMemCreate with CU_MEM_LOCATION_TYPE_HOST), so a context can be larger than VRAM. Attention kernels are
unchanged: they only touch host pages once a sequence is that deep. Output is bit-identical to an
all-device cache.

When an attention op's K/V range reaches the host tail, the used host rows are first copied into a VRAM
staging buffer with the copy engine (a second VMM alias maps the same VRAM head pages plus the staging
buffer), and the op reads the all-VRAM alias. Prefill reads the tail once per op instead of once per
query tile; decode moves it by DMA instead of SMs reading host memory from inside the kernel.
GGML_CUDA_KV_TIER_STAGING=0 disables staging.

Bonsai 2 27B on an RTX 4070 12 GB: q8_0 K/V at the full 262,144-token window (9.1 GB of KV), first ~94k
positions in VRAM. Staging: prefill in the tail 146 -> 280 tok/s at 180k, decode +28% at 180k.
Plumbed through llama_cparams to the standard KV cache and the hybrid (attention + recurrent) memory via a
trailing defaulted constructor parameter; other memory types are untouched.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 0b36705)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant