Repository navigation
llama/cuda: tiered KV cache (--kv-vram-cells): VRAM head, pinned-host tail in one VMM range - #319
Open
professorpalmer wants to merge 1 commit into
Open
professorpalmer wants to merge 1 commit into
professorpalmer wants to merge 1 commit into
Conversation
… tail in one VMM range --kv-vram-cells N (llama_context_params.n_kv_vram_cells) keeps cells [0, N) of each layer's K/V in device memory and the rest in pinned host memory, mapped into the same device virtual range with CUDA VMM (cuMemCreate with CU_MEM_LOCATION_TYPE_HOST), so a context can be larger than VRAM. Attention kernels are unchanged: they only touch host pages once a sequence is that deep. Output is bit-identical to an all-device cache. When an attention op's K/V range reaches the host tail, the used host rows are first copied into a VRAM staging buffer with the copy engine (a second VMM alias maps the same VRAM head pages plus the staging buffer), and the op reads the all-VRAM alias. Prefill reads the tail once per op instead of once per query tile; decode moves it by DMA instead of SMs reading host memory from inside the kernel. GGML_CUDA_KV_TIER_STAGING=0 disables staging. Bonsai 2 27B on an RTX 4070 12 GB: q8_0 K/V at the full 262,144-token window (9.1 GB of KV), first ~94k positions in VRAM. Staging: prefill in the tail 146 -> 280 tok/s at 180k, decode +28% at 180k. Plumbed through llama_cparams to the standard KV cache and the hybrid (attention + recurrent) memory via a trailing defaulted constructor parameter; other memory types are untouched. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> (cherry picked from commit 0b36705)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
One of the five pieces of #285, split per bri-prism's request on #317 so each can land on its own. Single commit on current
prism(eaecb50c7), applies without conflicts.What it does
llama_context_params.n_kv_vram_cells/--kv-vram-cells Nkeeps cells[0, N)of each layer's K/V in device memory and the rest in pinned host memory mapped into the same device range with CUDA VMM (cuMemCreatewithCU_MEM_LOCATION_TYPE_HOST);ggml_backend_cuda_tier_buffer_typeis exported through the registry. Kernels are unchanged and output is bit-identical to an all-device cache. When an attention op reaches the host tail, the used rows are copied into a VRAM staging buffer by the copy engine (a second VMM alias maps the same VRAM head plus the staging buffer) and the op reads the alias: prefill reads the tail once per op instead of once per query tile (146 -> 280 tok/s at 180k on the 4070), decode moves it by DMA (+28% at 180k). Plumbed throughllama_cparamsinto the standard KV cache and the hybrid memory with a trailing defaulted constructor parameter; other memory types untouched. Works on Windows/WDDM.Why: q8_0 at 262k is 9.1 GB of K/V for Bonsai 2 27B and does not fit next to the weights on 12 GB; #221 reached 262k with q4_0 (mean KLD 0.00218 vs f16 K/V, 97.9% top-1 agreement) where q8_0 measures 0.00017 / 99.4%.
Receipts (RTX 4070 12 GB, sm_89, PCIe 4.0 x16, GDDR6X +1500 MHz as in #221;
llama-server, one slot, greedy, MTP head, genuine prefill)The decode numbers above 32k include the MTP drafting-window commit (separate PR) and the FA GQA-decode commit that is being folded into #307; the tiered cache itself changes no arithmetic.
--kv-vram-cells 2048vs all-VRAM, q8_0, greedy, at 4096 / 20480 / 40960 tokens: 6/6 identical.test-backend-ops -b CUDA0on the stacked head and again on the rebased head:FLASH_ATTN_EXT2994/2994,MUL_MAT1516/1516,GATED_DELTA_NET39/39,GET_ROWS119/119 (MUL_MAT_ID1034/1035: the intermittentm=70, n=1case already inprism, fixed in cuda: bound the one-column MUL_MAT_ID mat-vec store by the expert row count #295).Serving recipe and all receipts: https://github.com/professorpalmer/bonsai-ada-surgery/blob/main/docs/Q8_FULL_CONTEXT.md