Repository navigation
kv-cache : limit seq_rm to the used cell range - #311
Conversation
seq_rm walked all cells of the cache on every call. Cells outside [used_min, used_max_p1) are empty, so skip them. With a large context and speculative decoding (seq_rm after rejected drafts) this showed up in the host profile. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
bri-prism
left a comment
There was a problem hiding this comment.
No blocking findings in the loop-bound change.
I compared the complete base/head seq_rm methods using the real KV-cell metadata type: 21,000 calls matched under AddressSanitizer and UndefinedBehaviorSanitizer across 1-4 streams, eight sequence IDs, shared cells, holes, empty caches, full clears, negative bounds, reversed intervals, and repeated removals. Cell metadata, sequence-position extrema, used ranges/counts, and allocation heads matched. The changed translation unit also passed a CPU syntax check.
The skipped cells have position -1, and p0 is normalized to 0, so they cannot match the original removal predicate. Capturing the upper bound before erasures preserves the walk while the used set changes.
This was focused metadata validation; I did not independently rerun real-model generation or reproduce the speed measurements.
Overview
llama_kv_cache::seq_rmwalked all cells of the cache on every call. Cells outside[used_min(), used_max_p1())are empty (pos == -1), so the loop now only covers that range. With a large context and speculative decoding (seq_rmafter rejected drafts) the full walk showed up in the host profile (0.9 % of the decode step at 114688 cells, 0.02 % at 8192).RTX 4070, Bonsai 2 27B PTQ1_0 + MTP head (n-max 2), server arguments as in an agent setup (context 114688, reasoning budget, thinking sampling), 8 interleaved A/B pairs: +0.77 % (95 % CI [+0.59, +0.96] %), outputs identical. Depth 0 / 16k / 65k / 100k: outputs identical, +0.4 to +1.8 %, never slower.
Additional information
src/llama-kv-cache.cpp, no behavior change: the skipped cells were empty.seq_rmtest in the repo.Test results
-DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=89.prismat88c4bc60b. The four commits since (SYCL, WebGPU andcuda: fused FWHT quantizer for 64-wide warps (#303)) do not touch the code paths of this PR; the branch is rebased on2459f68b5, builds, and thetest-backend-opsruns below were repeated on it.--spec-type draft-mtp --spec-draft-n-max 2, q4_0 K/V cache, one slot.--reasoning-format deepseek --reasoning-budget 16384, thinking sampling 1.0 / 0.95 / 20 / 0.05, fixed seed, 4 prompts x 256 tokens), 8 interleaved pairs: +0.77 % (95 % CI [+0.59, +0.96] %), outputs identical in all runs. This is the favourable case (almost empty cache in a 114688-cell context); the full walk costs 0.02 % of a decode step at 8192 cells and 0.9 % at 114688.seq_rmtest in the repository (the change is a loop bound; skipped cells havepos == -1and would not have been touched).Requirements