Skip to content

server : reuse the buffers of evicted prompt checkpoints - #312

Open
sb32445 wants to merge 2 commits into
PrismML-Eng:prismfrom
sb32445:pr/server-ckpt-buffer-reuse
Open

sb32445 wants to merge 2 commits into
PrismML-Eng:prismfrom
sb32445:pr/server-ckpt-buffer-reuse

Conversation

@sb32445

@sb32445 sb32445 commented Oct 4, 2026

Copy link
Copy Markdown

Overview

create_checkpoint() evicts old checkpoints and then emplaces a new one whose data vectors are resized from empty, so every prompt checkpoint (about 150 MiB for a hybrid model) is zero-filled and page-faulted again. The buffers of the last evicted checkpoint (data_tgt, data_dft) are handed to the next one instead (two local variables of create_checkpoint(), no new state).

Growing multi-turn conversation (44 turns, ~300 words of new text per turn, 8 greedy tokens, MTP n-max 2), 3 alternating A/B pairs of the same binary: 646.3 -> 626.3 ms per turn (-20 ms, -3.1 %), outputs identical. Decode speed is unchanged (+0.00 %, CI [-0.11, +0.13] %): the effect is per request, not per token.

Additional information

  • The branch has two commits: the first contains the environment switch LLAMA_CKPT_REUSE (=0 restores the old behavior) that was used for the measurements below, the last one removes it. To reproduce a measurement, build the first commit.
  • Matters for agent-style use with many short turns; irrelevant for long single answers.
  • Only tools/server/server-context.cpp changes.

Test results

  • Hardware / software: RTX 4070 12 GB (AD104, cc 8.9, 504 GB/s, 48 MB L2), Linux 6.18, NVIDIA driver 615.71, CUDA 13.4, GCC 16.2; Release build, -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=89.
  • Base: speed numbers were measured on prism at 88c4bc60b. The four commits since (SYCL, WebGPU and cuda: fused FWHT quantizer for 64-wide warps (#303)) do not touch the code paths of this PR; the branch is rebased on 2459f68b5, builds, and the test-backend-ops runs below were repeated on it.
  • Model: Ternary-Bonsai-2-27B (PTQ1_0) with a community MTP draft head (Q8_0), speculative decoding with --spec-type draft-mtp --spec-draft-n-max 2, q4_0 K/V cache, one slot.
  • Method: the same binary, one switch (first commit of the branch) flipped through an environment variable, alternating runs A B B A, paired differences with a 95 % bootstrap interval. Run-to-run noise is about 0.04 to 0.12 %, so effects above ~0.3 % are reliable.
  • Growing multi-turn conversation: 44 turns, ~300 words of new text per turn, reply capped at 8 greedy tokens, prefix cache on, 32 checkpoints, 3 alternating A/B pairs with LLAMA_CKPT_REUSE=0 / =1; median over turns 24 to 43: 646.3 -> 626.3 ms per turn (-20.0 ms, -3.1 %), the gain is the same in all three pairs (spread ~2 ms), outputs identical.
  • Why it helps: each prompt checkpoint is about 150 MiB for this hybrid model; a fresh vector was zero-filled and page-faulted for every one.
  • Decode speed (standard 4-prompt benchmark, 8 pairs, measured with the first version of this patch that kept the buffers in two members; the final version only moves them into local variables): 107.69 -> 107.72 tok/s, +0.00 % (95 % CI [-0.11, +0.13] %), i.e. no change in steady-state decode, outputs identical.
  • Functional test suite (80 cases, greedy, 128k context): 80 of 80 answers identical, same pass/fail verdict for each (this run covered the stack up to and including this change and the sampler change, not this commit alone).
  • Not tested: multiple slots, checkpoints of models without recurrent state.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: The patches were developed with Claude Code (Anthropic's coding agent): it wrote the code, the measurement scripts and the first drafts of the commit messages and PR texts. I decided what to work on (which kernels and host paths to optimise, based on profiles of my own decode setup). The measurements and checks listed in the PR texts were run in the Claude Code sessions; I did not re-run them independently. I will maintain the changes. Commits where Claude Code was used carry a Co-Authored-By trailer.

sb32445 and others added 2 commits October 4, 2026 14:21
create_checkpoint() evicts old checkpoints and then emplaces a new one
whose data vectors are resized from empty, so every checkpoint (about
150 MiB for a hybrid model) is zero-filled and page-faulted again. Hand
the buffers of the last evicted checkpoint to the next one instead.
LLAMA_CKPT_REUSE=0 turns this off.

On a growing multi-turn conversation this saves about 18 ms per turn
(645 -> 627 ms); decode speed and outputs are unchanged.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The environment switch of the previous commit was only there to measure the
change; evicted checkpoint buffers are now always handed to the next checkpoint.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@github-actions github-actions Bot added the server label Oct 4, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant