Conversation
create_checkpoint() evicts old checkpoints and then emplaces a new one whose data vectors are resized from empty, so every checkpoint (about 150 MiB for a hybrid model) is zero-filled and page-faulted again. Hand the buffers of the last evicted checkpoint to the next one instead. LLAMA_CKPT_REUSE=0 turns this off. On a growing multi-turn conversation this saves about 18 ms per turn (645 -> 627 ms); decode speed and outputs are unchanged. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The environment switch of the previous commit was only there to measure the change; evicted checkpoint buffers are now always handed to the next checkpoint. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
create_checkpoint()evicts old checkpoints and then emplaces a new one whose data vectors are resized from empty, so every prompt checkpoint (about 150 MiB for a hybrid model) is zero-filled and page-faulted again. The buffers of the last evicted checkpoint (data_tgt,data_dft) are handed to the next one instead (two local variables ofcreate_checkpoint(), no new state).Growing multi-turn conversation (44 turns, ~300 words of new text per turn, 8 greedy tokens, MTP n-max 2), 3 alternating A/B pairs of the same binary: 646.3 -> 626.3 ms per turn (-20 ms, -3.1 %), outputs identical. Decode speed is unchanged (+0.00 %, CI [-0.11, +0.13] %): the effect is per request, not per token.
Additional information
LLAMA_CKPT_REUSE(=0restores the old behavior) that was used for the measurements below, the last one removes it. To reproduce a measurement, build the first commit.tools/server/server-context.cppchanges.Test results
-DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=89.prismat88c4bc60b. The four commits since (SYCL, WebGPU andcuda: fused FWHT quantizer for 64-wide warps (#303)) do not touch the code paths of this PR; the branch is rebased on2459f68b5, builds, and thetest-backend-opsruns below were repeated on it.--spec-type draft-mtp --spec-draft-n-max 2, q4_0 K/V cache, one slot.LLAMA_CKPT_REUSE=0/=1; median over turns 24 to 43: 646.3 -> 626.3 ms per turn (-20.0 ms, -3.1 %), the gain is the same in all three pairs (spread ~2 ms), outputs identical.Requirements