Skip to content

llama-server: Use reference-counted checkpoints instead of deep-copies - #30137

Open
cwriter wants to merge 2 commits into
ggml-org:masterfrom
cwriter:feat/checkpoint-cache-optimization
Open

cwriter wants to merge 2 commits into
ggml-org:masterfrom
cwriter:feat/checkpoint-cache-optimization

Conversation

@cwriter

@cwriter cwriter commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Overview

llama-server currently creates and reuses context-checkpoints through a vector copy.
Since the checkpoints are immutable by design, this patch introduces reference-counted sharing instead.

This PR includes 3 changes to the system:

  • Use std::shared_ptr for the context checkpoints, allowing cheap (ref-counted) copies
  • Skip zero-initialization - this would be a security risk, but we ensure that the bytes are actually written
  • Reject broken requests before inserting into the prompt cache, avoiding leakages and thrashing. This is cheap because it's just an std::move

Additional information

This applies mainly to the host-side copies. The effects become larger as context grows. The effects are most visible when queuing events, as they need to wait on the serialized cache-checkpoint copies.
Tested on 3x Arc Pro B60, Threadripper 1950X x399 platform, 128GB RAM, qwen3.8-flash-next:UD-IQ4XS, RAM prompt cache 8GB, 16 checkpoints, checkpoint spacing 8192 tokens, no MTP, running on the SYCL backend (but should not affect this code).

Workload Metric Master This PR Reduction
Fresh history, 2-3 retained checkpoints Cache update 529.213 ms 250.798 ms 52.6%
Fresh history, 2-3 retained checkpoints First visible token 831.134 ms 542.500 ms 34.7%
Fresh history, 2-3 retained checkpoints Full 16-token request 1583.031 ms 1295.366 ms 18.2%
Grown history, 8-9 retained checkpoints Cache update 1039.985 ms 244.963 ms 76.4%
Grown history, 8-9 retained checkpoints First visible token 1345.784 ms 532.247 ms 60.5%
Grown history, 8-9 retained checkpoints Full 16-token request 2100.605 ms 1287.922 ms 38.7%

Memory savings:

Workload Metric (MiB) Base PR Saved
One slot, 2 checkpoints Steady RSS 1420.5 1421.4 -0.9 (-0.1%)
One slot, 2 checkpoints Peak RSS 1845.5 1621.1 224.4 (12.2%)
One slot, 8-9 checkpoints Steady RSS 2828.0 2829.2 -1.2 (-0.0%)
One slot, 8-9 checkpoints Peak RSS 4209.6 3197.8 1011.8 (24.0%)
Two slots, idle caching Steady RSS 4757.2 3181.3 1576.0 (33.1%)
Two slots, idle caching Peak RSS 4757.2 3181.3 1576.0 (33.1%)

Requirements

  • I have read and agree with the contributing guidelines - Yes
  • AI usage disclosure: Yes, Claude Code Opus 5.5 for issue identification and code draft, Opus Sol 6.1 for PR extraction

@cwriter
cwriter requested review from a team as code owners October 8, 2026 07:49
@github-actions github-actions Bot added the server label Oct 8, 2026
cwriter added 2 commits October 8, 2026 11:14
Assisted-by: Codex
Adapt the slot save/restore checkpoint appendix from ggml-org#26004 to the
shared_ptr checkpoint list and the common_state_data buffers.

Assisted-by: Claude Opus 5.5
@cwriter
cwriter force-pushed the feat/checkpoint-cache-optimization branch from d133b80 to d2d35a3 Compare October 8, 2026 11:14
@Davidyz

Davidyz commented Oct 9, 2026

Copy link
Copy Markdown

Would the shared references to the checkpoints also reduce the memory footprint of the prompt cache? If so, it might be useful to include numbers in the table.

@cwriter

cwriter commented Oct 9, 2026

Copy link
Copy Markdown
Contributor Author

Would the shared references to the checkpoints also reduce the memory footprint of the prompt cache?

No, this patch uses shared_ptr only to have an easy copy mechanism. It does not reduce the memory usage overall, except maybe transiently during copies.

@Davidyz

Davidyz commented Oct 9, 2026

Copy link
Copy Markdown

No, this patch uses shared_ptr only to have an easy copy mechanism. It does not reduce the memory usage overall, except maybe transiently during copies.

I see. But just for clarification, the scenario I had in mind was when 2 prompts share the same prefix (eg. system prompt+tool schema) but different user messages. With this pr, do they share the same checkpoint object for the matching system prompt? I'm curious because if the original behaviour was a vector copy, it sounds like it's going to store the same checkpoint twice.

@cwriter

cwriter commented Oct 9, 2026

Copy link
Copy Markdown
Contributor Author

Yes, sorry, you're right: If there are the same prefixes (a growing conversation), this saves (peak) RAM. It does not save RAM unrelated histories, even if they have the same input (there is no deduplication).

Added the Memory table

@Davidyz

Davidyz commented Oct 9, 2026

Copy link
Copy Markdown

Sweet! I have only a few GB of RAM left once I load the model/prompt cache, so even a few hundred MB of savings would be meaningful to me. Looking forward to seeing this in mainline.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants