Skip to content

server : shared context checkpoints, no zero-fill of state buffers (port of upstream #30137) - #70

Merged
routhjim merged 1 commit into
mainfrom
pr/server-ckpt-shared
Oct 9, 2026
Merged

routhjim merged 1 commit into
mainfrom
pr/server-ckpt-shared

Conversation

@routhjim

@routhjim routhjim commented Oct 9, 2026

Copy link
Copy Markdown
Owner

Port of upstream ggml-org/llama.cpp#30137.

  • Shared checkpoints: context checkpoints become shared_ptr<const>. server_prompt_cache::alloc() used to deep-copy every checkpoint of the slot on every save, about 150 MiB of 27B GDN state each.
  • No zero-fill of state buffers: a no-init allocator skips zero-filling hundreds of MiB per save, checkpoint and NVMe reload. That includes the lab's spill and reload buffers and the checkpoint sidecar.
  • A state read that comes back short now drops its entry.

It doesn't conflict with the lab's NVMe tier: checkpoints stay in RAM before and after. Upstream #26004 (checkpoints across slot save/restore) is already covered by the lab's KCPT sidecar, so it is not ported.

Measured with the 27B on both XTX:

  • Setup: one instance, np1, -cram 0 --cache-disk 16384 -ctxcp 4, prod-1006 (A) vs this (B), alternated ABAB.
  • Workload: three real trajectories of 11-15K tokens in rotation, so each turn after the first saves and reloads.
  • Metric: the server's own prompt cache update took X ms.
A1 A2 B1 B2
median update (ms) 2707 1880 1480 1429
  • Pooled (16 per arm): median 2550 to 1480 ms (−42%), mean 2878 to 1479 ms.
  • Spill write: median 1780 to 818 ms. NVMe reload is unchanged (277 vs 271 ms).
  • Cache hits: identical in every arm.
  • Caveat: the box had ~8 GB of free RAM during the runs, so writes were throttled. The absolute times are pessimistic; the direction is the point.

Assisted-by: Claude Opus 5.5

🤖 Generated with Claude Code

https://claude.ai/code/session_013XBY17SSRKGmDXc2YFsrmp

…ill of state buffers (port of upstream #30137)

Checkpoints become shared_ptr<const>, so saving a slot to the prompt cache no longer deep-copies
them; state buffers (checkpoints, cache entries, NVMe reload) use a no-init allocator; prompt_save
drops an entry whose state read came back short. Lab sidecar save/restore adapted.

Assisted-by: Claude Opus 5.5
@routhjim
routhjim merged commit aa9afe0 into main Oct 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant