Skip to content

common : take the top-k straight from the logits when top-k starts the chain - #313

Open
sb32445 wants to merge 2 commits into
PrismML-Eng:prismfrom
sb32445:pr/sampler-topk-from-logits
Open

sb32445 wants to merge 2 commits into
PrismML-Eng:prismfrom
sb32445:pr/sampler-topk-from-logits

Conversation

@sb32445

@sb32445 sb32445 commented Oct 4, 2026

Copy link
Copy Markdown

Overview

common_sampler_sample() builds a token array for the whole vocabulary (set_logits, 248k entries) and then lets the top-k sampler partial-sort it. When top-k is the first sampler that does anything (penalties, DRY and top-n-sigma off with the current settings, no logit bias, no mirostat, no backend sampling), no grammar is applied first and the reasoning budget is not forcing, the k largest logits (k <= 128) are now selected directly from the logits pointer. The selection repeats the heap steps of libstdc++ std::partial_sort (make_heap on the first k, replace the heap top by every larger logit in index order, sort_heap), so the entries and the order of equal logits are the same as before; a 16-wide block test skips blocks without a candidate. Everything else (the rest of the chain, cur_p) runs unchanged on the k entries.

With MTP speculative decoding the sampler runs on 3 positions per step on the host while the GPU is idle (~0.55 ms of a 24.6 ms step). RTX 4070, Bonsai 2 27B PTQ1_0 + MTP head (n-max 2), 8 interleaved A/B pairs of the same binary, outputs identical in all runs: greedy benchmark +1.92 % (CI [+1.85, +2.01] %), thinking sampling (1.0 / 0.95 / 20 / 0.05, seed 42, reasoning budget 16384, context 114688) +1.85 % (CI [+1.75, +1.94] %).

Additional information

  • The branch has two commits: the first contains the environment switch LLAMA_SAMPLER_FAST_TOPK (=0 restores the old path) that was used for the measurements below, the last one removes it. To reproduce a measurement, build the first commit. Penalties, logit bias, mirostat, backend sampling, k > 128 or a grammar applied first take the old path.
  • Checked against std::partial_sort on 120000 random arrays (k 1-128, n up to 248k, many ties, NaN, +-inf, sorted inputs): identical index and logit at every rank. Server outputs with the switch on and off are identical for tool calls with a grammar, thinking, presence/repeat penalty, top_k 1 and 200, and logprobs (tool-call ids ignored). The checks are not part of this PR.
  • It copies the sift-down/sift-up of libstdc++'s __adjust_heap. The identical tie order is therefore guaranteed (and tested) with libstdc++ only; with libc++ or MSVC STL the order of exactly equal logits may differ from what std::partial_sort gave before (the result is still a correct top-k). If you prefer not to depend on that, the same speed-up is possible with a plain selection that gives up the tie order.
  • common_sampler_clone / common_sampler_copy carry the new fast_topk member along.
  • Only common/sampling.cpp changes.

Test results

  • Hardware / software: RTX 4070 12 GB (AD104, cc 8.9), Linux 6.18, NVIDIA driver 615.71, CUDA 13.4, GCC 16.2 (libstdc++); Release build, -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=89.
  • Base: speed numbers were measured on prism at 88c4bc60b; the four commits since (SYCL, WebGPU and cuda: fused FWHT quantizer for 64-wide warps (#303)) do not touch the sampler. The branch is rebased on 2459f68b5 and builds.
  • Model: Ternary-Bonsai-2-27B (PTQ1_0) with a community MTP draft head (Q8_0), --spec-type draft-mtp --spec-draft-n-max 2, q4_0 K/V cache, one slot; vocabulary 248320 tokens.
  • Method: the same binary, one switch (LLAMA_SAMPLER_FAST_TOPK, first commit of the branch) flipped through an environment variable, alternating runs A B B A, 8 pairs, paired differences with a 95 % bootstrap interval; run-to-run noise about 0.04 to 0.12 %.
  • Speed, 4 greedy prompts x 256 tokens, depth 0: 107.8 -> 109.85 tok/s, +1.92 % (95 % CI [+1.85, +2.01] %), outputs identical in all 16 runs.
  • Agent-like setup (context 114688, --reasoning-format deepseek --reasoning-budget 16384, thinking sampling 1.0 / 0.95 / 20 / 0.05, fixed seed 42): 100.84 -> 102.71 tok/s, +1.85 % (CI [+1.75, +1.94] %), outputs identical, acceptance unchanged (61.6 %).
  • Where the time went: with MTP the sampler runs on 3 positions per step on the host while the GPU is idle, about 0.55 ms of a 24.6 ms step (set_logits 0.36 ms, top-k partial sort 0.18 ms; nsys with CPU sampling).
  • Correctness of the selection: compared with std::partial_sort on 120000 random arrays (k 1 to 128, n up to 248k, many ties, NaN, +-inf, ascending and descending inputs): identical index and logit at every rank (libstdc++, GCC 16). This test is not part of the PR; I can turn it into a test case in tests/ if you want it.
  • Server outputs with LLAMA_SAMPLER_FAST_TOPK=0 and =1 on the same build are identical for: tool calls with a grammar (sampling and greedy), thinking with sampling, presence_penalty 1.5 (old path), repeat_penalty 1.1 (old path), top_k 200 (old path) and top_k 1, and logprobs/top_logprobs (tool-call ids ignored).
  • test-sampling passes on the rebased branch (it exercises llama_sampler_*, not common_sampler, so it does not cover the new path).
  • Not tested: libc++ / MSVC STL (see above), several slots, top_k above 128 (old path), grammar-first sampling with an active grammar (old path by construction).

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: The patches were developed with Claude Code (Anthropic's coding agent): it wrote the code, the measurement scripts and the first drafts of the commit messages and PR texts. I decided what to work on (which kernels and host paths to optimise, based on profiles of my own decode setup). The measurements and checks listed in the PR texts were run in the Claude Code sessions; I did not re-run them independently. I will maintain the changes. Commits where Claude Code was used carry a Co-Authored-By trailer.

sb32445 and others added 2 commits October 4, 2026 14:21
…e chain

common_sampler_sample() built a token array for the whole vocabulary
(set_logits) and then let the top-k sampler partial-sort it. When every
sampler before top-k does nothing with the current settings (penalties,
dry and top-n-sigma off, no logit bias, no mirostat, no backend
sampling), no grammar applied first and the reasoning budget not
forcing, select the k largest logits directly. The selection repeats
the heap steps of libstdc++ std::partial_sort, so entries and the order
of equal logits are the same as before. Checked against std::partial_sort
on 120000 random arrays (ties, NaN, inf) and on server outputs (sampling,
tool calls, penalties, top_k 1/200, logprobs): identical.
LLAMA_SAMPLER_FAST_TOPK=0 turns it off.

Decode with MTP n-max 2: +1.9 % (greedy benchmark), +1.85 % with the
Hermes sampling settings (thinking, budget 16384, top_k 20).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The environment switch of the previous commit was only there to measure the
change; the direct top-k selection is now used whenever its conditions hold.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@sb32445 sb32445 changed the title server : reuse the buffers of evicted prompt checkpoints common : take the top-k straight from the logits when top-k starts the chain Oct 4, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant