Conversation
…e chain common_sampler_sample() built a token array for the whole vocabulary (set_logits) and then let the top-k sampler partial-sort it. When every sampler before top-k does nothing with the current settings (penalties, dry and top-n-sigma off, no logit bias, no mirostat, no backend sampling), no grammar applied first and the reasoning budget not forcing, select the k largest logits directly. The selection repeats the heap steps of libstdc++ std::partial_sort, so entries and the order of equal logits are the same as before. Checked against std::partial_sort on 120000 random arrays (ties, NaN, inf) and on server outputs (sampling, tool calls, penalties, top_k 1/200, logprobs): identical. LLAMA_SAMPLER_FAST_TOPK=0 turns it off. Decode with MTP n-max 2: +1.9 % (greedy benchmark), +1.85 % with the Hermes sampling settings (thinking, budget 16384, top_k 20). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The environment switch of the previous commit was only there to measure the change; the direct top-k selection is now used whenever its conditions hold. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
common_sampler_sample()builds a token array for the whole vocabulary (set_logits, 248k entries) and then lets the top-k sampler partial-sort it. When top-k is the first sampler that does anything (penalties, DRY and top-n-sigma off with the current settings, no logit bias, no mirostat, no backend sampling), no grammar is applied first and the reasoning budget is not forcing, the k largest logits (k <= 128) are now selected directly from the logits pointer. The selection repeats the heap steps of libstdc++std::partial_sort(make_heap on the first k, replace the heap top by every larger logit in index order, sort_heap), so the entries and the order of equal logits are the same as before; a 16-wide block test skips blocks without a candidate. Everything else (the rest of the chain,cur_p) runs unchanged on the k entries.With MTP speculative decoding the sampler runs on 3 positions per step on the host while the GPU is idle (~0.55 ms of a 24.6 ms step). RTX 4070, Bonsai 2 27B PTQ1_0 + MTP head (n-max 2), 8 interleaved A/B pairs of the same binary, outputs identical in all runs: greedy benchmark +1.92 % (CI [+1.85, +2.01] %), thinking sampling (1.0 / 0.95 / 20 / 0.05, seed 42, reasoning budget 16384, context 114688) +1.85 % (CI [+1.75, +1.94] %).
Additional information
LLAMA_SAMPLER_FAST_TOPK(=0restores the old path) that was used for the measurements below, the last one removes it. To reproduce a measurement, build the first commit. Penalties, logit bias, mirostat, backend sampling, k > 128 or a grammar applied first take the old path.std::partial_sorton 120000 random arrays (k 1-128, n up to 248k, many ties, NaN, +-inf, sorted inputs): identical index and logit at every rank. Server outputs with the switch on and off are identical for tool calls with a grammar, thinking, presence/repeat penalty, top_k 1 and 200, and logprobs (tool-call ids ignored). The checks are not part of this PR.__adjust_heap. The identical tie order is therefore guaranteed (and tested) with libstdc++ only; with libc++ or MSVC STL the order of exactly equal logits may differ from whatstd::partial_sortgave before (the result is still a correct top-k). If you prefer not to depend on that, the same speed-up is possible with a plain selection that gives up the tie order.common_sampler_clone/common_sampler_copycarry the newfast_topkmember along.common/sampling.cppchanges.Test results
-DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=89.prismat88c4bc60b; the four commits since (SYCL, WebGPU andcuda: fused FWHT quantizer for 64-wide warps (#303)) do not touch the sampler. The branch is rebased on2459f68b5and builds.--spec-type draft-mtp --spec-draft-n-max 2, q4_0 K/V cache, one slot; vocabulary 248320 tokens.LLAMA_SAMPLER_FAST_TOPK, first commit of the branch) flipped through an environment variable, alternating runs A B B A, 8 pairs, paired differences with a 95 % bootstrap interval; run-to-run noise about 0.04 to 0.12 %.--reasoning-format deepseek --reasoning-budget 16384, thinking sampling 1.0 / 0.95 / 20 / 0.05, fixed seed 42): 100.84 -> 102.71 tok/s, +1.85 % (CI [+1.75, +1.94] %), outputs identical, acceptance unchanged (61.6 %).set_logits0.36 ms, top-k partial sort 0.18 ms; nsys with CPU sampling).std::partial_sorton 120000 random arrays (k 1 to 128, n up to 248k, many ties, NaN, +-inf, ascending and descending inputs): identical index and logit at every rank (libstdc++, GCC 16). This test is not part of the PR; I can turn it into a test case intests/if you want it.LLAMA_SAMPLER_FAST_TOPK=0and=1on the same build are identical for: tool calls with a grammar (sampling and greedy), thinking with sampling,presence_penalty1.5 (old path),repeat_penalty1.1 (old path),top_k200 (old path) andtop_k1, andlogprobs/top_logprobs(tool-call ids ignored).test-samplingpasses on the rebased branch (it exercisesllama_sampler_*, notcommon_sampler, so it does not cover the new path).top_kabove 128 (old path), grammar-first sampling with an active grammar (old path by construction).Requirements