Repository navigation
Conversation
…ed K/V on Ada With quantized K/V on Ada the vector kernel was picked for 1 and 2 queries at every context length. At long context it is much slower than the tensor core kernel. The branch for F16 K/V already avoids the vector kernel when the GQA ratio is above 4; this applies the same ratio rule to the quantized branch. Ternary Bonsai 2 27B PTQ1_0 (head size 256, GQA 6, q4_0 K/V, mean-centered), RTX 4070 12 GB, greedy decode, one slot, tokens per second by filled context: depth no MTP: before -> after draft-mtp n-max 1: before -> after 32768 44.8 -> 54.2 53.3 -> 84.0 120000 25.3 -> 41.2 23.2 -> 68.5 draft-mtp n-max 2 (3 queries, already on the MMA kernel) and depth 0 are unchanged. llama-bench pp1/pp2 at depth 2048..16384 shows the MMA kernel is not slower from 2k keys on. test-backend-ops FLASH_ATTN_EXT passes (2994/2994). With 12 extra local cases for q4_0 and q8_0 at GQA 6, head size 256, 1 and 2 queries it passes on both paths (3006/3006). GGML_CUDA_FATTN_GQA_MMA=0 restores the old choice, GGML_CUDA_FATTN_VEC_MAXQ changes the query limit of the vector kernel (default 2). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
GGML_CUDA_FATTN_GQA_MMA and GGML_CUDA_FATTN_VEC_MAXQ of the previous commit were only there to measure the change. The vector kernel keeps the original limit of 2 queries and is used for GQA <= 4 only. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This was referenced Oct 5, 2026
Open
Open
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
On Ada, flash attention with quantized K/V takes the vector kernel for 1-2 queries regardless of the GQA ratio. With GQA above 4 the MMA kernel is faster (same rule that already exists for F16 K/V a few lines above). This makes the vector kernel apply only for GQA <= 4.
Ternary-Bonsai-2-27B (GQA 6, head size 256, q4_0 K/V), RTX 4070, 120k context, decode without speculative decoding: 25.3 -> 41.2 tok/s (+63 %). With MTP speculative decoding (n-max 2) 79.7 tok/s at 120k; depth 0 unchanged.
Additional information
GGML_CUDA_FATTN_GQA_MMA(=0restores the old choice) andGGML_CUDA_FATTN_VEC_MAXQ(query limit of the vector kernel, default 2) that were used for the measurements below, the last one removes both (the vector kernel keeps its original limit of 2 queries). To reproduce a measurement, build the first commit.K->ne[1] >= 8192). The MMA kernel was not slower than the vector kernel inllama-benchpp1/pp2 from 2k keys on, and decode at depth 0 is also slightly faster (+0.62 %), so I did not add one.Test results
-DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=89.prismat88c4bc60b; the four commits since (SYCL, WebGPU andcuda: fused FWHT quantizer for 64-wide warps (#303)) do not touch this code path. The branch is rebased on2459f68b5, builds, andtest-backend-opswas repeated on it.GGML_CUDA_FATTN_GQA_MMA=0against=1). Withdraft-mtp n-max 1(2 queries): 32768 keys 53.3 -> 84.0, 120000 keys 23.2 -> 68.5 tok/s. Withn-max 2(3 queries, already on the MMA kernel) and at depth 0 nothing changes.llama-benchpp1/pp2 at depth 2048 to 16384: the MMA kernel is not slower than the vector kernel from 2k keys on.test-backend-ops test -b CUDA0 -o FLASH_ATTN_EXT: 2994/2994 passed on the rebased branch (CUDA0 against CPU). On the earlier base I also ran 12 extra local cases (q4_0/q8_0, GQA 6, head size 256, 1-512 queries, 1024-16384 keys), 3006/3006 on both paths; they are not part of this PR.=0) against this PR (build without the tile change of the other PR): 3 of 4 texts diverge, the first difference at 13 %, 40 % and 77 % of the text, 1 of 4 is identical. I did not measure a task-level quality metric for the single-query path (perplexity does not exercise it; it runs batches of 512 queries).n-max 2, 3 queries) the output hashes are identical with and without this change at depth 16384 and 32768.cc >= GGML_CUDA_CC_ADA_LOVELACE, so the new rule applies to every NVIDIA GPU from cc 8.9 up (Ada, Hopper, Blackwell), not only to Ada; I could only test Ada and can restrict it to cc 8.9 if you prefer. HIP and MUSA are not affected by this condition.Requirements