Skip to content

cuda : backport GGML_CUDA_FA_QUANTS from upstream (ggml-org/llama.cpp#28079) - #317

Open
cheese-cakee wants to merge 1 commit into
PrismML-Eng:prismfrom
cheese-cakee:cuda-fa-quants-backport
Open

cheese-cakee wants to merge 1 commit into
PrismML-Eng:prismfrom
cheese-cakee:cuda-fa-quants-backport

Conversation

@cheese-cakee

Copy link
Copy Markdown

Overview

Fixes #267.

Backport of upstream ggml-org#28079 (5a4d0fe, by @pwilkin). With the default CUDA build, a K/V type pair without a compiled vector kernel (for example q8_0/q4_0) is now converted to f16 on the GPU with a one-time warning, instead of being rejected by the CUDA backend and running flash attention on the CPU without any message:

ggml_cuda_flash_attn_ext_vec: no FlashAttention vector kernel compiled for K/V types q8_0-q4_0, converting K and V to f16 instead (slow). Add "q8_0-q4_0" to GGML_CUDA_FA_QUANTS to compile it.

GGML_CUDA_FA_QUANTS selects which pairs get native kernels (default q4_0-q4_0;q8_0-q8_0;f16-f16;bf16-bf16, or all). GGML_CUDA_FA_ALL_QUANTS stays as a deprecated alias for all.

Additional information

The cherry-pick conflicted only in the docs/build.md options table (kept GGML_CUDA_PEER_MAX_BATCH_SIZE). fattn.cu merged without conflicts. The in-place q4_0/q8_0 MMA path from #221 still requires K->type == V->type, so mixed pairs take the f16 conversion path.

Testing on RTX 4050 Laptop (cc 8.9), CUDA 12.6, Linux (WSL2), -DCMAKE_CUDA_ARCHITECTURES=89:

llama-bench, Qwen3-0.6B-Q8_0, -ngl 99 -fa 1 -t 8 -r 3, q8_0 K / q4_0 V, t/s, median of 5 alternating rounds:

test prism default (FA on CPU) this PR, default (f16 conversion) prism + GGML_CUDA_FA_ALL_QUANTS=ON
pp512 888 13937 14378
pp512 @ d4096 47 8631 8881
tg128 78.7 190.7 186.5
tg128 @ d4096 17.3 99.6 143.3

The f16 conversion is much faster than the CPU fallback, but decode at long context is still slower than a compiled pair, which is what the warning says. For q8_0/q8_0 (already supported) the three builds are within run-to-run spread, except pp512 at depth 0, which varies by up to 15% between runs on this laptop.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES. AI was used for the backport and for running the tests; I checked and verified the change and the results afterwards.

…er what is compiled (ggml-org#28079)

* CUDA: add configurable FA quant combinations

Assisted-by: Codex

* remove all flags but , add runtime fallback with warning for uncompiled combination

* Update docs/build.md

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>

* apply code review comments

---------

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
@github-actions github-actions Bot added documentation Improvements or additions to documentation ggml CUDA labels Oct 5, 2026
@bri-prism

Copy link
Copy Markdown
Collaborator

@cheese-cakee @sb32445 @professorpalmer @Wyzix33, several open PRs touch flash attention with quantized K/V: #317 (upstream backport), #307 (MMA for GQA above 4), the FA commit inside #285, and #189. Our plan is to land #317 first, since it matches upstream.

@professorpalmer

Copy link
Copy Markdown

Agreed on the order: #317 first.

On #307 vs the FA commit in #285 (quantized-KV GQA decode on the in-place MMA kernel): one change, #307 with the ggml_cuda_fattn_mma_kv_native_supported(dst) guard, is fine by me, and I am not opening a separate PR for that commit. @sb32445, if you would rather take the guard as a patch against your branch than fold it in yourself, say so and I will send it.

On splitting #285: the branch was six commits, each of which applies cleanly on its own and on the current prism head (eaecb50c7), so #285 is now replaced by one PR per feature, each with the receipts that belong to it:

The FA GQA-decode commit stays out, per the above. On the two speculative-deferral defects from the #221 review: both are fixed in prism as merged (begin() clears catchup_failed with the other per-task state; draft() marks the stash failed on a failed decode instead of dropping it), so none of these PRs touch that path.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CUDA documentation Improvements or additions to documentation ggml

Projects

None yet

Development

Successfully merging this pull request may close these issues.

CUDA: mixed K/V cache types silently run flash attention on CPU (2x slower generation) unless built with GGML_CUDA_FA_ALL_QUANTS

4 participants