Skip to content

cuda: use the MMA flash attention kernel for GQA above 4 with quantized K/V on Ada - #307

Open
sb32445 wants to merge 2 commits into
PrismML-Eng:prismfrom
sb32445:pr/fattn-gqa-mma
Open

sb32445 wants to merge 2 commits into
PrismML-Eng:prismfrom
sb32445:pr/fattn-gqa-mma

cuda: remove the GQA/MMA and vector-kernel query-limit switches

069cd48
Select commit
Loading
Failed to load commit list.
Sign in for the full log view

Annotations

1 warning
labeler
succeeded Oct 4, 2026 in 11s