Skip to content

cuda: use the MMA flash attention kernel for GQA above 4 with quantized K/V on Ada - #307

Open
sb32445 wants to merge 3 commits into
PrismML-Eng:prismfrom
sb32445:pr/fattn-gqa-mma
Open

sb32445 wants to merge 3 commits into
PrismML-Eng:prismfrom
sb32445:pr/fattn-gqa-mma

Commits

Commits on Oct 4, 2026

Commits on Oct 7, 2026