Skip to content

Add NVFP4/FP8 Q/K/P/V + 2:4 attention quantization for MLA (vLLM TRITON_MLA) - #2244

Draft
kaix-nv wants to merge 5 commits into
mainfrom
kaix/mla_opt
Draft

Add NVFP4/FP8 Q/K/P/V + 2:4 attention quantization for MLA (vLLM TRITON_MLA)#2244
kaix-nv wants to merge 5 commits into
mainfrom
kaix/mla_opt