Skip to content

perf(kernel): optimize SM120 FP8 GEMM with FP16 accumulation - #1479

Merged
llmc-reviewer merged 7 commits into
mainfrom
yr/h3-fp8-f16-accum
Sep 10, 2026
Merged

perf(kernel): optimize SM120 FP8 GEMM with FP16 accumulation#1479
llmc-reviewer merged 7 commits into
mainfrom
yr/h3-fp8-f16-accum

Conversation

@STwangyingrui

Copy link
Copy Markdown
Contributor

Optimizes MiniMax-H3 inference on RTX 5090/SM120 with an FP8 GEMM path using FP16 accumulation. It includes exact-shape CUTLASS autotuning with persistent caching, qmax-aware checkpoint conversion and validation, and safe FP8-SGL fallback. The optimized path covers selected DiT and Video VAE decoder projections, with configuration and documentation included.

@llmc-reviewer
llmc-reviewer merged commit 05a81ec into main Sep 10, 2026
1 check passed
@llmc-reviewer
llmc-reviewer deleted the yr/h3-fp8-f16-accum branch September 10, 2026 05:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants