Skip to content

Support MiniMax-H3 Visual VAE quantization in ModelOpt #2137

Description

@baonudesifeizhai

vLLM-Omni is currently working on production VAE optimization for MiniMax-H3 in
vllm-project/vllm-omni#5948

As a follow-up, would the ModelOpt team be interested in adding quantization support for the MiniMax-H3 Visual VAE?

This could explore FP8, NVFP4, or mixed-precision quantization depending on quality and kernel support, and potentially integrate with the VAE optimization path in vLLM-Omni.

Thanks!

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions