MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video
4 hours ago
- MiniMax H3 is an open-weights omni-modal video model that can generate video up to 2K resolution and 15 seconds with native stereo audio.
- It supports text-to-video, image-to-video, first-and-last-frame control, and reference-to-video using images, video, or audio.
- The model integrates multimodal context understanding, collapsing multiple tasks into one by processing images, audio, and video together.
- Native stereo audio is generated in the same pass as the video, not as a post-process.
- Editing and motion transfer capabilities allow using a reference video for movement while applying different subject and style.
- The model is optimized for local inference in ComfyUI with a 66% memory reduction (from 123.6 GB to 42.5 GB), enabling it to run on consumer GPUs like the RTX 3060.
- Optimizations include pruning modulation weights, int8 convrot quantization, and custom kernels with dynamic VRAM offloading.
- Example outputs demonstrate various creative uses: comic-book style, product commercial, high-fashion editorial, and vibrant product ad.