๐ŸŒ Website  |  ๐Ÿ™ GitHub  |  ๐Ÿค— Hugging Face  |  ๐Ÿค– ModelScope  |  ๐ŸŽฎ Discord

MiniMax-H3-QuantFunc-4bit

4x compression, quality held.

QuantFunc INT4 cuts MiniMax H3's core weight precision from 16-bit to 4-bit. Character detail, style fidelity and fast motion stay clear and coherent, while H3's native video+audio generation is fully preserved.

In our internal FL2VA evaluation, QuantFunc INT4 vs the BF16 baseline (same prompt, same seed) measures ~23.7 dB PSNR.

Showcase

All clips below were generated by MiniMax-H3-QuantFunc-4bit โ€” click a player to watch.

Live-action performance Animated chase
High-speed car chase Stylized character

Poster frames are the first frame of each clip. Showcase spec: 896 ร— 1184, 124 frames, 24 FPS, ~5s.

Same 124 frames, up to 3.19x FP8's per-step speed

On an RTX 4090, 768 ร— 768, 5s, 124 frames:

Backend Per-step time Relative speed
QuantFunc INT4 3.2s โ€”
INT8 ConvRot 8.5s 2.66x faster
FP8 10.2s 3.19x faster

Per-step core-model inference time only โ€” excludes text encoder, VAE, audio processing and video saving. Actual speed varies with resolution, frame count, driver and software version.

Swap one loader, keep the rest of your workflow

  1. Install or update ComfyUI-QuantFunc.
  2. Download the FL2VA 4-step or Ref2VA 8-step weights.
  3. Swap your model loader for the QuantFunc loader and pick the matching weight file.

Prompts, reference images/videos, audio assets and the rest of your workflow nodes stay unchanged.

Both weight sets already have the acceleration LoRA, Token Refiner, INT4 Refiner and INT8 Conv Sidecar fused in โ€” no extra components to attach.

RTX 20-series through GB300, one build covers it all

Runs on every NVIDIA SM75+ GPU: RTX 20/30/40/50-series, A100, H100, H200, B100, B200, GB300.

4-bit weights significantly cut the weight-bandwidth cost of loading and inference, making MiniMax H3 much easier to run on consumer GPUs. Actual VRAM needs depend on resolution, frame count, reference-asset count and the rest of your workflow.

Choose a model

Variant File Size Recommended steps Use case
FL2VA minimax_h3_fl2va_4step_quantfunc_int4_r128.safetensors 12.37 GB 4 steps Text-to-audio/video, first-frame, last-frame, first+last-frame control
Ref2VA minimax_h3_ref2va_8steps_quantfunc_int4_r128.safetensors 12.37 GB 8 steps Multimodal reference generation from images, video and audio

FL2VA supports zero, one or two input images; Ref2VA targets more complex multimodal reference scenarios. See the official MiniMax H3 repo for details on both modes.

Use the matching scheduler for the model. Both weight sets already carry the acceleration LoRA fused in โ€” no need to load it separately.

Loading

Weights use QuantFunc's own sealed safetensors format (quantization parameters and metadata are sealed).

Load with ComfyUI-QuantFunc or the QuantFunc inference engine โ€” this is not a drop-in Diffusers checkpoint. (This repo declares library_name: diffusers purely so Hugging Face counts real downloads of the flat .safetensors files above โ€” see HF's download-stats docs โ€” not because the files load via the diffusers package.)

If the model isn't recognized or fails to load, update ComfyUI-QuantFunc to the latest version and restart ComfyUI.

Technical details

  • INT4 weights ร— INT4 activations (W4A4)
  • SVDQuant rank 128, group size 64
  • INT4 Token Refiner + INT4 Refiner
  • INT8 Conv Sidecar
  • FL2VA measured max relative output error from fp16 adaLN folding: ~7.58e-4
  • Minimum GPU architecture: NVIDIA SM75

Source & license

This repository distributes derived quantized weights produced from MiniMax H3.

MiniMax H3 and its derivative weights follow the MiniMax H3 Community License Agreement. QuantFunc's quantization tooling and implementation follow their own respective licenses. Please read and comply with the original model's license terms before use.

Community

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support