Instructions to use QuantFunc/Minimax-H3-Quantfunc-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use QuantFunc/Minimax-H3-Quantfunc-4bit with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("QuantFunc/Minimax-H3-Quantfunc-4bit", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
import torch
from diffusers import DiffusionPipeline
# switch to "mps" for apple devices
pipe = DiffusionPipeline.from_pretrained("QuantFunc/Minimax-H3-Quantfunc-4bit", dtype=torch.bfloat16, device_map="cuda")
prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k"
image = pipe(prompt).images[0]๐ Website | ๐ GitHub | ๐ค Hugging Face | ๐ค ModelScope | ๐ฎ Discord
MiniMax-H3-QuantFunc-4bit
4x compression, quality held.
QuantFunc INT4 cuts MiniMax H3's core weight precision from 16-bit to 4-bit. Character detail, style fidelity and fast motion stay clear and coherent, while H3's native video+audio generation is fully preserved.
In our internal FL2VA evaluation, QuantFunc INT4 vs the BF16 baseline (same prompt, same seed) measures ~23.7 dB PSNR.
Showcase
All clips below were generated by MiniMax-H3-QuantFunc-4bit โ click a player to watch.
| Live-action performance | Animated chase |
|---|---|
| High-speed car chase | Stylized character |
Poster frames are the first frame of each clip. Showcase spec: 896 ร 1184, 124 frames, 24 FPS, ~5s.
Same 124 frames, up to 3.19x FP8's per-step speed
On an RTX 4090, 768 ร 768, 5s, 124 frames:
| Backend | Per-step time | Relative speed |
|---|---|---|
| QuantFunc INT4 | 3.2s | โ |
| INT8 ConvRot | 8.5s | 2.66x faster |
| FP8 | 10.2s | 3.19x faster |
Per-step core-model inference time only โ excludes text encoder, VAE, audio processing and video saving. Actual speed varies with resolution, frame count, driver and software version.
Swap one loader, keep the rest of your workflow
- Install or update ComfyUI-QuantFunc.
- Download the FL2VA 4-step or Ref2VA 8-step weights.
- Swap your model loader for the QuantFunc loader and pick the matching weight file.
Prompts, reference images/videos, audio assets and the rest of your workflow nodes stay unchanged.
Both weight sets already have the acceleration LoRA, Token Refiner, INT4 Refiner and INT8 Conv Sidecar fused in โ no extra components to attach.
RTX 20-series through GB300, one build covers it all
Runs on every NVIDIA SM75+ GPU: RTX 20/30/40/50-series, A100, H100, H200, B100, B200, GB300.
4-bit weights significantly cut the weight-bandwidth cost of loading and inference, making MiniMax H3 much easier to run on consumer GPUs. Actual VRAM needs depend on resolution, frame count, reference-asset count and the rest of your workflow.
Choose a model
| Variant | File | Size | Recommended steps | Use case |
|---|---|---|---|---|
| FL2VA | minimax_h3_fl2va_4step_quantfunc_int4_r128.safetensors |
12.37 GB | 4 steps | Text-to-audio/video, first-frame, last-frame, first+last-frame control |
| Ref2VA | minimax_h3_ref2va_8steps_quantfunc_int4_r128.safetensors |
12.37 GB | 8 steps | Multimodal reference generation from images, video and audio |
FL2VA supports zero, one or two input images; Ref2VA targets more complex multimodal reference scenarios. See the official MiniMax H3 repo for details on both modes.
Use the matching scheduler for the model. Both weight sets already carry the acceleration LoRA fused in โ no need to load it separately.
Loading
Weights use QuantFunc's own sealed safetensors format (quantization parameters and metadata are sealed).
Load with ComfyUI-QuantFunc or the QuantFunc inference engine โ this is not a drop-in Diffusers checkpoint. (This repo declares library_name: diffusers purely so Hugging Face counts real downloads of the flat .safetensors files above โ see HF's download-stats docs โ not because the files load via the diffusers package.)
If the model isn't recognized or fails to load, update ComfyUI-QuantFunc to the latest version and restart ComfyUI.
Technical details
- INT4 weights ร INT4 activations (W4A4)
- SVDQuant rank 128, group size 64
- INT4 Token Refiner + INT4 Refiner
- INT8 Conv Sidecar
- FL2VA measured max relative output error from fp16 adaLN folding: ~7.58e-4
- Minimum GPU architecture: NVIDIA SM75
Source & license
This repository distributes derived quantized weights produced from MiniMax H3.
MiniMax H3 and its derivative weights follow the MiniMax H3 Community License Agreement. QuantFunc's quantization tooling and implementation follow their own respective licenses. Please read and comply with the original model's license terms before use.
Community
- ๐ QuantFunc website
- ๐ ComfyUI-QuantFunc on GitHub
- ๐ฎ Discord
- Downloads last month
- -