Upload README.md with huggingface_hub
Browse files
README.md
ADDED
|
@@ -0,0 +1,89 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: other
|
| 3 |
+
license_name: minimax-h3-license
|
| 4 |
+
license_link: https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE
|
| 5 |
+
base_model: MiniMaxAI/MiniMax-H3
|
| 6 |
+
tags:
|
| 7 |
+
- svdquant
|
| 8 |
+
- w4a4
|
| 9 |
+
- int4
|
| 10 |
+
- video
|
| 11 |
+
- text-to-video
|
| 12 |
+
- quantized
|
| 13 |
+
---
|
| 14 |
+
|
| 15 |
+
# MiniMax-H3 · SVDQuant W4A4 (int4, rank 32, GPTQ)
|
| 16 |
+
|
| 17 |
+
4-bit **weights and activations** for the MiniMax-H3 31B video+audio
|
| 18 |
+
transformer — true [SVDQuant](http://arxiv.org/abs/2411.05007) (ICLR 2025
|
| 19 |
+
Spotlight): activation outliers absorbed into a 16-bit rank-32 low-rank
|
| 20 |
+
branch, the residual GPTQ-rounded to int4, activations quantized to int4
|
| 21 |
+
per-token at runtime, executed on fused CUTLASS tensor-core kernels. This
|
| 22 |
+
is not weight-only quantization.
|
| 23 |
+
|
| 24 |
+
The text conditioner (Qwen3-VL 31B) ships W4A16+GPTQ in the same release:
|
| 25 |
+
**every stored weight is int4**; only norms, embeddings, and the visual
|
| 26 |
+
tower stay bf16, matching MiniMax's own quantization recipe.
|
| 27 |
+
|
| 28 |
+
## Measured (NVIDIA A100 80GB, 124 frames @ 24fps, 960x544)
|
| 29 |
+
|
| 30 |
+
| | BF16 | this release | factor |
|
| 31 |
+
|---|---|---|---|
|
| 32 |
+
| DiT checkpoint | 61.7 GB | **19.6 GB** | 3.15x |
|
| 33 |
+
| generation wall-clock | 484 s | **369 s** | **1.31x faster** |
|
| 34 |
+
| vs unfused reference dequant | 2427 s | 369 s | 6.6x |
|
| 35 |
+
| quantized GEMM (layer level) | — | — | 1.37-1.38x |
|
| 36 |
+
|
| 37 |
+
**All-resident configuration** (this release's DiT + TE together,
|
| 38 |
+
no CPU offload — unreachable for BF16 on one 80GB card):
|
| 39 |
+
|
| 40 |
+
| | BF16 (offloaded) | all-int4 resident | factor |
|
| 41 |
+
|---|---|---|---|
|
| 42 |
+
| generation wall-clock | 484 s | **318 s** | **1.52x faster** |
|
| 43 |
+
| pipeline VRAM | 65 GB peak, offload churn | 48.9 GB steady, 54.3 peak | fits |
|
| 44 |
+
| DiT + TE weights on disk | 123.8 GB | 37.6 GB | 3.3x |
|
| 45 |
+
|
| 46 |
+
Conversion cost: 44 min for the DiT (18 calib + 26 GPTQ) on one A100.
|
| 47 |
+
Kernel outputs agree with the fp32 reference oracle to 1.5-2.1% (the
|
| 48 |
+
bf16-vs-fp32 activation-rounding delta) at every layer shape.
|
| 49 |
+
|
| 50 |
+
Quality: same-seed renders are visually indistinguishable from BF16
|
| 51 |
+
(samples in this repo). On the Z-Image anchor, the same pipeline's GPTQ
|
| 52 |
+
checkpoint scores **better LPIPS than the officially published nunchaku
|
| 53 |
+
checkpoint** (0.288 vs 0.334).
|
| 54 |
+
|
| 55 |
+
## Use
|
| 56 |
+
|
| 57 |
+
```bash
|
| 58 |
+
pip install git+https://github.com/ModelsLab/svdquant git+https://github.com/rootonchair/nunchaku-lite
|
| 59 |
+
```
|
| 60 |
+
|
| 61 |
+
```python
|
| 62 |
+
import svdquant
|
| 63 |
+
transformer = svdquant.load_model("minimax-h3-packed.safetensors") # this repo's file
|
| 64 |
+
# drop into the diffusers ModularPipeline in place of the BF16 transformer
|
| 65 |
+
```
|
| 66 |
+
|
| 67 |
+
Requires an int4-tensor-core GPU (sm_75-89: RTX 20/30/40, A100) and
|
| 68 |
+
torch >= 2.11. An NVFP4 sibling for RTX 50-series (sm_120 has no int4
|
| 69 |
+
path) is planned from a fresh BF16 pass — int4 and fp4 grids do not nest,
|
| 70 |
+
so transcoding is never used.
|
| 71 |
+
|
| 72 |
+
## Honest notes
|
| 73 |
+
|
| 74 |
+
1. Attention stays bf16 — at video sequence lengths it bounds the
|
| 75 |
+
end-to-end speedup (Amdahl); the 1.31x reflects that.
|
| 76 |
+
2. The `token_refiner` (2 blocks, ~2% of params) stays bf16, following
|
| 77 |
+
MiniMax's own int8 recipe.
|
| 78 |
+
3. The AWQ repack of the 50 modulation layers re-derives scales; groups
|
| 79 |
+
GPTQ pushed to -8 take one extra bounded rounding.
|
| 80 |
+
|
| 81 |
+
## Credits and license
|
| 82 |
+
|
| 83 |
+
Weights derive from [MiniMaxAI/MiniMax-H3](https://huggingface.co/MiniMaxAI/MiniMax-H3)
|
| 84 |
+
and inherit its license. Method: SVDQuant (Li et al., MIT HAN Lab).
|
| 85 |
+
Kernels and packed layout: [nunchaku](https://github.com/nunchaku-ai/nunchaku),
|
| 86 |
+
[nunchaku-lite](https://github.com/rootonchair/nunchaku-lite) and
|
| 87 |
+
[diffuse-compressor](https://github.com/rootonchair/diffuse-compressor)
|
| 88 |
+
by rootonchair (Apache-2.0, vendored with attribution). Quantized with
|
| 89 |
+
[svdquant](https://github.com/ModelsLab/svdquant) by ModelsLab.
|