adhikjoshi commited on
Commit
0076fee
·
verified ·
1 Parent(s): aea0a02

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +89 -0
README.md ADDED
@@ -0,0 +1,89 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: other
3
+ license_name: minimax-h3-license
4
+ license_link: https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE
5
+ base_model: MiniMaxAI/MiniMax-H3
6
+ tags:
7
+ - svdquant
8
+ - w4a4
9
+ - int4
10
+ - video
11
+ - text-to-video
12
+ - quantized
13
+ ---
14
+
15
+ # MiniMax-H3 · SVDQuant W4A4 (int4, rank 32, GPTQ)
16
+
17
+ 4-bit **weights and activations** for the MiniMax-H3 31B video+audio
18
+ transformer — true [SVDQuant](http://arxiv.org/abs/2411.05007) (ICLR 2025
19
+ Spotlight): activation outliers absorbed into a 16-bit rank-32 low-rank
20
+ branch, the residual GPTQ-rounded to int4, activations quantized to int4
21
+ per-token at runtime, executed on fused CUTLASS tensor-core kernels. This
22
+ is not weight-only quantization.
23
+
24
+ The text conditioner (Qwen3-VL 31B) ships W4A16+GPTQ in the same release:
25
+ **every stored weight is int4**; only norms, embeddings, and the visual
26
+ tower stay bf16, matching MiniMax's own quantization recipe.
27
+
28
+ ## Measured (NVIDIA A100 80GB, 124 frames @ 24fps, 960x544)
29
+
30
+ | | BF16 | this release | factor |
31
+ |---|---|---|---|
32
+ | DiT checkpoint | 61.7 GB | **19.6 GB** | 3.15x |
33
+ | generation wall-clock | 484 s | **369 s** | **1.31x faster** |
34
+ | vs unfused reference dequant | 2427 s | 369 s | 6.6x |
35
+ | quantized GEMM (layer level) | — | — | 1.37-1.38x |
36
+
37
+ **All-resident configuration** (this release's DiT + TE together,
38
+ no CPU offload — unreachable for BF16 on one 80GB card):
39
+
40
+ | | BF16 (offloaded) | all-int4 resident | factor |
41
+ |---|---|---|---|
42
+ | generation wall-clock | 484 s | **318 s** | **1.52x faster** |
43
+ | pipeline VRAM | 65 GB peak, offload churn | 48.9 GB steady, 54.3 peak | fits |
44
+ | DiT + TE weights on disk | 123.8 GB | 37.6 GB | 3.3x |
45
+
46
+ Conversion cost: 44 min for the DiT (18 calib + 26 GPTQ) on one A100.
47
+ Kernel outputs agree with the fp32 reference oracle to 1.5-2.1% (the
48
+ bf16-vs-fp32 activation-rounding delta) at every layer shape.
49
+
50
+ Quality: same-seed renders are visually indistinguishable from BF16
51
+ (samples in this repo). On the Z-Image anchor, the same pipeline's GPTQ
52
+ checkpoint scores **better LPIPS than the officially published nunchaku
53
+ checkpoint** (0.288 vs 0.334).
54
+
55
+ ## Use
56
+
57
+ ```bash
58
+ pip install git+https://github.com/ModelsLab/svdquant git+https://github.com/rootonchair/nunchaku-lite
59
+ ```
60
+
61
+ ```python
62
+ import svdquant
63
+ transformer = svdquant.load_model("minimax-h3-packed.safetensors") # this repo's file
64
+ # drop into the diffusers ModularPipeline in place of the BF16 transformer
65
+ ```
66
+
67
+ Requires an int4-tensor-core GPU (sm_75-89: RTX 20/30/40, A100) and
68
+ torch >= 2.11. An NVFP4 sibling for RTX 50-series (sm_120 has no int4
69
+ path) is planned from a fresh BF16 pass — int4 and fp4 grids do not nest,
70
+ so transcoding is never used.
71
+
72
+ ## Honest notes
73
+
74
+ 1. Attention stays bf16 — at video sequence lengths it bounds the
75
+ end-to-end speedup (Amdahl); the 1.31x reflects that.
76
+ 2. The `token_refiner` (2 blocks, ~2% of params) stays bf16, following
77
+ MiniMax's own int8 recipe.
78
+ 3. The AWQ repack of the 50 modulation layers re-derives scales; groups
79
+ GPTQ pushed to -8 take one extra bounded rounding.
80
+
81
+ ## Credits and license
82
+
83
+ Weights derive from [MiniMaxAI/MiniMax-H3](https://huggingface.co/MiniMaxAI/MiniMax-H3)
84
+ and inherit its license. Method: SVDQuant (Li et al., MIT HAN Lab).
85
+ Kernels and packed layout: [nunchaku](https://github.com/nunchaku-ai/nunchaku),
86
+ [nunchaku-lite](https://github.com/rootonchair/nunchaku-lite) and
87
+ [diffuse-compressor](https://github.com/rootonchair/diffuse-compressor)
88
+ by rootonchair (Apache-2.0, vendored with attribution). Quantized with
89
+ [svdquant](https://github.com/ModelsLab/svdquant) by ModelsLab.