AbsorbQuant β W4A4 NVFP4 diffusion-transformer checkpoints
Pre-quantized 4-bit (weights and activations) checkpoints produced by
AbsorbQuant: decompose-first H-metric low-rank quantization with a
rank-32 fp16 branch on the raw input and a GPTQ-quantized NVFP4 residual,
running on the real nunchaku gemm_w4a4 kernel. Code, method, and theory:
https://github.com/chenjiaj109550158/AbsorbQuant.
| file | model | release | frozen config | content digest (verify.py) | size |
|---|---|---|---|---|---|
| pixart-sigma/absorbquant-pixart-sigma-nvfp4.pt | PixArt-Sigma-XL-2-1024 | v2 | Ξ»=0.1 + S_amax@0.25 | 4a712360β¦ | 325 MB |
| sana-1.6b/absorbquant-sana-1.6b-nvfp4.pt | SANA-1.6B | v2 | Ξ»=0.003 + S_amax@0.75 | c4fde5b3β¦ | 822 MB |
| sdxl-turbo/absorbquant-sdxl-turbo-nvfp4.pt | SDXL-Turbo | v2 | Ξ»=0.1 | 136050d3β¦ | 1.7 GB |
| sdxl-base/absorbquant-sdxl-base-nvfp4.pt | SDXL-Base-1.0 | v2 | λ=0.01 + S_amax@0.5 | d81a86fe⦠| 1.7 GB |
| flux-schnell/absorbquant-flux.1-schnell-nvfp4.safetensors | FLUX.1-schnell | v2 | λ=0.003 + S_amax@0.25 | 7a4f524c⦠| 6.6 GB |
| flux-dev/absorbquant-flux.1-dev-nvfp4.safetensors | FLUX.1-dev | v2 | λ=0.3 + S_amax@0.5 | 240bbc8e⦠| 6.6 GB |
| pixart-sigma-mx/absorbquant-pixart-sigma-mxfp4.pt | PixArt-Sigma (OCP MXFP4, Triton runtime) | v1 | Ξ»=0.3 + S_rms@0.25 | β | 308 MB |
Release lam0.1-v4 β search-free recipe (2026-09-23)
Six additional checkpoints built with protocol v4 and the search-free recipe of
docs/ALPHA_SELECTION_SSE.md: H-SVD damping Ξ» fixed at 0.1 and the smoothing strength Ξ± chosen per layer type from the
qdiff-128 layer-SSE tables (config smooth: {family: amax, rule: sse-per-type, tau: 0.0}), i.e. no Algorithm-1 image
generation at all. Calibration inputs: the clean-room v4 qdiff-128 caches and covariances (machine #3, 2026-09-21).
Each folder also carries PROVENANCE.md, the per-hook alpha_map_per_type.json, the selection record select_sse.json
and the builder sidecars.
| file | model | recipe (Ξ± per layer type) | content digest | size |
|---|---|---|---|---|
| sdxl-turbo-lam0.1-v4/absorbquant-sdxl-turbo-nvfp4-lam0.1-v4.pt | SDXL-Turbo | Ξ»=0.1; attn1_out=0.25, attn1_qkv=0.5, attn2_out=0.25, attn2_q=0.75, ffn_down=0.25, ffn_up=0.5 | e1db3587β¦ | 1.77 GB |
| pixart-sigma-lam0.1-v4/absorbquant-pixart-sigma-nvfp4-lam0.1-v4.pt | PixArt-Sigma-XL-2-1024 | λ=0.1; attn1_out=0.5, attn1_qkv=0.5, attn2_out=0.25, attn2_q=0.5, ffn_down=0.75, ffn_up=0.5 | b0d20f3b⦠| 340 MB |
| sana-1.6b-lam0.1-v4/absorbquant-sana-1.6b-nvfp4-lam0.1-v4.pt | SANA-1.6B | λ=0.1; attn1_out=0.25, attn1_qkv=0.5, attn2_out=0.5, attn2_q=0.75, ffn_down=0.25, ffn_up=0.5 | b1169e2d⦠| 861 MB |
| sdxl-base-lam0.1-v4/absorbquant-sdxl-base-nvfp4-lam0.1-v4.pt | SDXL-Base-1.0 | λ=0.1; attn1_out=0.25, attn1_qkv=0.5, attn2_out=0.25, attn2_q=0.75, ffn_down=0.25, ffn_up=0.5 | 8aaf5f5f⦠| 1.77 GB |
| flux-schnell-lam0.1-v4/absorbquant-flux.1-schnell-nvfp4-lam0.1-v4.safetensors | FLUX.1-schnell | Ξ»=0.1; double.mlp_context_fc1=0.5, double.mlp_context_fc2=0.5, double.mlp_fc1=0.5, double.mlp_fc2=0.25, double.out_proj=0.25, double.out_proj_context=0.25, double.qkv_proj=0.5, double.qkv_proj_context=0.5, single.mlp_fc1=0.5, single.mlp_fc2=0.5, single.out_proj=0.25, single.qkv_proj=0.5 | b62c9ae6β¦ | 7.02 GB |
| flux-dev-lam0.1-v4/absorbquant-flux.1-dev-nvfp4-lam0.1-v4.safetensors | FLUX.1-dev | λ=0.1; double.mlp_context_fc1=0.5, double.mlp_context_fc2=0.5, double.mlp_fc1=0.5, double.mlp_fc2=0.25, double.out_proj=0.25, double.out_proj_context=0.25, double.qkv_proj=0.5, double.qkv_proj_context=0.5, single.mlp_fc1=0.5, single.mlp_fc2=0.5, single.out_proj=0.5, single.qkv_proj=0.5 | c564ca3f⦠| 7.04 GB |
Final 5000-image evaluation (MJHQ-5000 and sDCI-5000, one library each; baselines = the archived SVDQuant+GPTQ / official SVDQuant libraries on the same 16-bit references; S2 minus strongest baseline, PSNR, and wins on PSNR/LPIPS/SSIM/FID-ref/FID-GT):
| model | MJHQ-5000 | sDCI-5000 |
|---|---|---|
| SDXL-Turbo | +0.48 dB, βββββ | +0.53 dB, βββββ |
| PixArt-Ξ£ | +0.16 dB, βββββ | +0.18 dB, βββββ |
| SANA-1.6B | +0.15 dB, βββββ | +0.18 dB, βββββ |
| SDXL-Base | +1.42 dB, βββββ | +1.31 dB, βββββ |
| FLUX.1-schnell (vs official SVDQuant) | +0.09 dB, βββββ | β0.04 dB, βββββ (tie) |
| FLUX.1-dev | pending | pending |
Full tables, the development-set study (MJHQ-500 / sDCI-500, which also chose this recipe) and caveats are in
docs/ALPHA_SELECTION_SSE.md. Download and verify:
python scripts/download.py --model pixart-sigma-lam0.1-v4
python scripts/verify.py --digest workdir/pixart-sigma-lam0.1-v4/checkpoints/absorbquant-pixart-sigma-nvfp4-lam0.1-v4.pt
# rebuild from scratch (metric-level reproduction; bit-exact from the same calibration inputs):
bash scripts/from_scratch/sse.sh pixart-sigma
Release v2 (2026-09-09). The selection algorithm's smoothing gate was
corrected (per-hook gains measured at the built strength Ξ± with both
activation and weight NVFP4-quantized; amax closed-form family; gate
threshold 0) and every recipe was re-derived from scratch. The four v2
files above replace their v1 versions; SANA-1.6B's v2 file followed on
2026-09-10 (re-derived from scratch on an independent RTX 5090, bit-identical
when rebuilt on the release machine), and FLUX.1-schnell's on 2026-09-10
(re-derived from scratch on that independent machine, Algorithm-1 included:
same recipe and same content digest as the release machine's product). v1 recipes are preserved at the repository's git tag v1; v1
files are not reproducible from the current main code.
Usage
git clone https://github.com/chenjiaj109550158/AbsorbQuant && cd AbsorbQuant
pip install -e ".[models]" # plus the nunchaku wheel for your torch/CUDA
python scripts/download.py --model pixart-sigma
python scripts/generate.py --model pixart-sigma --prompt "a corgi in space"
Provenance / verification
Every checkpoint here is the byte-for-byte output of the repository's
from-scratch pipeline (scripts/calibrate.py + scripts/build.py) under
the pinned environment in the repository README: calibration uses only the
128 fixed COCO-derived prompts shipped in the repo β no SVDQuant calibration
artifacts. scripts/verify.py re-checks tensor-level bit-identity between a
rebuild and these files and prints a canonical content digest; the digests
for these files are recorded in the repository's configs/*.yaml.
Licenses
The quantization code is Apache-2.0. Each checkpoint is a derivative of its base model and inherits that model's license β notably FLUX.1-dev (Black Forest Labs Non-Commercial License) and SDXL (CreativeML Open RAIL++-M). Check the base model licenses before commercial use.
Model tree for chenjiaj109550158/AbsorbQuant-NVFP4
Unable to build the model tree, the base model loops to the model itself. Learn more.