AbsorbQuant β€” W4A4 NVFP4 diffusion-transformer checkpoints

Pre-quantized 4-bit (weights and activations) checkpoints produced by AbsorbQuant: decompose-first H-metric low-rank quantization with a rank-32 fp16 branch on the raw input and a GPTQ-quantized NVFP4 residual, running on the real nunchaku gemm_w4a4 kernel. Code, method, and theory: https://github.com/chenjiaj109550158/AbsorbQuant.

file model release frozen config content digest (verify.py) size
pixart-sigma/absorbquant-pixart-sigma-nvfp4.pt PixArt-Sigma-XL-2-1024 v2 Ξ»=0.1 + S_amax@0.25 4a712360… 325 MB
sana-1.6b/absorbquant-sana-1.6b-nvfp4.pt SANA-1.6B v2 Ξ»=0.003 + S_amax@0.75 c4fde5b3… 822 MB
sdxl-turbo/absorbquant-sdxl-turbo-nvfp4.pt SDXL-Turbo v2 Ξ»=0.1 136050d3… 1.7 GB
sdxl-base/absorbquant-sdxl-base-nvfp4.pt SDXL-Base-1.0 v2 Ξ»=0.01 + S_amax@0.5 d81a86fe… 1.7 GB
flux-schnell/absorbquant-flux.1-schnell-nvfp4.safetensors FLUX.1-schnell v2 Ξ»=0.003 + S_amax@0.25 7a4f524c… 6.6 GB
flux-dev/absorbquant-flux.1-dev-nvfp4.safetensors FLUX.1-dev v2 Ξ»=0.3 + S_amax@0.5 240bbc8e… 6.6 GB
pixart-sigma-mx/absorbquant-pixart-sigma-mxfp4.pt PixArt-Sigma (OCP MXFP4, Triton runtime) v1 Ξ»=0.3 + S_rms@0.25 β€” 308 MB

Release lam0.1-v4 β€” search-free recipe (2026-09-23)

Six additional checkpoints built with protocol v4 and the search-free recipe of docs/ALPHA_SELECTION_SSE.md: H-SVD damping Ξ» fixed at 0.1 and the smoothing strength Ξ± chosen per layer type from the qdiff-128 layer-SSE tables (config smooth: {family: amax, rule: sse-per-type, tau: 0.0}), i.e. no Algorithm-1 image generation at all. Calibration inputs: the clean-room v4 qdiff-128 caches and covariances (machine #3, 2026-09-21). Each folder also carries PROVENANCE.md, the per-hook alpha_map_per_type.json, the selection record select_sse.json and the builder sidecars.

file model recipe (Ξ± per layer type) content digest size
sdxl-turbo-lam0.1-v4/absorbquant-sdxl-turbo-nvfp4-lam0.1-v4.pt SDXL-Turbo Ξ»=0.1; attn1_out=0.25, attn1_qkv=0.5, attn2_out=0.25, attn2_q=0.75, ffn_down=0.25, ffn_up=0.5 e1db3587… 1.77 GB
pixart-sigma-lam0.1-v4/absorbquant-pixart-sigma-nvfp4-lam0.1-v4.pt PixArt-Sigma-XL-2-1024 Ξ»=0.1; attn1_out=0.5, attn1_qkv=0.5, attn2_out=0.25, attn2_q=0.5, ffn_down=0.75, ffn_up=0.5 b0d20f3b… 340 MB
sana-1.6b-lam0.1-v4/absorbquant-sana-1.6b-nvfp4-lam0.1-v4.pt SANA-1.6B Ξ»=0.1; attn1_out=0.25, attn1_qkv=0.5, attn2_out=0.5, attn2_q=0.75, ffn_down=0.25, ffn_up=0.5 b1169e2d… 861 MB
sdxl-base-lam0.1-v4/absorbquant-sdxl-base-nvfp4-lam0.1-v4.pt SDXL-Base-1.0 Ξ»=0.1; attn1_out=0.25, attn1_qkv=0.5, attn2_out=0.25, attn2_q=0.75, ffn_down=0.25, ffn_up=0.5 8aaf5f5f… 1.77 GB
flux-schnell-lam0.1-v4/absorbquant-flux.1-schnell-nvfp4-lam0.1-v4.safetensors FLUX.1-schnell Ξ»=0.1; double.mlp_context_fc1=0.5, double.mlp_context_fc2=0.5, double.mlp_fc1=0.5, double.mlp_fc2=0.25, double.out_proj=0.25, double.out_proj_context=0.25, double.qkv_proj=0.5, double.qkv_proj_context=0.5, single.mlp_fc1=0.5, single.mlp_fc2=0.5, single.out_proj=0.25, single.qkv_proj=0.5 b62c9ae6… 7.02 GB
flux-dev-lam0.1-v4/absorbquant-flux.1-dev-nvfp4-lam0.1-v4.safetensors FLUX.1-dev Ξ»=0.1; double.mlp_context_fc1=0.5, double.mlp_context_fc2=0.5, double.mlp_fc1=0.5, double.mlp_fc2=0.25, double.out_proj=0.25, double.out_proj_context=0.25, double.qkv_proj=0.5, double.qkv_proj_context=0.5, single.mlp_fc1=0.5, single.mlp_fc2=0.5, single.out_proj=0.5, single.qkv_proj=0.5 c564ca3f… 7.04 GB

Final 5000-image evaluation (MJHQ-5000 and sDCI-5000, one library each; baselines = the archived SVDQuant+GPTQ / official SVDQuant libraries on the same 16-bit references; S2 minus strongest baseline, PSNR, and wins on PSNR/LPIPS/SSIM/FID-ref/FID-GT):

model MJHQ-5000 sDCI-5000
SDXL-Turbo +0.48 dB, βœ“βœ“βœ“βœ“βœ“ +0.53 dB, βœ“βœ“βœ“βœ“βœ—
PixArt-Ξ£ +0.16 dB, βœ“βœ“βœ“βœ—βœ— +0.18 dB, βœ“βœ“βœ“βœ“βœ“
SANA-1.6B +0.15 dB, βœ“βœ“βœ“βœ—βœ“ +0.18 dB, βœ“βœ“βœ“βœ“βœ“
SDXL-Base +1.42 dB, βœ“βœ“βœ“βœ“βœ“ +1.31 dB, βœ“βœ“βœ“βœ“βœ“
FLUX.1-schnell (vs official SVDQuant) +0.09 dB, βœ“βœ“βœ“βœ“βœ— βˆ’0.04 dB, βœ—βœ—βœ—βœ“βœ— (tie)
FLUX.1-dev pending pending

Full tables, the development-set study (MJHQ-500 / sDCI-500, which also chose this recipe) and caveats are in docs/ALPHA_SELECTION_SSE.md. Download and verify:

python scripts/download.py --model pixart-sigma-lam0.1-v4
python scripts/verify.py --digest workdir/pixart-sigma-lam0.1-v4/checkpoints/absorbquant-pixart-sigma-nvfp4-lam0.1-v4.pt
# rebuild from scratch (metric-level reproduction; bit-exact from the same calibration inputs):
bash scripts/from_scratch/sse.sh pixart-sigma

Release v2 (2026-09-09). The selection algorithm's smoothing gate was corrected (per-hook gains measured at the built strength Ξ± with both activation and weight NVFP4-quantized; amax closed-form family; gate threshold 0) and every recipe was re-derived from scratch. The four v2 files above replace their v1 versions; SANA-1.6B's v2 file followed on 2026-09-10 (re-derived from scratch on an independent RTX 5090, bit-identical when rebuilt on the release machine), and FLUX.1-schnell's on 2026-09-10 (re-derived from scratch on that independent machine, Algorithm-1 included: same recipe and same content digest as the release machine's product). v1 recipes are preserved at the repository's git tag v1; v1 files are not reproducible from the current main code.

Usage

git clone https://github.com/chenjiaj109550158/AbsorbQuant && cd AbsorbQuant
pip install -e ".[models]"   # plus the nunchaku wheel for your torch/CUDA
python scripts/download.py --model pixart-sigma
python scripts/generate.py --model pixart-sigma --prompt "a corgi in space"

Provenance / verification

Every checkpoint here is the byte-for-byte output of the repository's from-scratch pipeline (scripts/calibrate.py + scripts/build.py) under the pinned environment in the repository README: calibration uses only the 128 fixed COCO-derived prompts shipped in the repo β€” no SVDQuant calibration artifacts. scripts/verify.py re-checks tensor-level bit-identity between a rebuild and these files and prints a canonical content digest; the digests for these files are recorded in the repository's configs/*.yaml.

Licenses

The quantization code is Apache-2.0. Each checkpoint is a derivative of its base model and inherits that model's license β€” notably FLUX.1-dev (Black Forest Labs Non-Commercial License) and SDXL (CreativeML Open RAIL++-M). Check the base model licenses before commercial use.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for chenjiaj109550158/AbsorbQuant-NVFP4

Unable to build the model tree, the base model loops to the model itself. Learn more.