Instructions to use ProCreations/Image-2.1-Turbo-Calibrated-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use ProCreations/Image-2.1-Turbo-Calibrated-NVFP4 with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("ProCreations/Image-2.1-Turbo-Calibrated-NVFP4", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
Image 2.1 Turbo Calibrated NVFP4
Built with Qwen. An independently calibrated NVFP4 W4A4 quantization (with BF16 rank-128 corrections) of the image transformer from Qwen-Image-2.1-Turbo, with an accelerated runtime. It generates 1024 × 1024 images in 0.96 s end to end on an RTX PRO 6000 Blackwell. Research/evaluation only under the included Qwen Research License. This is not an official Qwen release.
This repository contains the quantized transformer (4.87 GB), an explicit loader and CLI, the full calibration/evaluation source, paired comparisons against BF16 Turbo, raw benchmark and profiler evidence, and a 30-second real-time demonstration. The BF16 text encoder, VAE, processor and the checkpoint's 8-step schedule load from the pinned upstream revision d65dbc9a7e8f6b5479e33dee6030eaab2a906509.
Default: 7 of 8 steps in NVFP4, step 1 in calibrated FP8
By default, step 1 of the 8-step schedule runs the calibrated FP8 transformer from Image-2.1-Turbo-Calibrated-FP8. That step also fills the prompt/reference KV cache. Steps 2–8 run NVFP4. A few-step model fixes its composition early, so this one higher-precision step brings NVFP4 much closer to BF16 (LPIPS 0.123 → 0.086, latent cosine 0.966 → 0.983) at the same speed. FP8's cheaper prefix pass offsets its slower GEMMs. The FP8 transformer (7.26 GB) downloads once from its pinned revision and stays resident, which adds about 7 GB of VRAM. --fp8-first-steps 0 gives pure NVFP4; --fp8-first-steps 2 is closer still (LPIPS 0.063–0.071) at about 1.0 s.
Speed
RTX PRO 6000 Blackwell (SM120, 96 GB), batch 1, all 8 saved-schedule steps, CFG 1, prefix KV cache.
| Runtime | 1024 × 1024 | 2048 × 2048 |
|---|---|---|
| BF16 Turbo (upstream Diffusers, eager) | 2.05 s | 11.45 s |
| NVFP4, eager, pure | 1.31 s | 8.40 s |
| NVFP4 + FP8 step 1, compiled, BF16 attention | 1.03 s | 6.80 s |
| Pure NVFP4, compiled + SageAttention2 | 0.96 s | 5.33 s |
| Release default (NVFP4 + FP8 step 1, compiled + SageAttention2) | 0.96 s | 5.38 s |
CUDA-synchronized end-to-end wall time: prompt encoding, all 8 transformer passes and VAE decode. Excluded: model load (about 4.4 s from local NVMe), first-use compilation in a fresh process (about 11 s at 1024, 15 s at 2048; once per resolution), and PNG writing. Five measured runs at 1024 and three at 2048 (five for the rows measured at both sizes in one process), after one warmup. Release default at 1024: mean 0.960 s, min 0.958 s. Stage breakdown: encoder 16 ms, step 1 (FP8) 138 ms, cached NVFP4 steps 101 ms each, VAE 93 ms. Peak GPU allocation: 33.7 GB (1024) and 42.8 GB (2048) with the FP8 step-1 model resident; 26.4 GB and 35.5 GB for pure NVFP4. Raw numbers: reports/benchmarks. With the 2× faster GEMMs, attention is the largest single cost at 2048 (16,384 image tokens).
Quality against BF16 Turbo
18 held-out cases (16 text-to-image, 2 edits) that were not used for calibration, with matched seeds and resolutions (four at 2048², two non-square, the rest 1024²). Images are composited over white and resized to 512 px for LPIPS (AlexNet) and SSIM. Final latents are compared at full resolution.
| Runtime | Mean LPIPS ↓ | Mean SSIM ↑ | Mean PSNR | Latent cosine |
|---|---|---|---|---|
| Release default (FP8 step 1 + NVFP4) | 0.086 | 0.915 | 25.1 dB | 0.9827 |
| Pure NVFP4, compiled + Sage | 0.123 | 0.887 | 23.5 dB | 0.9660 |
| Pure NVFP4, eager | 0.134 | 0.878 | 23.1 dB | 0.9612 |
| FP8 release, for reference | 0.055 | 0.942 | 28.6 dB | 0.9891 |
All 18 pairs were reviewed visually, along with all 31 images in the demo video. The English and Chinese titles stay correct, transparent RGBA output and both edits work, and fur, feathers, glass and hands look natural. No artifacts were observed. Outputs are not identical to BF16: differences show up as composition, pose or identity changes rather than lost detail. With pure NVFP4, a "young man at a pottery wheel" became a different young man and a gouache harbour got a new layout; the FP8 first step restores both. The largest default-mode pair is a macro beetle (LPIPS 0.24, same subject with a shifted fern). SageAttention measured within run-to-run noise (paired runs 0.084–0.089 with Sage vs 0.082–0.087 without). See comparisons (BF16 | release default | pure NVFP4 | FP8 release) and reports/quality_metrics.json. All trials are in reports/trials. Keeping every attn.to_out.0 in FP8 was tested too, but only improved LPIPS 0.134 → 0.130, so it was not adopted. This is a finite fidelity check, not a human-preference study.
Precision and runtime
224 projections (q/k/v/out and SwiGLU gate/proj/out in all 32 blocks) run as NVFP4 W4A4, which works like this:
- Weights: a per-layer SmoothQuant factor is folded into the weights, a BF16 rank-128 SVD correction is removed from the smoothed weight, and the residual is packed as E2M1 with E4M3 block scales over 16 elements plus an FP32 global scale.
- Activations: smoothed and quantized per call with a dynamic tensor-wide range, recomputed every call because static ranges clip rare outliers, plus dynamic per-16 block scales.
- GEMM: native FlashInfer 0.7.1 SM120 CuTe-DSL block-scaled FP4 kernel, with the low-rank BF16 branch fused into the FP32 accumulator epilogue.
Conditioning, input/output projections, norms, the text encoder and the VAE stay BF16.
Shared low-rank GEMM: projections that read the same input (q/k/v, gate/proj) compute their rank-128 down-projections in one concatenated GEMM.
Compiled decode: cached-prefix blocks run under
torch.compile(dynamic shapes, emulated BF16 cast boundaries).Split prefill: step 1's prompt/reference prefix runs alone and fills the KV cache; the image tokens then take the compiled cached path. Prefix tokens never attend to image tokens, so the attention structure is unchanged. The prefix and the image tokens each get their own dynamic activation range.
VAE: single-frame decode skips Wan's unused video frame cache (bit-identical) and is compiled: 172 → 93 ms at 1024.
Attention:
--attention sageuses SageAttention2 (INT8 QK, FP8 PV, FP32 accumulation) in the cached passes;--attention bf16keeps native BF16 SDPA. The defaultautouses Sage when installed.Not used: step skipping, feature caching, reduced resolution, distillation or LoRA.
Kernel evidence for one 1024² default generation, from the profiler: 1,568 native SM120 block-scaled FP4 GEMM launches (224 × 7 NVFP4 steps); 448 FP8 GEMMs for step 1 (prefix + target passes); 256 SageAttention qk_int_sv_f8_attn_kernel launches. One cached NVFP4 step takes 99 ms of GPU time: FP4 GEMMs 52%, the SageAttention kernel 14%, everything else small fused kernels. See reports/kernel_evidence.json.
This is a custom Diffusers/FlashInfer format that requires the included loader. It is not a drop-in Transformers, ComfyUI, Nunchaku, vLLM or TensorRT checkpoint, and needs a GPU with SM120 block-scaled FP4 (RTX PRO 6000 / RTX 50-series Blackwell; only SM120 was tested). Do not cast the loaded transformer. Tested with Torch 2.14.0+cu130, Diffusers 0.41.0, Transformers 5.17.0, flashinfer-python 0.7.1 and driver 615.71.09 (environment).
Calibration
- Trajectories: 64 deterministic BF16 Turbo trajectories on the checkpoint's 8-step schedule: 56 generations and 8 edits, 15 at 2048², 47 at 1024², and two 16:9/9:16. They cover photographs, portraits, textures, illustration, six languages, typography, transparency and compositions.
- Activation rows: 4 token rows per layer at every step, 2,048 rows per layer: 512 fit and 512 disjoint diagnostic rows.
- Second moments: full-token input second moments come from all 3.62 M tokens of the same trajectories.
- Search per layer: five smoothing exponents × three residual-weight rounding methods: plain FlashInfer quantizer, activation-weighted 15-point block-scale search, and GPTQ error feedback (natural column order, so block scales stay contiguous; per-group scale search). Every candidate is scored with the actual native W4A4 kernel and dynamic range.
- Result: GPTQ won for all 224 layers. Mean diagnostic output NRMSE is 5.03% (max 10.7%, mid-block
attn.to_out.0). At the same smoothing, plain rounding averaged 6.40% and block-scale search 5.97% fit NRMSE. The September Image 2.1 (40-step) NVFP4 release was at 5.80%.
Everything is in calibration_search.json, quantization_summary.json and calibration_manifest.json. This is post-training calibration only.
Run
Linux, Python 3.12, an SM120 Blackwell GPU and a CUDA 13 driver.
hf download ProCreations/Image-2.1-Turbo-Calibrated-NVFP4 --local-dir Image-2.1-Turbo-Calibrated-NVFP4
cd Image-2.1-Turbo-Calibrated-NVFP4
python -m venv .venv && source .venv/bin/activate
pip install torch==2.14.0 torchvision==0.29.0 --index-url https://download.pytorch.org/whl/cu130
pip install -r requirements.txt
# optional, faster attention (see requirements.txt): SageAttention 2.2.0 built from source
python generate.py --prompt "A kingfisher above a forest stream, wildlife photography" --width 1024 --height 1024 --output image.png
The first run downloads the step-1 FP8 transformer from ProCreations/Image-2.1-Turbo-Calibrated-FP8 at its pinned revision. Use --fp8-quant /path/to/transformer for a local copy, or --fp8-first-steps 0 for pure NVFP4 with no FP8 download. The 8-step schedule comes from the checkpoint. Defaults are 2048 × 2048; use --width/--height (multiples of 16; upstream presets 2048², 2400×1792, 1792×2400, 2528×1696, 1696×2528, 2752×1536, 1536×2752). --warmup reports first-use cost separately. --prompts-json list.json serves many prompts from one loaded pipeline; --cache-prompts reuses exact encoder outputs for repeated text-only prompts (0.93 s per 1024² image on a cache hit). --eager disables acceleration. --base /path/to/Qwen-Image-2.1-Turbo uses a local upstream snapshot.
Library use: load_pipeline(base, "transformer") from nvfp4_runtime.py, then accelerate_pipeline(pipe, attention="sage") from acceleration.py, then optionally enable_fp8_first_steps(pipe, fp8_dir, 1, accelerate=True, attention="sage").
Transparent PNGs, editing and multiple references
python generate.py --prompt "A cute mint-green baby dragon mascot, full body" --transparent --width 1024 --height 1024 --output dragon.png
python generate.py --image input.png --prompt "Change the jacket to blue, keep everything else" --width 1024 --height 1024 --output edited.png
python generate.py --image examples/ref-dragon-white.png examples/ref-whale-white.png \
--prompt "Image 1 shows a mint-green baby dragon. Image 2 shows a blue cartoon whale. Create one new picture with the dragon standing next to the whale, both fully visible." \
--width 1024 --height 1024 --output together.png
--transparent adds a transparency instruction and keeps the native alpha channel; no background removal is applied (examples/nvfp4-dragon.png, 74% transparent pixels). References keep command-line order. Transparent RGBA references with a short "place X next to Y" prompt lost the second character in BF16 Turbo as well, so this is upstream behaviour. White-background references with an explicit description worked in BF16, FP8 and NVFP4 (examples/nvfp4-together-rgb.png).
Real-time video
demo/realtime-30s-nvfp4.mp4: 30 s, 1024², 30 fps, image only, release default runtime. One completed warmup image is shown at t = 0; after that, each change happens when an actual generation finishes, with all waits kept at 1× speed. The video shows 31 new images in 30 seconds (mean 0.963 s each, including the Python loop). Maximum capture lag is 2.2 ms, and every sampled decoded frame matches the logged image (min PSNR 34.6 dB). See verification, timestamps and the contact sheet. The prompts are disjoint from calibration and evaluation.
Reproduce
source/ holds the experiment scripts in run order: verify_base.py, calibrate.py, hessians.py, quantize_nvfp4.py (--shard i/n workers, then one assembly run; needs weight_quant.py, gptq.py and dynamic_scale.py), evaluate.py, hybrid.py (step-hybrid trial), assemble_mixed.py (FP8 to_out trial), metrics.py, speedlab.py, benchmark.py, video.py, verify_video.py. common.py holds the workstation paths. Calibration activations and second moments (about 30 GB) are regenerated locally and are not needed for inference. Unit checks: test_gptq.py (the GPTQ packer reproduces the reference FP4 quantizer bit-exactly when H = I, and the native kernel decodes it), test_vae.py and test_split.py.
License and modifications
The original model is under the Qwen Research License: non-commercial research and evaluation, subject to its terms. Keep Notice and the license when redistributing. The modified weights carry a modification notice in their safetensors metadata and in transformer/quantization_config.json. ProCreations, 2026-10-09.
- Downloads last month
- -