Instructions to use QuantFunc/Qwen-Image-2.1-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use QuantFunc/Qwen-Image-2.1-4bit with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("QuantFunc/Qwen-Image-2.1-4bit", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
๐ Website | ๐ GitHub | ๐ค Hugging Face | ๐ค ModelScope | ๐ฎ Discord
Qwen-Image-2.1-QuantFunc-4bit
4x compression, quality held.
QuantFunc INT4 compresses Qwen-Image-2.1 to about a quarter of its 16-bit size. In our visual comparisons, composition, text rendering, detail and style stay close to the 16-bit baseline.
Showcase
All images below were generated by Qwen-Image-2.1-QuantFunc-4bit.
Swap one loader, keep your workflow
- Install or update ComfyUI-QuantFunc.
- Download the r128 or r32 weights.
- Swap your model loader for the QuantFunc loader and pick the matching weight file.
Every other node, connection and generation parameter stays as-is.
RTX 20-series through GB300, one build covers it all
Runs on every NVIDIA SM75+ GPU: RTX 20/30/40/50-series, A100, H100, H200, B100, B200, GB300.
4-bit weights significantly cut the weight-bandwidth cost of loading and inference, making Qwen-Image-2.1 much easier to run on consumer GPUs. Actual VRAM needs depend on resolution, batch size and the rest of your workflow.
Choose a model
| Variant | File | Size | Use case |
|---|---|---|---|
| r128 | qwen_image_21_base_quantfunc_int4_r128.safetensors |
~4.3 GB | Quality-first, recommended |
| r32 | qwen_image_21_base_quantfunc_int4_r32.safetensors |
~4.0 GB | Size/VRAM-first |
Loading
Weights use QuantFunc's own safetensors format โ load with ComfyUI-QuantFunc or the QuantFunc inference engine, not as a drop-in Diffusers checkpoint. (This repo declares library_name: diffusers purely so Hugging Face counts real downloads of the flat .safetensors files above โ see HF's download-stats docs โ not because the files load via the diffusers package.)
If the model isn't recognized or fails to load, update ComfyUI-QuantFunc to the latest version and restart ComfyUI.
Source & license
This repository distributes derived quantized weights produced from Qwen/Qwen-Image-2.1, released by the Qwen team.
Qwen-Image-2.1 and its derivative weights follow the Qwen Research License (license text). QuantFunc's quantization tooling and implementation follow their own respective licenses. Please read and comply with the original model's license terms before use.
Community
- ๐ QuantFunc website
- ๐ ComfyUI-QuantFunc on GitHub
- ๐ฎ Discord
- Downloads last month
- 242
Model tree for QuantFunc/Qwen-Image-2.1-4bit
Base model
Qwen/Qwen-Image-2.1



