Instructions to use ixim/Image21-INT4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use ixim/Image21-INT4 with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("ixim/Image21-INT4", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
Download cards/modelscope.md from ixim/Image21-INT4: direct link, hf CLI and curl.
- Browser
- Download file 3.73 kB
-
https://huggingface.co/ixim/Image21-INT4/resolve/main/cards/modelscope.md
- Command line
-
hf download hf://ixim/Image21-INT4/cards/modelscope.md
-
curl -L -o modelscope.md https://huggingface.co/ixim/Image21-INT4/resolve/main/cards/modelscope.md
license: other
license_name: qwen-research
license_link: LICENSE
base_model: Qwen/Qwen-Image-2.1
base_model_relation: quantized
library_name: diffusers
pipeline_tag: text-to-image
language:
- zh
- en
tags:
- diffusers
- sdnq
- int4
- uint4
- image-generation
- image-editing
- apple-silicon
Image21-INT4
Built with Qwen。 这是 Qwen-Image-2.1 的独立 4-bit 衍生模型,由 iximbox 制作,不是 Qwen 官方发布。
权重格式是 SDNQ UINT4(无符号 4-bit 整数)加秩为 32 的 SVD 残差。仓库名中的 INT4 指这份 4-bit 整数权重,不是 bitsandbytes NF4,也不是 GPTQ,更不是全整数流水线。 量化矩阵乘法保持关闭,因此 CUDA 与 Apple Silicon 都用普通 PyTorch 算子做反量化。
能在哪里运行
- 8GB 显存的 CUDA: 作者已在一台配备 8GB 显存的 Windows 笔记本电脑上运行本 INT4 权重。 分组卸载使 Transformer 一次只将一个 block 放上 GPU,文本编码器一次只放一个线性层; VAE 编码和解码在 CPU 上进行,以降低显存占用。下方 1024×1024、40 步的耗时与显存数据 来自在 32GB RTX 5090 上将 PyTorch 分配器限制为 7.2 GiB 的独立测试,并非该笔记本的测量值。 笔记本显卡仍须受到所安装 PyTorch 版本的支持。
- 更大的 CUDA 显卡: 加载器按组件做 model CPU offload。下方的成对对比使用这条路径。
- Apple Silicon: 作者已通过 MPS 在一台 M4 芯片的 MacBook 上运行本 INT4 权重。
此处尚未提供 MPS 耗时、峰值内存或最低统一内存需求的测量数据。若某个 MPS 算子不接受
bfloat16,可设置
IMAGE21_DTYPE=float16。
量化
基座版本:b3179ad355be050328e483a9dfdd9e60cd62adfa。
Transformer 与 Qwen3-VL 文本编码器中符合条件的线性层转为 UINT4。敏感投影、归一化、
嵌入、视觉塔、输出头和 VAE 保持浮点。具体模块见两个组件的量化 JSON。没有校准数据集,
也没有微调。
安装
pip install torch torchvision
pip install -r requirements.txt
NVIDIA 机器请先安装与显卡匹配的 PyTorch 2.10 CUDA 轮子,再安装 requirements.txt。
本次转换使用 torch 2.10.0+cu128。
import torch
import sys
from modelscope import snapshot_download
model_dir = snapshot_download("iximbox/Image21-INT4")
sys.path.insert(0, model_dir)
from scripts.runtime import load_pipeline
pipe = load_pipeline(model_dir, local_files_only=True)
image = pipe(
prompt="雨夜霓虹灯,灯牌上写着 CREATE WITH LIGHT",
width=1024, height=1024, num_inference_steps=40,
true_cfg_scale=1.0, use_kv_cache=True,
generator=torch.Generator("cpu").manual_seed(42),
).images[0]
image.save("output.png")
load_pipeline 会选择 CUDA、MPS 或 CPU。8GB 级别的 CUDA 会自动使用分组卸载。
不要在加载时再次量化这些文件。
编辑时,编辑种子必须和生成原图的种子不同。已发布的编辑对使用原图种子 42,编辑种子 1000042 和 1000123。需要透明通道时保存 PNG。
如需可视化操作,可使用另行开源的 IMGEN Web 应用; 它支持 BF16、INT8 和 INT4 模型的生图与改图。安装与使用说明见其 GitHub 仓库。
非正式评测
下方对比基于 RTX 5090 上少量自选样例;耗时与 CUDA 显存数据不适用于上述笔记本或 MPS 运行记录。
{{DETAILS}}
成对图片
{{SAMPLES}}
Qwen Research License 将本衍生模型限制在非商业研究与评测。仓库内附有许可证、Notice 和 CHANGES.md。