Qwen-Image-2.1 — OpenVINO INT4
將 Qwen/Qwen-Image-2.1 轉換為 OpenVINO INT4 權重,並以 純 CPU 實測出圖。
optimum-intel 匯出 → NNCF weight-only INT4 量化(耗時 115 s)→ FP16 43.6 GB → INT4 13.12 GB(↓ ~70%,約 3.3×)。
⚠️ 環境需求與其他 repo 不同
QwenImage21Pipeline在 diffusers 0.40.0 之後才加入,因此diffusers==0.37.1無法使用。 本 repo 需要 diffusers 主分支版本,且需對transformers與optimum-intel套用兩處小幅度修改 (詳見下方安裝說明與REPORT.md)。
目錄
快速資訊
| 項目 | 內容 |
|---|---|
| 基座模型 | Qwen/Qwen-Image-2.1 |
| Pipeline | QwenImage21Pipeline |
| Scheduler | FlowMatchEulerDiscreteScheduler(shift=1.0、shift_terminal=0.02、use_dynamic_shifting=true) |
| 量化格式 | transformer、text_encoder、text_encoder_i2i → INT4;其餘 → INT8 |
| 模型大小 | FP16 43.6 GB → INT4 13.12 GB(↓ ~70%,約 3.3×) |
| 推論裝置 | CPU(OpenVINO CPU plugin,無需 GPU) |
| 推薦參數 | num_inference_steps=40、true_cfg_scale=1.0、1024×1024 |
| 轉換工具 | optimum-intel 2.2.x / optimum 2.3.0 / OpenVINO 2026.4.0 / NNCF 3.4.0 |
特色
- 三個元件 INT4:Transformer 與兩個 Qwen3-VL text encoder 全部 INT4,
text_encoder_i2i是編輯功能用的第二個 text encoder。 - ⚠️ 需要 diffusers 主分支:
QwenImage21Pipeline需要diffusers==0.41.0.dev0與transformers==5.10.4,與同批其他 repo 的版本需求不同。 - **體積 ↓ 70%**:FP16 43.6 GB → INT4 13.12 GB(約 3.3×)。
- 支援編輯能力:保留了 image-to-image 的
text_encoder_i2i與vision_encoder。 - ⚠️ 速度最慢:1024×1024 / 40 steps 平均 870 s / 張(約 14.5 分鐘)。
- 記憶體需求最高:行程結束時 RSS 達 40.2 GB。
安裝
pip install diffusers==0.41.0.dev0 transformers==5.10.4 tokenizers==0.22.2 huggingface-hub==1.33.0 optimum==2.3.0 optimum-intel==2.3.0.dev0+32a317a openvino==2026.4.0 nncf==3.4.0 torch==2.14.0 pillow==12.3.0 psutil==7.2.2
nncf 僅重新量化時需要。
快速開始
import torch
from optimum.intel import OVDiffusionPipeline
pipe = OVDiffusionPipeline.from_pretrained("HelloSun/Qwen-Image-2.1-OpenVINO-INT4", compile=True)
image = pipe(
prompt="Astronaut in a jungle, cold color palette, muted colors, "
"detailed, 8k, photorealistic, cinematic lighting",
height=1024, width=1024,
num_inference_steps=40,
true_cfg_scale=1.0,
generator=torch.Generator().manual_seed(43),
).images[0]
image.save("out.png")
完整可執行範例:inference_int4.py 批次生成 + benchmark:generate5.py
推論參數建議
| 參數 | 建議值 | 說明 |
|---|---|---|
num_inference_steps |
40 |
蒸餾 / 推薦步數。請勿隨意增加。 |
true_cfg_scale |
1.0 |
CFG 設定;Turbo / 蒸餾模型通常為 0.0 或 1.0(即不啟用)。 |
shift |
取自 scheduler/scheduler_config.json |
不需手動傳入,載入時自動套用。 |
height / width |
1024 |
實測解析度。 |
compile=True |
開啟 | 編譯模型以取得較佳效能。 |
範例結果
全部為 40 steps / true_cfg_scale 1.0 / 1024×1024 / CPU,seed 42–46 固定,可完全重現。
另附 512px 對照圖(
outputs/*_512.png)。
⚠️ 512px 為縮圖:這批 512px 圖是 1024px 輸出的 LANCZOS 縮圖,不是重新以 512px 推理的結果,因此沒有對應的獨立耗時資料。
01_hanfu — seed 42 — 980.8 s
Young Chinese woman in red Hanfu, intricate embroidery, impeccable makeup, red floral forehead pattern, elaborate high bun, golden phoenix headdress, soft-lit outdoor night background, silhouetted tiered pagoda, blurred colorful distant lights, photorealistic, ultra detailed, 8k
02_astronaut — seed 43 — 865.7 s
Astronaut in a jungle, cold color palette, muted colors, detailed, 8k, photorealistic, cinematic lighting
03_taipei — seed 44 — 871.5 s
Cyberpunk street in Taipei at night, heavy rain, neon signs with text 'TAIPEI' and Chinese characters '台北', reflections on wet asphalt, crowded night market, cinematic, ultra detailed
04_shiba — seed 45 — 832.1 s
Cute Shiba Inu wearing a tiny astronaut helmet, sitting in a field of sunflowers under a starry sky, dreamy illustration, vibrant colors, high quality
05_ink — seed 46 — 798.9 s
Traditional Chinese ink wash landscape, misty mountains, a small pagoda on a cliff, cranes flying, minimalist, elegant, high aesthetic quality
效能實測摘要
完整逐 step 數據見 REPORT.md 與 outputs/benchmark.json。
測試環境
| 項目 | 內容 |
|---|---|
| CPU | Intel(R) Xeon(R) Platinum 8559C |
| 拓撲 | 2 sockets × 48 cores × 2 threads/core = 192 vCPU(96 實體核心) |
| RAM | 2.0 TiB |
| 虛擬化 | KVM(完整虛擬化) |
| OpenVINO | CPU only,2026.4.0(build 2026.4.0-22959-99c81491cc3-releases/2026/4) |
| 設定 | num_inference_steps=40、true_cfg_scale=1.0、1024×1024 |
總結
| 指標 | 數值 |
|---|---|
| 解析度 | 1024×1024 |
| 平均總耗時 | 869.80 s / 張 |
| 平均單步耗時 | 21.57 s |
| 最快 / 最慢 | 798.94 s / 980.80 s |
| 總計(5 張) | 4349.00 s |
| 記憶體高水位 | 40,211 MB |
- 512px 沒有獨立耗時資料:
outputs/*_512.png是 1024px 輸出的縮圖,非重新推理。 - 文字編碼與 VAE decode 的時間已包含在總耗時內。
逐張結果(1024×1024)
| # | Prompt | Seed | 總耗時 (s) | 平均單步 (s) | true_cfg_scale |
|---|---|---|---|---|---|
| 01_hanfu | 42 | 980.80 | 24.257 | 1.0 | |
| 02_astronaut | 43 | 865.67 | 21.497 | 1.0 | |
| 03_taipei | 44 | 871.46 | 21.645 | 1.0 | |
| 04_shiba | 45 | 832.12 | 20.592 | 1.0 | |
| 05_ink | 46 | 798.94 | 19.846 | 1.0 | |
| 平均 | 869.80 | 21.568 |
模型大小
以下為 repo 內 openvino_model.bin 的實際位元組數(Git LFS 記錄值)。
| 元件 | 位元組 | 大小 | 精度 |
|---|---|---|---|
text_encoder |
4,256,132,893 | 4.26 GB | INT4 |
text_encoder_i2i |
4,256,132,733 | 4.26 GB | INT4 |
transformer |
3,696,679,634 | 3.70 GB | INT4 |
vision_encoder |
577,780,620 | 0.58 GB | — |
vae_decoder |
253,235,856 | 0.25 GB | — |
vae_encoder |
78,002,698 | 0.08 GB | — |
| 合計 | 13,117,964,434 | 13.12 GB |
FP16 匯出模型約 43.6 GB(transformer 14.23 GB + text_encoder 15.14 GB + text_encoder_i2i 15.14 GB + vision_encoder 1.1 GB + vae_decoder 484 MB + vae_encoder 150 MB)——取自原始轉換紀錄。
INT8 部分(vision encoder + VAE)為 0.88 GB,佔 INT4 總量的 7%。
text_encoder 與 text_encoder_i2i 各自 INT4 後仍有 4.26 GB,是整體壓縮率(~70%)的主要限制來源。
檔案結構
.
├── README.md # 本文件
├── REPORT.md # 完整轉換 + 實測報告
├── model_index.json # diffusers pipeline 索引(QwenImage21Pipeline)
├── openvino_config.json # OpenVINO 量化設定
├── inference_int4.py # 單張推論範例
├── generate5.py # 5 組 prompt 批次生成 + benchmark
├── quantize_int4.py # FP16 OV → INT4 OV 量化腳本
├── transformer/ # INT4 QwenImage21Transformer2DModel(32 層)
├── text_encoder/ # INT4 Qwen3VLForConditionalGeneration
├── text_encoder_i2i/ # INT4 第二 text encoder(編輯用)
├── vision_encoder/ # INT8 Qwen3VLVisionModel
├── processor/ # Qwen3VLProcessor + tokenizer
├── vae_encoder/ # INT8 VAE encoder
├── vae_decoder/ # INT8 VAE decoder
├── scheduler/ # FlowMatchEulerDiscreteScheduler 設定
├── examples/ # 5 張展示圖(與 outputs 1024 相同)
└── outputs/ # 10 張實測圖 + 3 個資料檔
├── *_1024.png # 主測組(5 張)
├── *_512.png # 對照組(5 張,為 1024 縮圖)
├── benchmark.json # 逐 step 耗時 + 記憶體 + 系統資訊
├── benchmark_quantization.json # 量化設定與耗時
└── prompts.txt # 5 組 prompt 與 seed
從零復現
# 1. 匯出 FP16 OpenVINO 模型
optimum-cli export openvino \
-m Qwen/Qwen-Image-2.1 \
--task text-to-image \
--library diffusers \
--weight-format fp16 \
./Qwen-Image-2.1-ov-fp16
# 2. NNCF weight-only INT4 量化
python quantize_int4.py --fp16-dir ./Qwen-Image-2.1-ov-fp16 \
--int4-dir ./Qwen-Image-2.1-ov-int4
# 3. 單張推論
python inference_int4.py
# 4. 批次生成 5 組 + benchmark
python generate5.py --outdir outputs
量化設定:
from optimum.intel.openvino.configuration import (
OVConfig, OVWeightQuantizationConfig, OVPipelineQuantizationConfig,
)
int4_config = OVWeightQuantizationConfig(
bits=4, sym=False, group_size=128,
group_size_fallback="adjust", ratio=1.0,
)
pipeline_config = OVPipelineQuantizationConfig(
quantization_configs={
"transformer": int4_config,
"text_encoder": int4_config,
"text_encoder_i2i": int4_config,
},
default_config=OVWeightQuantizationConfig(bits=8),
)
已知限制
CPU 速度最慢:1024×1024 / 40 steps 平均 870 s / 張(約 14.5 分鐘),是這批 repo 中最慢的,純 CPU 實務上不太實用。
記憶體需求最高:INT4 模型 13.12 GB,實測行程結束時 RSS 達 40.2 GB,建議至少預留 48 GB 可用記憶體。
需要 diffusers 主分支:
diffusers==0.41.0.dev0(含QwenImage21Pipeline,v0.40.0 之後才加入)。optimum-intel需使用2.3.0.dev0+32a317a。需要對上游套件套用兩處修改:
transformers的相依套件上限(hub cap<1.0→<2.0)
optimum-intel的modeling_visual_language.py(transformers ≥ 5 的VisionRotaryEmbedding別名)
512px 對照為縮圖:
outputs/*_512.png是 1024px 輸出的 LANCZOS 縮圖,非重新推理。**
model_index.json未列出text_encoder_i2i與vision_encoder**:這兩個目錄存在於 repo 中但未在索引中宣告,屬 optimum-intel 匯出的已知落差。10 張風格展示圖無實測紀錄:只有 5 張主測圖有 benchmark 資料。
授權與出處
來源模型:
Qwen/Qwen-Image-2.1授權:
qwen-research(license: other)完整條款:LICENSE
轉換:僅做格式轉換與權重量化,模型權重來自來源模型
使用本模型時請遵守來源模型的授權條款。
Made with OpenVINO + optimum-intel + NNCF
Model tree for HelloSun/Qwen-Image-2.1-OpenVINO-INT4
Base model
Qwen/Qwen-Image-2.1



