Instructions to use HelloSun/Z-Image-Turbo-OpenVINO-INT4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use HelloSun/Z-Image-Turbo-OpenVINO-INT4 with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("HelloSun/Z-Image-Turbo-OpenVINO-INT4", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
Z-Image-Turbo — OpenVINO INT4
將 Tongyi-MAI/Z-Image-Turbo(6B S3-DiT,Turbo 蒸餾)轉換為 OpenVINO INT4 權重,
並以 純 CPU 實測 1024×1024 出圖。
optimum-intel 匯出 → NNCF weight-only INT4 量化 → FP16 20.32 GB → INT4 5.52 GB(↓ ~73%)。
步數說明:設定
num_inference_steps=9,scheduler 實際執行 9 個 step。 來源模型標示的「8 NFE」是指 DiT forward 次數的官方說法,與本 repo 量測到的 9 個 scheduler step 並不完全對應,實測數據以examples/benchmark.json為準。
目錄
快速資訊
| 項目 | 內容 |
|---|---|
| 基座模型 | Tongyi-MAI/Z-Image-Turbo(6B S3-DiT,Turbo 蒸餾) |
| Pipeline | ZImagePipeline(diffusers 格式) |
| Scheduler | FlowMatchEulerDiscreteScheduler(shift=3.0,由 scheduler_config.json 自動載入) |
| 量化格式 | transformer / text_encoder → INT4;VAE → INT8 |
| 模型大小 | FP16 20.32 GB → INT4 5.52 GB(↓ ~73%) |
| 推論裝置 | CPU(OpenVINO CPU plugin,無需 GPU) |
| 推薦參數 | num_inference_steps=9、guidance_scale=0.0、1024×1024 |
| 記憶體需求 | 實測推論期間 RSS 約 16.3 GB |
| 轉換工具 | optimum 2.3.0 / optimum-intel 2.2.0 / OpenVINO 2026.4.0 / NNCF 3.4.0 |
特色
- **體積 ↓ 73%**:FP16 20.32 GB → INT4 5.52 GB。
- 純 CPU 可用:
guidance_scale=0.0,每步只需跑一次 Transformer。 - diffusers 目錄結構:可直接
from_pretrained()。 - 完整記憶體量測:本 repo 記錄了逐張的推論前後 RSS,是同批中記憶體資料最完整的。
- 附 15 張實測圖:5 組 prompt × (1024px + 512px 對照組),另加 10 張風格展示。
- ⚠️ 記憶體需求偏高:INT4 模型僅 5.52 GB,但推論時 RSS 達 16.3 GB。
安裝
pip install diffusers==0.37.1 transformers==4.57.6 tokenizers==0.22.0 huggingface-hub==0.35.1 optimum==2.3.0 optimum-intel==2.2.0 openvino==2026.4.0 nncf==3.4.0 torch pillow psutil
僅推論時不需要
nncf;nncf為重新量化所需。
快速開始
import torch
from optimum.intel import OVDiffusionPipeline
pipe = OVDiffusionPipeline.from_pretrained(
"HelloSun/Z-Image-Turbo-OpenVINO-INT4", compile=True)
image = pipe(
prompt="Young Chinese woman in red Hanfu, intricate embroidery, "
"photorealistic, ultra detailed, 8k",
height=1024, width=1024,
num_inference_steps=9, guidance_scale=0.0,
generator=torch.Generator().manual_seed(42),
).images[0]
image.save("out.png")
完整可執行範例:inference_int4.py 批次生成 + benchmark:generate5.py
推論參數建議
| 參數 | 建議值 | 說明 |
|---|---|---|
num_inference_steps |
9 |
Turbo 蒸餾設定。請勿增加。 |
guidance_scale |
0.0 |
Turbo 模型不啟用 CFG。 |
shift |
3.0 |
不需手動傳入,scheduler/scheduler_config.json 已內建。 |
height / width |
1024 |
模型原生解析度。 |
compile=True |
開啟 | 編譯模型以取得較佳效能。 |
範例結果
全部為 9 steps / guidance 0.0 / 1024×1024 / CPU,seed 42–46 固定,可完全重現。
每組另附 512px 對照組(
examples/*_512.png)。
檔名說明:主測組檔名為
*_8step_*、風格展示為*_9step_*,但兩者實際都以num_inference_steps=9生成,檔名中的步數前綴並不可靠。
01_hanfu_portrait — seed 42 — 116.2 s
Young Chinese woman in red Hanfu, intricate embroidery, impeccable makeup, red floral forehead pattern, elaborate high bun, golden phoenix headdress, soft-lit outdoor night background, silhouetted tiered pagoda, blurred colorful distant lights, photorealistic, ultra detailed, 8k
02_astronaut_jungle — seed 43 — 110.8 s
Astronaut in a jungle, cold color palette, muted colors, detailed, 8k, photorealistic, cinematic lighting
03_taipei_cyberpunk — seed 44 — 110.3 s
Cyberpunk street in Taipei at night, heavy rain, neon signs with text 'TAIPEI' and Chinese characters '台北', reflections on wet asphalt, crowded night market, cinematic, ultra detailed
04_shiba_astronaut — seed 45 — 109.3 s
Cute Shiba Inu wearing a tiny astronaut helmet, sitting in a field of sunflowers under a starry sky, dreamy illustration, vibrant colors, high quality
05_ink_landscape — seed 46 — 109.3 s
Traditional Chinese ink wash landscape, misty mountains, a small pagoda on a cliff, cranes flying, minimalist, elegant, high aesthetic quality
風格展示
以下 10 張為 10 類風格 × 同一組 5 組基礎 prompt(seed 42–51), 使用與上方基準測試完全相同的推論設定。
注意:這批圖片沒有對應的
examples/benchmark.json紀錄,也沒有腳本可重現, 完整 prompt 亦未收錄於 repo(早期說明檔中的版本已被截斷,無法還原)。
traditional_painting · 傳統繪畫
油畫、版畫、水彩等傳統媒材
anime_manga · 動漫繪師
吉卜力 × 新海誠動畫風
digital_3d · 數位 3D
UE5 寫實渲染、CG 質感
photography_cinema · 攝影電影
電影感人像攝影
cultural_regional · 文化地域
敦煌壁畫、水墨等東方美學
material_craft · 材質工藝
絲線刺繡、陶瓷、木雕等
design_commercial · 平面商業
商業插畫、扁平幾何設計
abstract_generative · 抽象生成
分形、生成藝術
era_subculture · 復古次文化
80s 霓虹、賽博龐克
render_technique · 厚塗渲染
油畫厚塗、概念設定
效能實測摘要
完整逐 step 數據見 REPORT.md 與 examples/benchmark.json。
測試環境
| 項目 | 內容 |
|---|---|
| CPU | Intel(R) Xeon(R) Platinum 8559C |
| 拓撲 | 2 sockets × 48 cores × 2 threads/core = 192 vCPU(96 實體核心) |
| RAM | 2.0 TiB |
| 虛擬化 | KVM(完整虛擬化) |
| OpenVINO | CPU only,2026.4.0(build 2026.4.0-22959-99c81491cc3-releases/2026/4) |
| 設定 | num_inference_steps=9、guidance_scale=0.0、1024×1024 + 512×512 對照 |
總結
| 指標 | 主測組 | 對照組 |
|---|---|---|
| 解析度 | 1024×1024 | 512×512 |
| 平均總耗時 | 111.18 s / 張 | 33.46 s / 張 |
| 平均單步耗時 | 12.15 s | 3.67 s |
| 最快 / 最慢 | 109.26 s / 116.19 s | 31.05 s / 38.72 s |
| 對照組倍數 | — | 3.32× |
- 對照組為同一 prompt/seed 的 512×512 重新推理(非縮圖)。
- 文字編碼與 VAE decode 的時間已包含在總耗時內,每張約多出 1.7–1.9 s。
逐張結果(1024×1024 / 9 steps)
| # | Prompt | Seed | 總耗時 (s) | 平均單步 (s) | 512px 總耗時 (s) | 512px 單步 (s) | 推論後 RSS (GB) |
|---|---|---|---|---|---|---|---|
| 01_hanfu_portrait | 42 | 116.19 | 12.696 | 38.72 | 4.236 | 16.31 | |
| 02_astronaut_jungle | 43 | 110.82 | 12.112 | 33.04 | 3.617 | 16.31 | |
| 03_taipei_cyberpunk | 44 | 110.31 | 12.066 | 31.81 | 3.489 | 16.32 | |
| 04_shiba_astronaut | 45 | 109.30 | 11.943 | 31.05 | 3.396 | 16.32 | |
| 05_ink_landscape | 46 | 109.26 | 11.941 | 32.69 | 3.587 | 16.32 | |
| 平均 | 111.18 | 12.152 | 33.46 | 3.665 | 16.32 |
模型大小
以下為 repo 內 openvino_model.bin 的實際位元組數(Git LFS 記錄值)。
| 元件 | 位元組 | 大小 | 精度 |
|---|---|---|---|
transformer |
3,199,343,400 | 3.20 GB | INT4 |
text_encoder |
2,236,949,085 | 2.24 GB | INT4 |
vae_decoder |
49,612,856 | 0.05 GB | INT8 |
vae_encoder |
34,325,660 | 0.03 GB | INT8 |
| 合計 | 5,520,231,001 | 5.52 GB |
FP16 匯出模型約 20.32 GB(transformer 12.31 GB + text_encoder 7.84 GB + vae_decoder 99 MB + vae_encoder 69 MB)——取自 examples/benchmark.json 的 sysinfo.model_sizes_GB。
⚠️ 早期說明中的 19 GB / 5.2 GB 與實測不符:實際為 20.32 GB / 5.52 GB,本表為準。
INT4 後的 text_encoder 仍有 2.24 GB,因為它是 Qwen3-4B 級別的語言模型,即使 4-bit 量化體積依然可觀。
檔案結構
.
├── README.md # 本文件
├── REPORT.md # 完整轉換 + 實測報告
├── model_index.json # diffusers pipeline 索引(ZImagePipeline)
├── openvino_config.json # OpenVINO 量化設定
├── inference_int4.py # 單張推論範例
├── generate5.py # 5 組 prompt 批次生成 + benchmark
├── quantize_int4.py # FP16 OV → INT4 OV 量化腳本
├── transformer/ # INT4 ZImageTransformer2DModel(6B)
├── text_encoder/ # INT4 Qwen3 語言模型
├── tokenizer/ # Qwen2Tokenizer
├── vae_encoder/ # INT8 VAE encoder
├── vae_decoder/ # INT8 VAE decoder
├── scheduler/ # FlowMatchEulerDiscreteScheduler (shift=3.0)
└── examples/ # 20 張實測圖 + benchmark.json + prompts.txt
├── *_8step_1024.png # 主測組(5 張)
├── *_8step_512.png # 對照組(5 張)
├── *_9step_1024.png # 風格展示(10 張)
├── benchmark.json # 逐 step 耗時 + 記憶體 + 完整系統資訊
└── prompts.txt # 5 組 prompt 與 seed
從零復現
# 1. 匯出 FP16 OpenVINO 模型
optimum-cli export openvino \
-m Tongyi-MAI/Z-Image-Turbo \
--task text-to-image \
--library diffusers \
--weight-format fp16 \
./z-image-turbo-ov-fp16
# 2. NNCF weight-only INT4 量化
python quantize_int4.py --fp16-dir ./z-image-turbo-ov-fp16 \
--int4-dir ./z-image-turbo-ov-int4
# 3. 單張推論
python inference_int4.py
# 4. 批次生成 5 組 + 512px 對照 + benchmark
python generate5.py --outdir examples
量化設定:
from optimum.intel.openvino.configuration import (
OVConfig, OVWeightQuantizationConfig, OVPipelineQuantizationConfig,
)
int4_config = OVWeightQuantizationConfig(
bits=4, sym=False, group_size=128,
group_size_fallback="adjust", ratio=1.0,
)
pipeline_config = OVPipelineQuantizationConfig(
quantization_configs={
"transformer": int4_config,
"text_encoder": int4_config,
},
default_config=OVWeightQuantizationConfig(bits=8, sym=True),
)
已知限制
- 記憶體需求高:INT4 模型僅 5.52 GB,但實測推論期間 RSS 達 16.3 GB。這是 OpenVINO CPU 在 INT4 執行時的即時解量化與中間張量所致,建議至少預留 20 GB 可用記憶體。
- CPU 速度有限:1024×1024 平均 111 s / 張,遠慢於同批的 SDXL 蒸餾模型。
- 僅為 weight-only 量化:首次載入 + 編譯約 7.7 s。
- CPU 專用:GPU 使用請改用原模型或自行轉換 OpenVINO GPU。
- 9 steps 蒸餾設定:
num_inference_steps > 9不會變好。 - **
model_index.json宣告單一vae,但實際目錄為vae_decoder/+vae_encoder/**:這是 optimum-intel 匯出的常見情況,OVDiffusionPipeline可正常載入。 - 風格展示圖無實測紀錄:10 張風格圖未納入 benchmark,完整 prompt 亦未收錄。
授權與出處
授權:
apache-2.0(沿用來源模型授權)轉換:僅做格式轉換與權重量化,模型權重來自來源模型
使用本模型時請一併遵守來源模型的授權條款。
Made with OpenVINO + optimum-intel + NNCF
- Downloads last month
- 85
Model tree for HelloSun/Z-Image-Turbo-OpenVINO-INT4
Base model
Tongyi-MAI/Z-Image-Turbo













