How to use from the
Use from the
Diffusers library
pip install -U diffusers transformers accelerate
import torch
from diffusers import DiffusionPipeline

# switch to "mps" for apple devices
pipe = DiffusionPipeline.from_pretrained("HelloSun/Z-Image-Turbo-OpenVINO-INT4", dtype=torch.bfloat16, device_map="cuda")

prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k"
image = pipe(prompt).images[0]

Z-Image-Turbo — OpenVINO INT4

將 Tongyi-MAI/Z-Image-Turbo(6B S3-DiT,Turbo 蒸餾)轉換為 OpenVINO INT4 權重, 並以 純 CPU 實測 1024×1024 出圖。

optimum-intel 匯出 → NNCF weight-only INT4 量化 → FP16 20.32 GB → INT4 5.52 GB(↓ ~73%)。

步數說明:設定 num_inference_steps=9,scheduler 實際執行 9 個 step。 來源模型標示的「8 NFE」是指 DiT forward 次數的官方說法,與本 repo 量測到的 9 個 scheduler step 並不完全對應,實測數據以 examples/benchmark.json 為準。


目錄


快速資訊

項目 內容
基座模型 Tongyi-MAI/Z-Image-Turbo(6B S3-DiT,Turbo 蒸餾)
Pipeline ZImagePipeline(diffusers 格式)
Scheduler FlowMatchEulerDiscreteScheduler(shift=3.0,由 scheduler_config.json 自動載入)
量化格式 transformer / text_encoder → INT4;VAE → INT8
模型大小 FP16 20.32 GB → INT4 5.52 GB(↓ ~73%)
推論裝置 CPU(OpenVINO CPU plugin,無需 GPU)
推薦參數 num_inference_steps=9、guidance_scale=0.0、1024×1024
記憶體需求 實測推論期間 RSS 約 16.3 GB
轉換工具 optimum 2.3.0 / optimum-intel 2.2.0 / OpenVINO 2026.4.0 / NNCF 3.4.0

特色

  • **體積 ↓ 73%**:FP16 20.32 GB → INT4 5.52 GB。
  • 純 CPU 可用:guidance_scale=0.0,每步只需跑一次 Transformer。
  • diffusers 目錄結構:可直接 from_pretrained()。
  • 完整記憶體量測:本 repo 記錄了逐張的推論前後 RSS,是同批中記憶體資料最完整的。
  • 附 15 張實測圖:5 組 prompt × (1024px + 512px 對照組),另加 10 張風格展示。
  • ⚠️ 記憶體需求偏高:INT4 模型僅 5.52 GB,但推論時 RSS 達 16.3 GB。

安裝

pip install diffusers==0.37.1 transformers==4.57.6 tokenizers==0.22.0 huggingface-hub==0.35.1 optimum==2.3.0 optimum-intel==2.2.0 openvino==2026.4.0 nncf==3.4.0 torch pillow psutil

僅推論時不需要 nncf;nncf 為重新量化所需。

快速開始

import torch
from optimum.intel import OVDiffusionPipeline

pipe = OVDiffusionPipeline.from_pretrained(
    "HelloSun/Z-Image-Turbo-OpenVINO-INT4", compile=True)

image = pipe(
    prompt="Young Chinese woman in red Hanfu, intricate embroidery, "
             "photorealistic, ultra detailed, 8k",
    height=1024, width=1024,
    num_inference_steps=9, guidance_scale=0.0,
    generator=torch.Generator().manual_seed(42),
).images[0]

image.save("out.png")

完整可執行範例:inference_int4.py 批次生成 + benchmark:generate5.py

推論參數建議

參數 建議值 說明
num_inference_steps 9 Turbo 蒸餾設定。請勿增加。
guidance_scale 0.0 Turbo 模型不啟用 CFG。
shift 3.0 不需手動傳入,scheduler/scheduler_config.json 已內建。
height / width 1024 模型原生解析度。
compile=True 開啟 編譯模型以取得較佳效能。

範例結果

全部為 9 steps / guidance 0.0 / 1024×1024 / CPU,seed 42–46 固定,可完全重現。

每組另附 512px 對照組(examples/*_512.png)。

檔名說明:主測組檔名為 *_8step_*、風格展示為 *_9step_*,但兩者實際都以 num_inference_steps=9 生成,檔名中的步數前綴並不可靠。

01_hanfu_portrait — seed 42 — 116.2 s

Young Chinese woman in red Hanfu, intricate embroidery, impeccable makeup, red floral forehead pattern, elaborate high bun, golden phoenix headdress, soft-lit outdoor night background, silhouetted tiered pagoda, blurred colorful distant lights, photorealistic, ultra detailed, 8k

01_hanfu_portrait

02_astronaut_jungle — seed 43 — 110.8 s

Astronaut in a jungle, cold color palette, muted colors, detailed, 8k, photorealistic, cinematic lighting

02_astronaut_jungle

03_taipei_cyberpunk — seed 44 — 110.3 s

Cyberpunk street in Taipei at night, heavy rain, neon signs with text 'TAIPEI' and Chinese characters '台北', reflections on wet asphalt, crowded night market, cinematic, ultra detailed

03_taipei_cyberpunk

04_shiba_astronaut — seed 45 — 109.3 s

Cute Shiba Inu wearing a tiny astronaut helmet, sitting in a field of sunflowers under a starry sky, dreamy illustration, vibrant colors, high quality

04_shiba_astronaut

05_ink_landscape — seed 46 — 109.3 s

Traditional Chinese ink wash landscape, misty mountains, a small pagoda on a cliff, cranes flying, minimalist, elegant, high aesthetic quality

05_ink_landscape

風格展示

以下 10 張為 10 類風格 × 同一組 5 組基礎 prompt(seed 42–51), 使用與上方基準測試完全相同的推論設定。

注意:這批圖片沒有對應的 examples/benchmark.json 紀錄,也沒有腳本可重現, 完整 prompt 亦未收錄於 repo(早期說明檔中的版本已被截斷,無法還原)。

traditional_painting · 傳統繪畫

油畫、版畫、水彩等傳統媒材

01_traditional_painting

anime_manga · 動漫繪師

吉卜力 × 新海誠動畫風

02_anime_manga

digital_3d · 數位 3D

UE5 寫實渲染、CG 質感

03_digital_3d

photography_cinema · 攝影電影

電影感人像攝影

04_photography_cinema

cultural_regional · 文化地域

敦煌壁畫、水墨等東方美學

05_cultural_regional

material_craft · 材質工藝

絲線刺繡、陶瓷、木雕等

06_material_craft

design_commercial · 平面商業

商業插畫、扁平幾何設計

07_design_commercial

abstract_generative · 抽象生成

分形、生成藝術

08_abstract_generative

era_subculture · 復古次文化

80s 霓虹、賽博龐克

09_era_subculture

render_technique · 厚塗渲染

油畫厚塗、概念設定

10_render_technique

效能實測摘要

完整逐 step 數據見 REPORT.md 與 examples/benchmark.json。

測試環境

項目 內容
CPU Intel(R) Xeon(R) Platinum 8559C
拓撲 2 sockets × 48 cores × 2 threads/core = 192 vCPU(96 實體核心)
RAM 2.0 TiB
虛擬化 KVM(完整虛擬化)
OpenVINO CPU only,2026.4.0(build 2026.4.0-22959-99c81491cc3-releases/2026/4)
設定 num_inference_steps=9、guidance_scale=0.0、1024×1024 + 512×512 對照

總結

指標 主測組 對照組
解析度 1024×1024 512×512
平均總耗時 111.18 s / 張 33.46 s / 張
平均單步耗時 12.15 s 3.67 s
最快 / 最慢 109.26 s / 116.19 s 31.05 s / 38.72 s
對照組倍數 — 3.32×
  • 對照組為同一 prompt/seed 的 512×512 重新推理(非縮圖)。
  • 文字編碼與 VAE decode 的時間已包含在總耗時內,每張約多出 1.7–1.9 s。

逐張結果(1024×1024 / 9 steps)

# Prompt Seed 總耗時 (s) 平均單步 (s) 512px 總耗時 (s) 512px 單步 (s) 推論後 RSS (GB)
01_hanfu_portrait 42 116.19 12.696 38.72 4.236 16.31
02_astronaut_jungle 43 110.82 12.112 33.04 3.617 16.31
03_taipei_cyberpunk 44 110.31 12.066 31.81 3.489 16.32
04_shiba_astronaut 45 109.30 11.943 31.05 3.396 16.32
05_ink_landscape 46 109.26 11.941 32.69 3.587 16.32
平均 111.18 12.152 33.46 3.665 16.32

模型大小

以下為 repo 內 openvino_model.bin 的實際位元組數(Git LFS 記錄值)。

元件 位元組 大小 精度
transformer 3,199,343,400 3.20 GB INT4
text_encoder 2,236,949,085 2.24 GB INT4
vae_decoder 49,612,856 0.05 GB INT8
vae_encoder 34,325,660 0.03 GB INT8
合計 5,520,231,001 5.52 GB

FP16 匯出模型約 20.32 GB(transformer 12.31 GB + text_encoder 7.84 GB + vae_decoder 99 MB + vae_encoder 69 MB)——取自 examples/benchmark.json 的 sysinfo.model_sizes_GB。

⚠️ 早期說明中的 19 GB / 5.2 GB 與實測不符:實際為 20.32 GB / 5.52 GB,本表為準。

INT4 後的 text_encoder 仍有 2.24 GB,因為它是 Qwen3-4B 級別的語言模型,即使 4-bit 量化體積依然可觀。

檔案結構

.
├── README.md                  # 本文件
├── REPORT.md                  # 完整轉換 + 實測報告
├── model_index.json           # diffusers pipeline 索引(ZImagePipeline)
├── openvino_config.json       # OpenVINO 量化設定
├── inference_int4.py          # 單張推論範例
├── generate5.py               # 5 組 prompt 批次生成 + benchmark
├── quantize_int4.py           # FP16 OV → INT4 OV 量化腳本
├── transformer/               # INT4 ZImageTransformer2DModel(6B)
├── text_encoder/              # INT4 Qwen3 語言模型
├── tokenizer/                 # Qwen2Tokenizer
├── vae_encoder/               # INT8 VAE encoder
├── vae_decoder/               # INT8 VAE decoder
├── scheduler/                 # FlowMatchEulerDiscreteScheduler (shift=3.0)
└── examples/                  # 20 張實測圖 + benchmark.json + prompts.txt
    ├── *_8step_1024.png       # 主測組(5 張)
    ├── *_8step_512.png        # 對照組(5 張)
    ├── *_9step_1024.png       # 風格展示(10 張)
    ├── benchmark.json         # 逐 step 耗時 + 記憶體 + 完整系統資訊
    └── prompts.txt            # 5 組 prompt 與 seed

從零復現

# 1. 匯出 FP16 OpenVINO 模型
optimum-cli export openvino \
  -m Tongyi-MAI/Z-Image-Turbo \
  --task text-to-image \
  --library diffusers \
  --weight-format fp16 \
  ./z-image-turbo-ov-fp16

# 2. NNCF weight-only INT4 量化
python quantize_int4.py --fp16-dir ./z-image-turbo-ov-fp16 \
                          --int4-dir  ./z-image-turbo-ov-int4

# 3. 單張推論
python inference_int4.py

# 4. 批次生成 5 組 + 512px 對照 + benchmark
python generate5.py --outdir examples

量化設定:

from optimum.intel.openvino.configuration import (
    OVConfig, OVWeightQuantizationConfig, OVPipelineQuantizationConfig,
)

int4_config = OVWeightQuantizationConfig(
    bits=4, sym=False, group_size=128,
    group_size_fallback="adjust", ratio=1.0,
)

pipeline_config = OVPipelineQuantizationConfig(
    quantization_configs={
        "transformer":  int4_config,
        "text_encoder": int4_config,
    },
    default_config=OVWeightQuantizationConfig(bits=8, sym=True),
)

已知限制

  • 記憶體需求高:INT4 模型僅 5.52 GB,但實測推論期間 RSS 達 16.3 GB。這是 OpenVINO CPU 在 INT4 執行時的即時解量化與中間張量所致,建議至少預留 20 GB 可用記憶體。
  • CPU 速度有限:1024×1024 平均 111 s / 張,遠慢於同批的 SDXL 蒸餾模型。
  • 僅為 weight-only 量化:首次載入 + 編譯約 7.7 s。
  • CPU 專用:GPU 使用請改用原模型或自行轉換 OpenVINO GPU。
  • 9 steps 蒸餾設定:num_inference_steps > 9 不會變好。
  • **model_index.json 宣告單一 vae,但實際目錄為 vae_decoder/ + vae_encoder/**:這是 optimum-intel 匯出的常見情況,OVDiffusionPipeline 可正常載入。
  • 風格展示圖無實測紀錄:10 張風格圖未納入 benchmark,完整 prompt 亦未收錄。

授權與出處

  • 來源模型:Tongyi-MAI/Z-Image-Turbo

  • 授權:apache-2.0(沿用來源模型授權)

  • 轉換:僅做格式轉換與權重量化,模型權重來自來源模型

使用本模型時請一併遵守來源模型的授權條款。


Made with OpenVINO + optimum-intel + NNCF

Downloads last month
85
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for HelloSun/Z-Image-Turbo-OpenVINO-INT4

Finetuned
(160)
this model