FLUX.2-klein-4B — OpenVINO INT4

將 black-forest-labs/FLUX.2-klein-4B 轉換為 OpenVINO INT4 權重,並以 純 CPU 實測出圖。

optimum-intel 匯出 → NNCF weight-only INT4 量化(耗時 45 s)→ FP16 15.79 GB → INT4 4.39 GB(↓ ~72%,約 3.6×)。


目錄


快速資訊

項目 內容
基座模型 black-forest-labs/FLUX.2-klein-4B
Pipeline Flux2KleinPipeline
Scheduler FlowMatchEulerDiscreteScheduler(shift=3.0、use_dynamic_shifting=true)
量化格式 transformer、text_encoder → INT4;其餘 → INT8
模型大小 FP16 15.79 GB → INT4 4.39 GB(↓ ~72%,約 3.6×)
推論裝置 CPU(OpenVINO CPU plugin,無需 GPU)
推薦參數 num_inference_steps=4、guidance_scale=1.0、1024×1024
轉換工具 optimum-intel 2.2.x / optimum 2.3.0 / OpenVINO 2026.4.0 / NNCF 3.4.0

特色

  • 4 step 出圖:FLUX.2-klein-4B 為蒸餾模型,num_inference_steps=4。
  • **體積 ↓ 72%**:FP16 15.79 GB → INT4 4.39 GB(約 3.6×)。
  • 記憶體實測完整:逐 step 記憶體取樣 + 行程結束時 RSS 皆有記錄。
  • 純 CPU 可用:不依賴 GPU / CUDA。
  • diffusers 目錄結構:可直接 from_pretrained()。
  • ⚠️ 512px 對照為縮圖:不是重新推理,無對應耗時資料。

安裝

pip install diffusers==0.37.1 transformers==4.57.6 tokenizers==0.22.0 huggingface-hub==0.35.1 optimum==2.3.0 optimum-intel==2.2.0 openvino==2026.4.0 nncf==3.4.0 torch pillow psutil

nncf 僅重新量化時需要。

快速開始

import torch
from optimum.intel import OVDiffusionPipeline

pipe = OVDiffusionPipeline.from_pretrained("HelloSun/FLUX.2-klein-4B-OpenVINO-INT4", compile=True)

image = pipe(
    prompt="Astronaut in a jungle, cold color palette, muted colors, "
             "detailed, 8k, photorealistic, cinematic lighting",
    height=1024, width=1024,
    num_inference_steps=4,
    guidance_scale=1.0,
    generator=torch.Generator().manual_seed(43),
).images[0]

image.save("out.png")

完整可執行範例:inference_int4.py 批次生成 + benchmark:generate5.py

推論參數建議

參數 建議值 說明
num_inference_steps 4 蒸餾 / 推薦步數。請勿隨意增加。
guidance_scale 1.0 CFG 設定;Turbo / 蒸餾模型通常為 0.0 或 1.0(即不啟用)。
shift 取自 scheduler/scheduler_config.json 不需手動傳入,載入時自動套用。
height / width 1024 實測解析度。
compile=True 開啟 編譯模型以取得較佳效能。

範例結果

全部為 4 steps / guidance_scale 1.0 / 1024×1024 / CPU,seed 42–46 固定,可完全重現。

另附 512px 對照圖(outputs/*_512.png)。

⚠️ 512px 為縮圖:這批 512px 圖是 1024px 輸出的 LANCZOS 縮圖,不是重新以 512px 推理的結果,因此沒有對應的獨立耗時資料。

01_hanfu — seed 42 — 39.8 s

Young Chinese woman in red Hanfu, intricate embroidery, impeccable makeup, red floral forehead pattern, elaborate high bun, golden phoenix headdress, soft-lit outdoor night background, silhouetted tiered pagoda, blurred colorful distant lights, photorealistic, ultra detailed, 8k

01_hanfu

02_astronaut — seed 43 — 34.7 s

Astronaut in a jungle, cold color palette, muted colors, detailed, 8k, photorealistic, cinematic lighting

02_astronaut

03_taipei — seed 44 — 34.0 s

Cyberpunk street in Taipei at night, heavy rain, neon signs with text 'TAIPEI' and Chinese characters '台北', reflections on wet asphalt, crowded night market, cinematic, ultra detailed

03_taipei

04_shiba — seed 45 — 34.2 s

Cute Shiba Inu wearing a tiny astronaut helmet, sitting in a field of sunflowers under a starry sky, dreamy illustration, vibrant colors, high quality

04_shiba

05_ink — seed 46 — 34.0 s

Traditional Chinese ink wash landscape, misty mountains, a small pagoda on a cliff, cranes flying, minimalist, elegant, high aesthetic quality

05_ink

風格展示

以下 10 張為 10 類風格 × 同一組 5 組基礎 prompt(seed 42–51), 使用與上方基準測試完全相同的推論設定。

注意:這批圖片沒有對應的 outputs/benchmark.json 紀錄,也沒有腳本可重現, 完整 prompt 亦未收錄於 repo(早期說明檔中的版本已被截斷,無法還原)。

traditional_painting · 傳統繪畫

油畫、版畫、水彩等傳統媒材

01_traditional_painting

anime_manga · 動漫繪師

吉卜力 × 新海誠動畫風

02_anime_manga

digital_3d · 數位 3D

UE5 寫實渲染、CG 質感

03_digital_3d

photography_cinema · 攝影電影

電影感人像攝影

04_photography_cinema

cultural_regional · 文化地域

敦煌壁畫、水墨等東方美學

05_cultural_regional

material_craft · 材質工藝

絲線刺繡、陶瓷、木雕等

06_material_craft

design_commercial · 平面商業

商業插畫、扁平幾何設計

07_design_commercial

abstract_generative · 抽象生成

分形、生成藝術

08_abstract_generative

era_subculture · 復古次文化

80s 霓虹、賽博龐克

09_era_subculture

render_technique · 厚塗渲染

油畫厚塗、概念設定

10_render_technique

效能實測摘要

完整逐 step 數據見 REPORT.md 與 outputs/benchmark.json。

測試環境

項目 內容
CPU Intel(R) Xeon(R) Platinum 8559C
拓撲 2 sockets × 48 cores × 2 threads/core = 192 vCPU(96 實體核心)
RAM 2.0 TiB
虛擬化 KVM(完整虛擬化)
OpenVINO CPU only,2026.4.0(build 2026.4.0-22959-99c81491cc3-releases/2026/4)
設定 num_inference_steps=4、guidance_scale=1.0、1024×1024

總結

指標 數值
解析度 1024×1024
平均總耗時 35.36 s / 張
平均單步耗時 7.87 s
最快 / 最慢 34.02 s / 39.83 s
總計(5 張) 176.81 s
記憶體高水位 14,636 MB
  • 512px 沒有獨立耗時資料:outputs/*_512.png 是 1024px 輸出的縮圖,非重新推理。
  • 文字編碼與 VAE decode 的時間已包含在總耗時內。

逐張結果(1024×1024)

# Prompt Seed 總耗時 (s) 平均單步 (s) guidance_scale
01_hanfu 42 39.83 8.166 1.0
02_astronaut 43 34.72 7.833 1.0
03_taipei 44 34.02 7.814 1.0
04_shiba 45 34.21 7.780 1.0
05_ink 46 34.02 7.731 1.0
平均 35.36 7.865

模型大小

以下為 repo 內 openvino_model.bin 的實際位元組數(Git LFS 記錄值)。

元件 位元組 大小 精度
text_encoder 2,289,393,757 2.29 GB INT4
transformer 2,013,667,900 2.01 GB INT4
vae_decoder 49,702,491 0.05 GB —
vae_encoder 34,489,244 0.03 GB —
合計 4,387,253,392 4.39 GB

FP16 匯出模型約 15.79 GB(transformer 7.75 GB + text_encoder 8.04 GB + vae_decoder 95 MB + vae_encoder 66 MB)——取自原始轉換紀錄。

⚠️ 早期說明中的 4.2 GB 與實測不符:repo 內四個 openvino_model.bin 相加為 4.39 GB,本表為準。

壓縮比按位元組計算約 3.6×;分元件看 transformer 為 3.85×、text_encoder 為 3.51×。

text_encoder INT4 後仍有 2.29 GB(Qwen3 級別語言模型),是整體壓縮率的主要限制。

檔案結構

.
├── README.md                  # 本文件
├── REPORT.md                  # 完整轉換 + 實測報告
├── model_index.json           # diffusers pipeline 索引(Flux2KleinPipeline)
├── openvino_config.json       # OpenVINO 量化設定
├── inference_int4.py          # 單張推論範例
├── generate5.py               # 5 組 prompt 批次生成 + benchmark
├── quantize_int4.py           # FP16 OV → INT4 OV 量化腳本
├── transformer/               # INT4 Flux2Transformer2DModel(4B)
├── text_encoder/              # INT4 Qwen3 語言模型
├── tokenizer/                 # Qwen2Tokenizer
├── vae_encoder/               # INT8 VAE encoder
├── vae_decoder/               # INT8 VAE decoder
├── scheduler/                 # FlowMatchEulerDiscreteScheduler 設定
├── examples/                  # 15 張展示圖(含 10 張風格圖)
└── outputs/                   # 10 張實測圖 + 3 個資料檔
    ├── *_1024.png             # 主測組(5 張)
    ├── *_512.png              # 對照組(5 張,為 1024 縮圖)
    ├── benchmark.json         # 逐 step 耗時 + 記憶體 + 系統資訊
    ├── benchmark_quantization.json  # 量化設定與耗時
    └── prompts.txt            # 5 組 prompt 與 seed

從零復現

# 1. 匯出 FP16 OpenVINO 模型
optimum-cli export openvino \
  -m black-forest-labs/FLUX.2-klein-4B \
  --task text-to-image \
  --library diffusers \
  --weight-format fp16 \
  ./FLUX.2-klein-4B-ov-fp16

# 2. NNCF weight-only INT4 量化
python quantize_int4.py --fp16-dir ./FLUX.2-klein-4B-ov-fp16 \
                          --int4-dir  ./FLUX.2-klein-4B-ov-int4

# 3. 單張推論
python inference_int4.py

# 4. 批次生成 5 組 + benchmark
python generate5.py --outdir outputs

量化設定:

from optimum.intel.openvino.configuration import (
    OVConfig, OVWeightQuantizationConfig, OVPipelineQuantizationConfig,
)

int4_config = OVWeightQuantizationConfig(
    bits=4, sym=False, group_size=128,
    group_size_fallback="adjust", ratio=1.0,
)

pipeline_config = OVPipelineQuantizationConfig(
    quantization_configs={
        "transformer":  int4_config,
        "text_encoder": int4_config,
    },
    default_config=OVWeightQuantizationConfig(bits=8),
)

已知限制

  • 512px 對照非重新推理:outputs/*_512.png 是 1024px 輸出的 LANCZOS 縮圖,若要評估 512px 的真實速度需自行重新推理。
  • 記憶體需求高:INT4 模型 4.39 GB,但實測推論期間 RSS 達 14.6 GB。
  • CPU 速度有限:1024×1024 平均 35.4 s / 張。
  • 首次生成有額外開銷:01_hanfu 為 39.83 s,其後四張穩定在 34.0–34.2 s,差距來自 text encoder 首次執行與 warmup。
  • 10 張風格展示圖無實測紀錄:只有 5 張主測圖有 benchmark 資料,完整 prompt 亦未收錄於 outputs/prompts.txt。
  • **outputs/benchmark_quantization.json 的 export_time_seconds 為 0.0**:代表 FP16 匯出是在另一次執行中完成的,該腳本本身未做匯出,因此 repo 內沒有 FP16 的機器紀錄。

授權與出處

使用本模型時請遵守來源模型的授權條款。


Made with OpenVINO + optimum-intel + NNCF

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for HelloSun/FLUX.2-klein-4B-OpenVINO-INT4

Finetuned
(64)
this model