Image-2.1 — CMF

Qwen-Image-2.1 packaged as a single CMF file: the 7B image transformer, the Qwen3-VL-8B text encoder with its vision tower, and the RGBA VAE together, with the generation recipe stored inside. It runs with cortiq, a Rust engine with no Python: Apple silicon through Metal, NVIDIA GPUs through Vulkan, and a CPU fallback everywhere. Text-to-image, native transparency (RGBA), and generation from condition images. Built with Qwen.

A neon shop sign that reads CORTIQ, rainy night
A neon shop sign that reads "CORTIQ", rainy night, reflections on wet pavement
A transparent cartoon dragon sticker
RGBA: a cute cartoon dragon sticker, transparent background
A red fox in fresh snow at golden hour
A red fox sitting in fresh snow at golden hour, photorealistic
A vintage travel poster of Lake Baikal
A vintage travel poster of Lake Baikal, the title "BAIKAL"
Condition image and the edit
--image (left, the fox above) + "Change the background to a sunset beach, keep the fox unchanged" (right)

1024×1024, 40 steps, default settings, generated by cortiq (Vulkan, RTX PRO 4000); click for the full PNG.

English · Русский · 中文

Quick start

cargo install cortiq-cli          # 0.8.2 or later; prebuilt binaries: github.com/infosave2007/cmf/releases
hf download infosave/Image-2.1-cmf qwen-image-2.1.cmf qwen-image-2.1.cmf.sha256 --local-dir .
cortiq imagine qwen-image-2.1.cmf --prompt "A neon shop sign that reads \"CORTIQ\", rainy night, reflections on wet pavement"

No flags needed: the file carries the recipe (1024×1024, 40 steps, no CFG). Set XDG_RUNTIME_DIR=/tmp on a headless Linux box.

Transparent images — ask for them in the prompt, PNG keeps the alpha channel:

cortiq imagine qwen-image-2.1.cmf --out sticker.png --prompt "This is an RGBA image with transparency. A cute cartoon dragon sticker. The image has alpha channel and the background is transparent."

From condition images (repeat --image for several):

cortiq imagine qwen-image-2.1.cmf --image photo.png --prompt "Change the background to a sunset beach"

File

file size inside
qwen-image-2.1.cmf 12.7 GB transformer 4-bit (q4tp), Qwen3-VL-8B text encoder and vision tower 8-bit (q8_2f), VAE f16

sha256 b94309e4ed716c7a224c3dc88d583d3095d6e34edd76a76c01b04b9123a40bce (qwen-image-2.1.cmf.sha256 is next to it). The text encoder is 8-bit by measurement: at 4 bits its output moves 40 % against the fp32 reference, at 8 bits 11 %. With bf16 weights the engine matches the diffusers fp32 reference to 1e-5.

Performance

RTX PRO 4000 Blackwell (24 GB), Vulkan, one image per process, default settings (40 steps):

task size time to image DiT steps prompt encode VAE (load + decode)
text-to-image 512×512 15 s 9.6 s (0.24 s a step) 2.4 s 1.4 s
text-to-image 1024×1024 48 s 41.6 s (1.04 s a step) 2.2 s 1.8 s
text-to-image 2048×2048 246 s 238 s (6.0 s a step) 2.0 s 3.2 s
one 1024² condition image 1024×1024 122 s 47 s (1.18 s a step) 60 s 1.7 s

Peak VRAM 17.5 GB at 1024² (24 GB cards). With several GPUs the engine takes the one with tensor cores and the most memory; CMF_GPU_ADAPTER=<n> (see cortiq gpu) pins another.

Mac mini M4 (10-core GPU, 24 GB), Metal:

task size time to image DiT steps VAE (load + decode)
text-to-image 512×512 3.9 min 5.5 s a step 7.6 s
text-to-image 1024×1024 20 min 25–30 s a step 34 s

About 9 GB of memory. The GEMMs run at ≈ 83 % of the M4 GPU's matrix peak, so the time is the model's arithmetic; --steps 20 halves it. Measured with cortiq 0.8.2.

Options

option meaning
--width, --height multiples of 32 (default 1024×1024, up to 2048² and beyond on 24 GB; with --image, the last image's aspect)
--steps, --seed, --num-images N image i uses seed + i
--image PATH condition image, repeatable
--reference-size S condition images are resized to S² area (default 1024)
--cfg G, --negative-prompt true CFG when G > 1 and a negative prompt is given (the model samples without guidance)
--out .png (default out.png, RGBA), .jpg, .ppm

CMF_QI21_PROF=1 prints stage times. CMF_GPU=0 runs everything on the CPU (very slow).

License

Qwen-Image-2.1 is released under the Qwen Research License: non-commercial use only (research and evaluation). This repository redistributes it modified — quantized and repackaged into CMF by infosave — under the same agreement; see LICENSE and NOTICE. Commercial use needs a separate license from the Qwen team. Built with Qwen.

Verify

sha256sum -c qwen-image-2.1.cmf.sha256    # macOS: shasum -a 256 -c
cortiq info qwen-image-2.1.cmf

Документация на русском

Qwen-Image-2.1 одним файлом CMF: трансформер 7B, текстовый энкодер Qwen3-VL-8B с vision-башней и RGBA-VAE вместе, рецепт генерации записан в файл. Запускается движком cortiq на Rust без Python: Apple silicon через Metal, NVIDIA через Vulkan, везде есть путь на CPU. Генерация по тексту, нативная прозрачность (RGBA), генерация по картинкам-условиям.

cargo install cortiq-cli          # 0.8.2 или новее; готовые бинарники: github.com/infosave2007/cmf/releases
hf download infosave/Image-2.1-cmf qwen-image-2.1.cmf qwen-image-2.1.cmf.sha256 --local-dir .
cortiq imagine qwen-image-2.1.cmf --prompt "Лиса сидит в свежем снегу, золотой час, фотореализм"
cortiq imagine qwen-image-2.1.cmf --image photo.png --prompt "Замени фон на закат на пляже"

Флаги не нужны: 1024×1024, 40 шагов, без CFG. Прозрачный фон — попросите в промпте, PNG сохраняет альфа-канал. На Linux без дисплея задайте XDG_RUNTIME_DIR=/tmp.

Файл: qwen-image-2.1.cmf, 12.7 ГБ — трансформер 4 бит (q4tp), энкодер и vision-башня 8 бит (q8_2f), VAE f16; sha256 в qwen-image-2.1.cmf.sha256.

Скорость (один процесс на картинку, настройки по умолчанию, 40 шагов):

устройство 512² 1024² 2048² 1024² с картинкой-условием
RTX PRO 4000 Blackwell, Vulkan 15 с 48 с (шаг 1.04 с) 246 с (шаг 6.0 с) 122 с (энкодер 60 с)
Mac mini M4 24 ГБ, Metal 3.9 мин 20 мин (шаг 25–30 с) — —

Пик видеопамяти 17.5 ГБ (карты на 24 ГБ); из нескольких GPU берётся карта с тензорными ядрами и большей памятью, CMF_GPU_ADAPTER=<n> выбирает другую. На Mac около 9 ГБ памяти, GEMM идут на ≈ 83 % матричного пика M4; --steps 20 вдвое быстрее.

Опции: --width/--height (кратны 32), --steps, --seed, --num-images N, --image (повторяемый), --reference-size S, --cfg G с --negative-prompt, --out (.png с альфой, .jpg, .ppm). CMF_QI21_PROF=1 печатает время стадий.

Лицензия: Qwen Research License — только некоммерческое использование; файлы изменены (квантованы и упакованы в CMF), см. LICENSE и NOTICE. Built with Qwen.


中文文档

Qwen-Image-2.1 打包为单个 CMF 文件:7B 图像 Transformer、Qwen3-VL-8B 文本编码器及其视觉塔、 RGBA VAE 合在一起,生成配方也写在文件里。由 Rust 引擎 cortiq 运行,无需 Python:Apple silicon 走 Metal,NVIDIA 走 Vulkan,任何机器都有 CPU 后备路径。支持文生图、原生透明(RGBA)和以图为条件的生成。

cargo install cortiq-cli          # 0.8.2 或更新
hf download infosave/Image-2.1-cmf qwen-image-2.1.cmf qwen-image-2.1.cmf.sha256 --local-dir .
cortiq imagine qwen-image-2.1.cmf --prompt "雨夜里写着 \"CORTIQ\" 的霓虹招牌,湿润路面上的倒影"

无需参数:1024×1024,40 步,无 CFG。透明背景请在提示词中说明,PNG 保留 alpha 通道。

文件: qwen-image-2.1.cmf,12.7 GB —— Transformer 4 位(q4tp)、文本编码器与视觉塔 8 位(q8_2f)、VAE f16。

性能(每进程一张图,默认设置,40 步):

设备 512² 1024² 2048² 1024² 带条件图
RTX PRO 4000 Blackwell,Vulkan 15 s 48 s(每步 1.04 s) 246 s(每步 6.0 s) 122 s(编码 60 s)
Mac mini M4 24 GB,Metal 3.9 分钟 20 分钟(每步 25–30 s) — —

显存峰值 17.5 GB(24 GB 显卡);多卡时自动选择带张量核心、显存最大的卡,CMF_GPU_ADAPTER=<n> 可指定。Mac 约占 9 GB 内存,--steps 20 可减半耗时。

许可: Qwen Research License —— 仅限非商业用途;文件经过修改(量化并打包为 CMF),见 LICENSE 与 NOTICE。Built with Qwen。

Downloads last month
3
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for infosave/Image-2.1-cmf

Quantized
(91)
this model