File size: 9,369 Bytes
4f03424
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
"""Render bilingual model cards from actual evaluation records."""
import json
from pathlib import Path

def main():
    summary=json.loads(Path('artifacts/eval/summary.json').read_text())
    review=json.loads(Path('artifacts/eval/visual-review.json').read_text())
    if review.get('approved_for_research_release') is not True:
        raise ValueError('Complete the paired visual review before rendering release cards')
    conversion=json.loads(Path('models/Image21-MLX-8bit/conversion.json').read_text())
    gib=sum(x['tensor_bytes'] for x in conversion['components'].values())/2**30
    header='''---
license: other
license_name: qwen-research
license_link: LICENSE
base_model: Qwen/Qwen-Image-2.1
library_name: mlx
pipeline_tag: text-to-image
tags:
- mlx
- mlx-vlm
- apple-silicon
- quantized
- image-to-image
- rgba
---
'''
    common=f'''
## Reproducible inference

```bash
uv venv --python 3.13 .venv
uv pip install --python .venv/bin/python -r requirements.lock.txt
.venv/bin/python -m scripts.infer --model . \\
  --prompt 'A natural portrait in soft window light' --output outputs/portrait.png
```

```bash
.venv/bin/python -m scripts.infer --model . --input input.png \\
  --prompt 'Change only the blue sweater to a red sweater. Preserve the person and background.' \\
  --source-seed 42 --seed 1000042 --output outputs/edit.png
```

Transparent generation prompt:
`This is an RGBA image with transparency. A cute cartoon dragon sticker. The image has alpha channel and the background is transparent.`

## Measurements / 实测

Apple M4 Max, 40 GPU cores, 128 GiB unified memory, macOS 26.6.2, MLX 0.32.2.
Both models use the same pinned MLX runtime and internal SSD; 1024×1024, 40 steps,
CFG=1, no VAE tiling, one full warm-up per process, seven cases, one seed per case.
Sequential component loading is enabled in both models. Per-image time includes
component loading, prompt encoding, denoising and VAE decoding; PNG writing is excluded.
This is a desktop session with other applications open. BF16 ran before Q8;
there were no repeated or interleaved trials to control order and thermal effects.

| Case | BF16 seconds | 8-bit seconds | BF16 peak GiB | 8-bit peak GiB |
|---|---:|---:|---:|---:|
'''
    for r in summary['rows']:
        common+=f'| {r["case_id"]} | {r["bf16_seconds"]:.2f} | {r["q8_seconds"]:.2f} | {r["bf16_peak_gib"]:.2f} | {r["q8_peak_gib"]:.2f} |\n'
    relative=100*(summary['q8_t2i_mean_seconds']/summary['bf16_t2i_mean_seconds']-1)
    common+=f'\nSix-case text-to-image mean / 六类文生图平均:BF16 **{summary["bf16_t2i_mean_seconds"]/60:.2f} min**, Q8 **{summary["q8_t2i_mean_seconds"]/60:.2f} min**. Q8 generation time relative to BF16 / Q8 相对耗时:**{relative:+.1f}%** in this run.\n'
    common+='''
Peak figures measure MLX allocations, not minimum physical RAM. Full raw records,
original RGBA samples and the visual review are in [evaluation](evaluation/report.md).
These are informal measurements, not an official benchmark. The BF16 baseline is
the same MLX implementation; cross-runtime CUDA equivalence is not claimed.
One seed per case is insufficient to establish statistical quality equivalence.
There is no claim of lossless quantization or a guaranteed speedup.

![All seven paired samples](evaluation/comparison.png)

## Provenance

- Source: `Qwen/Qwen-Image-2.1@b3179ad355be050328e483a9dfdd9e60cd62adfa`.
- Runtime: `Blaizzy/mlx-vlm@95b01ccad2d9f65a9e87f6a87bd1c5df69626261`.
- Native MLX affine packed weights; not CUDA bitsandbytes INT8, FP8, GGUF or an MFLUX checkpoint.
- All three components passed exact tensor round-trip verification.
- [conversion.json](conversion.json), [modifications](CHANGES.md), [Notice](Notice), [license](LICENSE), [file hashes](MANIFEST.json).

To rebuild, first download the original source snapshot at the revision above to
`source-bf16` (requires additional disk space), then run from this repository:

```bash
.venv/bin/python -m scripts.audit --source source-bf16 --manifest evaluation/source-files.json
.venv/bin/python -m scripts.convert --source source-bf16 --bits 8 --output rebuilt-8bit
```
The audit verifies the original source file hashes before conversion. Rebuilding
uses the original floating-point weights, never a dequantized CUDA INT8 checkpoint.

To repeat the paired evaluation after the source audit, create a native baseline
and run each precision in its own process. This is a lengthy research workflow
and needs extra memory and disk space beyond ordinary inference:

```bash
.venv/bin/python -m scripts.convert --source source-bf16 --bits 16 --output rebuilt-bf16
.venv/bin/python -m scripts.benchmark --model rebuilt-bf16 --output artifacts/eval/bf16
.venv/bin/python -m scripts.benchmark --model . --output artifacts/eval/8bit --edit-input artifacts/eval/bf16/portrait-s42.png
.venv/bin/python -m scripts.report
```
'''
    en=f'''# Image21-MLX-8bit

**Built with Qwen.** An independent native MLX quantization of Qwen-Image-2.1 for
Apple Silicon, by ixim / iximbox. **Non-commercial research and evaluation only**
under the original Qwen Research License. This is not an official Qwen release.

The complete checkpoint is approximately **{gib:.2f} GiB**. DiT attention/MLP and
language-encoder attention/MLP linears use **8-bit affine weights, group size 64**,
with BF16 activations. The full vision tower, token embeddings, language head,
norms and DiT input/output/timestep/modulation layers retain floating-point precision.
The VAE retains its original **FP32** weights. See the exact 476 quantized modules in conversion.json.

Supports text-to-image, native reference-image editing and RGBA transparency through
the included scripts. The wrapper preserves the sampler's alpha channel, which the
pinned upstream text-to-image convenience method otherwise slices away.
The runtime also enables the upstream fixed-prefix KV cache for text-to-image.
Components are loaded and released by phase to reduce unified-memory use.
See scripts/mlx_pipeline.py and its retained MIT attribution.

With the included phase-loading runtime, plan for **32GB unified memory as a starting
budget; 48GB or more gives more room for other applications** at 1024px with one reference.
These are capacity estimates, not verified minimums: only the 128GB M4 Max was tested.
The measured Q8 MLX allocation peak is about **16.20 GiB**, excluding OS/driver overhead.
16/24GB Macs are not the target for this default 1024px configuration. 2048px and
10-reference editing have not been benchmarked. Allow roughly **50GB free disk space**
for ordinary inference; rebuilding also needs the original source and baseline checkpoints.
'''
    zh=f'''# Image21-MLX-8bit

**Built with Qwen。** ixim / iximbox 基于 Qwen-Image-2.1 原始 BF16 权重制作的
Apple Silicon 原生 MLX 量化版。**仅限非商业研究与评估**,沿用 Qwen Research License;
本项目为独立衍生模型,不是官方发布。

完整权重约 **{gib:.2f} GiB**。DiT 和文本编码器中的注意力/MLP 线性层采用
**8-bit affine、group size 64、BF16 激活**。整个视觉编码器、词嵌入、语言输出头、
归一化及 DiT 输入/输出、时间嵌入和 modulation 保留浮点精度;VAE 保留原始 **FP32**。
全部 476 个量化模块详见 conversion.json。本模型从原始权重直接转换,没有从 CUDA INT8 二次量化。

通过配套脚本支持文生图、参考图指令编辑和 RGBA 透明图;适配器保留原生 sampler 的 alpha,
避免固定版本上游文生图便捷方法丢弃透明通道。
运行时还将上游固定前缀 KV 缓存启用于文生图;完整改动与 MIT 归属保留在 scripts 中。
配套脚本默认按阶段加载与释放组件,组件加载开销计入下表时间。

使用配套分阶段加载运行时,1024px、单参考图建议按 **32GB 起、48GB 以上更有余量**规划。
这是容量估算,未验证最低内存边界;本次只实测了 128GB M4 Max。量化版 MLX 分配峰值约
**16.20 GiB**,还需为操作系统、驱动和其他软件留空间。16/24GB 不是默认 1024px 配置的目标。
2048px 与十张参考图的编辑尚未测评。普通推理建议预留约 **50GB 可用磁盘**;
重建量化与 BF16 基线还需额外存放原始及基线权重。

下表为非正式、同后端、单种子对照,不能证明无损,也不能与 CUDA 时间直接横向比较。
MLX 分配峰值不等于整机内存需求。参考图编辑共用同一张 BF16 原图,编辑种子与原图生成种子不同。
'''
    Path('cards').mkdir(exist_ok=True)
    for platform,body in [('huggingface',en),('modelscope',zh)]:
        command=(
            'uvx --from huggingface-hub==1.33.0 hf download ixim/Image21-MLX-8bit --local-dir Image21-MLX-8bit'
            if platform=='huggingface' else
            'uvx --from modelscope-hub==0.4.5 modelscope download iximbox/Image21-MLX-8bit --local-dir Image21-MLX-8bit'
        )
        download=f'\n## Download / 下载\n\nOn an Apple Silicon Mac with [uv](https://docs.astral.sh/uv/) installed:\n\n```bash\n{command}\ncd Image21-MLX-8bit\n```\n'
        quality='\n## Visual findings / 视觉检查\n\n'+review['summary_en' if platform=='huggingface' else 'summary_zh']+'\n'
        Path(f'cards/{platform}.md').write_text(header+body+quality+download+common)

if __name__=='__main__': main()