studio-mlx / README.md
kruatech's picture
Upload folder using huggingface_hub
6a10500 verified
|
Raw History Blame Contribute Delete
10.5 kB
---
license: other
license_name: mixed-apache-2.0-and-mit
license_link: https://huggingface.co/kruatech/studio-mlx/blob/main/LICENSES.md
library_name: mlx
tags:
- mlx
- mlx-swift
- apple-silicon
- text-to-image
- image-editing
- text-to-music
- quantized
---
<div align="center">
# studio-mlx
**Image and music models converted for native MLX inference on Apple silicon**
![MLX](https://img.shields.io/badge/MLX-0.32-1f6feb) ![platform](https://img.shields.io/badge/platform-Apple%20silicon-black) ![precision](https://img.shields.io/badge/precision-bf16%20and%204--bit-6f42c1) [![licence](https://img.shields.io/badge/licence-Apache--2.0%20and%20MIT-brightgreen)](LICENSES.md)
</div>
Three models and the text encoder the two image models share. Every folder is a self-contained bundle with its own card, so you download only what you need.
Weights are not retrained. Tensor layouts were adapted for MLX and the 4-bit folders are quantized, so parameters are mathematically equivalent to upstream rather than byte-identical.
## Quick start
Pick one line. The image models need the shared encoder, the music model does not.
```bash
pip install "huggingface_hub[hf_xet]"
```
The `-q4` folders sit at the top level rather than inside the bf16 ones, so an include pattern fetches one precision without dragging in the other.
## What is inside
| folder | what it does | precision | size | licence |
|----------------------------------------------------------------------|------------------------------------------|-----------|----------|------------|
| [`qwen3-tts-12hz-1.7b-base`](qwen3-tts-12hz-1.7b-base) | Voice cloning from a reference recording | `bf16` | 4.23 GiB | Apache-2.0 |
| [`qwen3-tts-12hz-1.7b-customvoice`](qwen3-tts-12hz-1.7b-customvoice) | Nine preset voices, ten languages | `bf16` | 4.21 GiB | Apache-2.0 |
| [`qwen3-tts-12hz-1.7b-voicedesign`](qwen3-tts-12hz-1.7b-voicedesign) | Describe a voice in words and get it | `bf16` | 4.21 GiB | Apache-2.0 |
## bf16 or 4-bit
**Use the 4-bit folders.** Quantization is MLX affine: 4 bits per weight with a bf16 scale and bias for every group of 64, about 4.5 bits per weight and 28-30% of the bf16 size. Normalizations, modulations and the encoder's embedding table stay in bf16, and matrix multiplication still runs in bf16 - the gain is memory and load time, not integer arithmetic.
| | bf16 | 4-bit |
|-----------------------------------|-----------|----------|
| Z-Image DiT | 11.46 GiB | 3.40 GiB |
| klein DiT | 7.22 GiB | 2.03 GiB |
| shared encoder | 7.51 GiB | 2.58 GiB |
| klein 512 px, 4 steps, wall clock | 170 s | 73 s |
Output quality was indistinguishable: the same prompt and seed gave two clean images that differ the way two seeds differ. The bf16 folders exist as the accuracy reference a port can be checked against.
## Bundle layout
```
<folder>/
β”œβ”€β”€ manifest.json SHA-256, byte sizes and tensor counts for every file
β”œβ”€β”€ config/ component configs copied from upstream
β”œβ”€β”€ tokenizer/ BPE plus chat_format.json, where applicable
└── weights/ safetensors in MLX tensor layout
```
`manifest.json` records the quantization settings, so a loader rebuilds the exact layout without being told. Verify integrity before first use: the manifest carries SHA-256 per file, which catches corruption that size checks miss.
`tokenizer/chat_format.json` holds the Qwen chat template already rendered into a prefix and a suffix for each pipeline, so a runtime needs no Jinja at all.
<details>
<summary><b>Tensor layout</b></summary>
MLX convolutions expect channels last, PyTorch expects channels first, so convolution weights are permuted during conversion. Everything else keeps its upstream shape.
| kind | PyTorch | MLX |
|-----------------------------|---------------------|---------------------|
| `Conv2d.weight` | `(out, in, kH, kW)` | `(out, kH, kW, in)` |
| `Linear`, norms, embeddings | `(out, in)` | unchanged |
No `weight_norm` and no 3-D convolutions appear in these models, so no folding was needed. Parameter names match the upstream checkpoints, so weights load without remapping.
</details>
<details>
<summary><b>Verification numbers</b></summary>
Every module was compared against the upstream reference on fixed inputs in float32. The metric is `rel_max = max|a-b| / max|a|`.
| module | rel_max | reference |
|-------------------------------------------------|-----------|---------------------------------|
| Qwen3 encoder, hidden_states[-2] | 2.7e-07 | transformers |
| Qwen3 encoder, layers 9/18/27 with padding mask | 6.0e-07 | transformers |
| Z-Image DiT | 4.2e-06 | diffusers |
| Flux2 DiT | 3.7e-07 | diffusers |
| VAE decoder | 1.3e-05 | diffusers |
| VAE encoder | 5.3e-06 | diffusers |
| VAE tiled decode | 5.1e-06 | diffusers tiled_decode |
| sigma schedule, static shift | 3.2e-08 | FlowMatchEulerDiscreteScheduler |
| sigma schedule, exponential dynamic shift | 7.7e-08 | FlowMatchEulerDiscreteScheduler |
| position ids, latent packing, patchify | bit-exact | pipeline helpers |
A Swift MLX implementation was then checked against the Python one:
| module | rel_max | note |
|------------------------------|-----------------|------------------------------------------------|
| tokenization, both pipelines | bit-exact | same ids |
| Qwen3 encoder, bf16 | 2.2e-04 | Z-Image branch |
| Qwen3 encoder, bf16 | 4.6e-03 | klein branch, 512 tokens with mask |
| Qwen3 encoder, 4-bit | 3.0e-04 | proves the quantized layout is rebuilt exactly |
| Z-Image DiT, float32 | 3.7e-07 | 8 layers |
| Flux2 DiT, float32 | 1.0e-06 | 2 double + 2 single blocks |
| VAE decode / tiled / encode | 1e-05 or better | both VAE classes |
| VAE BatchNorm statistics | 0 | exact |
The bf16 figures for a full-depth DiT are larger - 2e-02 for both models - and that is rounding order, not a defect: at float32 the same code agrees to 1e-06, and the deviation grows with depth from a bf16-level 1e-04 per block. Comparing two bf16 implementations below 1e-02 is not meaningful for a 30-block network.
</details>
<details>
<summary><b>Performance on an M1 Max</b></summary>
Mac Studio, Apple M1 Max, 10-core CPU, 24-core GPU, 32 GB, macOS 26.5. Release build, 4-bit weights, batch 1, warm page cache.
| run | per step | VAE decode |
|--------------------------------------|----------|------------|
| Z-Image, 512 px, 4 steps | 2.6 s | 1.5 s |
| klein, 512 px, 4 steps | 2.3 s | 0.9 s |
| klein, 1024 px, 4 steps | 8.0 s | 0.2 s |
| klein, 512 px + one 512 px reference | 4.1 s | 0.9 s |
Reading the weights dominates a cold run and depends entirely on the storage: on an external volume delivering about 0.2 GiB/s, the 4-bit encoder took 22-32 s and a 4-bit DiT 18-35 s, against 60-107 s for the same weights in bf16. On internal storage expect these to be several times shorter. Compute is unaffected: the numbers above are steady state.
Stages run one at a time and release their weights, so peak memory is set by the largest single component rather than their sum. The 4-bit DiT is 2.03 GiB for klein and 3.40 GiB for Z-Image; activations at 1024 px add to that, and no separate peak measurement was made.
</details>
<details>
<summary><b>Limitations</b></summary>
- Seeds are not compatible with the upstream pipelines, which use `torch.Generator`. The same prompt gives comparable images, never the same file.
- Both models are distilled: CFG is not applied, and step counts are low by design (8 for Z-Image, 4 for klein).
- Z-Image's Omni mode is not covered: it needs a SigLIP encoder that is not part of this bundle.
- Tiled VAE decode is available for high resolutions and gives a result that differs slightly from a single pass, exactly as it does in diffusers.
- Batches larger than one are not implemented.
</details>
## Licences
Derivatives of three upstream models under two licences. Each folder carries the licence of its upstream model.
| folder | upstream | licence |
|-----------------------------------|-----------------------------------------------------------------------------------------------------|-----------------------------------------------------------|
| `qwen3-tts-12hz-1.7b-base` | [Qwen/Qwen3-TTS-12Hz-1.7B-Base](https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-Base) | [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0) |
| `qwen3-tts-12hz-1.7b-customvoice` | [Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice](https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice) | [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0) |
| `qwen3-tts-12hz-1.7b-voicedesign` | [Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign](https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign) | [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0) |
The shared text encoder is redistributed from inside the two image repositories, both Apache-2.0. Conversion tooling and these cards: MIT.
The upstream cards state usage restrictions and responsible-use commitments that redistribution does not repeal. See in particular the out-of-scope use section of the FLUX.2 [klein] card.