studio-mlx

Image and music models converted for native MLX inference on Apple silicon

MLX platform precision licence

Three models and the text encoder the two image models share. Every folder is a self-contained bundle with its own card, so you download only what you need.

Weights are not retrained. Tensor layouts were adapted for MLX and the 4-bit folders are quantized, so parameters are mathematically equivalent to upstream rather than byte-identical.

Quick start

Pick one line. The image models need the shared encoder, the music model does not.

pip install "huggingface_hub[hf_xet]"

The -q4 folders sit at the top level rather than inside the bf16 ones, so an include pattern fetches one precision without dragging in the other.

What is inside

folder what it does precision size licence
qwen3-tts-12hz-1.7b-base Voice cloning from a reference recording bf16 4.23 GiB Apache-2.0
qwen3-tts-12hz-1.7b-customvoice Nine preset voices, ten languages bf16 4.21 GiB Apache-2.0
qwen3-tts-12hz-1.7b-voicedesign Describe a voice in words and get it bf16 4.21 GiB Apache-2.0

bf16 or 4-bit

Use the 4-bit folders. Quantization is MLX affine: 4 bits per weight with a bf16 scale and bias for every group of 64, about 4.5 bits per weight and 28-30% of the bf16 size. Normalizations, modulations and the encoder's embedding table stay in bf16, and matrix multiplication still runs in bf16 - the gain is memory and load time, not integer arithmetic.

bf16 4-bit
Z-Image DiT 11.46 GiB 3.40 GiB
klein DiT 7.22 GiB 2.03 GiB
shared encoder 7.51 GiB 2.58 GiB
klein 512 px, 4 steps, wall clock 170 s 73 s

Output quality was indistinguishable: the same prompt and seed gave two clean images that differ the way two seeds differ. The bf16 folders exist as the accuracy reference a port can be checked against.

Bundle layout

<folder>/
β”œβ”€β”€ manifest.json     SHA-256, byte sizes and tensor counts for every file
β”œβ”€β”€ config/           component configs copied from upstream
β”œβ”€β”€ tokenizer/        BPE plus chat_format.json, where applicable
└── weights/          safetensors in MLX tensor layout

manifest.json records the quantization settings, so a loader rebuilds the exact layout without being told. Verify integrity before first use: the manifest carries SHA-256 per file, which catches corruption that size checks miss.

tokenizer/chat_format.json holds the Qwen chat template already rendered into a prefix and a suffix for each pipeline, so a runtime needs no Jinja at all.

Tensor layout

MLX convolutions expect channels last, PyTorch expects channels first, so convolution weights are permuted during conversion. Everything else keeps its upstream shape.

kind PyTorch MLX
Conv2d.weight (out, in, kH, kW) (out, kH, kW, in)
Linear, norms, embeddings (out, in) unchanged

No weight_norm and no 3-D convolutions appear in these models, so no folding was needed. Parameter names match the upstream checkpoints, so weights load without remapping.

Verification numbers

Every module was compared against the upstream reference on fixed inputs in float32. The metric is rel_max = max|a-b| / max|a|.

module rel_max reference
Qwen3 encoder, hidden_states[-2] 2.7e-07 transformers
Qwen3 encoder, layers 9/18/27 with padding mask 6.0e-07 transformers
Z-Image DiT 4.2e-06 diffusers
Flux2 DiT 3.7e-07 diffusers
VAE decoder 1.3e-05 diffusers
VAE encoder 5.3e-06 diffusers
VAE tiled decode 5.1e-06 diffusers tiled_decode
sigma schedule, static shift 3.2e-08 FlowMatchEulerDiscreteScheduler
sigma schedule, exponential dynamic shift 7.7e-08 FlowMatchEulerDiscreteScheduler
position ids, latent packing, patchify bit-exact pipeline helpers

A Swift MLX implementation was then checked against the Python one:

module rel_max note
tokenization, both pipelines bit-exact same ids
Qwen3 encoder, bf16 2.2e-04 Z-Image branch
Qwen3 encoder, bf16 4.6e-03 klein branch, 512 tokens with mask
Qwen3 encoder, 4-bit 3.0e-04 proves the quantized layout is rebuilt exactly
Z-Image DiT, float32 3.7e-07 8 layers
Flux2 DiT, float32 1.0e-06 2 double + 2 single blocks
VAE decode / tiled / encode 1e-05 or better both VAE classes
VAE BatchNorm statistics 0 exact

The bf16 figures for a full-depth DiT are larger - 2e-02 for both models - and that is rounding order, not a defect: at float32 the same code agrees to 1e-06, and the deviation grows with depth from a bf16-level 1e-04 per block. Comparing two bf16 implementations below 1e-02 is not meaningful for a 30-block network.

Performance on an M1 Max

Mac Studio, Apple M1 Max, 10-core CPU, 24-core GPU, 32 GB, macOS 26.5. Release build, 4-bit weights, batch 1, warm page cache.

run per step VAE decode
Z-Image, 512 px, 4 steps 2.6 s 1.5 s
klein, 512 px, 4 steps 2.3 s 0.9 s
klein, 1024 px, 4 steps 8.0 s 0.2 s
klein, 512 px + one 512 px reference 4.1 s 0.9 s

Reading the weights dominates a cold run and depends entirely on the storage: on an external volume delivering about 0.2 GiB/s, the 4-bit encoder took 22-32 s and a 4-bit DiT 18-35 s, against 60-107 s for the same weights in bf16. On internal storage expect these to be several times shorter. Compute is unaffected: the numbers above are steady state.

Stages run one at a time and release their weights, so peak memory is set by the largest single component rather than their sum. The 4-bit DiT is 2.03 GiB for klein and 3.40 GiB for Z-Image; activations at 1024 px add to that, and no separate peak measurement was made.

Limitations
  • Seeds are not compatible with the upstream pipelines, which use torch.Generator. The same prompt gives comparable images, never the same file.
  • Both models are distilled: CFG is not applied, and step counts are low by design (8 for Z-Image, 4 for klein).
  • Z-Image's Omni mode is not covered: it needs a SigLIP encoder that is not part of this bundle.
  • Tiled VAE decode is available for high resolutions and gives a result that differs slightly from a single pass, exactly as it does in diffusers.
  • Batches larger than one are not implemented.

Licences

Derivatives of three upstream models under two licences. Each folder carries the licence of its upstream model.

folder upstream licence
qwen3-tts-12hz-1.7b-base Qwen/Qwen3-TTS-12Hz-1.7B-Base Apache-2.0
qwen3-tts-12hz-1.7b-customvoice Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice Apache-2.0
qwen3-tts-12hz-1.7b-voicedesign Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign Apache-2.0

The shared text encoder is redistributed from inside the two image repositories, both Apache-2.0. Conversion tooling and these cards: MIT.

The upstream cards state usage restrictions and responsible-use commitments that redistribution does not repeal. See in particular the out-of-scope use section of the FLUX.2 [klein] card.

Downloads last month

-

Downloads are not tracked for this model. How to track
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support