Instructions to use kruatech/studio-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use kruatech/studio-mlx with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir studio-mlx kruatech/studio-mlx
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Three models and the text encoder the two image models share. Every folder is a self-contained bundle with its own card, so you download only what you need.
Weights are not retrained. Tensor layouts were adapted for MLX and the 4-bit folders are quantized, so parameters are mathematically equivalent to upstream rather than byte-identical.
Quick start
Pick one line. The image models need the shared encoder, the music model does not.
pip install "huggingface_hub[hf_xet]"
The -q4 folders sit at the top level rather than inside the bf16 ones, so an include pattern fetches one precision without dragging in the other.
What is inside
| folder | what it does | precision | size | licence |
|---|---|---|---|---|
qwen3-tts-12hz-1.7b-base |
Voice cloning from a reference recording | bf16 |
4.23 GiB | Apache-2.0 |
qwen3-tts-12hz-1.7b-customvoice |
Nine preset voices, ten languages | bf16 |
4.21 GiB | Apache-2.0 |
qwen3-tts-12hz-1.7b-voicedesign |
Describe a voice in words and get it | bf16 |
4.21 GiB | Apache-2.0 |
bf16 or 4-bit
Use the 4-bit folders. Quantization is MLX affine: 4 bits per weight with a bf16 scale and bias for every group of 64, about 4.5 bits per weight and 28-30% of the bf16 size. Normalizations, modulations and the encoder's embedding table stay in bf16, and matrix multiplication still runs in bf16 - the gain is memory and load time, not integer arithmetic.
| bf16 | 4-bit | |
|---|---|---|
| Z-Image DiT | 11.46 GiB | 3.40 GiB |
| klein DiT | 7.22 GiB | 2.03 GiB |
| shared encoder | 7.51 GiB | 2.58 GiB |
| klein 512 px, 4 steps, wall clock | 170 s | 73 s |
Output quality was indistinguishable: the same prompt and seed gave two clean images that differ the way two seeds differ. The bf16 folders exist as the accuracy reference a port can be checked against.
Bundle layout
<folder>/
βββ manifest.json SHA-256, byte sizes and tensor counts for every file
βββ config/ component configs copied from upstream
βββ tokenizer/ BPE plus chat_format.json, where applicable
βββ weights/ safetensors in MLX tensor layout
manifest.json records the quantization settings, so a loader rebuilds the exact layout without being told. Verify integrity before first use: the manifest carries SHA-256 per file, which catches corruption that size checks miss.
tokenizer/chat_format.json holds the Qwen chat template already rendered into a prefix and a suffix for each pipeline, so a runtime needs no Jinja at all.
Tensor layout
MLX convolutions expect channels last, PyTorch expects channels first, so convolution weights are permuted during conversion. Everything else keeps its upstream shape.
| kind | PyTorch | MLX |
|---|---|---|
Conv2d.weight |
(out, in, kH, kW) |
(out, kH, kW, in) |
Linear, norms, embeddings |
(out, in) |
unchanged |
No weight_norm and no 3-D convolutions appear in these models, so no folding was needed. Parameter names match the upstream checkpoints, so weights load without remapping.
Verification numbers
Every module was compared against the upstream reference on fixed inputs in float32. The metric is rel_max = max|a-b| / max|a|.
| module | rel_max | reference |
|---|---|---|
| Qwen3 encoder, hidden_states[-2] | 2.7e-07 | transformers |
| Qwen3 encoder, layers 9/18/27 with padding mask | 6.0e-07 | transformers |
| Z-Image DiT | 4.2e-06 | diffusers |
| Flux2 DiT | 3.7e-07 | diffusers |
| VAE decoder | 1.3e-05 | diffusers |
| VAE encoder | 5.3e-06 | diffusers |
| VAE tiled decode | 5.1e-06 | diffusers tiled_decode |
| sigma schedule, static shift | 3.2e-08 | FlowMatchEulerDiscreteScheduler |
| sigma schedule, exponential dynamic shift | 7.7e-08 | FlowMatchEulerDiscreteScheduler |
| position ids, latent packing, patchify | bit-exact | pipeline helpers |
A Swift MLX implementation was then checked against the Python one:
| module | rel_max | note |
|---|---|---|
| tokenization, both pipelines | bit-exact | same ids |
| Qwen3 encoder, bf16 | 2.2e-04 | Z-Image branch |
| Qwen3 encoder, bf16 | 4.6e-03 | klein branch, 512 tokens with mask |
| Qwen3 encoder, 4-bit | 3.0e-04 | proves the quantized layout is rebuilt exactly |
| Z-Image DiT, float32 | 3.7e-07 | 8 layers |
| Flux2 DiT, float32 | 1.0e-06 | 2 double + 2 single blocks |
| VAE decode / tiled / encode | 1e-05 or better | both VAE classes |
| VAE BatchNorm statistics | 0 | exact |
The bf16 figures for a full-depth DiT are larger - 2e-02 for both models - and that is rounding order, not a defect: at float32 the same code agrees to 1e-06, and the deviation grows with depth from a bf16-level 1e-04 per block. Comparing two bf16 implementations below 1e-02 is not meaningful for a 30-block network.
Performance on an M1 Max
Mac Studio, Apple M1 Max, 10-core CPU, 24-core GPU, 32 GB, macOS 26.5. Release build, 4-bit weights, batch 1, warm page cache.
| run | per step | VAE decode |
|---|---|---|
| Z-Image, 512 px, 4 steps | 2.6 s | 1.5 s |
| klein, 512 px, 4 steps | 2.3 s | 0.9 s |
| klein, 1024 px, 4 steps | 8.0 s | 0.2 s |
| klein, 512 px + one 512 px reference | 4.1 s | 0.9 s |
Reading the weights dominates a cold run and depends entirely on the storage: on an external volume delivering about 0.2 GiB/s, the 4-bit encoder took 22-32 s and a 4-bit DiT 18-35 s, against 60-107 s for the same weights in bf16. On internal storage expect these to be several times shorter. Compute is unaffected: the numbers above are steady state.
Stages run one at a time and release their weights, so peak memory is set by the largest single component rather than their sum. The 4-bit DiT is 2.03 GiB for klein and 3.40 GiB for Z-Image; activations at 1024 px add to that, and no separate peak measurement was made.
Limitations
- Seeds are not compatible with the upstream pipelines, which use
torch.Generator. The same prompt gives comparable images, never the same file. - Both models are distilled: CFG is not applied, and step counts are low by design (8 for Z-Image, 4 for klein).
- Z-Image's Omni mode is not covered: it needs a SigLIP encoder that is not part of this bundle.
- Tiled VAE decode is available for high resolutions and gives a result that differs slightly from a single pass, exactly as it does in diffusers.
- Batches larger than one are not implemented.
Licences
Derivatives of three upstream models under two licences. Each folder carries the licence of its upstream model.
| folder | upstream | licence |
|---|---|---|
qwen3-tts-12hz-1.7b-base |
Qwen/Qwen3-TTS-12Hz-1.7B-Base | Apache-2.0 |
qwen3-tts-12hz-1.7b-customvoice |
Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice | Apache-2.0 |
qwen3-tts-12hz-1.7b-voicedesign |
Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign | Apache-2.0 |
The shared text encoder is redistributed from inside the two image repositories, both Apache-2.0. Conversion tooling and these cards: MIT.
The upstream cards state usage restrictions and responsible-use commitments that redistribution does not repeal. See in particular the out-of-scope use section of the FLUX.2 [klein] card.
Quantized