Instructions to use kruatech/studio-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use kruatech/studio-mlx with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir studio-mlx kruatech/studio-mlx
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
|
Download README.md from kruatech/studio-mlx: direct link, hf CLI and curl.
- Browser
- Download file 10.5 kB
-
https://huggingface.co/kruatech/studio-mlx/resolve/main/README.md
- Command line
-
hf download hf://kruatech/studio-mlx/README.md
-
curl -L -o README.md https://huggingface.co/kruatech/studio-mlx/resolve/main/README.md
10.5 kB
| license: other | |
| license_name: mixed-apache-2.0-and-mit | |
| license_link: https://huggingface.co/kruatech/studio-mlx/blob/main/LICENSES.md | |
| library_name: mlx | |
| tags: | |
| - mlx | |
| - mlx-swift | |
| - apple-silicon | |
| - text-to-image | |
| - image-editing | |
| - text-to-music | |
| - quantized | |
| <div align="center"> | |
| # studio-mlx | |
| **Image and music models converted for native MLX inference on Apple silicon** | |
|    [](LICENSES.md) | |
| </div> | |
| Three models and the text encoder the two image models share. Every folder is a self-contained bundle with its own card, so you download only what you need. | |
| Weights are not retrained. Tensor layouts were adapted for MLX and the 4-bit folders are quantized, so parameters are mathematically equivalent to upstream rather than byte-identical. | |
| ## Quick start | |
| Pick one line. The image models need the shared encoder, the music model does not. | |
| ```bash | |
| pip install "huggingface_hub[hf_xet]" | |
| ``` | |
| The `-q4` folders sit at the top level rather than inside the bf16 ones, so an include pattern fetches one precision without dragging in the other. | |
| ## What is inside | |
| | folder | what it does | precision | size | licence | | |
| |----------------------------------------------------------------------|------------------------------------------|-----------|----------|------------| | |
| | [`qwen3-tts-12hz-1.7b-base`](qwen3-tts-12hz-1.7b-base) | Voice cloning from a reference recording | `bf16` | 4.23 GiB | Apache-2.0 | | |
| | [`qwen3-tts-12hz-1.7b-customvoice`](qwen3-tts-12hz-1.7b-customvoice) | Nine preset voices, ten languages | `bf16` | 4.21 GiB | Apache-2.0 | | |
| | [`qwen3-tts-12hz-1.7b-voicedesign`](qwen3-tts-12hz-1.7b-voicedesign) | Describe a voice in words and get it | `bf16` | 4.21 GiB | Apache-2.0 | | |
| ## bf16 or 4-bit | |
| **Use the 4-bit folders.** Quantization is MLX affine: 4 bits per weight with a bf16 scale and bias for every group of 64, about 4.5 bits per weight and 28-30% of the bf16 size. Normalizations, modulations and the encoder's embedding table stay in bf16, and matrix multiplication still runs in bf16 - the gain is memory and load time, not integer arithmetic. | |
| | | bf16 | 4-bit | | |
| |-----------------------------------|-----------|----------| | |
| | Z-Image DiT | 11.46 GiB | 3.40 GiB | | |
| | klein DiT | 7.22 GiB | 2.03 GiB | | |
| | shared encoder | 7.51 GiB | 2.58 GiB | | |
| | klein 512 px, 4 steps, wall clock | 170 s | 73 s | | |
| Output quality was indistinguishable: the same prompt and seed gave two clean images that differ the way two seeds differ. The bf16 folders exist as the accuracy reference a port can be checked against. | |
| ## Bundle layout | |
| ``` | |
| <folder>/ | |
| βββ manifest.json SHA-256, byte sizes and tensor counts for every file | |
| βββ config/ component configs copied from upstream | |
| βββ tokenizer/ BPE plus chat_format.json, where applicable | |
| βββ weights/ safetensors in MLX tensor layout | |
| ``` | |
| `manifest.json` records the quantization settings, so a loader rebuilds the exact layout without being told. Verify integrity before first use: the manifest carries SHA-256 per file, which catches corruption that size checks miss. | |
| `tokenizer/chat_format.json` holds the Qwen chat template already rendered into a prefix and a suffix for each pipeline, so a runtime needs no Jinja at all. | |
| <details> | |
| <summary><b>Tensor layout</b></summary> | |
| MLX convolutions expect channels last, PyTorch expects channels first, so convolution weights are permuted during conversion. Everything else keeps its upstream shape. | |
| | kind | PyTorch | MLX | | |
| |-----------------------------|---------------------|---------------------| | |
| | `Conv2d.weight` | `(out, in, kH, kW)` | `(out, kH, kW, in)` | | |
| | `Linear`, norms, embeddings | `(out, in)` | unchanged | | |
| No `weight_norm` and no 3-D convolutions appear in these models, so no folding was needed. Parameter names match the upstream checkpoints, so weights load without remapping. | |
| </details> | |
| <details> | |
| <summary><b>Verification numbers</b></summary> | |
| Every module was compared against the upstream reference on fixed inputs in float32. The metric is `rel_max = max|a-b| / max|a|`. | |
| | module | rel_max | reference | | |
| |-------------------------------------------------|-----------|---------------------------------| | |
| | Qwen3 encoder, hidden_states[-2] | 2.7e-07 | transformers | | |
| | Qwen3 encoder, layers 9/18/27 with padding mask | 6.0e-07 | transformers | | |
| | Z-Image DiT | 4.2e-06 | diffusers | | |
| | Flux2 DiT | 3.7e-07 | diffusers | | |
| | VAE decoder | 1.3e-05 | diffusers | | |
| | VAE encoder | 5.3e-06 | diffusers | | |
| | VAE tiled decode | 5.1e-06 | diffusers tiled_decode | | |
| | sigma schedule, static shift | 3.2e-08 | FlowMatchEulerDiscreteScheduler | | |
| | sigma schedule, exponential dynamic shift | 7.7e-08 | FlowMatchEulerDiscreteScheduler | | |
| | position ids, latent packing, patchify | bit-exact | pipeline helpers | | |
| A Swift MLX implementation was then checked against the Python one: | |
| | module | rel_max | note | | |
| |------------------------------|-----------------|------------------------------------------------| | |
| | tokenization, both pipelines | bit-exact | same ids | | |
| | Qwen3 encoder, bf16 | 2.2e-04 | Z-Image branch | | |
| | Qwen3 encoder, bf16 | 4.6e-03 | klein branch, 512 tokens with mask | | |
| | Qwen3 encoder, 4-bit | 3.0e-04 | proves the quantized layout is rebuilt exactly | | |
| | Z-Image DiT, float32 | 3.7e-07 | 8 layers | | |
| | Flux2 DiT, float32 | 1.0e-06 | 2 double + 2 single blocks | | |
| | VAE decode / tiled / encode | 1e-05 or better | both VAE classes | | |
| | VAE BatchNorm statistics | 0 | exact | | |
| The bf16 figures for a full-depth DiT are larger - 2e-02 for both models - and that is rounding order, not a defect: at float32 the same code agrees to 1e-06, and the deviation grows with depth from a bf16-level 1e-04 per block. Comparing two bf16 implementations below 1e-02 is not meaningful for a 30-block network. | |
| </details> | |
| <details> | |
| <summary><b>Performance on an M1 Max</b></summary> | |
| Mac Studio, Apple M1 Max, 10-core CPU, 24-core GPU, 32 GB, macOS 26.5. Release build, 4-bit weights, batch 1, warm page cache. | |
| | run | per step | VAE decode | | |
| |--------------------------------------|----------|------------| | |
| | Z-Image, 512 px, 4 steps | 2.6 s | 1.5 s | | |
| | klein, 512 px, 4 steps | 2.3 s | 0.9 s | | |
| | klein, 1024 px, 4 steps | 8.0 s | 0.2 s | | |
| | klein, 512 px + one 512 px reference | 4.1 s | 0.9 s | | |
| Reading the weights dominates a cold run and depends entirely on the storage: on an external volume delivering about 0.2 GiB/s, the 4-bit encoder took 22-32 s and a 4-bit DiT 18-35 s, against 60-107 s for the same weights in bf16. On internal storage expect these to be several times shorter. Compute is unaffected: the numbers above are steady state. | |
| Stages run one at a time and release their weights, so peak memory is set by the largest single component rather than their sum. The 4-bit DiT is 2.03 GiB for klein and 3.40 GiB for Z-Image; activations at 1024 px add to that, and no separate peak measurement was made. | |
| </details> | |
| <details> | |
| <summary><b>Limitations</b></summary> | |
| - Seeds are not compatible with the upstream pipelines, which use `torch.Generator`. The same prompt gives comparable images, never the same file. | |
| - Both models are distilled: CFG is not applied, and step counts are low by design (8 for Z-Image, 4 for klein). | |
| - Z-Image's Omni mode is not covered: it needs a SigLIP encoder that is not part of this bundle. | |
| - Tiled VAE decode is available for high resolutions and gives a result that differs slightly from a single pass, exactly as it does in diffusers. | |
| - Batches larger than one are not implemented. | |
| </details> | |
| ## Licences | |
| Derivatives of three upstream models under two licences. Each folder carries the licence of its upstream model. | |
| | folder | upstream | licence | | |
| |-----------------------------------|-----------------------------------------------------------------------------------------------------|-----------------------------------------------------------| | |
| | `qwen3-tts-12hz-1.7b-base` | [Qwen/Qwen3-TTS-12Hz-1.7B-Base](https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-Base) | [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0) | | |
| | `qwen3-tts-12hz-1.7b-customvoice` | [Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice](https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice) | [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0) | | |
| | `qwen3-tts-12hz-1.7b-voicedesign` | [Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign](https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign) | [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0) | | |
| The shared text encoder is redistributed from inside the two image repositories, both Apache-2.0. Conversion tooling and these cards: MIT. | |
| The upstream cards state usage restrictions and responsible-use commitments that redistribution does not repeal. See in particular the out-of-scope use section of the FLUX.2 [klein] card. | |