Instructions to use kruatech/Stable-Audio-3-Medium-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use kruatech/Stable-Audio-3-Medium-MLX with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir Stable-Audio-3-Medium-MLX kruatech/Stable-Audio-3-Medium-MLX
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
You need to agree to share your contact information to access this model
This repository is publicly accessible, but you have to accept the conditions to access its files and content.
These weights are a converted version of stable-audio-3-medium and are governed by the Stability AI Community License Agreement, a copy of which is in this repository. The checkpoint also contains Google's T5Gemma text encoder, redistributed under the Gemma Terms of Use, which apply in addition - including the use restrictions in their section 3.2. By requesting access you agree to be bound by both. In particular: the Stability licence is free for research and non-commercial use, and for commercial use only while you and your affiliates generate less than US $1,000,000 in annual revenue - above that figure the licence terminates and you must request one from Stability; commercial use below the threshold still requires registration with Stability; you may not use the model, a derivative, or its output to create or improve any foundational generative AI model; and if you redistribute you must pass on both sets of terms, carry the required notices, state what you changed, and display "Powered by Stability AI".
Log in or Sign Up to review the conditions and access this model content.
Stable-Audio-3-Medium-MLX
stable-audio-3-medium converted for Apple MLX, runnable by MLXBundle without Python at runtime.
This Stability AI Model is licensed under the Stability AI Community License, Copyright © Stability AI Ltd. All Rights Reserved
Powered by Stability AI
Modified: this is a modified version of stable-audio-3-medium.
Why this repository is gated
Two sets of terms apply, and both have to reach you. Stability's licence requires that a redistributor provide a copy of the agreement (§4a). Google's Gemma Terms go further for the T5Gemma encoder inside the checkpoint: a copy of the terms must be given to recipients, the section 3.2 use restrictions must be passed on as an enforceable provision, modified files must be marked, and a notice must travel with the model. Stability themselves gate the original for the same reason.
What you are agreeing to
LICENSE.md and LICENSE_GEMMA.md are in this repository. The parts that bite:
Revenue threshold terminates the licence (§3). "If at any time You or Your Affiliate(s)... generate more than USD $1,000,000 in annual revenue... any licenses granted to You under this Agreement shall terminate as of such date." Not "you need a different licence" - the grant ends, and §4f then requires deleting the materials. A licence must be requested from stability.ai/enterprise, which Stability may grant at its discretion.
Commercial use below the threshold requires registration at stability.ai/community-license (§3). Research and non-commercial use does not. "Commercial Purpose" includes hosted services, APIs, and your organisation's internal operations.
No training other foundation models (§4b). You may not use the model, a derivative, or its output to create or improve any foundational generative AI model. Stability's Acceptable Use Policy is part of the licence.
If you redistribute (§4a). Provide the agreement; carry the notice above in
a Notice file; display "Powered by Stability AI" prominently on a related
website, user interface, blog post, about page, or product documentation; and
state in the Notice file that the model was changed and how.
Gemma Terms of Use. Apply to the T5Gemma encoder in addition to everything above, including the section 3.2 use restrictions, and carry their own obligations when you pass the model on. Satisfying Stability's licence does not satisfy Google's.
Governing law is California. No trademark licence beyond the §4a attribution.
What conversion changed
- Tensor layouts transformed for MLX's convolution convention.
- Components split into separate files with an index.
- Everything kept in float32 except the T5Gemma encoder, which is bf16 as upstream. Nothing quantized, so no lossy step here.
No training, fine-tuning, distillation or merging, and no learned parameters added. The bundle carries the DiT, the SAME autoencoder, the T5Gemma encoder, the duration conditioner and the tokenizer, and is self-contained at 9.7 GiB.
Verified against Stability's own stable_audio_tools rather than against
diffusers, which does not reproduce this model correctly: the DiT agrees to
1.066e-05 with a correlation of 1.000000, the decoder to a correlation of
0.99999, the text encoder to 6.5e-03 on real tokens at a correlation of
0.999900. End to end, a prompt asking for 120 bpm comes back at a 0.50 s
autocorrelation period with a peak of 0.94.
Using it
mlxbundle-cli audio ~/models/Stable-Audio-3-Medium-MLX \
~/models/Stable-Audio-3-Medium-MLX/tokenizer \
~/models/Stable-Audio-3-Medium-MLX/config/scheduler_config.json \
"solo grand piano playing a slow melody in C major" piano.wav 10
Note on the scheduler: Stability publish no diffusers-format scheduler config
for this checkpoint, so config/scheduler_config.json in this bundle stands in
for one. It records where each of its three values came from - two read from the
reference implementation, one chosen and then confirmed by measurement.
Downloading a gated repository needs a Hugging Face token with read access; see the MLXBundle documentation.
Disclaimer
An independent conversion. No warranty of any kind, and no representation that any particular use of it complies with either set of terms or with any law. The obligations above are yours. If your use is commercial, read the agreements rather than this summary.
Quantized