Instructions to use mlx-community/Step-Audio-EditX-bf16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use mlx-community/Step-Audio-EditX-bf16 with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download mlx-community/Step-Audio-EditX-bf16 --local-dir Step-Audio-EditX-bf16
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Step-Audio-EditX β bf16 MLX bundle
Step-Audio-EditX (StepFun, Apache-2.0): a 3 B audio LLM that re-delivers an existing take β emotion, speaking style, inserted paralinguistics (laughter, sighs, breaths), denoise, silence trim β in the same voice, plus zero-shot cloning. This is the whole pipeline in one directory of MLX safetensors, bf16 throughout, no quantisation, converted from the stock checkpoint:
| file | component | source weights |
|---|---|---|
model.safetensors (7.06 GB) + config.json |
step1 LM β 32 Γ 3072, 48 heads / 4 KV groups, sqrt-ALiBi | stepfun-ai/Step-Audio-EditX |
vq02.safetensors + vq02-config.json |
Paraformer-large zh-cantonese-en online encoder (the 16.7 Hz linguistic tokenizer) | FunASR dengcunqin/β¦-vocab8501-online (FunASR model licence) |
vq06.safetensors + vq06-config.json |
S3 v1 semantic tokenizer (25 Hz, 4 096 codes) | stepfun-ai/Step-Audio-Tokenizer |
step-audio-tokenizer-assets.safetensors |
the k-means linguistic codebook + CMVN | stepfun-ai/Step-Audio-Tokenizer |
flow-model.safetensors, flow-conditioner.safetensors + configs |
upsample conformer + DiT flow matching (24 kHz mel) | Step-Audio-EditX/CosyVoice-300M-25Hz/flow.pt (StepFun-trained) |
hift.safetensors + hift-config.json |
HiFT vocoder | β¦/hift.pt (StepFun-trained) |
campplus.safetensors + campplus-config.json |
CAM++ speaker embedding | β¦/campplus.onnx (Alibaba, Apache-2.0) |
tokenizer.json, tokenizer_config.json |
the sentencepiece Unigram text tokenizer | derived from tokenizer.model |
frontend-config.json |
the 24 kHz mel front end | β |
Layout and conversion: appautomaton/mlx-speech (MIT)
scripts/convert/step_audio_editx.py, run with a --no-quantize patch so every component stays bf16; tokenizer.json
was added from tokenizer.model (LlamaTokenizerFast).
Consumers
- Swift: xocialize/mlx-step-audio-editx-swift β
EditXPipeline.load(bundle: EditXBundle(root: dir)); quantises the LM to int8 at load when asked. Every stage is parity-locked against StepFun's torch code (itsPORTING-SPEC.md). - Python:
mlx-speechβStepAudioEditXModel.from_dir(dir)(its resampler should be a windowed sinc; the shippednp.interpperturbs the 16 kHz tokenizers β see the Swift package's notes).
Notes
- The Paraformer front end in the upstream code dithers with an unseeded RNG at inference (only ~96 % of its own codes agree run to run). Both consumers above run it deterministically (dither off).
- Measured on an M5 Max with the Swift consumer: β 9.4 GB resident, β 1.9 GB per edit, real-time factor 0.72β0.75 (0.49β0.58 with the LM at int8). Content, voice and quality of the edits equal the torch reference's on the evaluation set (mlxengine-audio E19).
Licence
Apache-2.0 (StepFun) for the LM, flow, HiFT and the S3 tokenizer; the Paraformer encoder under the FunASR model licence; CAM++ Apache-2.0 (Alibaba / 3D-Speaker). Converted and re-hosted by xocialize for mlx-community.
- Downloads last month
- -
Quantized
Model tree for mlx-community/Step-Audio-EditX-bf16
Base model
stepfun-ai/Step-Audio-EditX