Step-Audio-EditX Technical Report
Paper • 2511.03601 • Published • 30
How to use appautomaton/step-audio-editx-8bit-mlx with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download appautomaton/step-audio-editx-8bit-mlx --local-dir step-audio-editx-8bit-mlx
This repository contains a self-contained pure-MLX int8 conversion of
Step-Audio-EditX for local voice cloning and expressive audio editing on
Apple Silicon. All pipeline components are stored as .safetensors — no
PyTorch, ONNX, or NumPy files are required at inference time.
stepfun-ai/Step-Audio-EditXThis bundle is self-contained — all weights are packaged in one repository.
| File | Component | Format | Size |
|---|---|---|---|
model.safetensors |
Step1 LM (3.5B params) | int8 | 3.5 GB |
flow-model.safetensors |
Flow model (DiT + conformer) | int8 | 181 MB |
vq02.safetensors |
VQ02 audio tokenizer | int8 | 162 MB |
vq06.safetensors |
VQ06 audio tokenizer | bf16 | 249 MB |
hift.safetensors |
HiFT vocoder | bf16 | 40 MB |
campplus.safetensors |
CampPlus speaker embedding | bf16 | 13 MB |
flow-conditioner.safetensors |
Flow conditioner | bf16 | 2.5 MB |
config.json |
Step1 LM config + quantization | JSON | — |
flow-model-config.json |
Flow model config | JSON | — |
vq02-config.json, vq06-config.json |
Tokenizer configs | JSON | — |
step-audio-tokenizer-assets.safetensors |
VQ02 codebook + CMVN | FP32 | ~2 MB |
step-audio-tokenizer-config.json |
Tokenizer runtime config | JSON | — |
frontend-config.json |
Prompt mel frontend config | JSON | — |
hift-config.json, campplus-config.json, flow-conditioner-config.json |
Component configs | JSON | — |
tokenizer.json, tokenizer_config.json |
Step1 tokenizer | JSON | — |
Download the bundle:
hf download appautomaton/step-audio-editx-8bit-mlx \
--local-dir models/stepfun/step_audio_editx/mlx-int8
Voice cloning:
python scripts/generate/step_audio_editx.py \
--prompt-audio reference.wav \
--prompt-text "Transcript of reference audio." \
-o cloned.wav \
clone --target-text "New speech in the cloned voice."
Audio editing (change emotion):
python scripts/generate/step_audio_editx.py \
--prompt-audio input.wav \
--prompt-text "Transcript of input audio." \
-o happy.wav \
edit --edit-type emotion --edit-info happy
| Edit type | Description | --edit-info examples |
|---|---|---|
emotion |
Change the emotion of speech | happy, sad, angry, surprised |
style |
Change speaking style | whispering, broadcasting, formal |
speed |
Change speaking speed | fast, slow |
denoise |
Remove noise from audio | not used |
vad |
Remove silences from audio | not used |
paralinguistic |
Add non-verbal sounds | requires --target-text |
Five-stage pipeline, all running pure MLX with bf16 activations:
The VQ02 and VQ06 tokenizers encode reference audio into dual codebook tokens consumed by Step1.
mlx-speechstepfun-ai/Step-Audio-EditXApache 2.0 — following the upstream license published with
stepfun-ai/Step-Audio-EditX.
Quantized
Base model
stepfun-ai/Step-Audio-EditX