Step-Audio-EditX β€” bf16 MLX bundle

Step-Audio-EditX (StepFun, Apache-2.0): a 3 B audio LLM that re-delivers an existing take β€” emotion, speaking style, inserted paralinguistics (laughter, sighs, breaths), denoise, silence trim β€” in the same voice, plus zero-shot cloning. This is the whole pipeline in one directory of MLX safetensors, bf16 throughout, no quantisation, converted from the stock checkpoint:

file component source weights
model.safetensors (7.06 GB) + config.json step1 LM β€” 32 Γ— 3072, 48 heads / 4 KV groups, sqrt-ALiBi stepfun-ai/Step-Audio-EditX
vq02.safetensors + vq02-config.json Paraformer-large zh-cantonese-en online encoder (the 16.7 Hz linguistic tokenizer) FunASR dengcunqin/…-vocab8501-online (FunASR model licence)
vq06.safetensors + vq06-config.json S3 v1 semantic tokenizer (25 Hz, 4 096 codes) stepfun-ai/Step-Audio-Tokenizer
step-audio-tokenizer-assets.safetensors the k-means linguistic codebook + CMVN stepfun-ai/Step-Audio-Tokenizer
flow-model.safetensors, flow-conditioner.safetensors + configs upsample conformer + DiT flow matching (24 kHz mel) Step-Audio-EditX/CosyVoice-300M-25Hz/flow.pt (StepFun-trained)
hift.safetensors + hift-config.json HiFT vocoder …/hift.pt (StepFun-trained)
campplus.safetensors + campplus-config.json CAM++ speaker embedding …/campplus.onnx (Alibaba, Apache-2.0)
tokenizer.json, tokenizer_config.json the sentencepiece Unigram text tokenizer derived from tokenizer.model
frontend-config.json the 24 kHz mel front end β€”

Layout and conversion: appautomaton/mlx-speech (MIT) scripts/convert/step_audio_editx.py, run with a --no-quantize patch so every component stays bf16; tokenizer.json was added from tokenizer.model (LlamaTokenizerFast).

Consumers

  • Swift: xocialize/mlx-step-audio-editx-swift β€” EditXPipeline.load(bundle: EditXBundle(root: dir)); quantises the LM to int8 at load when asked. Every stage is parity-locked against StepFun's torch code (its PORTING-SPEC.md).
  • Python: mlx-speech β€” StepAudioEditXModel.from_dir(dir) (its resampler should be a windowed sinc; the shipped np.interp perturbs the 16 kHz tokenizers β€” see the Swift package's notes).

Notes

  • The Paraformer front end in the upstream code dithers with an unseeded RNG at inference (only ~96 % of its own codes agree run to run). Both consumers above run it deterministically (dither off).
  • Measured on an M5 Max with the Swift consumer: β‰ˆ 9.4 GB resident, β‰ˆ 1.9 GB per edit, real-time factor 0.72–0.75 (0.49–0.58 with the LM at int8). Content, voice and quality of the edits equal the torch reference's on the evaluation set (mlxengine-audio E19).

Licence

Apache-2.0 (StepFun) for the LM, flow, HiFT and the S3 tokenizer; the Paraformer encoder under the FunASR model licence; CAM++ Apache-2.0 (Alibaba / 3D-Speaker). Converted and re-hosted by xocialize for mlx-community.

Downloads last month
-
Safetensors
Model size
4B params
Tensor type
BF16
Β·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for mlx-community/Step-Audio-EditX-bf16

Finetuned
(5)
this model