BS-RoFormer for Core AI (.aimodel)

Apple Core AI conversion of BS-RoFormer β€” Wei-Tsung Lu, Ju-Chiang Wang, Qiuqiang Kong, Yun-Ning Hung, Music Source Separation with Band-Split RoPE Transformer (ByteDance, ICASSP 2024) β€” in lucidrains' implementation (MIT), with the four-stem weights model_bs_roformer_ep_17_sdr_9.6568.ckpt trained and released by ZFTurbo (Music-Source-Separation-Training, MIT): MUSDB18 test average 9.65 dB SDR (drums 11.61, vocals 11.08, bass 8.48, other 7.44). 132M parameters, drums, bass, other and vocals.

stems-bs_roformer-float32.aimodel (528 MB) holds the band-split transformer between the spectrogram and the mask at the model's chunk: input x [1, 1101, 4100], the STFT of 485,100 samples laid out per frame as (bin, channel, real/imaginary); output mask [1, 4, 1101, 4100]. The STFT (n_fft 2048, hop 441, periodic Hann, unnormalised), the complex mask, the DC zeroing, the inverse and the chunked inference with linear fades run in the host β€” swift-vocal-isolation (MIT), exposed by the stems CLI as --engine roformer. Nothing was re-authored, retrained or pruned.

Faithfulness

  • The exported network is asserted equal to model(chunk) before export.
  • Spectrogram layout and masked inverse: 151–155 dB PSNR against torch.
  • Each of 72 chunks (a 10-second clip and a 6-minute mix) through Core AI on the GPU: 74–110 dB PSNR against upstream's output; whole clips end to end 97–114 dB on every stem. Lower than a convolutional model's 140 dB because eight transformer layers of fp32 attention accumulate GPU-versus-CPU rounding β€” 1e-4 relative on the worst chunk, 1e-5 typical, repeatable run to run, far below anything audible.
  • The host runs every chunk twice and settles a mismatch with a third run, because Core AI's GPU was measured to return a slightly wrong result now and then on other models.

Use

hf download arraypress/stems-roformer --local-dir models
stems model install models/stems-bs_roformer-float32.aimodel
stems song.wav --engine roformer      # song-drums.wav, -bass, -other, -vocals

Requirements: macOS 27, Apple silicon. Reproduce with uv run Tools/export_roformer.py in the library repo, given ZFTurbo's config and checkpoint.

Licence and citation

MIT, as the implementation and the weights. Please cite:

W.-T. Lu, J.-C. Wang, Q. Kong, Y.-N. Hung. "Music Source Separation with Band-Split RoPE Transformer." ICASSP 2024.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support