|
Download README.md from arraypress/stems-roformer: direct link, hf CLI and curl.
- Browser
- Download file 2.75 kB
-
https://huggingface.co/arraypress/stems-roformer/resolve/main/README.md
- Command line
-
hf download hf://arraypress/stems-roformer/README.md
-
curl -L -o README.md https://huggingface.co/arraypress/stems-roformer/resolve/main/README.md
2.75 kB
| license: mit | |
| tags: | |
| - music-source-separation | |
| - stems | |
| - bs-roformer | |
| - core-ai | |
| - aimodel | |
| - apple-silicon | |
| - macos | |
| library_name: swift-vocal-isolation | |
| # BS-RoFormer for Core AI (`.aimodel`) | |
| Apple Core AI conversion of **BS-RoFormer** β Wei-Tsung Lu, Ju-Chiang Wang, Qiuqiang Kong, | |
| Yun-Ning Hung, *Music Source Separation with Band-Split RoPE Transformer* (ByteDance, ICASSP | |
| 2024) β in lucidrains' implementation (MIT), with the four-stem weights | |
| `model_bs_roformer_ep_17_sdr_9.6568.ckpt` trained and released by ZFTurbo | |
| ([Music-Source-Separation-Training](https://github.com/ZFTurbo/Music-Source-Separation-Training), | |
| MIT): MUSDB18 test average 9.65 dB SDR (drums 11.61, vocals 11.08, bass 8.48, other 7.44). | |
| 132M parameters, drums, bass, other and vocals. | |
| `stems-bs_roformer-float32.aimodel` (528 MB) holds the band-split transformer between the | |
| spectrogram and the mask at the model's chunk: input `x` `[1, 1101, 4100]`, the STFT of 485,100 | |
| samples laid out per frame as (bin, channel, real/imaginary); output `mask` `[1, 4, 1101, 4100]`. | |
| The STFT (n_fft 2048, hop 441, periodic Hann, unnormalised), the complex mask, the DC zeroing, the | |
| inverse and the chunked inference with linear fades run in the host β | |
| [swift-vocal-isolation](https://github.com/arraypress/swift-vocal-isolation) (MIT), exposed by the | |
| `stems` CLI as `--engine roformer`. Nothing was re-authored, retrained or pruned. | |
| ## Faithfulness | |
| - The exported network is asserted equal to `model(chunk)` before export. | |
| - Spectrogram layout and masked inverse: 151β155 dB PSNR against torch. | |
| - Each of 72 chunks (a 10-second clip and a 6-minute mix) through Core AI on the GPU: 74β110 dB | |
| PSNR against upstream's output; whole clips end to end 97β114 dB on every stem. Lower than a | |
| convolutional model's 140 dB because eight transformer layers of fp32 attention accumulate | |
| GPU-versus-CPU rounding β 1e-4 relative on the worst chunk, 1e-5 typical, repeatable run to | |
| run, far below anything audible. | |
| - The host runs every chunk twice and settles a mismatch with a third run, because Core AI's GPU | |
| was measured to return a slightly wrong result now and then on other models. | |
| ## Use | |
| ```sh | |
| hf download arraypress/stems-roformer --local-dir models | |
| stems model install models/stems-bs_roformer-float32.aimodel | |
| stems song.wav --engine roformer # song-drums.wav, -bass, -other, -vocals | |
| ``` | |
| Requirements: macOS 27, Apple silicon. Reproduce with `uv run Tools/export_roformer.py` in the | |
| library repo, given ZFTurbo's config and checkpoint. | |
| ## Licence and citation | |
| MIT, as the implementation and the weights. Please cite: | |
| > W.-T. Lu, J.-C. Wang, Q. Kong, Y.-N. Hung. "Music Source Separation with Band-Split RoPE | |
| > Transformer." ICASSP 2024. | |