model card
Browse files
README.md
ADDED
|
@@ -0,0 +1,58 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: mit
|
| 3 |
+
tags:
|
| 4 |
+
- music-source-separation
|
| 5 |
+
- stems
|
| 6 |
+
- bs-roformer
|
| 7 |
+
- core-ai
|
| 8 |
+
- aimodel
|
| 9 |
+
- apple-silicon
|
| 10 |
+
- macos
|
| 11 |
+
library_name: swift-vocal-isolation
|
| 12 |
+
---
|
| 13 |
+
|
| 14 |
+
# BS-RoFormer for Core AI (`.aimodel`)
|
| 15 |
+
|
| 16 |
+
Apple Core AI conversion of **BS-RoFormer** — Wei-Tsung Lu, Ju-Chiang Wang, Qiuqiang Kong,
|
| 17 |
+
Yun-Ning Hung, *Music Source Separation with Band-Split RoPE Transformer* (ByteDance, ICASSP
|
| 18 |
+
2024) — in lucidrains' implementation (MIT), with the four-stem weights
|
| 19 |
+
`model_bs_roformer_ep_17_sdr_9.6568.ckpt` trained and released by ZFTurbo
|
| 20 |
+
([Music-Source-Separation-Training](https://github.com/ZFTurbo/Music-Source-Separation-Training),
|
| 21 |
+
MIT): MUSDB18 test average 9.65 dB SDR (drums 11.61, vocals 11.08, bass 8.48, other 7.44).
|
| 22 |
+
132M parameters, drums, bass, other and vocals.
|
| 23 |
+
|
| 24 |
+
`stems-bs_roformer-float32.aimodel` (528 MB) holds the band-split transformer between the
|
| 25 |
+
spectrogram and the mask at the model's chunk: input `x` `[1, 1101, 4100]`, the STFT of 485,100
|
| 26 |
+
samples laid out per frame as (bin, channel, real/imaginary); output `mask` `[1, 4, 1101, 4100]`.
|
| 27 |
+
The STFT (n_fft 2048, hop 441, periodic Hann, unnormalised), the complex mask, the DC zeroing, the
|
| 28 |
+
inverse and the chunked inference with linear fades run in the host —
|
| 29 |
+
[swift-vocal-isolation](https://github.com/arraypress/swift-vocal-isolation) (MIT), exposed by the
|
| 30 |
+
`stems` CLI as `--engine roformer`. Nothing was re-authored, retrained or pruned.
|
| 31 |
+
|
| 32 |
+
## Faithfulness
|
| 33 |
+
|
| 34 |
+
- The exported network is asserted equal to `model(chunk)` before export.
|
| 35 |
+
- Spectrogram layout and masked inverse: 151–155 dB PSNR against torch.
|
| 36 |
+
- Each chunk through Core AI on the GPU: 96–110 dB PSNR against upstream's output; whole clips
|
| 37 |
+
end to end 97–107 dB. Lower than a convolutional model's 140 dB because eight transformer
|
| 38 |
+
layers of fp32 attention accumulate GPU-versus-CPU rounding, still far below anything audible.
|
| 39 |
+
- The host runs every chunk twice and settles a mismatch with a third run, because Core AI's GPU
|
| 40 |
+
was measured to return a slightly wrong result now and then on other models.
|
| 41 |
+
|
| 42 |
+
## Use
|
| 43 |
+
|
| 44 |
+
```sh
|
| 45 |
+
hf download arraypress/stems-roformer --local-dir models
|
| 46 |
+
stems model install models/stems-bs_roformer-float32.aimodel
|
| 47 |
+
stems song.wav --engine roformer # song-drums.wav, -bass, -other, -vocals
|
| 48 |
+
```
|
| 49 |
+
|
| 50 |
+
Requirements: macOS 27, Apple silicon. Reproduce with `uv run Tools/export_roformer.py` in the
|
| 51 |
+
library repo, given ZFTurbo's config and checkpoint.
|
| 52 |
+
|
| 53 |
+
## Licence and citation
|
| 54 |
+
|
| 55 |
+
MIT, as the implementation and the weights. Please cite:
|
| 56 |
+
|
| 57 |
+
> W.-T. Lu, J.-C. Wang, Q. Kong, Y.-N. Hung. "Music Source Separation with Band-Split RoPE
|
| 58 |
+
> Transformer." ICASSP 2024.
|