File size: 2,748 Bytes
288f41e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
57c4e4d
 
 
 
 
288f41e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
---
license: mit
tags:
  - music-source-separation
  - stems
  - bs-roformer
  - core-ai
  - aimodel
  - apple-silicon
  - macos
library_name: swift-vocal-isolation
---

# BS-RoFormer for Core AI (`.aimodel`)

Apple Core AI conversion of **BS-RoFormer** — Wei-Tsung Lu, Ju-Chiang Wang, Qiuqiang Kong,
Yun-Ning Hung, *Music Source Separation with Band-Split RoPE Transformer* (ByteDance, ICASSP
2024) — in lucidrains' implementation (MIT), with the four-stem weights
`model_bs_roformer_ep_17_sdr_9.6568.ckpt` trained and released by ZFTurbo
([Music-Source-Separation-Training](https://github.com/ZFTurbo/Music-Source-Separation-Training),
MIT): MUSDB18 test average 9.65 dB SDR (drums 11.61, vocals 11.08, bass 8.48, other 7.44).
132M parameters, drums, bass, other and vocals.

`stems-bs_roformer-float32.aimodel` (528 MB) holds the band-split transformer between the
spectrogram and the mask at the model's chunk: input `x` `[1, 1101, 4100]`, the STFT of 485,100
samples laid out per frame as (bin, channel, real/imaginary); output `mask` `[1, 4, 1101, 4100]`.
The STFT (n_fft 2048, hop 441, periodic Hann, unnormalised), the complex mask, the DC zeroing, the
inverse and the chunked inference with linear fades run in the host —
[swift-vocal-isolation](https://github.com/arraypress/swift-vocal-isolation) (MIT), exposed by the
`stems` CLI as `--engine roformer`. Nothing was re-authored, retrained or pruned.

## Faithfulness

- The exported network is asserted equal to `model(chunk)` before export.
- Spectrogram layout and masked inverse: 151–155 dB PSNR against torch.
- Each of 72 chunks (a 10-second clip and a 6-minute mix) through Core AI on the GPU: 74–110 dB
  PSNR against upstream's output; whole clips end to end 97–114 dB on every stem. Lower than a
  convolutional model's 140 dB because eight transformer layers of fp32 attention accumulate
  GPU-versus-CPU rounding — 1e-4 relative on the worst chunk, 1e-5 typical, repeatable run to
  run, far below anything audible.
- The host runs every chunk twice and settles a mismatch with a third run, because Core AI's GPU
  was measured to return a slightly wrong result now and then on other models.

## Use

```sh
hf download arraypress/stems-roformer --local-dir models
stems model install models/stems-bs_roformer-float32.aimodel
stems song.wav --engine roformer      # song-drums.wav, -bass, -other, -vocals
```

Requirements: macOS 27, Apple silicon. Reproduce with `uv run Tools/export_roformer.py` in the
library repo, given ZFTurbo's config and checkpoint.

## Licence and citation

MIT, as the implementation and the weights. Please cite:

> W.-T. Lu, J.-C. Wang, Q. Kong, Y.-N. Hung. "Music Source Separation with Band-Split RoPE
> Transformer." ICASSP 2024.