arraypress commited on
Commit
288f41e
·
verified ·
1 Parent(s): db4e66d

model card

Browse files
Files changed (1) hide show
  1. README.md +58 -0
README.md ADDED
@@ -0,0 +1,58 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: mit
3
+ tags:
4
+ - music-source-separation
5
+ - stems
6
+ - bs-roformer
7
+ - core-ai
8
+ - aimodel
9
+ - apple-silicon
10
+ - macos
11
+ library_name: swift-vocal-isolation
12
+ ---
13
+
14
+ # BS-RoFormer for Core AI (`.aimodel`)
15
+
16
+ Apple Core AI conversion of **BS-RoFormer** — Wei-Tsung Lu, Ju-Chiang Wang, Qiuqiang Kong,
17
+ Yun-Ning Hung, *Music Source Separation with Band-Split RoPE Transformer* (ByteDance, ICASSP
18
+ 2024) — in lucidrains' implementation (MIT), with the four-stem weights
19
+ `model_bs_roformer_ep_17_sdr_9.6568.ckpt` trained and released by ZFTurbo
20
+ ([Music-Source-Separation-Training](https://github.com/ZFTurbo/Music-Source-Separation-Training),
21
+ MIT): MUSDB18 test average 9.65 dB SDR (drums 11.61, vocals 11.08, bass 8.48, other 7.44).
22
+ 132M parameters, drums, bass, other and vocals.
23
+
24
+ `stems-bs_roformer-float32.aimodel` (528 MB) holds the band-split transformer between the
25
+ spectrogram and the mask at the model's chunk: input `x` `[1, 1101, 4100]`, the STFT of 485,100
26
+ samples laid out per frame as (bin, channel, real/imaginary); output `mask` `[1, 4, 1101, 4100]`.
27
+ The STFT (n_fft 2048, hop 441, periodic Hann, unnormalised), the complex mask, the DC zeroing, the
28
+ inverse and the chunked inference with linear fades run in the host —
29
+ [swift-vocal-isolation](https://github.com/arraypress/swift-vocal-isolation) (MIT), exposed by the
30
+ `stems` CLI as `--engine roformer`. Nothing was re-authored, retrained or pruned.
31
+
32
+ ## Faithfulness
33
+
34
+ - The exported network is asserted equal to `model(chunk)` before export.
35
+ - Spectrogram layout and masked inverse: 151–155 dB PSNR against torch.
36
+ - Each chunk through Core AI on the GPU: 96–110 dB PSNR against upstream's output; whole clips
37
+ end to end 97–107 dB. Lower than a convolutional model's 140 dB because eight transformer
38
+ layers of fp32 attention accumulate GPU-versus-CPU rounding, still far below anything audible.
39
+ - The host runs every chunk twice and settles a mismatch with a third run, because Core AI's GPU
40
+ was measured to return a slightly wrong result now and then on other models.
41
+
42
+ ## Use
43
+
44
+ ```sh
45
+ hf download arraypress/stems-roformer --local-dir models
46
+ stems model install models/stems-bs_roformer-float32.aimodel
47
+ stems song.wav --engine roformer # song-drums.wav, -bass, -other, -vocals
48
+ ```
49
+
50
+ Requirements: macOS 27, Apple silicon. Reproduce with `uv run Tools/export_roformer.py` in the
51
+ library repo, given ZFTurbo's config and checkpoint.
52
+
53
+ ## Licence and citation
54
+
55
+ MIT, as the implementation and the weights. Please cite:
56
+
57
+ > W.-T. Lu, J.-C. Wang, Q. Kong, Y.-N. Hung. "Music Source Separation with Band-Split RoPE
58
+ > Transformer." ICASSP 2024.