arraypress commited on
Commit
cc01d84
·
verified ·
1 Parent(s): 540aae1

model card

Browse files
Files changed (1) hide show
  1. README.md +55 -0
README.md ADDED
@@ -0,0 +1,55 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: mit
3
+ tags:
4
+ - music-source-separation
5
+ - stems
6
+ - demucs
7
+ - core-ai
8
+ - aimodel
9
+ - apple-silicon
10
+ - macos
11
+ library_name: swift-vocal-isolation
12
+ ---
13
+
14
+ # Hybrid Transformer Demucs for Core AI (`.aimodel`)
15
+
16
+ Apple Core AI conversion of **HTDemucs** (`htdemucs`) — Simon Rouard, Francisco Massa, Alexandre
17
+ Défossez, *Hybrid Transformers for Music Source Separation*, ICASSP 2023 — from
18
+ [github.com/facebookresearch/demucs](https://github.com/facebookresearch/demucs) (MIT, 42M
19
+ parameters). Four stems: drums, bass, other, vocals.
20
+
21
+ `stems-htdemucs-float32.aimodel` (168 MB) holds the network between the complex spectrogram and
22
+ the mask at the model's 7.8-second training segment. Inputs: the normalised mix `[1, 2, 343980]`,
23
+ its complex-as-channels spectrogram `[1, 4, 2048, 336]` and the four per-segment statistics the
24
+ model normalises by. Outputs: the spectral stems `[1, 16, 2048, 336]` and the time-branch stems
25
+ `[1, 8, 343980]`. Core AI has no STFT and no variance op, so those run in the host — in
26
+ [swift-vocal-isolation](https://github.com/arraypress/swift-vocal-isolation) (MIT), together with
27
+ upstream's chunking, overlap-add weights, centred padding and global normalisation, line for line.
28
+ Nothing was re-authored, retrained or pruned. The `stems` CLI exposes it as `--engine demucs`.
29
+
30
+ ## Faithfulness
31
+
32
+ Held to upstream's Python with `shifts=0` (random shifts make upstream itself non-deterministic):
33
+
34
+ - The exported network is asserted equal to `model(mix)` before export.
35
+ - One training segment through Core AI on the GPU: 136–147 dB PSNR against upstream's output.
36
+ - Every one of 63 chunks of a 6-minute mix: worst stem 110 dB.
37
+ - Whole clips end to end (10 s and 6 min): 127–148 dB on every stem.
38
+
39
+ ## Use
40
+
41
+ ```sh
42
+ hf download arraypress/stems-demucs --local-dir models
43
+ stems model install models/stems-htdemucs-float32.aimodel
44
+ stems song.wav --engine demucs # song-drums.wav, -bass, -other, -vocals
45
+ ```
46
+
47
+ Requirements: macOS 27, Apple silicon. Reproduce with `uv run Tools/export_demucs.py` in the
48
+ library repo (fetches the checkpoint through the `demucs` package). `htdemucs_ft` (a bag of four)
49
+ and `htdemucs_6s` (adds guitar and piano) export the same way but are not verified here.
50
+
51
+ ## Licence and citation
52
+
53
+ MIT, as upstream. Please cite:
54
+
55
+ > S. Rouard, F. Massa, A. Défossez. "Hybrid Transformers for Music Source Separation." ICASSP 2023.