|
Download README.md from arraypress/stems-demucs: direct link, hf CLI and curl.
- Browser
- Download file 2.66 kB
-
https://huggingface.co/arraypress/stems-demucs/resolve/main/README.md
- Command line
-
hf download hf://arraypress/stems-demucs/README.md
-
curl -L -o README.md https://huggingface.co/arraypress/stems-demucs/resolve/main/README.md
2.66 kB
| license: mit | |
| tags: | |
| - music-source-separation | |
| - stems | |
| - demucs | |
| - core-ai | |
| - aimodel | |
| - apple-silicon | |
| - macos | |
| library_name: swift-vocal-isolation | |
| # Hybrid Transformer Demucs for Core AI (`.aimodel`) | |
| Apple Core AI conversion of **HTDemucs** (`htdemucs`) — Simon Rouard, Francisco Massa, Alexandre | |
| Défossez, *Hybrid Transformers for Music Source Separation*, ICASSP 2023 — from | |
| [github.com/facebookresearch/demucs](https://github.com/facebookresearch/demucs) (MIT, 42M | |
| parameters). Four stems: drums, bass, other, vocals. | |
| `stems-htdemucs-float32.aimodel` (168 MB) holds the network between the complex spectrogram and | |
| the mask at the model's 7.8-second training segment. Inputs: the normalised mix `[1, 2, 343980]`, | |
| its complex-as-channels spectrogram `[1, 4, 2048, 336]` and the four per-segment statistics the | |
| model normalises by. Outputs: the spectral stems `[1, 16, 2048, 336]` and the time-branch stems | |
| `[1, 8, 343980]`. Core AI has no STFT and no variance op, so those run in the host — in | |
| [swift-vocal-isolation](https://github.com/arraypress/swift-vocal-isolation) (MIT), together with | |
| upstream's chunking, overlap-add weights, centred padding and global normalisation, line for line. | |
| Nothing was re-authored, retrained or pruned. The `stems` CLI exposes it as `--engine demucs`. | |
| ## Faithfulness | |
| Held to upstream's Python with `shifts=0` (random shifts make upstream itself non-deterministic): | |
| - The exported network is asserted equal to `model(mix)` before export. | |
| - One training segment through Core AI on the GPU: 136–147 dB PSNR against upstream's output. | |
| - Every one of 63 chunks of a 6-minute mix: worst stem 110 dB. | |
| - Whole clips end to end (10 s and 6 min): 127–148 dB on every stem. | |
| - One measured wrinkle in the runtime, not the model: the GPU returned a slightly wrong segment | |
| about 2% of the time (77–105 dB), never the same one twice. A correct run is bit-for-bit | |
| repeatable, so the host runs every segment twice and settles a mismatch with a third run. | |
| ## Use | |
| ```sh | |
| hf download arraypress/stems-demucs --local-dir models | |
| stems model install models/stems-htdemucs-float32.aimodel | |
| stems song.wav --engine demucs # song-drums.wav, -bass, -other, -vocals | |
| ``` | |
| Requirements: macOS 27, Apple silicon. Reproduce with `uv run Tools/export_demucs.py` in the | |
| library repo (fetches the checkpoint through the `demucs` package). `htdemucs_ft` (a bag of four) | |
| and `htdemucs_6s` (adds guitar and piano) export the same way but are not verified here. | |
| ## Licence and citation | |
| MIT, as upstream. Please cite: | |
| > S. Rouard, F. Massa, A. Défossez. "Hybrid Transformers for Music Source Separation." ICASSP 2023. | |