File size: 2,663 Bytes
cc01d84
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
5b5bafc
 
 
cc01d84
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
---
license: mit
tags:
  - music-source-separation
  - stems
  - demucs
  - core-ai
  - aimodel
  - apple-silicon
  - macos
library_name: swift-vocal-isolation
---

# Hybrid Transformer Demucs for Core AI (`.aimodel`)

Apple Core AI conversion of **HTDemucs** (`htdemucs`) — Simon Rouard, Francisco Massa, Alexandre
Défossez, *Hybrid Transformers for Music Source Separation*, ICASSP 2023 — from
[github.com/facebookresearch/demucs](https://github.com/facebookresearch/demucs) (MIT, 42M
parameters). Four stems: drums, bass, other, vocals.

`stems-htdemucs-float32.aimodel` (168 MB) holds the network between the complex spectrogram and
the mask at the model's 7.8-second training segment. Inputs: the normalised mix `[1, 2, 343980]`,
its complex-as-channels spectrogram `[1, 4, 2048, 336]` and the four per-segment statistics the
model normalises by. Outputs: the spectral stems `[1, 16, 2048, 336]` and the time-branch stems
`[1, 8, 343980]`. Core AI has no STFT and no variance op, so those run in the host — in
[swift-vocal-isolation](https://github.com/arraypress/swift-vocal-isolation) (MIT), together with
upstream's chunking, overlap-add weights, centred padding and global normalisation, line for line.
Nothing was re-authored, retrained or pruned. The `stems` CLI exposes it as `--engine demucs`.

## Faithfulness

Held to upstream's Python with `shifts=0` (random shifts make upstream itself non-deterministic):

- The exported network is asserted equal to `model(mix)` before export.
- One training segment through Core AI on the GPU: 136–147 dB PSNR against upstream's output.
- Every one of 63 chunks of a 6-minute mix: worst stem 110 dB.
- Whole clips end to end (10 s and 6 min): 127–148 dB on every stem.
- One measured wrinkle in the runtime, not the model: the GPU returned a slightly wrong segment
  about 2% of the time (77–105 dB), never the same one twice. A correct run is bit-for-bit
  repeatable, so the host runs every segment twice and settles a mismatch with a third run.

## Use

```sh
hf download arraypress/stems-demucs --local-dir models
stems model install models/stems-htdemucs-float32.aimodel
stems song.wav --engine demucs      # song-drums.wav, -bass, -other, -vocals
```

Requirements: macOS 27, Apple silicon. Reproduce with `uv run Tools/export_demucs.py` in the
library repo (fetches the checkpoint through the `demucs` package). `htdemucs_ft` (a bag of four)
and `htdemucs_6s` (adds guitar and piano) export the same way but are not verified here.

## Licence and citation

MIT, as upstream. Please cite:

> S. Rouard, F. Massa, A. Défossez. "Hybrid Transformers for Music Source Separation." ICASSP 2023.