|
Download README.md from TrevorJS/htdemucs-CoreAI: direct link, hf CLI and curl.
- Browser
- Download file 2.7 kB
-
https://huggingface.co/TrevorJS/htdemucs-CoreAI/resolve/main/README.md
- Command line
-
hf download hf://TrevorJS/htdemucs-CoreAI/README.md
-
curl -L -o README.md https://huggingface.co/TrevorJS/htdemucs-CoreAI/resolve/main/README.md
2.7 kB
| license: mit | |
| tags: | |
| - audio | |
| - source-separation | |
| - music | |
| - demucs | |
| - core-ai | |
| - deprecated | |
| # htdemucs on Core AI | |
| > **Deprecated.** slurper no longer uses this model. It runs the Core ML conversion, [TrevorJS/htdemucs-CoreML](https://huggingface.co/TrevorJS/htdemucs-CoreML), which works on macOS 26 and later. This repository is kept for reference and is not updated. | |
| Meta's [Hybrid Transformer Demucs](https://github.com/facebookresearch/demucs) (`htdemucs`: drums, bass, other, vocals) converted to a Core AI model for macOS 27 on Apple Silicon. It is the drum, bass and other separator in slurper, a stem-splitting command-line tool. | |
| ## Graph | |
| `htdemucs_fp32.aimodel`, float32, one function `main`, fixed shapes for one 7.8 s segment at 44.1 kHz: | |
| | | Name | Shape | Contents | | |
| |---|---|---|---| | |
| | in | `mix` | `[1, 2, 343980]` | stereo audio | | |
| | in | `spec` | `[1, 4, 2048, 336]` | `HTDemucs._magnitude(HTDemucs._spec(mix))`: left real, left imaginary, right real, right imaginary | | |
| | out | `time` | `[1, 8, 343980]` | time branch, 4 sources × 2 channels | | |
| | out | `freq` | `[1, 16, 2048, 336]` | frequency branch, 4 sources × 2 channels × real/imaginary | | |
| The graph is `HTDemucs.forward` from its normalization to just before `_mask`; both outputs are denormalized. The complex STFT stays on the host, so a caller: | |
| 1. computes `spec` as demucs does: reflect-pad by 1536 samples plus the remainder of the last hop, take a centered, normalized STFT (4096-sample periodic Hann, hop 1024), and keep bins 0..<2048 of frames 2..<338; | |
| 2. runs `main`; | |
| 3. inverts each source and channel of `freq` with `HTDemucs._ispec` (zero Nyquist bin, two zero frames each side, normalized inverse STFT, trim) and adds `time`. | |
| For whole songs, split into 7.8 s segments overlapping by a quarter with triangular crossfades, as demucs's `apply_model` does. Source order is drums, bass, other, vocals. | |
| float16 overflows to NaN in the frequency branch. Running two inferences at once on one loaded model corrupted the outputs, so run segments one at a time. | |
| ## Verification | |
| - A synthetic segment through Core AI on the GPU against PyTorch `HTDemucs.forward`: 114 dB (drums), 129 dB (bass), 114 dB (other), 100 dB (vocals) SDR. | |
| - A 135 s song through slurper's Swift host against PyTorch `apply_model` (no shifts, overlap 0.25): 113-120 dB SDR per stem. | |
| ## Conversion | |
| `scripts/convert_htdemucs.py` in slurper: demucs 4.1.0, torch 2.13.0, coreai-torch 0.4.2 (coreai-core 1.0.0b2). It exports the core with `torch.export`, converts it with `TorchConverter`, and checks the Core AI output against PyTorch before saving. | |
| ## License | |
| MIT, as are the htdemucs weights in facebookresearch/demucs. | |