File size: 3,509 Bytes
b56f9e6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
---

license: mit
tags:
  - onnx
  - webnn
  - audio-source-separation
  - htdemucs
pipeline_tag: audio-to-audio
---


# Stem Separator (HTDemucs v4, fwd-only ONNX export) - WebNN

ONNX export of Meta's pretrained **HTDemucs v4** (`htdemucs.th`, base/non-fine-tuned checkpoint)
used by the [WebNN Developer Preview](https://github.com/microsoft/webnn-developer-preview) Stem
Separator demo. Splits a stereo mix into **drums / bass / other / vocals**, running entirely
client-side via ONNX Runtime Web with WebNN (NPU/GPU) acceleration, cascading to WebGPU/WASM.

## Files

| File | Size | Description |
|---|---|---|
| `onnx/htdemucs_fwd.onnx` | ~2.3 MB | Model graph |
| `onnx/htdemucs_fwd.onnx.data` | ~168 MB | External weights |

## Provenance

- **Source checkpoint**: Meta Research's pretrained `htdemucs.th` base checkpoint (signature
  `955717e8`, filename `955717e8-8726e21a.th`) from the
  [facebookresearch/demucs](https://github.com/facebookresearch/demucs) project (MIT license),
  [checkpoint link](https://dl.fbaipublicfiles.com/demucs/hybrid_transformer/955717e8-8726e21a.th).
  No training or fine-tuning was performed. This is a format conversion of an already-trained model.
- **Conversion**: exported via `convert_htdemucs_fwd_only.py`, a script modeled on the HTDemucs
  conversion utilities from the
  [RyanMetcalfeInt8/Music-Source-Separation-Training](https://github.com/RyanMetcalfeInt8/Music-Source-Separation-Training)
  fork, purpose-built to produce a clean, opset-pinned ONNX file for web/WebNN consumption rather
  than an OpenVINO IR intermediate.
- **Why fwd-only**: `torch.onnx.export` cannot trace the complex-tensor ops used by HTDemucs's
  STFT/iSTFT stages. Those stages are reimplemented client-side in JavaScript instead, matching the
  original PyTorch implementation within floating-point tolerance.
- **Conversion command**:
  ```

  python convert_htdemucs_fwd_only.py \

    --config configs/config_musdb18_htdemucs.yaml \

    --checkpoint 955717e8-8726e21a.th \

    --output htdemucs_fwd.onnx

  ```
  The exact script used is included in this repository at
  [`tools/convert_htdemucs_fwd_only.py`](./tools/convert_htdemucs_fwd_only.py) as reproducibility
  collateral (adapted from the fork above, MIT License).

## Input / Output shapes

| Tensor | Shape | Description |
|---|---|---|
| `x` (input) | `[1, 4, 2048, 336]` | Normalized CAC spectrogram (freq-domain branch) |
| `xt` (input) | `[1, 2, 343980]` | Normalized stereo waveform (time-domain branch), ~7.8s @ 44.1kHz |
| `x_out` (output) | `[1, 16, 2048, 336]` | 4 stems × 4-channel CAC spectrogram |
| `xt_out` (output) | `[1, 8, 343980]` | 4 stems × stereo waveform |

Segments are processed with 50% overlap (`SEGMENT_LENGTH=343980`, `OVERLAP=171990`) and
reassembled with triangular-weighted overlap-add.

## Validation

Compared against a Python fp32 reference pipeline (sine-wave input, every intermediate stage dumped):

| Stage | Max diff |
|---|---|
| STFT | 9.5e-7 |
| pre_forward | 7.6e-6 |

| ONNX `fwd` (fp16) | MAE 9e-4 |

| post_forward | 1.8e-7 |

## Usage

See the [Stem Separator demo](https://github.com/microsoft/webnn-developer-preview/tree/main/demos/stem-separator)
in the WebNN Developer Preview repository for the full client-side pre/post-processing pipeline.

## License

MIT - see [LICENSE.md](./LICENSE.md). Derived from
[facebookresearch/demucs](https://github.com/facebookresearch/demucs) (MIT).