spoti.pw Sing voice model
The vocal separator spoti.pw's Sing runs on the iPhone to turn a song's vocals down
while it plays. It is Mel-Band RoFormer with KimberleyJensen's vocal checkpoint, its spectral core
exported for Core ML with two-second windows: a float32 spectrum of shape [1, 2050, 201, 2] in, the
vocals_spectrum of the same shape out, the STFT around it done by the app. Normalization, attention,
softmax and matrix products are float32, the rest float16. The files are a compiled separator.mlmodelc
as they are: the app downloads each one and keeps it only if its size and SHA-256 are the ones it pins.
It needs iOS 27, and no audio leaves the phone.
| File | Bytes | SHA-256 |
|---|---|---|
weights/weight.bin |
488986336 | 970a99fb4b15724bf76d2918ceb177df592c69265d3e2fabaab6e5ba72738e62 |
model.mil |
669060 | c612ac3798e1ce0f89603453ff72ca78bf4f1b07701925afa726ffce4c6e3ecb |
metadata.json |
2416 | af098e7f6b0af360cb2d5df7c3059fe83e3e2f38c27e3f67c875417c8d9d4102 |
coremldata.bin |
507 | 5ff2c235fcf4e153f1ea47ee9939132ebad276dc7509fc1069b45f7ef8c0207d |
analytics/coremldata.bin |
243 | 9a949d4e0b28ab778d750f30ec3bb120b22a8c01ef8dce25f11b0832d6d5fe39 |
Provenance
- Checkpoint: KimberleyJSN/melbandroformer, MIT.
- Conversion: john-rocky/coreai-model-zoo.
- Reference implementation: Mel-Band-Roformer-Vocal-Model at revision
25f44ffb55ee3c301281bba21b2d6d311cb69ae2. - Exported by
harness/sing/export_coreml.pyin the spoti.pw repository.
Mel-Band RoFormer by Ju-Chiang Wang, Wei-Tsung Lu and Minz Won; the vocal checkpoint by KimberleyJensen;
lucidrains' BS-RoFormer implementation; ZFTurbo's training code. The MIT notices are in NOTICE.