spoti.pw Sing voice model

The vocal separator spoti.pw's Sing runs on the iPhone to turn a song's vocals down while it plays. It is Mel-Band RoFormer with KimberleyJensen's vocal checkpoint, its spectral core exported for Core ML with two-second windows: a float32 spectrum of shape [1, 2050, 201, 2] in, the vocals_spectrum of the same shape out, the STFT around it done by the app. Normalization, attention, softmax and matrix products are float32, the rest float16. The files are a compiled separator.mlmodelc as they are: the app downloads each one and keeps it only if its size and SHA-256 are the ones it pins. It needs iOS 27, and no audio leaves the phone.

File Bytes SHA-256
weights/weight.bin 488986336 970a99fb4b15724bf76d2918ceb177df592c69265d3e2fabaab6e5ba72738e62
model.mil 669061 966560ed5125174a98f19b94f5de04450a7112ade0e731f2236c202c0280a623
metadata.json 2431 52a8d5e3f09e33236d495dbed5bbce1c75bac6f2a6b6097637cf214f37c1de53
coremldata.bin 507 2090acaf7a6df72ec83857cb88d101654023a6baad222a25d0173827d3347e28
analytics/coremldata.bin 243 f7ee4ec9b5cc1c97171bd5aad93af61e183aa0db1451ef3b775bd4e77f0b7cfd

Provenance

Mel-Band RoFormer by Ju-Chiang Wang, Wei-Tsung Lu and Minz Won; the vocal checkpoint by KimberleyJensen; lucidrains' BS-RoFormer implementation; ZFTurbo's training code. The MIT notices are in NOTICE.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support