spoti.pw Sing voice model

The vocal separator spoti.pw's Sing runs on the iPhone to turn a song's vocals down while it plays. It is Mel-Band RoFormer with KimberleyJensen's vocal checkpoint, its spectral core exported for Core ML with two-second windows: a float32 spectrum of shape [1, 2050, 201, 2] in, the vocals_spectrum of the same shape out, the STFT around it done by the app. Normalization, attention, softmax and matrix products are float32, the rest float16. The files are a compiled separator.mlmodelc as they are: the app downloads each one and keeps it only if its size and SHA-256 are the ones it pins. It needs iOS 27, and no audio leaves the phone.

File Bytes SHA-256
weights/weight.bin 488986336 970a99fb4b15724bf76d2918ceb177df592c69265d3e2fabaab6e5ba72738e62
model.mil 669060 c612ac3798e1ce0f89603453ff72ca78bf4f1b07701925afa726ffce4c6e3ecb
metadata.json 2416 af098e7f6b0af360cb2d5df7c3059fe83e3e2f38c27e3f67c875417c8d9d4102
coremldata.bin 507 5ff2c235fcf4e153f1ea47ee9939132ebad276dc7509fc1069b45f7ef8c0207d
analytics/coremldata.bin 243 9a949d4e0b28ab778d750f30ec3bb120b22a8c01ef8dce25f11b0832d6d5fe39

Provenance

Mel-Band RoFormer by Ju-Chiang Wang, Wei-Tsung Lu and Minz Won; the vocal checkpoint by KimberleyJensen; lucidrains' BS-RoFormer implementation; ZFTurbo's training code. The MIT notices are in NOTICE.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support