LatentSlate Audio Analysis model bundle
All model files used by the Audio Analysis recipe of LatentSlate Engine: audio in, timing metadata out (sections, lyric lines, word timings, beats and bar starts).
Each component keeps its original license, which is in its folder. Nothing here was
retrained or modified. Where a component was only published as a Python pickle
checkpoint, it has been converted to safetensors so it can load without executing
pickled code. The converted tensors are bit-identical to the originals, which are kept
under original/ for verification.
| Folder | What it does | Source | License |
|---|---|---|---|
whisper-large-v3/ |
Transcribes what is sung | openai/whisper-large-v3 @ 06f233fe06e710322aca913c1bc4249a0d71fce1 |
Apache-2.0 |
wav2vec2-large-960h-lv60-self/ |
CTC forced alignment, English (default) | facebook/wav2vec2-large-960h-lv60-self @ 54074b1c16f4de6a5ad59affb4caa8f2ea03a119; pytorch_model.bin converted to model.safetensors |
Apache-2.0 |
mms-300m-1130-forced-aligner/ |
CTC forced alignment, multilingual (alternative) | MahmoudAshraf/mms-300m-1130-forced-aligner @ 49402e9577b1158620820667c218cd494cc44486 |
CC-BY-NC-4.0 (non-commercial) |
demucs-htdemucs-ft/ |
Separates the vocal from the mix | facebookresearch/demucs htdemucs_ft, from dl.fbaipublicfiles.com/demucs/hybrid_transformer/; converted to safetensors + config.json (constructor arguments) |
MIT |
beat-this/ |
Beat and downbeat tracking | CPJKU/beat_this final0 checkpoint; converted to safetensors + config.json (hyper-parameters) |
MIT |
Notes
htdemucs_ftis a bag of four specialist models (one per source, one-hot bag weights). LatentSlate only needs vocals and loads04573f0d-f3cf25b2alone. The Demucs filename suffix is the first eight hex digits of the original file's SHA-256.- The multilingual MMS aligner is included for completeness and is not for commercial use. LatentSlate's default configuration uses the Apache-2.0 English aligner.
Citations
- Radford et al., Robust Speech Recognition via Large-Scale Weak Supervision (Whisper), 2022.
- Baevski et al., wav2vec 2.0, 2020; Pratap et al., Scaling Speech Technology to 1,000+ Languages (MMS), 2023.
- Rouard, Massa, Défossez, Hybrid Transformers for Music Source Separation, ICASSP 2023.
- Foscarin, Schlüter, Widmer, Beat this! Accurate beat tracking without DBN postprocessing, ISMIR 2024.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support