LatentSlate Audio Analysis model bundle

All model files used by the Audio Analysis recipe of LatentSlate Engine: audio in, timing metadata out (sections, lyric lines, word timings, beats and bar starts).

Each component keeps its original license, which is in its folder. Nothing here was retrained or modified. Where a component was only published as a Python pickle checkpoint, it has been converted to safetensors so it can load without executing pickled code. The converted tensors are bit-identical to the originals, which are kept under original/ for verification.

Folder What it does Source License
whisper-large-v3/ Transcribes what is sung openai/whisper-large-v3 @ 06f233fe06e710322aca913c1bc4249a0d71fce1 Apache-2.0
wav2vec2-large-960h-lv60-self/ CTC forced alignment, English (default) facebook/wav2vec2-large-960h-lv60-self @ 54074b1c16f4de6a5ad59affb4caa8f2ea03a119; pytorch_model.bin converted to model.safetensors Apache-2.0
mms-300m-1130-forced-aligner/ CTC forced alignment, multilingual (alternative) MahmoudAshraf/mms-300m-1130-forced-aligner @ 49402e9577b1158620820667c218cd494cc44486 CC-BY-NC-4.0 (non-commercial)
demucs-htdemucs-ft/ Separates the vocal from the mix facebookresearch/demucs htdemucs_ft, from dl.fbaipublicfiles.com/demucs/hybrid_transformer/; converted to safetensors + config.json (constructor arguments) MIT
beat-this/ Beat and downbeat tracking CPJKU/beat_this final0 checkpoint; converted to safetensors + config.json (hyper-parameters) MIT

Notes

  • htdemucs_ft is a bag of four specialist models (one per source, one-hot bag weights). LatentSlate only needs vocals and loads 04573f0d-f3cf25b2 alone. The Demucs filename suffix is the first eight hex digits of the original file's SHA-256.
  • The multilingual MMS aligner is included for completeness and is not for commercial use. LatentSlate's default configuration uses the Apache-2.0 English aligner.

Citations

  • Radford et al., Robust Speech Recognition via Large-Scale Weak Supervision (Whisper), 2022.
  • Baevski et al., wav2vec 2.0, 2020; Pratap et al., Scaling Speech Technology to 1,000+ Languages (MMS), 2023.
  • Rouard, Massa, Défossez, Hybrid Transformers for Music Source Separation, ICASSP 2023.
  • Foscarin, Schlüter, Widmer, Beat this! Accurate beat tracking without DBN postprocessing, ISMIR 2024.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support