--- license: cc-by-nc-4.0 base_model: m-a-p/SheetSage2 tags: [coreml, music-transcription] --- # SheetSage2 on Core ML [SheetSage2](https://huggingface.co/m-a-p/SheetSage2) (revision `398b22834dac7dd05e09b9c4e40a39fc479ec502`) and its parent encoder [MERT-v2-FullSong](https://huggingface.co/m-a-p/MERT-v2-FullSong), converted to Core ML for [slurper](https://github.com/TrevorS/slurper) by `scripts/convert_sheetsage2.py`. The weights keep their CC BY-NC 4.0 license: non-commercial use only, with attribution to the SheetSage2 and MERT-v2 authors. - `encoder_fp32.mlpackage`: `samples[1,7200000]` (300 s of 24 kHz mono, zero padded) `-> cross[12,8,7500,64]`, the log-mel frontend, MERT-v2 with SheetSage2's adapters merged, the layer mix and projection, and each decoder layer's cross-attention keys (2i) and values (2i + 1). float32. - `decoder_fp16.mlpackage`: one step of the 6-layer BART decoder, `token[1,1] + position[1] + mask[1,1,1,L] -> logits[1,31678]`, with self- and cross-attention caches as states. `mask` is zeros of length position + 1. - `golden_audio.f32`, `golden_tokens.json`: 30 s of "Swansong" by Josh Woodward (CC BY 4.0) at 24 kHz mono, and the tokens PyTorch decodes from it, which these models reproduce exactly. Please cite the SheetSage2 technical report (Jiang et al., 2026) and MERT (Li et al., ICLR 2024).