scribe-piano / README.md
arraypress's picture
model card
159bd41 verified
|
Raw History Blame Contribute Delete
2.77 kB
metadata
license: apache-2.0
tags:
  - piano-transcription
  - audio-to-midi
  - music-transcription
  - core-ai
  - aimodel
  - apple-silicon
  - macos
library_name: swift-music-transcriber

Piano transcription with pedals for Core AI (.aimodel)

Apple Core AI conversion of High-resolution Piano Transcription with Pedals by Regressing Onset and Offset Times — Qiuqiang Kong, Bochen Li, Xuchen Song, Yuan Hou, Yuxuan Wang (ByteDance, 2020), checkpoint CRNN_note_F1=0.9677_pedal_F1=0.9186 from github.com/bytedance/piano_transcription (Apache 2.0, 43M parameters). The piano specialist: one instrument, but a velocity for every note and the sustain pedal, which general audio-to-MIDI models do not hear.

scribe-piano-float32.aimodel (136 MB) has two entry points:

  • trunk: log-mel [1, 1001, 229] (one 10-second segment at 100 fps; 229 Slaney mels 30–8000 Hz in dB) → the seven CRNN trunks' features [1, 1001, 768] (frame, onset, offset, velocity, pedal onset, pedal offset, pedal frame).
  • parameters: every GRU and head weight as one flat vector. Core AI has no recurrent op and a decomposed GRU unrolls to ~770k ops at this size, so the sixteen bidirectional recurrences run in the host (Swift, Accelerate) from these weights. Nothing was re-authored, retrained or pruned.

The host side — front end, segmentation with 5-second hop, stitching, the regression post-processor and the note/pedal state machines — lives in swift-music-transcriber (MIT), where the scribe CLI exposes it as --engine piano.

Faithfulness

Held to upstream's Python on real audio (a piano loop, and a 17-second three-segment piece):

  • The trunk is asserted equal to upstream's forward before export, and agrees at 155 dB PSNR through Core AI on the GPU.
  • The Swift GRU agrees with PyTorch's at 112 dB on upstream's own trunk activations.
  • The seven framewise outputs agree at 123–147 dB.
  • Note and pedal events are identical: every note, velocity and pedal event, with onset and offset times within 0.001 ms of upstream's.

Use

hf download arraypress/scribe-piano --local-dir models
scribe model install models/scribe-piano-float32.aimodel
scribe recording.wav --engine piano          # one piano track, velocities, controller 64 for sustain

Requirements: macOS 27, Apple silicon. Reproduce with uv run Tools/export_piano.py in the library repo (downloads the checkpoint from Zenodo).

Licence and citation

Apache 2.0, as upstream. Please cite:

Q. Kong, B. Li, X. Song, Y. Hou, Y. Wang. "High-resolution Piano Transcription with Pedals by Regressing Onset and Offset Times." arXiv:2010.01815, 2020.