File size: 2,773 Bytes
159bd41
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
---
license: apache-2.0
tags:
  - piano-transcription
  - audio-to-midi
  - music-transcription
  - core-ai
  - aimodel
  - apple-silicon
  - macos
library_name: swift-music-transcriber
---

# Piano transcription with pedals for Core AI (`.aimodel`)

Apple Core AI conversion of **High-resolution Piano Transcription with Pedals by Regressing
Onset and Offset Times** — Qiuqiang Kong, Bochen Li, Xuchen Song, Yuan Hou, Yuxuan Wang
(ByteDance, 2020), checkpoint `CRNN_note_F1=0.9677_pedal_F1=0.9186` from
[github.com/bytedance/piano_transcription](https://github.com/bytedance/piano_transcription)
(Apache 2.0, 43M parameters). The piano specialist: one instrument, but a velocity for every note
and the sustain pedal, which general audio-to-MIDI models do not hear.

`scribe-piano-float32.aimodel` (136 MB) has two entry points:

- `trunk`: log-mel `[1, 1001, 229]` (one 10-second segment at 100 fps; 229 Slaney mels
  30–8000 Hz in dB) → the seven CRNN trunks' features `[1, 1001, 768]` (frame, onset, offset,
  velocity, pedal onset, pedal offset, pedal frame).
- `parameters`: every GRU and head weight as one flat vector. Core AI has no recurrent op and a
  decomposed GRU unrolls to ~770k ops at this size, so the sixteen bidirectional recurrences run in
  the host (Swift, Accelerate) from these weights. Nothing was re-authored, retrained or pruned.

The host side — front end, segmentation with 5-second hop, stitching, the regression
post-processor and the note/pedal state machines — lives in
[swift-music-transcriber](https://github.com/arraypress/swift-music-transcriber) (MIT), where the
`scribe` CLI exposes it as `--engine piano`.

## Faithfulness

Held to upstream's Python on real audio (a piano loop, and a 17-second three-segment piece):

- The trunk is asserted equal to upstream's forward before export, and agrees at 155 dB PSNR
  through Core AI on the GPU.
- The Swift GRU agrees with PyTorch's at 112 dB on upstream's own trunk activations.
- The seven framewise outputs agree at 123–147 dB.
- **Note and pedal events are identical**: every note, velocity and pedal event, with onset and
  offset times within 0.001 ms of upstream's.

## Use

```sh
hf download arraypress/scribe-piano --local-dir models
scribe model install models/scribe-piano-float32.aimodel
scribe recording.wav --engine piano          # one piano track, velocities, controller 64 for sustain
```

Requirements: macOS 27, Apple silicon. Reproduce with `uv run Tools/export_piano.py` in the
library repo (downloads the checkpoint from Zenodo).

## Licence and citation

Apache 2.0, as upstream. Please cite:

> Q. Kong, B. Li, X. Song, Y. Hou, Y. Wang. "High-resolution Piano Transcription with Pedals by
> Regressing Onset and Offset Times." arXiv:2010.01815, 2020.