copisteria
The models behind KobiMusic's copisteria optical music recognition (OMR) reader: a scanned or rendered page of printed music in, MusicXML out. Instead of hand-written rules, a 2M-parameter transformer reads every symbol in the context of the whole page: the naturals that say a key is wrong, the bars that add up to three beats under a 4/4 sign, the dot the repetitions have, the voice that works as the first. Nobody wrote that evidence down; the model learned it from rendered pages where the truth is known.
copisteria is the evidence model and the reader around it; the symbols come from a copista model's detector, in two sizes: copista-28m (the 28.7M-parameter detector) and copista-2m (the 1.94M-parameter detector: about 4M parameters for the whole reader). Both read with the same evidence model.
| File | What it is |
|---|---|
evidence-2m.pt |
The evidence model (2.03M parameters), both sizes: a 6-layer transformer (width 160, 8 heads) over a window of systems. A PyTorch checkpoint, {"model": state_dict, "cfg": {...}}. |
v7_obj-recall-30m_best.pt |
copista-28m's symbol detector (28.7M parameters). It finds every notation symbol on the page (270 classes) together with the note heads' attributes (staff position, stem direction, dots, voice). Trained on 87K pages. |
v7_obj-recall_best.pt |
copista-2m's symbol detector (1.94M parameters), the same classes and attributes. |
Setup
The weights run with the copisteria reader. You need git and Python 3.12 or newer (tested on Linux with 3.12 and 3.14). On a desktop CPU a page takes about 15 seconds with copista-28m and 5 with copista-2m, start-up included; an NVIDIA GPU is used automatically when PyTorch sees one.
1. Get the reader and make a virtual environment for it:
git clone https://github.com/kobimusic/copisteria
cd copisteria
python3 -m venv .venv
source .venv/bin/activate
2. Install it. Without an NVIDIA GPU, take PyTorch's CPU build first (a much smaller download); with one, skip that line and the default build uses it.
pip install torch torchvision --index-url https://download.pytorch.org/whl/cpu # CPU only
pip install -e .
pip install huggingface_hub
3. Download the weights into the reader's models/ folder (about 131 MB: the evidence model in models/, both
detectors in models/small/), from the repository's root:
hf download kobimusic/copisteria evidence-2m.pt --local-dir models
hf download kobimusic/copisteria v7_obj-recall-30m_best.pt v7_obj-recall_best.pt --local-dir models/small
4. Read a page: a scan of printed music as PNG or JPEG (a 300 dpi scan is plenty; the reader scales the page itself), or a PDF. Run from the repository's root:
python -m copisteria.pipeline page.png --out out # copista-28m
python -m copisteria.pipeline page.png --out out --detector 2m # copista-2m
python -m copisteria.pipeline --pdf score.pdf --range 1-12 --out out
PDFs need poppler (sudo apt install poppler-utils, or brew install poppler).
5. Open the result. For each page, out/<page>.musicxml opens in MuseScore, Dorico, Finale or Sibelius;
out/<page>.html is a viewer (the scan with every symbol, the readings the context changed, the score engraved);
out/<page>.reading.json holds every symbol's reading; out/index.html lists the pages.
Good to know
- The detections are cached beside each image (
<page>.dets_v7_<size>.jsonfor copista-28m,<page>.dets_v7small_<size>.jsonfor copista-2m), so reading a page again is quicker. - Page text (title, composer, part names, tempo and expression words) comes from an OCR file beside the page,
<page>.texts.json(the format is in the reader'sdocs/ARCHITECTURE.md); without it the music is read the same and the text is left out. FileNotFoundError: models/...: run from the repository's root (the folder withpyproject.toml) and check stepCUDA out of memory: another program is using the GPU. Pick another withCUDA_VISIBLE_DEVICES=1, or run on the CPU withCUDA_VISIBLE_DEVICES=(empty) in front of the command.python -c "import torch; print(torch.cuda.is_available())"says whether PyTorch sees a GPU; if it should and does not, install the build for your CUDA version from pytorch.org.
How it works
How it reads. A small front end turns the detector's boxes into symbols on staves and in bars, from the page's ink, and decides no music. The model then reads each system with its neighbours. For every reading the detector has (is the symbol there, its class, dots, staff position, voice, grace) the model adds evidence to the detector's own log-probability, in nats, so with no evidence the detector stands. The rest it reads from context: each note's sounding alteration, chord, tie, tuplet and onset in its bar, and each bar's clef, key and meter. A writer decodes each voice's rhythm from the model's onset and duration distributions, carries each accidental along its line to the bar's end as notation defines it, and writes MusicXML, adding nothing the detector did not see. The code is the copisteria repository on GitHub, which says where to put these files and how to run each size.
Training. 31,800 pages rendered from 28,000 PDMX scores in random engraving styles with scan effects, each with labels that say what every symbol means in the score. The detector's readings were corrupted on the fly (classes masked or swapped, attributes nudged, symbols dropped, ghost symbols, confidence shifts, tuplet marks left out after the first), so the model learned how far to trust the detector against the context. 70,000 steps from scratch, then 25,000 with emphasis on tuplets, on copista-28m's detector.
Benchmark
The OMR-NED benchmark: each page's MusicXML is compared with the dataset's own by musicdiff, the edits are summed over a set and divided by the symbols of both scores. The result is the error rate, the share of the score read wrong: a page read entirely wrong scores 100 %.
Every figure in the table and the charts is an error rate: lower is better ↓. copista-28m is lowest on every set; both sizes are ahead of every other system on every set.
| System | Figures | Quartets, scans ↓ | Quartets, renders ↓ | Lieder, scans ↓ | Lieder, renders ↓ | Polish piano, scans* ↓ |
|---|---|---|---|---|---|---|
| copista-28m | KobiMusic | 8.7 % | 5.8 % | 17.7 % | 14.6 % | 34.4 % |
| copista-2m | KobiMusic | 11.2 % | 8.8 % | 17.8 % | 15.9 % | 38.7 % |
| Legato 2 (unreleased) | published | 31.6 % | 17.1 % | n/a ⁴ | 27.6 % | n/a |
| Legato | published | 58.2 % | 32.9 % | n/a ⁴ | 39.5 % | n/a ¹ |
| homr 0.7 | run by KobiMusic | piano only | piano only | 42.3 % | 38.0 % | 48.3 % |
| Audiveris 5.11 | run by KobiMusic | 66.9 % ² | 33.2 % | 51.8 % | 28.8 % | 62.6 % |
| Transcoda | run by KobiMusic | piano only | piano only | 48.4 % ³ | 42.0 % ³ | 55.1 % ³ |
Pages per set: quartets 252 scans and 252 renders, Lieder 55 scans and 64 renders, Polish 112. The Lieder scans are the dataset's 64 less 9 broken pages: their ground truth lacks the vocal staff the page shows, so every system loses most of them. copista-28m is 30.7M parameters in all, copista-2m 4.0M. "Published" figures are the ones in the Legato papers (Legato, Legato 2). Legato 2 is not released: its figures are from the paper and nobody can run it (hatched bars in the charts). The systems marked "run by KobiMusic" were run on the same pages with the same scorer. A page a system gives no output for counts as read entirely wrong, as on the IMSLP piano leaderboard.
* Polish scans: the ground truth holds notes, rests, beams and tuplets only: no slurs, pedal marks, dynamics, octave lines or text, and almost no articulations (0.4 per 100 notes, against 5-19 in the other sets). Whatever a reader reads of those counts against it, and copisteria reads them: about 2.7 points of its error. Scored on notes and rests only, copista-28m's error rate is 28.2 % and copista-2m's 33.4 %. The comparison with the other systems stands: they are scored against the same ground truth.
- Legato's Polish figure (86.7 %) comes from a different pipeline and carries an asterisk where it is published; it is left out.
- Audiveris gave no output on 54 of the 252 quartet scans. Its pages were upscaled for it (it refuses small staff spacing); our Audiveris runs score far better than the Audiveris figures published in the Legato 2 paper.
- Transcoda writes
**kern; it is scored against the same MusicXML ground truth as the others (against the dataset's kern ground truth its Polish figure is 51.6 %). - Legato and Legato 2 publish their Lieder scan figures over all 64 pages (44.9 % and 43.6 %), the 9 broken pages included, so they have no figure on these 55.
copisteria alone. Over the 735 pages of the five sets copista-28m reads 88.3 % of the score right (error rate 11.7 %), copista-2m 85.5 % (14.5 %).
The figures are the copisteria code's run with these weights and its default settings (--detector 2m for
copista-2m), with page text from KobiMusic's OCR.
Limitations
- Printed music only; handwriting is out of scope.
- Triplets a page does not mark (marked once, often pages earlier) can be read as plain notes: with every tuplet mark removed from rendered pages, 65 % of tuplet notes are written as tuplets (96 % with the marks).
- Staves, systems and bars come from the front end's hand-made rules; a page whose staff lines the line finder cannot follow falls back to the detector's measure boxes, which are less exact.
- copista-2m's evidence model was trained on copista-28m's detector; its small detector misses more small marks (rests, articulations, dots), which shows on the dense quartet and Polish pages.
- Lyrics are not written.
Links
- 🤗 The copista collection: the public copista-28m and copista-2m repos
- 🤗 kobimusic on Hugging Face
- copisteria: the reader
- kobi.music: KobiMusic, with the hosted reader
License
The weights in this repository are released under the Apache License 2.0, like the copisteria code. You may use, fine-tune and redistribute them, commercially too, as long as you keep the copyright notice and the NOTICE file that credits KobiMusic, and say what you changed.
Citation
If you use copisteria in research or in a product, please cite it:
@software{copisteria,
author = {{KobiMusic}},
title = {copisteria: optical music recognition with a learned evidence model},
year = {2026},
version = {1.3.0},
url = {https://github.com/kobimusic/copisteria},
license = {Apache-2.0}
}

