Image-to-Text
optical-music-recognition
omr
musicxml
sheet-music

copista-2m

Collection: copista Code: kobimusic/copisteria kobi.music

copista-2m is KobiMusic's optical music recognition (OMR) model: a scanned page of sheet music in, MusicXML out. A 1.94M-parameter symbol detector finds every notation symbol on the page (270 classes) together with the note heads' attributes, and the 2.03M-parameter evidence model of the copisteria reader reads every symbol in the context of the whole page: the naturals that say a key is wrong, the bars that add up to three beats under a 4/4 sign, the dot the repetitions have, the voice that works as the first. Nobody wrote that evidence down; the model learned it from rendered pages where the truth is known. About 4.0M parameters in all.

Made for real scans. copista is focused on real scans, not just synthetic renders with wrinkles. It is strong on authentic scans of 1800s and 1900s editions, where other systems collapse (the charts below). It works for some handwritten scores too; but if even you have trouble reading a page, the model will too.

Its sibling copista-28m reads with a 28.7M-parameter detector (30.7M parameters in all); both are in the copista collection.

Error rate on scanned pages, lower is better; Legato 2 is not released. String quartets: copista-28m 8.7 %, copista-2m 11.2 %, Legato 2 31.6 %, Legato 58.2 %, Audiveris 66.9 %. Lieder, 55 pages: copista-28m 17.7 %, copista-2m 17.8 %, homr 42.3 %, Transcoda 48.4 %, Audiveris 51.8 %. Polish piano scores (ground truth of notes only, see the note): copista-28m 34.4 %, copista-2m 38.7 %, homr 48.3 %, Transcoda 55.1 %, Audiveris 62.6 %

File What it is
small/v7_obj-recall_best.pt The symbol detector (1.94M parameters): the same 270 classes and attributes as copista-28m's.
evidence-2m.pt The evidence model (2.03M parameters, the same in both sizes): a 6-layer transformer (width 160, 8 heads) over a window of systems. A PyTorch checkpoint, {"model": state_dict, "cfg": {...}}.

Setup

The weights run with the copisteria reader. You need git and Python 3.12 or newer (tested on Linux with 3.12 and 3.14). On a CPU a page takes about 5 seconds and under 1.1 GB of memory, start-up included, so a laptop is enough; an NVIDIA GPU is used automatically when PyTorch sees one.

1. Get the reader and make a virtual environment for it:

git clone https://github.com/kobimusic/copisteria
cd copisteria
python3 -m venv .venv
source .venv/bin/activate

2. Install it. Without an NVIDIA GPU, take PyTorch's CPU build first (a much smaller download); with one, skip that line and the default build uses it.

pip install torch torchvision --index-url https://download.pytorch.org/whl/cpu    # CPU only
pip install -e .
pip install huggingface_hub

3. Download the weights into the reader's models/ folder (about 16 MB: models/evidence-2m.pt and models/small/v7_obj-recall_best.pt), from the repository's root:

hf download kobimusic/copista-2m --include "*.pt" --local-dir models

4. Read a page: a scan of printed music as PNG or JPEG (a 300 dpi scan is plenty; the reader scales the page itself), or a PDF. Run from the repository's root:

python -m copisteria.pipeline page.png --out out --detector 2m
python -m copisteria.pipeline --pdf score.pdf --range 1-12 --out out --detector 2m

PDFs need poppler (sudo apt install poppler-utils, or brew install poppler).

5. Open the result. For each page, out/<page>.musicxml opens in MuseScore, Dorico, Finale or Sibelius; out/<page>.html is a viewer (the scan with every symbol, the readings the context changed, the score engraved); out/<page>.reading.json holds every symbol's reading; out/index.html lists the pages.

Good to know

  • The detections are cached beside each image (<page>.dets_v7small_<size>.json), so reading a page again is quicker.
  • Page text (title, composer, part names, tempo and expression words) comes from an OCR file beside the page, <page>.texts.json (the format is in the reader's docs/ARCHITECTURE.md); without it the music is read the same and the text is left out.
  • FileNotFoundError: models/...: run from the repository's root (the folder with pyproject.toml) and check step
  • CUDA out of memory: another program is using the GPU. Pick another with CUDA_VISIBLE_DEVICES=1, or run on the CPU with CUDA_VISIBLE_DEVICES= (empty) in front of the command.
  • python -c "import torch; print(torch.cuda.is_available())" says whether PyTorch sees a GPU; if it should and does not, install the build for your CUDA version from pytorch.org.

How it works

How it reads. A small front end turns the detector's boxes into symbols on staves and in bars, from the page's ink, and decides no music. The evidence model then reads each system with its neighbours. For every reading the detector has (is the symbol there, its class, dots, staff position, voice, grace) the model adds evidence to the detector's own log-probability, in nats, so with no evidence the detector stands. The rest it reads from context: each note's sounding alteration, chord, tie, tuplet and onset in its bar, and each bar's clef, key and meter. A writer decodes each voice's rhythm from the model's onset and duration distributions, carries each accidental along its line to the bar's end as notation defines it, and writes MusicXML, adding nothing the detector did not see.

Training. The evidence model: 31,800 pages rendered from 28,000 PDMX scores in random engraving styles with scan effects, each with labels that say what every symbol means in the score. The detector's readings were corrupted on the fly (classes masked or swapped, attributes nudged, symbols dropped, ghost symbols, confidence shifts, tuplet marks left out after the first), so the model learned how far to trust the detector against the context: 70,000 steps from scratch, then 25,000 with emphasis on tuplets, on copista-28m's detector.

Benchmark

The OMR-NED benchmark: each page's MusicXML is compared with the dataset's own by musicdiff, the edits are summed over a set and divided by the symbols of both scores. The result is the error rate, the share of the score read wrong: a page read entirely wrong scores 100 %.

Every figure in the table and the charts is an error rate: lower is better ↓. copista-28m is lowest on every set; both sizes are ahead of every other system on every set.

System Figures Quartets, scans ↓ Quartets, renders ↓ Lieder, scans ↓ Lieder, renders ↓ Polish piano, scans* ↓
copista-28m KobiMusic 8.7 % 5.8 % 17.7 % 14.6 % 34.4 %
copista-2m KobiMusic 11.2 % 8.8 % 17.8 % 15.9 % 38.7 %
Legato 2 (unreleased) published 31.6 % 17.1 % n/a ⁴ 27.6 % n/a
Legato published 58.2 % 32.9 % n/a ⁴ 39.5 % n/a ¹
homr 0.7 run by KobiMusic piano only piano only 42.3 % 38.0 % 48.3 %
Audiveris 5.11 run by KobiMusic 66.9 % ² 33.2 % 51.8 % 28.8 % 62.6 %
Transcoda run by KobiMusic piano only piano only 48.4 % ³ 42.0 % ³ 55.1 % ³

Pages per set: quartets 252 scans and 252 renders, Lieder 55 scans and 64 renders, Polish 112. The Lieder scans are the dataset's 64 less 9 broken pages: their ground truth lacks the vocal staff the page shows, so every system loses most of them. copista-28m is 30.7M parameters in all, copista-2m 4.0M. "Published" figures are the ones in the Legato papers (Legato, Legato 2). Legato 2 is not released: its figures are from the paper and nobody can run it (hatched bars in the charts). The systems marked "run by KobiMusic" were run on the same pages with the same scorer. A page a system gives no output for counts as read entirely wrong, as on the IMSLP piano leaderboard.

* Polish scans: the ground truth holds notes, rests, beams and tuplets only: no slurs, pedal marks, dynamics, octave lines or text, and almost no articulations (0.4 per 100 notes, against 5-19 in the other sets). Whatever a reader reads of those counts against it, and copisteria reads them: about 2.7 points of its error. Scored on notes and rests only, copista-28m's error rate is 28.2 % and copista-2m's 33.4 %. The comparison with the other systems stands: they are scored against the same ground truth.

  1. Legato's Polish figure (86.7 %) comes from a different pipeline and carries an asterisk where it is published; it is left out.
  2. Audiveris gave no output on 54 of the 252 quartet scans. Its pages were upscaled for it (it refuses small staff spacing); our Audiveris runs score far better than the Audiveris figures published in the Legato 2 paper.
  3. Transcoda writes **kern; it is scored against the same MusicXML ground truth as the others (against the dataset's kern ground truth its Polish figure is 51.6 %).
  4. Legato and Legato 2 publish their Lieder scan figures over all 64 pages (44.9 % and 43.6 %), the 9 broken pages included, so they have no figure on these 55.

Error rate on rendered pages, lower is better; Legato 2 is not released. String quartets: copista-28m 5.8 %, copista-2m 8.8 %, Legato 2 17.1 %, Legato 32.9 %, Audiveris 33.2 %. Lieder: copista-28m 14.6 %, copista-2m 15.9 %, Legato 2 27.6 %, Audiveris 28.8 %, homr 38.0 %, Legato 39.5 %, Transcoda 42.0 %

copista-2m alone. Over the 735 pages of the five sets copista-2m reads 85.5 % of the score right (error rate 14.5 %).

The figures are the copisteria code's run with these weights and its default settings (--detector 2m), with page text from KobiMusic's OCR.

Limitations

  • Handwriting: some handwritten scores read well, but a page that is hard for a person to read is as hard for the model.
  • Triplets a page does not mark (marked once, often pages earlier) can be read as plain notes: with every tuplet mark removed from rendered pages, 65 % of tuplet notes are written as tuplets (96 % with the marks).
  • Staves, systems and bars come from the front end's hand-made rules; a page whose staff lines the line finder cannot follow falls back to the detector's measure boxes, which are less exact.
  • The evidence model was trained on copista-28m's detector; copista-2m's small detector misses more small marks (rests, articulations, dots), which shows on the dense quartet and Polish pages.
  • Lyrics are not written.

Links

License

The weights in this repository are released under the Apache License 2.0, like the copisteria code. You may use, fine-tune and redistribute them, commercially too, as long as you keep the copyright notice and the NOTICE file that credits KobiMusic, and say what you changed.

Citation

If you use copista-2m in research or in a product, please cite it:

@misc{copista2m,
  author    = {{KobiMusic}},
  title     = {copista-2m: optical music recognition, printed sheet music to MusicXML},
  year      = {2026},
  publisher = {Hugging Face},
  url       = {https://huggingface.co/kobimusic/copista-2m},
  license   = {Apache-2.0}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including kobimusic/copista-2m

Papers for kobimusic/copista-2m