We propose Legato, a new end-to-end model for optical music recognition (OMR), a task of converting music score images to machine-readable documents. Legato is the first large-scale pretrained OMR model capable of recognizing full-page or multi-page typeset music scores and the first to generate documents in ABC notation, a concise, human-readable format for symbolic music. Bringing together a pretrained vision encoder with an ABC decoder trained on a dataset of more than 214K images, our model exhibits the strong ability to generalize across various typeset scores. We conduct comprehensive experiments on a range of datasets and metrics and demonstrate that Legato outperforms the previous state of the art. On our most realistic dataset, we see a 68.2% and 48.1% absolute error reduction on the standard metrics TEDn and OMR-NED, respectively.
Figures & tables
Figure 1: Model architecture. The input image is first cropped into overlapping segments with an aspect ratio of 1:4 or less, then resized and divided into four patches per segment (§ 4.2.1 ). The image patches are fed into a vision encoder (§ 4.2.2 ; parameters are frozen during training). The resulting latent embeddings serve as cross-attention keys and values in a transformer decoder, which autoregressively generates ABC tokens (§ 4.2.3 ). Special tokens <B> , <I> , and <E> denote <|begin_of_abc|> , <|image|> , and <|end_of_abc|> , respectively. For better visualization, here we use “_” to represent whitespace.
Figure 2: An example of our canonical ABC representation (below) with a MusicXML-rendered image (above).
Figure 3: Example vocabulary items from tokenization.
Metric
n
GPT-5
SMT++
Legato
Legato small
1. Camera OpenScore String Quartets (252 pages)
TEDn
252
90.2
98.6
62.2
84.0
TEDn convert
18
89.5
79.9
59.6
82.7
OMR-NED
252
97.3
94.8
58.3
93.6
2. Rendered OpenScore String Quartets (252 pages)
TEDn
252
92.5
97.9
61.2
78.8
Table 1: Experimental results on various datasets and metrics. Lower is better for all metrics. n represents the number of samples tested for the metric. TEDn is the primary metric, requiring outputs to be converted to MusicXML. TEDn convert : evaluated only on instances where SMT++ produces outputs that can be successfully converted to MusicXML; Legato outputs always converted to MusicXML successfully. OMR-NED is format agnostic, built on extraction of symbols from any format. For OMR-NED, **kern outputs are automatically corrected for syntax errors, while ABC is first converted to MusicXML (again, always successful in practice) before symbol extraction. OpenScore String Quartets is the most challenging dataset, since it has much denser score images. All metrics are explained in § 5 .
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Architecture
SMT++
SMT++
SMT++ large
SMT++ large
Legato
Training Dataset
GrandStaff †
GrandStaff †
GrandStaff †
PDMX-Synth
PDMX-Synth
Format & Tokenizer
**kern
ABC
ABC
ABC
ABC
IMSLP Piano Scores (32 pages)
TEDn
97.7
74.6
73.0
80.5
29.5
OMR-NED
91.9
70.9
82.9
96.2
43.8
Appendix
Table 2: Ablation study on IMSLP Piano Scores. Lower is better for all metrics. All metrics are explained in § 5 . † : Following SMT++ style of data processing, i.e., syetem-level pre-training and curriculum fine-tuning.
Figure 4: Learning Curve of Legato. The validation loss and error rates are tested on the validation split of PDMX-Synth.
Dataset
Total Files
Failed Files
Failure Rate
Camera OpenScore String Quartets
252
234
92.9%
Rendered OpenScore String Quartets
252
224
88.9%
Camera OpenScore Lieder
64
60
93.8%
Rendered OpenScore Lieder
64
60
93.8%
IMSLP Piano Scores
32
29
90.6%
Appendix
Table 3: Failure rates of **kern-to-MusicXML conversion on SMT++’s outputs.
Model
GPT-5
Gemini 2.5 Pro
Output Format
ABC
ABC
MusicXML
ABC
ABC
MusicXML
In-Context Learning
✓
✓
IMSLP Piano Scores
TEDn
96.7
74.7
86.6
87.8
78.6
80.3
OMR-NED
98.7
93.3
94.4
95.4
93.7
94.3
Appendix
Table 4: General-purpose VLMs’ ability on IMSLP Piano Scores.
Metric
SMT++
Legato
Legato small
Rendered String Quartets Violin Parts (256 pages)
†CERkern
30.7
16.7
50.6
†SERkern
42.6
19.1
55.5
†LERkern
75.2
29.2
73.5
TEDn
95.0
07.7
35.0
TEDn convert
64.5
08.7
35.9
Appendix
Table 5: Results on OpenScore String Quartets violin parts. Lower is better for all metrics. TEDn is the primary metric, requiring outputs to be converted to MusicXML. TEDn convert : evaluated only on instances where SMT++ produces outputs that can be successfully converted to MusicXML; Legato outputs always converted to MusicXML successfully. OMR-NED is format agnostic, built on extraction of symbols from any format. For OMR-NED, **kern outputs are automatically corrected for syntax errors, while ABC is first converted to MusicXML (again, always successful in practice) before symbol extraction. † : 10.2% of Legato’s outputs and 27.9% of Legato small ’s outputs are not convertible to **kern, so an empty prediction is used. OpenScore String Quartets is the most challenging dataset, since it has much denser score images. All metrics are explained in § 5 .
Figure 5: ABC error rates on PDMX-Synth test input with different aspect ratios. Error rates are reported by averaging over each bin. Legato is capable of recognizing multi-page scores.
Figure 6: Example (first system of Duetto No. 1 in E minor by Bach, BWV 802) from IMSLP Piano Scores (top), with output from Legato (middle) and SMT++ (bottom). Errors are marked in red boxes.
Figure 7: Qualitative examples of piano scores. Since SMT++ is trained only on piano data, we illustrate results on well-known piano works. All score images are sourced from IMSLP (not rendered by MuseScore, abcm2ps, or Verovio) and cropped to a single system for clarity, though both SMT++ and Legato can process full pages. The examples progress from easy ( 7(a) ), to difficult ( 7(b) ), to ill-formed input ( 7(c) ). TEDn values are reported at system-level. Since the number of errors is too large for some of the examples, we did not mark them with red boxes to preserve readability.
Jun 12, 2025·Juan C. Martinez-Sevilla, Joan Cerveto-Serrano, Noelia Luna +4
Pattern Recognition and Artificial Intelligence Group, University of Alicante, Spain · Self-employed · Center for Computer Research in Music and Acoustics, Stanford University, USA +1