cs.CVJun 23, 2025

LEGATO: Large-scale End-to-end Generalizable Approach to Typeset OMR

Authors: Guang Yang, Victoria Ebert, Nazif Can Tamer, Brian Siyuan Zheng, Luiza Amador Pozzobon, Noah A. Smith

Organizations: Paul G. Allen School of Computer Science & Engineering, University of Washington · Allen Institute for AI

Abstract

We propose Legato, a new end-to-end model for optical music recognition (OMR), a task of converting music score images to machine-readable documents. Legato is the first large-scale pretrained OMR model capable of recognizing full-page or multi-page typeset music scores and the first to generate documents in ABC notation, a concise, human-readable format for symbolic music. Bringing together a pretrained vision encoder with an ABC decoder trained on a dataset of more than 214K images, our model exhibits the strong ability to generalize across various typeset scores. We conduct comprehensive experiments on a range of datasets and metrics and demonstrate that Legato outperforms the previous state of the art. On our most realistic dataset, we see a 68.2% and 48.1% absolute error reduction on the standard metrics TEDn and OMR-NED, respectively.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. LEGATO 2: Toward Multimodal Sheet Music Recognition and Understanding

    Jul 7, 2026Guang Yang, Brian Siyuan Zheng, Victoria Ebert +1Multi-Instrument Music TranscriptionSymbolic Music Generation

  2. Sheet Music Benchmark: Standardized Optical Music Recognition Evaluation

    Jun 12, 2025Juan C. Martinez-Sevilla, Joan Cerveto-Serrano, Noelia Luna +4