Transcribing music into a human-readable score requires a coherent understanding of rhythm, harmony, melody, and form. Two obstacles limit this goal: annotated recordings are scarce, and accurate local predictions can still produce inconsistent musical sequences. We present SheetSage2, a unified music transcription framework that combines synthetic data, task-specific structured decoding, and autoregressive distillation. Automatically annotated MIDI, rendered into audio, provides scalable supervision across music understanding tasks. Task-specific structured decoders integrate complementary musical cues and their temporal dependencies to produce musically coherent scores. Autoregressive distillation further retains transcription accuracy without task-specific dynamic programming at inference. Across eight benchmark collections, a single SheetSage2-AR model exceeds the listed prior systems on 12 of 15 benchmark--metric pairs in our evaluation, substantially improving over SheetSage1 and surpassing task-specific models on several benchmarks. Model weights and inference code are publicly available.
Figures & tables
Figure 1: SheetSage2-AR inference. Audio representations and preceding events condition a unified event sequence, which is converted into ABC notation and an editable lead sheet. The right panel aligns an illustrative score, its ABC notation, and its events. The excerpt is serialized independently, with the active key emitted at its first beat and omitted thereafter unless it changes. Musical units are used for readability.
Figure 2: Real Prober activations and structured decoding on RWC-MDB-P-2001 No. 1. (a) Rhythm, (b) meter and relative tempo, (c) chord, (d) melody, (e) selected section and boundary activations, and (f) key. Thin grid lines indicate beats and bold lines downbeats.
Table 2: Multitask transcription on real recordings. SheetSage2-Prober and SheetSage2-AR are compared with SheetSage1, madmom ( Böck et al., 2016a ) , and prior task-specific systems. Scores are percentages; higher is better. Bold and underlining mark the largest and second-largest point estimates in each benchmark–metric pair.
Task-training audio
Beat
Downbeat
Key
Chord
Structure
Melody
Synthetic only
79.52
76.64
75.66
76.99
56.76
54.77
Real only
85.43
82.06
75.69
77.65
74.88
74.68
Real + synthetic
87.60
85.77
75.46
79.22
76.46
74.69
Improvement by synth.
+2.17
+3.71
-0.24
+1.57
+1.58
+0.01
Table 3: Effect of task-training data. Each entry is the unweighted mean (%) of a task’s benchmark–metric pairs in Table 2 . Structure averages accuracy and both boundary F1 scores; melody averages RWC-Pop vocal/full F1 and Rock Corpus vocal F1. Bold marks the best task average.
Figure 3: Selected examples of structured decoding. (a,b) P-DBN (orange) changes tempo level over several beats, while Prober (blue) remains stable. Lines mark decoded beats and thick lines denote downbeats. (c,d) Pitch-marginal decoding introduces octave shifts that Prober avoids. Gray bars show reference notes.
osu2017
Hooktheory
Beat This!
Prober
P-DBN
AR
Beat This!
Prober
P-DBN
AR
Mixing ratio (%)
20.74
2.96
3.70
1.48
17.30
0.34
1.36
0.51
Table 4: Global tempo consistency. Mixing ratio (%) on 135 common osu2017 and 2,943 common Hooktheory recordings; lower is better. P-DBN applies madmom’s DBN to Prober activations.
SheetSage1
Prober
Prober (Pitch Marginal)
AR
Mixing ratio (%)
34.83
9.31
12.91
11.71
Table 5: Octave consistency on RWC full melody. Mixing ratio (%) on 333 contiguous instrument passages common to all systems; lower is better.
Table 6: Opening metadata and melody. RWC-MDB-P-2001 No. 1, 0.02–7.17 s. The task prefix and first four complete measures are shown. The first event establishes meter, section, key, chord, and melody source.
Table 7: A key change. RWC-MDB-P-2001 No. 1, 35.61–39.17 s. The key changes from G ♯ minor to C minor at 37.39 s. The key token updates tonal context at the start of measure 22.
Table 8: A temporary meter change. A song from Chords1217. The meter changes from 4/4 to 3/4 at 60.86 s and returns to 4/4 at 90.99 s. The two excerpts show both boundaries; the 24 intervening measures are omitted.
Training set
Songs
Audio hours
Audio type
LA / MALD ( Lev, 2024 ; Jiang, 2025 )
355,095
19,269.7
MIDI synthesis
SLMS ( Eldeeb & Malandro, 2025 )
6,130
366.4
MIDI synthesis
Expanded lead-sheet corpus ( Donahue et al., 2022 )
32,243
2,349.1
Recordings
HarmonixSet training set ( Nieto et al., 2019 )
512
31.5
Recordings
Total
393,980
22,016.7
Mixed
Appendix
Table 9: Training sets of the multitask Prober. Hours are audio durations; task labels may cover only part of a recording. LA: Los Angeles MIDI Dataset; MALD: MIDI AutoLabel Dataset, the task labels for LA; SLMS: Segmented Lakh MIDI Subset.
Non-melody Prober
Melody Prober
SheetSage2-AR
Head
Two-layer MLP, width 512
Six RoFormer layers, width 512, 8 heads
Six-layer BART decoder, width 512, 8 heads
Output
160 binary channels
926-class softmax
31,678-class softmax over event tokens
Training data
393,980 songs (Table 9 )
220,341 songs
441,094 pseudo-labeled recordings
Excerpts
300 s and 30 s, 25-Hz frames
30 s, 384 grid positions
300 s, up to 5,120 tokens
Pitch augmentation
Continuous, [−6,6) semitones
Continuous, [−12,12) semitones
None
Global batch
72
128
32
Appendix
Table 10: Configurations of the evaluated models.
Figure 4: SheetSage2-Prober and structured decoding.
Model
Task
Channels
Content
Non-melody
Chord
48
Root, chroma, extension, and bass, 12 pitch classes each
Key
24
12 major and 12 minor keys
Rhythm
64
Downbeat, quarter-note, and eighth-note events, four 100-Hz slots each; two meter denominators; 25 median-tempo and 25 relative-tempo bins
Structure
24
23 section classes and one boundary
Melody
Melody
926
128×7 pitch–interval categories, 24 durations, one rest, and five auxiliary slots
Appendix
Table 11: Prober output channels.
Quality
0
1
2
3
4
5
6
7
8
9
10
11
Other bass
maj
∙
∙
∙
/2 , /3 , /5
min
∙
∙
∙
/2 , /b3 , /5
dim
∙
∙
∙
—
aug
∙
∙
∙
—
maj7
∙
∙
∙
∙
/3 , /5 , /7
min7
∙
∙
∙
∙
/b3 , /5 , /b7
Appendix
Table 12: Chord qualities in the decoder vocabulary. Labels follow the Harte syntax ( Harte et al., 2005 ) . Filled cells mark each quality’s chroma template, in semitones above the root. Every quality is used with all 12 roots in root position; the last column lists the additional bass options. As in the mir_eval encoding ( Raffel et al., 2014 ) , the bass note is always part of the chroma, so the /2 states also contain the second; for example, C:maj/2 has chroma C, D, E, G and bass D.
Semitone offset
Major key
Minor key
0
Perfect unison
Perfect unison
1
Augmented unison
Minor second
2
Major second
Major second
3
Augmented second
Minor third
4
Major third
Major third
5
Perfect fourth
Perfect fourth
Appendix
Table 13: Melody pitch spelling by key. Semitone offsets are measured from the tonic modulo 12. Each entry gives the interval spelling selected by the fixed lookup; alternative enharmonic spellings are not selected by this rule.
Figure 5: Tone cost for chord spelling in G major. Notes are ordered on the line of fifths. The seven notes of the G major key signature (shaded) have zero cost, and the cost increases by one per fifth beyond either edge; it continues linearly outside the plotted range.
Spelling
Chord tones (1, ♭ 3, ♭ 5, ♭ 7)
Tone costs
Total
C ♯ :hdim7
C ♯ , E, G, B
1, 0, 0, 0
2
D ♭ :hdim7
D ♭ , F ♭ , A ♭♭ , C ♭
5, 8, 11, 7
36
B x :hdim7
B x , D x , F x , A x
13, 10, 7, 11
54
Appendix
Table 14: Spelling C#:hdim7 in G major. Tone costs follow Figure 5 ; the total counts the root cost twice. The selected spelling is C ♯ :hdim7.
Relation
Example
Score
Same
C major vs. C major
1.0
Fifth
C major vs. G major
0.5
Relative
C major vs. A minor
0.3
Parallel
C major vs. C minor
0.2
Other
C major vs. D major
0.0
Appendix
Table 15: Weighted key score. Fifth errors in either direction receive 0.5.
osu2017
Chords1217
JAAH
Metric
CF
Prober
AR
CF
Prober
AR
CF
Prober
AR
Root
87.00
91.18
90.77
84.65
86.13
85.69
60.43
66.62
67.55
Thirds
85.72
90.21
89.91
81.79
83.04
82.55
57.56
63.02
64.14
Maj/min
86.55
90.43
90.08
83.94
84.29
83.81
59.45
62.94
64.50
Triads
83.96
88.26
87.99
77.73
78.80
78.53
56.65
60.27
61.81
Sevenths
76.19
74.64
75.58
72.39
71.60
72.07
46.49
45.15
47.89
Appendix
Table 16: Chord metrics (%). CF: ChordFormer, the strongest prior chord system in Table 2 ; on Chords1217 it uses five-fold cross-validation. Bold marks the best system for each benchmark and metric.
Benchmark
Metric
SheetSage1
MuScriptor
YourMT3+
Demucs+ ROSVOT
SheetSage2
Prober
AR
RWC-Pop
Vocal F1 ↑
62.71
47.04
49.12
26.09
83.08
82.51
Full F1 ↑
64.02
38.71
42.35
22.30
75.00
75.29
Rock Corpus
Vocal F1 ↑
49.19
37.36
36.05
19.86
65.98
67.08
Appendix
Table 17: External vocal-melody transcription systems. Benchmarks, metrics, and recordings match Table 2 ; external systems use their vocal output only. The Full row is grayed out because the external systems do not support melody reduction for non-vocal instruments. Scores are percentages; higher is better. Bold and underlining mark the largest and second-largest scores in each row.
Figure 6: First sung phrase of RWC-MDB-P-2001 No. 1 (9.8–24.8 s). Each panel shows one system of Table 17 . Gray bars are reference notes, hollow where missed; blue and orange bars are predictions that do and do not match a reference note. The reference is shifted by whole octaves to best match each system.
Task
Benchmark
Metric
Task-training audio
Real only
Synthetic only
Real + synthetic
Beat
GTZAN
F1 ↑
80.08
73.05
82.93
osu2017
90.79
85.98
92.28
Downbeat
GTZAN
F1 ↑
75.51
67.23
78.74
osu2017
88.60
86.05
92.79
Key
GiantSteps
Score ↑
75.93
75.05
78.29
Appendix
Table 18: Complete Prober data-source ablation. Rows and metrics match Table 2 ; real + synthetic reproduces its Prober results. Scores are percentages; higher is better. Bold marks the largest point estimate in each row.
Figure 7: Opening vocal melody in RWC-MDB-P-2001 No. 1. In each panel, colored Prober predictions overlay gray ground-truth notes over 10.0–16.8 s. All traces are the outputs scored in Table 18 . For comparison, the ground truth is shifted down one octave in the real-only and mixed panels and two octaves in the synthetic-only panel.
Figure 8: Tempo ratio composition. Each vertical slice is one recording, sorted independently per strip by mean level; colors give the duration share at each tempo level relative to the reference. A recording mixes tempo levels when at least two colors each cover at least 5% of its slice.
Figure 9: Octave-offset composition. Each vertical slice is one unit (a passage or a recording), sorted independently per strip by mean offset; colors give the share of matched notes at each octave offset from the reference. A unit mixes octaves when at least two colors each cover at least 5% of its slice.
Task
Benchmark
Metric
Prober
P-DBN
Pitch-marginal
Direct-pitch
Beat
GTZAN
F1 ↑
82.93
84.72
—
—
osu2017
92.28
91.55
—
—
Downbeat
GTZAN
F1 ↑
78.74
79.12
—
—
osu2017
92.79
90.65
—
—
Melody
RWC-Pop
Vocal F1 ↑
83.08
—
83.09
83.13
Full F1 ↑
75.00
—
75.02
74.85
Appendix
Table 19: F1 of the consistency ablations. Benchmarks, metrics, and recordings match Table 2 , whose Prober scores are reproduced. Scores are percentages; higher is better. Bold marks the largest point estimate in each row; an em dash marks an ablation that does not apply.
Figure 10: SheetSage2-AR lead sheet for RWC-MDB-P-2001 No. 1, mm. 14–37 and 74–91 of 118.
Figure 11: SheetSage1 lead sheet for RWC-MDB-P-2001 No. 1, mm. 14–37 and 74–91 of 114.
Figure 12: SheetSage2-AR lead sheet for RWC-MDB-P-2001 No. 2, mm. 1–18 and 40–57 of 93.
Figure 13: SheetSage1 lead sheet for RWC-MDB-P-2001 No. 2, mm. 1–18 and 40–57 of 90.
Existing methods for automatic music transcription are often limited to single-instrument recordings or fail on complex, real music mixes. Although previous work utilizes synthetic training data, the resulting models generalize poorly, leading to largely unusable transcription output in realistic, multi-instrument settings. In this work, we analyze the effectiveness of synthetic data for pre-training while combining it with fine-tuning on real music audio and post-training using reinforcement learning. We further introduce conditioning on instrument presence to customize transcriptions. Finally, we release MuScriptor, an open-weight multi-instrument music transcription model that works on real-world music recordings from across a diverse range of musical genres.
Simon Rouard, Michael Krause, Axel Roebel +2
Kyutai · Mirelo AI · UMR STMS, IRCAM-CNRS Sorbonne Univ.
Competitive music transcription models require large amounts of paired audio-score data, which is scarce due to collection costs, alignment difficulty, and copyright restrictions. Meanwhile, vast quantities of unpaired audio recordings and symbolic scores are freely available but have gone unused. We adopt a cycle-consistent translation framework in which a small amount of paired data acts as a minimal anchor, unlocking the full potential of the unpaired pool. We find that: unpaired data yields surprisingly large gains, especially under limited supervision; unpaired audio contributes more than unpaired scores; incorporating unlabeled audio from a new instrument during training improves transcription for that instrument without any paired supervision. Together, these results suggest that scaling unpaired data offers a practical path toward high-quality transcription for instruments where labeled data remains scarce.
Optical Music Recognition (OMR), the task of transcribing sheet music into a structured textual representation, is currently bottlenecked by a lack of large-scale, annotated datasets of real scans. This forces models to rely on either few-shot transfer or synthetic training pipelines that remain overly simplistic. A secondary challenge is encoding non-uniqueness: in the popular Humdrum **kern format for transcribing music, multiple different text encodings can render into the same visual sheet music. This one-to-many mapping creates a harder learning task and introduces high uncertainty during decoding. We propose Transcoda, an OMR system built on (i) an advanced synthetic data generation pipeline, (ii) a normalization of the **kern encoding to enforce a unique normal form and (iii) grammar-based decoding to ensure the syntactic correctness of the output. This approach allows us to train a compact 59M-parameter model in just 6 hours on a single GPU that outperforms billion-parameter baselines. Transcoda achieves the best score among state of the art baselines on a newly curated benchmark of synthetically rendered scores at 18.46% OMR-NED (compared to 43.91% for the next-best system, Legato) and reduces the error rate on historical Polish scans to 63.97% OMR-NED (down from 80.16% for SMT++).