Machine learning tools have significantly aided automatic piano music transcription; however, this domain has focused primarily on accurately predicting the pitches and timings of played notes. To produce sheet music for the piano, notes must be separated into two staves, one for each hand, and good sheet music often contains fingering annotations to guide the player when sight-reading or learning fast or complex pieces. We propose 8 statistical approaches for combined hand and fingering annotation of transcribed piano notes, including baseline hidden Markov models, rule-based methods, a synthesis of existing approaches, and N-gram language models. Furthermore, we develop a pipeline for complete transcription from piano audio to fingering-annotated sheet music. Evaluations with the PIG dataset demonstrate that our Synthesis model achieves a hand separation accuracy of 90.8% and a joint hand and finger annotation accuracy of 56.6%. These approaches serve as a new baseline for further research into this problem, while our pipeline demonstrates the feasibility of a combined system for automated note transcription, hand separation, and fingering annotation.
Figures & tables
Hand
Joint
Runtime
Model
Mean
Worst
Best
Mean
Worst
Best
GM (s)
Baseline-M 1st-order
73.9
41.0
90.5
28.0
12.1
53.1
4.29
Baseline-M 2nd-order
71.2
39.0
87.1
25.1
11.7
45.7
3.93
Baseline-M 3rd-order
66.4
19.7
90.1
20.7
8.9
32.3
6.11
Baseline-S 1st-order
86.7
60.2
98.4
39.7
14.5
62.4
6.19
Baseline-S 2nd-order
86.8
60.2
98.2
38.7
13.3
58.1
6.85
Table 1: Mean, worst-piece, and best-piece hand and joint annotation accuracy for the joint and hierarchical first-, second-, and third-order HMM baselines. Runtime is the geometric mean of training plus evaluation time for each 10-piece test split.
Figure 1: Diagram of our complete transcription pipeline outlining the flow of data at each stage.
Figure 2: Sample of sheet music for Bach’s Fugue No. 2 BWV 847 in C minor generated entirely by our pipeline with 3-gram-R for hand separation and fingering annotation.
Hand
Joint
Runtime
Model
Mean
Q1
Q3
Mean
Q1
Q3
GM (s)
Nakamura-Merged
67.4
50.0
87.3
39.7
24.7
56.7
498.02
Baseline-M
73.6
69.1
79.3
26.8
21.3
31.2
3.32
Baseline-S
86.1
82.6
91.6
37.9
31.8
43.3
5.18
Novel-B
88.7
85.8
93.6
41.4
36.0
46.1
16.06
Novel-V
86.8
83.7
92.1
33.1
27.5
38.3
9.17
Table 2: Mean model-only accuracy with the first and third quartiles of piece-level scores; runtime is the geometric mean across splits.
Evaluation
Metric
Comparison
Δ %
95% CI
Lower %
Upper %
Model
Hand
Synthesis – 3-gram-R
0.4
0.0
0.9
Model
Joint
Baseline-S – 3-gram-M
−0.5
−1.1
0.2
Pipeline
Hand
Novel-B – 3-gram-M
0.1
−0.2
0.4
Pipeline
Hand
Novel-V – 3-gram-M
−0.1
−0.5
0.3
Pipeline
Joint
Nakamura-Merged – Baseline-S
0.3
−0.5
1.1
Table 3: Pairwise comparisons where the 95% bootstrap confidence interval includes zero. Mean difference and upper and lower bound differences are reported in percentage points as the first model minus the second.
Hand
Joint
Runtime
Model
Mean
Q1
Q3
Mean
Q1
Q3
GM (s)
Nakamura-Merged
48.0
26.4
72.2
25.5
10.3
41.8
538.31
Baseline-M
54.4
44.2
71.0
18.2
13.3
25.0
40.77
Baseline-S
62.0
52.6
79.2
25.2
17.6
33.5
42.81
Novel-B
62.5
53.4
81.0
25.5
19.5
32.8
53.94
Novel-V
62.3
53.3
81.3
21.7
16.3
28.5
47.12
Table 4: Mean pipeline recall with the first and third quartiles of piece-level scores; runtime is the geometric mean across splits.
Model
Mean (%)
Performance-MIDI transcription
Audio-to-Score
90.6
Ours
89.3
Hand assignment on shared MIDI
Audio-to-Score
90.3
Synthesis
87.7
Table 5: Comparison with Audio-to-Score on the released MAPS subset.
Figure 3: Sample of sheet music annotated with the Synthesis model. Dashed red circles indicate the same finger being used for different sequential notes in quick succession. Solid red circles indicate unplayable finger crossing. Solid yellow circles indicate hand labels that violate the ground truth.
Figure 4: Sample of sheet music annotated with the 3-gram-R model.
Piano fingering shapes how a passage can be played, yet it is difficult to label after a performance. An annotator must decide which finger produced each note while reconciling the score, timing, video, and hand motion. We present PiAnnotate, a web-based pipeline for adding expert fingering annotations to the FurElise performance dataset. The tool brings together a piano-roll view, performance video, and a 3D MANO hand mesh so that reviewers can inspect each assignment in musical and physical context. Rather than storing only the final answer, PiAnnotate keeps paired rule-based and human-edited fingering tracks. These paired tracks make the annotation history auditable by showing where a geometric rule was sufficient, where experts intervened, and how labels changed across review passes. As a final diagnostic, we train a small Transformer probe on the paired tracks. The probe improves on the rule baseline on held-out pieces while remaining conservative about changing labels that were already correct, suggesting that the edited labels contain learnable structure rather than only isolated fixes.
Joonhyung Bae, Kirak Kim, Hyeyoon Cho +8
Korea Advanced Institute of Science and Technology (KAIST), Daejeon, South Korea · Seoul National University, Seoul, South Korea · YAMAHA, Hamamatsu, Japan
This paper describes a novel paradigm that formalizes automatic piano transcription (APT) as an optimal transport (OT) problem, not as a frame-level multi-label binary classification problem. Our method learns to minimize the cost of transporting a predicted distribution of note events to the ground-truth distribution over time and frequency. The OT loss can thus accommodate temporal misalignment, leading to perceptually relevant optimization. We also propose a convolutional recurrent neural network (CRNN) with a harmonics-aware attention mechanism to capture the spectro-temporal dependencies inherent in music.Our experiments using the MAESTRO dataset showed that our method attained a state-of-the-art performance in onset detection. We confirmed the versatility of the OT loss in application to existing models.
Weixing Wei, Raynaldi Lalang, Dichucheng Li +1
Graduate School of Informatics, Kyoto University, Japan · Graduate School of Engineering, Kyoto University, Japan · Independent Researcher, Hong Kong, China
We consider the conversion of musical recordings into human-readable sheet music annotated with timestamps. Such output lets a listener clearly visualize rubato (temporally expressive playing), a learner diagnose ensemble precision and timing choices against the written music, and a musicology scholar compare performance styles across recordings of the same work. We introduce (1) a prompt-conditioned encoder-decoder model, named Rubato, trained to output (2) a new textual representation for polyphonic music, named InterMo, which we designed for compatibility with sequence-to-sequence training. Our experiments demonstrate that Rubato produces timestamped piano sheet music from audio with higher notational accuracy than the best existing approaches, which are based on cascades. We find that even if the cascade is given ground-truth MIDI instead of audio, Rubato performs better, suggesting that the ceiling of existing approaches is primarily representational, not acoustic. Further, because Rubato is trained on several related tasks (with prompts), it competes with or outperforms the best single-task systems on related but simpler tasks like MIDI note grounding and beat/downbeat detection. A demo is available at https://nctamer.github.io/rubato-transcription .
Nazif Can Tamer, Victoria Ebert, Guang Yang +1
Paul G. Allen School of Computer Science & Engineering, University of Washington · Allen Institute for AI