Toward Part-Aware Choral Transcription with singing voice assignment
Authors: Hanyu Meng, Zhanhong He, Zixun Guo, Yaolong Ju
Organizations: The University of New South Wales, Sydney, Australia · Great Bay University, Guangdong, China · The University of Western Australia, Perth, Australia · Center for Digital Music (C4DM), Queen Mary University of London, London, United Kingdom
Single-instrument automatic music transcription (AMT) has advanced substantially, yet choral applications require soprano, alto, tenor, and bass (SATB) to be transcribed as separate parts. Recent note-level choral AMT instead produces a single merged note track, limiting rehearsal, education, and score reconstruction. To address this limitation, we introduce Part-aware Choral Transcription (PawCT), to our knowledge the first end-to-end neural framework that identifies active SATB parts from choral audio and transcribes each into a separate note-level track. PawCT combines part-specific onset, offset, and frame prediction with part-presence estimation, union-level supervision, and structured training targets using a range prior (RP) based on SATB pitch ranges and its ordered-continuity (OC) extension, which adds within-part melodic continuity and cross-part pitch ordering. On YouChorale, PawCT-RP-OC achieves a macro part-aware note F1 of 0.225 at a 50-ms onset tolerance, outperforming an adapted choral baseline (0.165) by 36.4% relative and a two-stage post-hoc assignment pipeline (0.175). Its part-agnostic variant, PagCT, achieves a 50-ms onset F1 of 0.382, compared with 0.237 for the previous state-of-the-art choral AMT model. Cross-dataset evaluations on CSD and Cantoria further assess performance under dataset shift. These results demonstrate the benefit of jointly modeling note transcription and vocal-part assignment. Code and demos are available at https://hanyu-meng.github.io/Paw_Choral_AMT_Demo/.
Figures & tables
Figure 1: Part-agnostic versus part-aware choral transcription. PagCT produces a single merged track, whereas PawCT jointly identifies active SATB parts (green check marks) and transcribes them into separate note-level MIDI tracks.
Figure 2: Two strategies for part-aware choral transcription considered in this paper. (a) The proposed PawCT jointly transcribes mixed choral audio into SATB-specific note tracks. (b)–(c) The two-stage alternative first uses PagCT to produce a merged note track and then applies Post-VA, a symbolic voice-assignment model, to assign each decoded note to a vocal part.
Model
#Params
Frame
Onset (50 ms)
Onset (100 ms)
Precision
Recall
F1
Precision
Recall
F1
Precision
Recall
F1
Onsets and Frames [ 12 ]
6.79M
0.806
0.326
0.428
0.450
0.178
0.242
0.688
0.248
0.344
MT3 [ 11 ]
44.7M
0.590
0.243
0.344
0.117
0.148
0.127
0.200
0.255
0.217
YourMT3+ [ 2 ]
45.8M
0.529
0.423
0.457
0.145
0.119
0.127
0.252
0.207
0.221
MuScriptor [ 28 ]
307M
0.720
0.745
0.730
0.086
0.083
0.084
0.236
0.228
0.231
Yu et al. [ 31 ]
4.23M
0.678
0.619
0.647
0.251
0.232
0.237
0.366
0.341
0.347
Table 1: Part-agnostic transcription on the YouChorale test set. Following [ 31 ] , we report frame and note precision, recall, and F1, with 50- and 100-ms onset tolerances for note metrics. “w/o Aug.” denotes no pitch-transposition augmentation. Blue rows indicate proposed systems; bold and underlined values denote the best and second-best results, respectively.
Model
Soprano
Alto
Tenor
Bass
Average
VA Rate (%)
Frame
Note
Frame
Note
Frame
Note
Frame
Note
Frame
Note
Frame
Note
YourMT3+ ⋆ [ 2 ]
0.371
0.083
0.297
0.068
0.294
0.073
0.371
0.093
0.333
0.079
72.87
62.20
Yu et al. ⋆ [ 31 ]
0.528
0.183
0.451
0.151
0.436
0.138
0.560
0.186
0.494
0.165
76.35
69.62
PagCT + Post-VA
0.501
0.180
0.439
0.145
0.447
0.156
0.568
0.219
0.489
0.175
66.26
45.81
PawCT w/o union-loss
0.417
0.192
0.338
0.169
0.343
0.178
0.419
0.222
0.379
0.190
51.36
49.74
PawCT
0.530
0.228
0.430
0.191
0.429
0.197
0.552
0.252
0.485
0.217
65.72
56.81
Table 2: Part-aware SATB transcription performance on the YouChorale test set. For each voice part: Soprano, Alto, Tenor, and Bass, we report frame F1 and note F1 with a 50-ms onset tolerance, together with their macro-average. VA Rate is defined as F1avg/F1agn×100% . All systems use pitch-transposition augmented training data. ⋆ denotes baselines adapted for part-aware transcription.
Dataset
Model
Soprano
Alto
Tenor
Bass
Avg.
CSD
PagCT + Post-VA
0.073
0.122
0.131
0.143
0.118
PawCT
0.140
0.175
0.183
0.145
0.161
PawCT-RP-OC
0.145
0.223
0.199
0.130
0.174
Cantoría
PagCT + Post-VA
0.179
0.115
0.085
0.084
0.115
PawCT
0.240
0.239
0.275
0.328
0.271
PawCT-RP-OC
0.329
0.273
0.300
0.348
0.313
Table 3: Cross-dataset evaluation of part-aware note F1 with a 50-ms onset tolerance on CSD and Cantoría. Models are trained only on YouChorale. Bold denotes the best result.
High-quality singing annotations are fundamental to modern Singing Voice Synthesis (SVS) systems. However, obtaining these annotations at scale through manual labeling is unrealistic due to the substantial labor and musical expertise required, making automatic annotation highly necessary. Despite their utility, current automatic transcription systems face significant challenges: they often rely on complex multi-stage pipelines, struggle to recover text-note alignments, and exhibit poor generalization to out-of-distribution (OOD) singing data. To alleviate these issues, we present VocalParse, a unified singing voice transcription (SVT) model built upon a Large Audio Language Model (LALM). Specifically, our novel contribution is to introduce an interleaved prompting formulation that jointly models lyrics, melody, and word-note correspondence, yielding a generated sequence that directly maps to a structured musical score. Furthermore, we propose a Chain-of-Thought (CoT) style prompting strategy, which decodes lyrics first as a semantic scaffold, significantly mitigating the context disruption problem while preserving the structural benefits of interleaved generation. Experiments demonstrate that VocalParse achieves state-of-the-art SVT performance on multiple singing datasets. The source code and checkpoint are available at https://github.com/pymaster17/VocalParse.
Yukun Chen, Tianrui Wang, Zhaoxi Mu +2
Xi’an Jiaotong University · Nanyang Technological University · Tianjin University +2
Transcribing music into a human-readable score requires a coherent understanding of rhythm, harmony, melody, and form. Two obstacles limit this goal: annotated recordings are scarce, and accurate local predictions can still produce inconsistent musical sequences. We present SheetSage2, a unified music transcription framework that combines synthetic data, task-specific structured decoding, and autoregressive distillation. Automatically annotated MIDI, rendered into audio, provides scalable supervision across music understanding tasks. Task-specific structured decoders integrate complementary musical cues and their temporal dependencies to produce musically coherent scores. Autoregressive distillation further retains transcription accuracy without task-specific dynamic programming at inference. Across eight benchmark collections, a single SheetSage2-AR model exceeds the listed prior systems on 12 of 15 benchmark--metric pairs in our evaluation, substantially improving over SheetSage1 and surpassing task-specific models on several benchmarks. Model weights and inference code are publicly available.
Precise note-level annotations are critical for training automatic music transcription (AMT) systems, in particular note-onset labels, which form a core component of many recent AMT systems. However, high-quality annotations for real-world recordings are scarce. Sequence-level score--audio alignment methods such as dynamic time warping provide only coarse correspondence, making a local refinement step necessary. This refinement step, known as snapping, adjusts aligned score onsets using peaks in a neural onset posteriorgram and often determines whether weakly aligned score--audio pairs become usable training data at all. Despite its practical importance, snapping is typically treated as a simple post-processing heuristic and implemented with greedy local decisions. We present a systematic analysis of snapping strategies for training instrument-agnostic transcribers, demonstrating that snapping is essential for learning from weakly aligned data. Building on this, we formulate snapping as a per-pitch assignment problem and solve it via bipartite graph matching, yielding context-aware onset decisions under overlapping refinement windows and uncertain initial alignments. Extensive cross-dataset experiments across piano, chamber, and orchestral recordings show improved onset alignment and transcription accuracy over greedy snapping, with gains increasing for wider snapping windows and coarser initial alignments. Qualitative examples are provided on our project page: https://abhirupsaha8.github.io