Toward Part-Aware Choral Transcription with singing voice assignment
Authors: Hanyu Meng, Zhanhong He, Zixun Guo, Yaolong Ju
Organizations: The University of New South Wales, Sydney, Australia · Great Bay University, Guangdong, China · The University of Western Australia, Perth, Australia · Center for Digital Music (C4DM), Queen Mary University of London, London, United Kingdom
Single-instrument automatic music transcription (AMT) has advanced substantially, yet choral applications require soprano, alto, tenor, and bass (SATB) to be transcribed as separate parts. Recent note-level choral AMT instead produces a single merged note track, limiting rehearsal, education, and score reconstruction. To address this limitation, we introduce Part-aware Choral Transcription (PawCT), to our knowledge the first end-to-end neural framework that identifies active SATB parts from choral audio and transcribes each into a separate note-level track. PawCT combines part-specific onset, offset, and frame prediction with part-presence estimation, union-level supervision, and structured training targets using a range prior (RP) based on SATB pitch ranges and its ordered-continuity (OC) extension, which adds within-part melodic continuity and cross-part pitch ordering. On YouChorale, PawCT-RP-OC achieves a macro part-aware note F1 of 0.225 at a 50-ms onset tolerance, outperforming an adapted choral baseline (0.165) by 36.4% relative and a two-stage post-hoc assignment pipeline (0.175). Its part-agnostic variant, PagCT, achieves a 50-ms onset F1 of 0.382, compared with 0.237 for the previous state-of-the-art choral AMT model. Cross-dataset evaluations on CSD and Cantoria further assess performance under dataset shift. These results demonstrate the benefit of jointly modeling note transcription and vocal-part assignment. Code and demos are available at https://hanyu-meng.github.io/Paw_Choral_AMT_Demo/.
Figures & tables
Figure 1: Part-agnostic versus part-aware choral transcription. PagCT produces a single merged track, whereas PawCT jointly identifies active SATB parts (green check marks) and transcribes them into separate note-level MIDI tracks.
Figure 2: Two strategies for part-aware choral transcription considered in this paper. (a) The proposed PawCT jointly transcribes mixed choral audio into SATB-specific note tracks. (b)–(c) The two-stage alternative first uses PagCT to produce a merged note track and then applies Post-VA, a symbolic voice-assignment model, to assign each decoded note to a vocal part.
Model
#Params
Frame
Onset (50 ms)
Onset (100 ms)
Precision
Recall
F1
Precision
Recall
F1
Precision
Recall
F1
Onsets and Frames [ 12 ]
6.79M
0.806
0.326
0.428
0.450
0.178
0.242
0.688
0.248
0.344
MT3 [ 11 ]
44.7M
0.590
0.243
0.344
0.117
0.148
0.127
0.200
0.255
0.217
YourMT3+ [ 2 ]
45.8M
0.529
0.423
0.457
0.145
0.119
0.127
0.252
0.207
0.221
MuScriptor [ 28 ]
307M
0.720
0.745
0.730
0.086
0.083
0.084
0.236
0.228
0.231
Yu et al. [ 31 ]
4.23M
0.678
0.619
0.647
0.251
0.232
0.237
0.366
0.341
0.347
Table 1: Part-agnostic transcription on the YouChorale test set. Following [ 31 ] , we report frame and note precision, recall, and F1, with 50- and 100-ms onset tolerances for note metrics. “w/o Aug.” denotes no pitch-transposition augmentation. Blue rows indicate proposed systems; bold and underlined values denote the best and second-best results, respectively.
Model
Soprano
Alto
Tenor
Bass
Average
VA Rate (%)
Frame
Note
Frame
Note
Frame
Note
Frame
Note
Frame
Note
Frame
Note
YourMT3+ ⋆ [ 2 ]
0.371
0.083
0.297
0.068
0.294
0.073
0.371
0.093
0.333
0.079
72.87
62.20
Yu et al. ⋆ [ 31 ]
0.528
0.183
0.451
0.151
0.436
0.138
0.560
0.186
0.494
0.165
76.35
69.62
PagCT + Post-VA
0.501
0.180
0.439
0.145
0.447
0.156
0.568
0.219
0.489
0.175
66.26
45.81
PawCT w/o union-loss
0.417
0.192
0.338
0.169
0.343
0.178
0.419
0.222
0.379
0.190
51.36
49.74
PawCT
0.530
0.228
0.430
0.191
0.429
0.197
0.552
0.252
0.485
0.217
65.72
56.81
Table 2: Part-aware SATB transcription performance on the YouChorale test set. For each voice part: Soprano, Alto, Tenor, and Bass, we report frame F1 and note F1 with a 50-ms onset tolerance, together with their macro-average. VA Rate is defined as F1avg/F1agn×100% . All systems use pitch-transposition augmented training data. ⋆ denotes baselines adapted for part-aware transcription.
Dataset
Model
Soprano
Alto
Tenor
Bass
Avg.
CSD
PagCT + Post-VA
0.073
0.122
0.131
0.143
0.118
PawCT
0.140
0.175
0.183
0.145
0.161
PawCT-RP-OC
0.145
0.223
0.199
0.130
0.174
Cantoría
PagCT + Post-VA
0.179
0.115
0.085
0.084
0.115
PawCT
0.240
0.239
0.275
0.328
0.271
PawCT-RP-OC
0.329
0.273
0.300
0.348
0.313
Table 3: Cross-dataset evaluation of part-aware note F1 with a 50-ms onset tolerance on CSD and Cantoría. Models are trained only on YouChorale. Bold denotes the best result.