Multi-pitch estimation in vocal ensembles is challenging because singers occupy overlapping pitch ranges and often sing at closely spaced fundamental frequencies, causing their harmonics to overlap in time-frequency representations. Existing models commonly use harmonic constant-Q transform (HCQT)-based representations to provide frequency-adaptive resolution, at the cost of expensive feature extraction when training mixtures are generated on the fly. We revisit this design and compare HCQT with a linear short-time Fourier transform (STFT), whose frequency bins are directly provided as model inputs. Despite its fixed frequency resolution and the absence of a pitch-aligned input grid, the linear STFT outperforms HCQT while substantially reducing feature-extraction cost. Further analysis shows that a longer analysis window or broader spectral coverage provides no additional improvement, while restricting the input to the predicted pitch range reduces the advantage of the linear STFT. These results suggest that finer frequency resolution does not necessarily improve vocal-ensemble MPE, and that shorter analysis windows can be more effective for time-varying vocal pitches.
Figures & tables
Figure 1: The two representations compared, on a four-second a cappella recording in jaCappella [ 20 ] : (a) the linear STFT with 2048-sample window and (b) the harmonic CQT. Reference fundamental frequencies are shown in white, with h=1,2,3 shown in red, green, and blue, respectively, in (b). The dashed line in (a) marks C7, the upper bound of the predicted pitch range.
Figure 2: Model architecture. The first 360 bins of a 2048-sample STFT (0–7.7 kHz) are processed by a convolutional stem and an encoder, followed by a frame-wise prediction head over 360 pitch bins (C1–C7).
jaCappella [ 20 ]
ACappellaSet [ 22 ]
Korean [ 18 ]
Cantoría [ 6 ]
ChoralSynth ∗ [ 21 ]
Model
20 cent
50 cent
20 cent
50 cent
20 cent
50 cent
20 cent
50 cent
20 cent
50 cent
HCQT [ 3 ]
Late/Deep CNN [ 9 ]
.387 ± .039
.804 ± .033
.334 ± .026
.716 ± .042
.309 ± .008
.791 ± .013
.545 ± .012
.863 ± .010
.626 ± .020
.932 ± .014
VoasCNN [ 8 ]
.291 ± .015
.747 ± .040
.208 ± .030
.513 ± .074
.255 ± .009
.715 ± .020
.358 ± .010
.791 ± .022
.469 ± .027
.904 ± .018
VoasCLSTM [ 8 ]
.299 ± .018
.731 ± .040
.248 ± .018
.573 ± .055
.269 ± .008
.745 ± .017
.416 ± .011
.819 ± .013
.527 ± .029
.894 ± .022
ConvNeXt [ 14 ]
.648 ± .036
.884 ± .019
.492 ± .027
.828 ± .023
.584 ± .016
.880 ± .009
.561 ± .017
.875 ± .012
.638 ± .020
.934 ± .016
Table 1: A cappella ensembles and choirs. Each evaluation song is scored independently with multipitch accuracy, TP/(TP+FP+FN) , on the 360-bin, 20-cent pitch grid: exact-bin matching (20 cent) and a 50-cent tolerance. The activation threshold is selected separately for each model to maximize the corresponding metric. Intervals denote the half-width of the 95% bootstrap confidence interval over songs. ∗ ChoralSynth is excluded from the training data of both our models and the published baselines.
Training
Validation
Dataset
Language
Songs
Hours
Songs
Hours
Korean [ 18 ]
Korean
51
50.9
48
14.5
jaCappella [ 20 ]
Japanese
32
40.3
8
4.8
ACappellaSet [ 22 ]
Chinese
66
32.8
17
6.4
Cantoría [ 6 ]
Latin
11
8.2
3
0.9
Dagstuhl [ 25 ]
German
6
0.8
2
0.2
Table 2: Training datasets. Training and validation splits are defined separately for each dataset. Hours denote the duration of songs.
Dagstuhl [ 25 ]
ESMUC [ 7 ]
20 cent
50 cent
20 cent
50 cent
HCQT
Late/Deep CNN
.470 ± .012
.904 ± .004
.564 ± .028
.910 ± .005
VoasCNN
.374 ± .015
.871 ± .003
.381 ± .018
.862 ± .007
VoasCLSTM
.394 ± .013
.858 ± .003
.406 ± .024
.841 ± .010
ConvNeXt
.617 ± .009
.907 ± .010
.475 ± .038
.826 ± .011
Table 3: Choirs recorded in one room. Dagstuhl ChoirSet and ESMUC are the only evaluation datasets whose singers were recorded together, and both are included in the training data of the Late/Deep CNN and our models.
Figure 3: F1 and computational cost. Time per epoch re-timed on one A100. All models are trained and evaluated on 30 s segments. InceptionNeXt coincides with ConvNeXt.
Figure 4: Analysis window. Linear STFT with the ConvNeXt encoder across analysis-window sizes, scored against the reference F0 before 20-cent quantization. Dashed lines: HCQT with the same encoder. The top panel reports F1 at three pitch tolerances, and the bottom panel reports the pitch error of matched pitches.