Speaker verification (SV) performance degrades under language mismatch due to the entanglement of speaker identity with language-specific acoustic cues. To address this problem, we leverage large-scale self-supervised pre-trained models (PTMs) to learn language-agnostic speaker representations. We utilize PTMs as robust front-end feature extractors, capitalizing on their rich acoustic and linguistic knowledge acquired from vast, diverse audio data. These generalized features are then used to train a downstream speaker embedding network, effectively disentangling speaker identity from language-specific characteristics. We validate our approach on the TidyVoice2026 benchmark, which benchmarks SV under language mismatch. Our proposed system (team T02) achieves equal error rates (EERs) of 2.21% on tv26_eval-A and 2.99% on tv26_eval-U.
Figures & tables
Figure 1: Model architecture and training paradigm of the speaker verification system based on self-supervised pre-trained models.
System
Params
GMACs
TidyVoiceX_Dev
tv26_eval-A
tv26_eval-U
EER(%) / minDCF
EER(%) / minDCF
EER(%) / minDCF
ReDimNet-B5 SF2-C
9.2M
9.87
1.1353 / 0.6024
-
-
+QMF
-
3.4699 / 0.2562
4.8389 / 0.2944
ReDimNet-B6 SF2-C
15M
20.27
0.9791 / 0.5887
-
-
+QMF
-
2.5163 / 0.1892
3.4344 / 0.2237
w2v-BERT 2.0 based SV AAM
580+13.7M
58.43
1.0255 / 0.6085
3.3382 / 0.2328
4.5895 / 0.2676
Table 1: Evaluation results of developed models on TidyVoiceX_Dev, tv26_eval-A and tv26_eval-U, reported as EER(%) / minDCF. Channel Concatenation is used for multi-scale feature aggregation. GMACs is measured on 2-second segments using the THOP library.
PTM
Params
#Langs
TidyVoiceX_Dev
EER(%) / minDCF
Whisper Large-v3 [ 21 ]
635M
99
3.1468 / 0.7352
XLS-R-300m [ 10 ]
317M
128
2.9394 / 0.6481
XLS-R-1b [ 10 ]
965M
128
3.0987 / 0.5663
XLS-R-2b [ 10 ]
2.16B
128
3.1022 / 0.6598
MMS-300m [ 11 ]
317M
1406
3.2738 / 0.6676
Table 2: Performance comparison of different pre-trained models for speaker verification on the TidyVoiceX_Dev set without score calibration, including their parameter sizes and the number of languages in their pre-training corpora.
Aggregation Method
TidyVoiceX_Dev
EER(%) / minDCF
Mean Fusion
2.2134 / 0.7129
Layer-wise Weighted Fusion
2.1856 / 0.7095
Channel-wise Weighted Fusion
2.1771 / 0.7080
Channel Concatenation
2.0927 / 0.7063
Temporal Concatenation
2.3210 / 0.7184
Table 3: Ablation study on multi-scale feature aggregation for w2v-BERT 2.0 based speaker verification models.
Training Stage
Training Data
TidyVoiceX_Dev
EER(%) / minDCF
Pre-training
Full dataset
1.5411 / 0.6636
+ Fine-tuning
Full dataset
1.1380 / 0.6250
++ LM-FT
TidyVoiceX_Train
1.0255 / 0.6085
Table 4: Ablation study of the w2v-BERT 2.0 based speaker verification models on TidyVoiceX_Dev without score calibration.