Speaker verification (SV) performance degrades under language mismatch due to the entanglement of speaker identity with language-specific acoustic cues. To address this problem, we leverage large-scale self-supervised pre-trained models (PTMs) to learn language-agnostic speaker representations. We utilize PTMs as robust front-end feature extractors, capitalizing on their rich acoustic and linguistic knowledge acquired from vast, diverse audio data. These generalized features are then used to train a downstream speaker embedding network, effectively disentangling speaker identity from language-specific characteristics. We validate our approach on the TidyVoice2026 benchmark, which benchmarks SV under language mismatch. Our proposed system (team T02) achieves equal error rates (EERs) of 2.21% on tv26_eval-A and 2.99% on tv26_eval-U.
Figures & tables
Figure 1: Model architecture and training paradigm of the speaker verification system based on self-supervised pre-trained models.
System
Params
GMACs
TidyVoiceX_Dev
tv26_eval-A
tv26_eval-U
EER(%) / minDCF
EER(%) / minDCF
EER(%) / minDCF
ReDimNet-B5 SF2-C
9.2M
9.87
1.1353 / 0.6024
-
-
+QMF
-
3.4699 / 0.2562
4.8389 / 0.2944
ReDimNet-B6 SF2-C
15M
20.27
0.9791 / 0.5887
-
-
+QMF
-
2.5163 / 0.1892
3.4344 / 0.2237
w2v-BERT 2.0 based SV AAM
580+13.7M
58.43
1.0255 / 0.6085
3.3382 / 0.2328
4.5895 / 0.2676
Table 1: Evaluation results of developed models on TidyVoiceX_Dev, tv26_eval-A and tv26_eval-U, reported as EER(%) / minDCF. Channel Concatenation is used for multi-scale feature aggregation. GMACs is measured on 2-second segments using the THOP library.
PTM
Params
#Langs
TidyVoiceX_Dev
EER(%) / minDCF
Whisper Large-v3 [ 21 ]
635M
99
3.1468 / 0.7352
XLS-R-300m [ 10 ]
317M
128
2.9394 / 0.6481
XLS-R-1b [ 10 ]
965M
128
3.0987 / 0.5663
XLS-R-2b [ 10 ]
2.16B
128
3.1022 / 0.6598
MMS-300m [ 11 ]
317M
1406
3.2738 / 0.6676
Table 2: Performance comparison of different pre-trained models for speaker verification on the TidyVoiceX_Dev set without score calibration, including their parameter sizes and the number of languages in their pre-training corpora.
Aggregation Method
TidyVoiceX_Dev
EER(%) / minDCF
Mean Fusion
2.2134 / 0.7129
Layer-wise Weighted Fusion
2.1856 / 0.7095
Channel-wise Weighted Fusion
2.1771 / 0.7080
Channel Concatenation
2.0927 / 0.7063
Temporal Concatenation
2.3210 / 0.7184
Table 3: Ablation study on multi-scale feature aggregation for w2v-BERT 2.0 based speaker verification models.
Training Stage
Training Data
TidyVoiceX_Dev
EER(%) / minDCF
Pre-training
Full dataset
1.5411 / 0.6636
+ Fine-tuning
Full dataset
1.1380 / 0.6250
++ LM-FT
TidyVoiceX_Train
1.0255 / 0.6085
Table 4: Ablation study of the w2v-BERT 2.0 based speaker verification models on TidyVoiceX_Dev without score calibration.
Cross-lingual mismatch remains a key source of overall degradation in modern speaker verification. The TidyVoice2026 Challenge targets this setting with text-independent verification, comprising 3,666 training and 808 development speakers in 40 languages and 2,200 evaluation speakers in 38 unseen languages, without language labels at test time. Starting from the official SimAM-ResNet34 baseline pretrained on VoxBlink2 and VoxCeleb2 and fine-tuned on TidyVoice, we revisit Nuisance Attribute Projection (NAP) as a simple language-normalization step in the embedding space. We estimate a compact language subspace from cross-language same-speaker differences and project embeddings onto its orthogonal complement before cosine scoring with Adaptive Symmetric score normalization. This reduces development EER from 2.97% with cosine and 2.70% with AS-Norm to 2.18% and yields a Codabench evaluation score of 8.40, showing that simple back-end language normalization can rival more complex systems.
Nina Hosseini-Kivanani
University of Luxembourg & Radio T´el´evisioun L¨etzebuerg (RTL), Luxembourg
Multilingual speaker verification remains challenging because language-dependent acoustic variability causes speaker identity to become entangled with linguistic characteristics, degrading generalization across languages. In multilingual training, embeddings often encode language cues with speaker identity, causing speakers to form language-specific clusters. We propose L-Proto, a language-aware episodic prototypical training strategy that constructs language-consistent episodes. By sampling speakers from a single language per episode, L-Proto reduces language-driven variation during training and encourages embeddings to focus more directly on speaker identity. Experiments on the TidyVoice Challenge benchmark demonstrate consistent performance improvements over conventional fine-tuning and random episodic sampling across multiple backbone architectures.
Hyung-Seok Oh, Deok-Hyeon Cho, Seung-Bin Kim +1
Department of Artificial Intelligence, Korea University, Seoul, Korea
Cross-lingual speaker verification (SV) systems typically exhibit performance degradation when enrollment and test utterances are spoken in different languages. However, standard evaluation protocols confound language mismatch with inter-speaker variability, as evaluation is generally performed with different speakers across languages. In this work, we introduce a bilingual same-speaker evaluation set for five Iberian languages, enabling analysis of cross-lingual SV under constant speaker identity. We apply this setup to a HuBERT-based SV system previously shown to exhibit strong language dependence, and analyze results using the Cross-Lingual Transfer Matrix (CLTM) to study pairwise cross-lingual transfer. Our results show that speaker-related variability accounts for part of the observed degradation, but language mismatch remains the main driver of cross-lingual performance loss. These findings provide a more precise characterization of language dependence in cross-lingual SV.
Pol Buitrago, Javier Hernando
1Barcelona Supercomputing Center (BSC), Spain · 2Universitat Polit`ecnica de Catalunya (UPC), Spain