Automatic speaker verification must remain reliable across devices, rooms, and compression pipelines. We present ReDimNet2+, which scales training of the compact ReDimNet2 backbone across seven public corpora (63,934 speakers, about 8,675 hours). Analysis of a VoxBlink2 subset reveals a shift in predicted spectral coloration, motivating codec and waveform augmentation alongside this multi-corpus training, large-margin fine-tuning (LMFT), and graph-based retrieval reranking. With random 4-second evaluation windows for all models, ReDimNet2+ LMFT reduces pooled VoxCeleb1 EER from 2.42% to 0.82% and a 26-condition robustness stress-test EER from 7.21% to 1.99%. Under this shared local protocol, it reaches 0.35% EER on VoxCeleb1-O versus 0.787% for the best evaluated WeSpeaker checkpoint. On a VoxBlink2 retrieval subset, reranking improves the final model's Pr@k from 0.7413 to 0.7687.
Figures & tables
Figure 1: Training-time data pipeline. Audio is indexed through Parquet metadata, decoded as a random window, augmented at waveform and codec level, batched as waveforms, and converted to spectrogram features on the GPU.
Figure 2: Examples of codec transformations used to model spectral coloration and bandwidth limitation: the same utterance after different codec presets.
System
EER o
EER e
EER h
EER p
EER ph
Pr@ k
ReDimNet2 baseline
1.601
1.725
3.039
2.420
7.214
0.5978
ReDimNet2 (FT)
2.414
2.337
4.260
3.408
5.599
0.4890
ReDimNet2 (cosine margin schedule)
2.143
2.236
4.122
3.271
5.502
0.4813
ReDimNet2 (VC2, 6 epochs)
1.734
1.645
3.010
2.404
4.624
0.6770
ReDimNet2 (multi-domain, 6 epochs)
1.095
1.067
1.997
1.587
3.702
0.6906
ReDimNet2 (TidyVoice added)
0.989
1.022
1.934
1.537
3.628
0.6936
Table 1: Main ReDimNet2+ development path on VoxCeleb1 (4-second evaluation windows; EER, %). EER p pools O/E/H scores; EER ph is the 26-condition codec/waveform stress test; Pr@ k is the VoxBlink2-subset retrieval metric before reranking.
System
Source
Params
EER o
EER e
EER h
ECAPA-TDNN C=512 [ 5 ]
published
6.2M
1.01
1.24
2.32
ECAPA-TDNN C=1024 [ 5 ]
published
14.7M
0.87
1.12
2.12
CAM++ [ 28 ]
published
7.2M
0.71
0.85
1.66
ECAPA2 [ 23 ]
published
27.0M
0.34
0.52
0.99
WavLM [ 28 ]
published
324M
0.52
0.63
1.34
W2V-BERT 2.0 [ 28 ]
published
587M
0.38
0.51
1.06
Table 2: VoxCeleb1-O/E/H EER (%) comparison with other models. “Published” values use their source protocols (VoxCeleb2-dev training); “local, 4 s” values use our shared 4-second window (Sec. 4.1 ). Bold: best local value among the systems in this table.
Method
Pr@1
Pr@10
Pr@45
VB2
Baseline
99.9577
99.8611
99.2404
0.7024
Mean chain
99.9577
99.9075
99.5299
0.7220
Graph rerank
99.9492
99.8798
99.5671
0.7369
Mean chain + graph
99.9511
99.9003
99.6992
0.7391
Table 3: Reranking ablation on ReDimNet2+ pretrained embeddings on two retrieval datasets. Pr@1/10/45 are measured on VoxCeleb1 (%); VB2 is Pr@ k on the VoxBlink2 subset (fraction). Bold: best per column.
Embedding model
Before
After
Δ
ReDimNet2 baseline
0.5978
0.6906
+0.0928
ReDimNet2+ pretrained
0.7024
0.7391
+0.0367
ReDimNet2+ LMFT
0.7413
0.7687
+0.0274
Table 4: VoxBlink2-subset retrieval Pr@ k before and after the full reranking pipeline (mean chain + graph, same settings) for three embedding models.
Corpus
Speakers
Languages
Utterances
Hours
Main domain
VoxBlink2 [ 13 ]
11,053
∼ 18
673,277
1,457.77
YouTube / in-the-wild
VoxCeleb2 [ 3 ]
5,994
Multi
1,092,009
2,442.00
YouTube / interviews
3D-Speaker [ 30 ]
10,000
Mandarin dialects
579,013
1,124.52
Multi-device, multi-distance
CN-Celeb [ 6 ]
997
Mandarin
130,109
273.73
Multi-genre web video
CN-Celeb2 [ 10 ]
1,996
Mandarin
529,485
1,090.00
Web-crawled speech
TidyVoice [ 7 ]
6,657
62
527,484
744.60
Read multilingual speech
Table 5: Training Corpora Used by ReDimNet2+
Duration bucket
Train
Evaluation
<2 s
11.0%
2.5%
2–3 s
11.7%
7.2%
3–5 s
22.4%
21.2%
5–10 s
31.7%
39.4%
10–20 s
16.7%
21.8%
20–30 s
3.9%
5.0%
Table 6: Duration Distribution in the VoxBlink2 Subset
Statistic
Train
Evaluation
Median bytes/s
21,040
14,021
p05 bytes/s
16,283
8,989
p95 bytes/s
26,860
20,831
Files with COL <2.0
41,599 (6.2%)
55,820 (41.4%)
Median COL gap
−0.977
Table 7: Byte-Rate and Coloration Shift Between Train and the Pr@ k Evaluation Subset
Component
Baseline
Optimized
Speedup
Random audio window
5.38 ms
2.06 ms
2.6 ×
Speed perturbation
455 ms
0.16 ms
>2800×
PCM A-law codec
14.33 ms
8.77 ms
1.63 ×
G.722 codec
14.74 ms
9.52 ms
1.55 ×
MP3 8 kHz codec
16.39 ms
10.69 ms
1.53 ×
Spectrograms, 4 workers
1029 samp/s
1570 samp/s
1.53 ×
Table 8: Preprocessing and Feature-Extraction Speedups
Stage
Probability
Transform family
Top-level keep
0.5
No waveform corruption
Top-level codec
0.2
Codec branch only
Top-level augment
0.2
Waveform branch only
Top-level codec+augment
0.1
Both branches
Filters
0.3
Band-pass, band-stop, high-pass, low-pass
Additive noise
0.3
Colored noise, MUSAN, RawBoost
Table 9: Augmentation Policy Used During Training
Figure 3: Examples of waveform-level transformations: clean mel reference, colored noise, RawBoost, band-pass filtering, and speed perturbation.
Run
Batch/GPU
GPUs
Accum.
Epochs
LR
Samples
ReDimNet2 (FT)
64
6
4
22
5⋅10−5
32,200
ReDimNet2 (cosine margin schedule)
64
6
4
10
10−4
32,200
ReDimNet2 (VC2, 6 epochs)
64
6
4
6
9⋅10−5
32,200
ReDimNet2 (multi-domain, 6 epochs)
48
6
5
6
9⋅10−5
48,300
ReDimNet2 (multi-domain, 10 epochs)
48
6
5
10
9⋅10−5
48,300
ReDimNet2 (TidyVoice added)
48
6
5
6
9⋅10−5
48,300
Table 10: Training Hyperparameters for the Main Runs
System
Training data
Params
EER o
EER e
EER h
ECAPA-TDNN C=512 [ 5 ]
VoxCeleb2-dev
6.2M
1.01
1.24
2.32
ECAPA-TDNN C=1024 [ 5 ]
VoxCeleb2-dev
14.7M
0.87
1.12
2.12
CAM++ [ 28 ]
VoxCeleb2-dev
7.2M
0.71
0.85
1.66
ECAPA2 [ 23 ]
VoxCeleb2-dev
27.0M
0.34
0.52
0.99
WavLM [ 28 ]
VoxCeleb2-dev
324M
0.52
0.63
1.34
W2V-BERT 2.0 [ 28 ]
VoxCeleb2-dev
587M
0.38
0.51
1.06
Table 11: Published VoxCeleb1 results under their source protocols. ReDimNet2-B6 uses full utterances; our results use 4-second windows. The local checkpoint comparison is given in Table 2 .
Figure 4: Data scaling trends for VoxCeleb1-O, VoxCeleb1-E, and VoxCeleb1-H as a function of training speakers and training utterances.
System
Runtime
EER o (%)
EER e (%)
EER h (%)
ReDimNet2+ LMFT
Torch
0.351
0.523
1.055
WeSpeaker CAM++
ONNX/CUDA
0.787
0.928
1.824
WeSpeaker ResNet34-LM
ONNX/CUDA
0.814
0.933
1.679
WeSpeaker ECAPA512-LM
ONNX/CUDA
0.877
1.071
1.968
Table 12: Local VoxCeleb1-O/E/H Comparison with 4-Second Windows
Implementation
xRTF
Batch
Workers
EER o
EER p
EER ph
Baseline ONNX FP32
707.61
32
4
14.647
17.197
41.634
Ours Torch FP32
59.09
32
4
0.489
1.076
8.995
Ours Torch BF16
82.53
32
4
0.484
1.072
8.968
Ours ONNX FP32
61.87
32
4
0.489
1.075
9.000
Ours ONNX FP16
113.66
32
4
0.489
1.076
8.984
Ours TensorRT FP32
116.43
32
4
0.489
1.076
9.000
Table 13: Inference Speed and Quality on Fixed 6-Second Segments. xRTF is audio duration divided by processing time; higher is faster
Modern speaker verification (SV) systems rely on speaker embeddings that are effective but difficult to interpret or query in natural language. Most existing speech-text corpora target controllable synthesis or utterance-level captioning, offering limited speaker-level supervision for in-the-wild speaker recognition. This paper introduces SpeakerCard-1M, a bilingual speaker resource for evidence-grounded SV, derived from VoxCeleb1/2 and CN-Celeb1/2, where the ``-1M'' suffix refers to the 1.78M utterance-level captions contained in the release. We adopt a tool-first, LLM-last approach in which ten acoustic probes produce field-level evidence, the evidence is aggregated into speaker profiles under a schema that separates relatively stable traits from utterance-level states, and bilingual Speaker Cards are rendered by a constrained LLM that sees only the structured fields. The release includes 56.7k Speaker Card records over 10.2k speakers, 1.78M utterance-level captions, and speaker-ID-disjoint hard-negative triplets. We further define two SV-oriented cross-modal protocols, bidirectional Speaker-Text Retrieval (T2S-R / S2T-R) and Attribute-Conditioned Verification (AC-Verify), and compare a dual-encoder baseline against recent audio language models under a zero-shot forced-choice setting. Joint audio-text training costs only 0.31% absolute EER on VoxCeleb1-O relative to the audio-only baseline. Under a style-symmetric LLM-generated counterfactual protocol, eight recent audio language models (7B-30B+ parameters, both open- and closed-source) score 49-77% on pitch-level AC-Verify in a 2-way forced-choice setting, compared with 88.66% for our dual encoder.
Junyi Peng, Oldřich Plchot, Xiao Song +9
Brno University of Technology, Czechia · Peking University, China · Xiaomi, China +3
Speaker verification is a task of confirming an individual's identity through the analysis of their voice. Whispered speech differs from phonated speech in acoustic characteristics, which degrades the performance of speaker verification systems in real-life scenarios, including avoiding fully phonated speech to protect privacy, disrupt others, or when the lack of full vocalization is dictated by a disease. In this paper we propose a model with a training recipe to obtain more robust representations against whispered speech hindrances. The proposed system employs an encoder--decoder structure built atop a fine-tuned speaker verification backbone, optimized jointly using cosine similarity--based classification and triplet loss. We gain relative improvement of 22.26% compared to the baseline (baseline 6.77% vs ours 5.27%) in normal vs whispered speech trials, achieving AUC of 98.16%. In tests comparing whispered to whispered, our model attains an EER of 1.88% with AUC equal to 99.73%, which represents a 15% relative enhancement over the prior leading ReDimNet-B2. We also offer a summary of the most popular and state-of-the-art speaker verification models in terms of their performance with whispered speech. Additionally, we evaluate how these models perform under noisy audios, obtaining that generally the same relative level of noise degrades the performance of speaker verification more significantly on whispered speech than on normal speech.
Magdalena Gołębiowska, Piotr Syga
Department of Artificial Intelligence, Wroclaw University of Science and Technology, Wybrzeze Wyspianskiego 27, Wroclaw, 50-370, Poland
In this paper, we propose MECT, a speaker verification model that integrates the Mixture-of-Experts (MoE) mechanism into a CNN-Transformer backbone with optimized block structure and stacking scheme. Specifically, we investigated four MoE variants that span utterance-level and frame-level granularity with dense and sparse routing strategies. The MoE mechanism proves to be effective over the baseline without MoE with only a small increase in parameters. We further scale MECT to a series of model sizes, all maintaining compact parameters and low computational complexity. In particular, MECT-B2 achieves state-of-the-art performance on VoxCeleb1 and delivers strong results on CN-Celeb, demonstrating its effectiveness across diverse datasets. In addition, we establish a streaming inference paradigm through causal retraining, which maintains strong performance at a chunk size of 100ms.