ReDimNet2+: Multi-Corpus Data Scaling for Robust Speaker Verification
Organizations: lab260, Yerevan, Armenia · BitmanagerAI, Dubai, UAE · MTUCI, Moscow, Russia
Abstract
Automatic speaker verification must remain reliable across devices, rooms, and compression pipelines. We present ReDimNet2+, which scales training of the compact ReDimNet2 backbone across seven public corpora (63,934 speakers, about 8,675 hours). Analysis of a VoxBlink2 subset reveals a shift in predicted spectral coloration, motivating codec and waveform augmentation alongside this multi-corpus training, large-margin fine-tuning (LMFT), and graph-based retrieval reranking. With random 4-second evaluation windows for all models, ReDimNet2+ LMFT reduces pooled VoxCeleb1 EER from 2.42% to 0.82% and a 26-condition robustness stress-test EER from 7.21% to 1.99%. Under this shared local protocol, it reaches 0.35% EER on VoxCeleb1-O versus 0.787% for the best evaluated WeSpeaker checkpoint. On a VoxBlink2 retrieval subset, reranking improves the final model's Pr@k from 0.7413 to 0.7687.
Figures & tables
| System | EER o | EER e | EER h | EER p | EER ph | Pr@ |
|---|---|---|---|---|---|---|
| ReDimNet2 baseline | 1.601 | 1.725 | 3.039 | 2.420 | 7.214 | 0.5978 |
| ReDimNet2 (FT) | 2.414 | 2.337 | 4.260 | 3.408 | 5.599 | 0.4890 |
| ReDimNet2 (cosine margin schedule) | 2.143 | 2.236 | 4.122 | 3.271 | 5.502 | 0.4813 |
| ReDimNet2 (VC2, 6 epochs) | 1.734 | 1.645 | 3.010 | 2.404 | 4.624 | 0.6770 |
| ReDimNet2 (multi-domain, 6 epochs) | 1.095 | 1.067 | 1.997 | 1.587 | 3.702 | 0.6906 |
| ReDimNet2 (TidyVoice added) | 0.989 | 1.022 | 1.934 | 1.537 | 3.628 | 0.6936 |
| System | Source | Params | EER o | EER e | EER h |
|---|---|---|---|---|---|
| ECAPA-TDNN C=512 [ 5 ] | published | 6.2M | 1.01 | 1.24 | 2.32 |
| ECAPA-TDNN C=1024 [ 5 ] | published | 14.7M | 0.87 | 1.12 | 2.12 |
| CAM++ [ 28 ] | published | 7.2M | 0.71 | 0.85 | 1.66 |
| ECAPA2 [ 23 ] | published | 27.0M | 0.34 | 0.52 | 0.99 |
| WavLM [ 28 ] | published | 324M | 0.52 | 0.63 | 1.34 |
| W2V-BERT 2.0 [ 28 ] | published | 587M | 0.38 | 0.51 | 1.06 |
| Method | Pr@1 | Pr@10 | Pr@45 | VB2 |
|---|---|---|---|---|
| Baseline | 99.9577 | 99.8611 | 99.2404 | 0.7024 |
| Mean chain | 99.9577 | 99.9075 | 99.5299 | 0.7220 |
| Graph rerank | 99.9492 | 99.8798 | 99.5671 | 0.7369 |
| Mean chain + graph | 99.9511 | 99.9003 | 99.6992 | 0.7391 |
| Embedding model | Before | After | |
|---|---|---|---|
| ReDimNet2 baseline | 0.5978 | 0.6906 | +0.0928 |
| ReDimNet2+ pretrained | 0.7024 | 0.7391 | +0.0367 |
| ReDimNet2+ LMFT | 0.7413 | 0.7687 | +0.0274 |
| Corpus | Speakers | Languages | Utterances | Hours | Main domain |
|---|---|---|---|---|---|
| VoxBlink2 [ 13 ] | 11,053 | 18 | 673,277 | 1,457.77 | YouTube / in-the-wild |
| VoxCeleb2 [ 3 ] | 5,994 | Multi | 1,092,009 | 2,442.00 | YouTube / interviews |
| 3D-Speaker [ 30 ] | 10,000 | Mandarin dialects | 579,013 | 1,124.52 | Multi-device, multi-distance |
| CN-Celeb [ 6 ] | 997 | Mandarin | 130,109 | 273.73 | Multi-genre web video |
| CN-Celeb2 [ 10 ] | 1,996 | Mandarin | 529,485 | 1,090.00 | Web-crawled speech |
| TidyVoice [ 7 ] | 6,657 | 62 | 527,484 | 744.60 | Read multilingual speech |
| Duration bucket | Train | Evaluation |
|---|---|---|
| s | 11.0% | 2.5% |
| 2–3 s | 11.7% | 7.2% |
| 3–5 s | 22.4% | 21.2% |
| 5–10 s | 31.7% | 39.4% |
| 10–20 s | 16.7% | 21.8% |
| 20–30 s | 3.9% | 5.0% |
| Statistic | Train | Evaluation |
|---|---|---|
| Median bytes/s | 21,040 | 14,021 |
| p05 bytes/s | 16,283 | 8,989 |
| p95 bytes/s | 26,860 | 20,831 |
| Files with COL | 41,599 (6.2%) | 55,820 (41.4%) |
| Median COL gap | ||
| Component | Baseline | Optimized | Speedup |
|---|---|---|---|
| Random audio window | 5.38 ms | 2.06 ms | 2.6 |
| Speed perturbation | 455 ms | 0.16 ms | |
| PCM A-law codec | 14.33 ms | 8.77 ms | 1.63 |
| G.722 codec | 14.74 ms | 9.52 ms | 1.55 |
| MP3 8 kHz codec | 16.39 ms | 10.69 ms | 1.53 |
| Spectrograms, 4 workers | 1029 samp/s | 1570 samp/s | 1.53 |
| Stage | Probability | Transform family |
| Top-level keep | 0.5 | No waveform corruption |
| Top-level codec | 0.2 | Codec branch only |
| Top-level augment | 0.2 | Waveform branch only |
| Top-level codec+augment | 0.1 | Both branches |
| Filters | 0.3 | Band-pass, band-stop, high-pass, low-pass |
| Additive noise | 0.3 | Colored noise, MUSAN, RawBoost |
| Run | Batch/GPU | GPUs | Accum. | Epochs | LR | Samples |
|---|---|---|---|---|---|---|
| ReDimNet2 (FT) | 64 | 6 | 4 | 22 | 32,200 | |
| ReDimNet2 (cosine margin schedule) | 64 | 6 | 4 | 10 | 32,200 | |
| ReDimNet2 (VC2, 6 epochs) | 64 | 6 | 4 | 6 | 32,200 | |
| ReDimNet2 (multi-domain, 6 epochs) | 48 | 6 | 5 | 6 | 48,300 | |
| ReDimNet2 (multi-domain, 10 epochs) | 48 | 6 | 5 | 10 | 48,300 | |
| ReDimNet2 (TidyVoice added) | 48 | 6 | 5 | 6 | 48,300 |
| System | Training data | Params | EER o | EER e | EER h |
|---|---|---|---|---|---|
| ECAPA-TDNN C=512 [ 5 ] | VoxCeleb2-dev | 6.2M | 1.01 | 1.24 | 2.32 |
| ECAPA-TDNN C=1024 [ 5 ] | VoxCeleb2-dev | 14.7M | 0.87 | 1.12 | 2.12 |
| CAM++ [ 28 ] | VoxCeleb2-dev | 7.2M | 0.71 | 0.85 | 1.66 |
| ECAPA2 [ 23 ] | VoxCeleb2-dev | 27.0M | 0.34 | 0.52 | 0.99 |
| WavLM [ 28 ] | VoxCeleb2-dev | 324M | 0.52 | 0.63 | 1.34 |
| W2V-BERT 2.0 [ 28 ] | VoxCeleb2-dev | 587M | 0.38 | 0.51 | 1.06 |
| System | Runtime | EER o (%) | EER e (%) | EER h (%) |
|---|---|---|---|---|
| ReDimNet2+ LMFT | Torch | 0.351 | 0.523 | 1.055 |
| WeSpeaker CAM++ | ONNX/CUDA | 0.787 | 0.928 | 1.824 |
| WeSpeaker ResNet34-LM | ONNX/CUDA | 0.814 | 0.933 | 1.679 |
| WeSpeaker ECAPA512-LM | ONNX/CUDA | 0.877 | 1.071 | 1.968 |
| Implementation | xRTF | Batch | Workers | EER o | EER p | EER ph |
|---|---|---|---|---|---|---|
| Baseline ONNX FP32 | 707.61 | 32 | 4 | 14.647 | 17.197 | 41.634 |
| Ours Torch FP32 | 59.09 | 32 | 4 | 0.489 | 1.076 | 8.995 |
| Ours Torch BF16 | 82.53 | 32 | 4 | 0.484 | 1.072 | 8.968 |
| Ours ONNX FP32 | 61.87 | 32 | 4 | 0.489 | 1.075 | 9.000 |
| Ours ONNX FP16 | 113.66 | 32 | 4 | 0.489 | 1.076 | 8.984 |
| Ours TensorRT FP32 | 116.43 | 32 | 4 | 0.489 | 1.076 | 9.000 |