cs.SDSep 29, 2026

ReDimNet2+: Multi-Corpus Data Scaling for Robust Speaker Verification

Authors: Kirill Borodin, Vasilii Kudryavtsev, Maxim Maslov, Grach Mkrtchian

Organizations: lab260, Yerevan, Armenia · BitmanagerAI, Dubai, UAE · MTUCI, Moscow, Russia

Abstract

Automatic speaker verification must remain reliable across devices, rooms, and compression pipelines. We present ReDimNet2+, which scales training of the compact ReDimNet2 backbone across seven public corpora (63,934 speakers, about 8,675 hours). Analysis of a VoxBlink2 subset reveals a shift in predicted spectral coloration, motivating codec and waveform augmentation alongside this multi-corpus training, large-margin fine-tuning (LMFT), and graph-based retrieval reranking. With random 4-second evaluation windows for all models, ReDimNet2+ LMFT reduces pooled VoxCeleb1 EER from 2.42% to 0.82% and a 26-condition robustness stress-test EER from 7.21% to 1.99%. Under this shared local protocol, it reaches 0.35% EER on VoxCeleb1-O versus 0.787% for the best evaluated WeSpeaker checkpoint. On a VoxBlink2 retrieval subset, reranking improves the final model's Pr@k from 0.7413 to 0.7687.

Figures & tables

Explore similar work

CardsList
  1. SpeakerCard-1M: An Evidence-Grounded Corpus for In-the-Wild Speaker Verification

    Jun 2, 2026Junyi Peng, Oldřich Plchot, Xiao Song +9Automatic Speaker VerificationSpeaker

  2. Enhancing Speaker Verification with Whispered Speech via Post-Processing

    Apr 22, 2026Magdalena Gołębiowska, Piotr SygaAutomatic Speaker VerificationSpeaker

  3. MECT: Mixture of Experts with CNN-Transformer Network for Speaker verification

    Sep 21, 2026Yu Zheng, Jinghan Peng, ChangHao Zhang +2SpeakerCnn-Transformer Tradeoff