Revisiting Label-Free Speaker Embedding Enhancement with vMF Profile Likelihood
Authors: Seunghwan Kim, Jinyong Kim, Sooyoung Yang, Youngjin Ko, Myungjoo Kang
Organizations: Interdisciplinary Program in Artificial Intelligence, Seoul National University, South Korea · Department of Mathematical Sciences, Seoul National University, South Korea · Research Institute of Mathematics, South Korea
Embedding enhancement improves speaker verification under acoustic mismatch without modifying a frozen backbone. Recent work has established a practical label-free setting for this task, but often adopts increasingly structured formulations. Here, the clean target is directly observed during training, making enhancement a matching problem on the unit hypersphere. We model the clean target with a von Mises--Fisher (vMF) likelihood and profile out a sample-wise concentration parameter, yielding a simple closed-form objective with adaptive weighting. Across VoxCeleb1, VoxSRC23, CN-Celeb, VOiCES, and VC-Mix, the proposed method largely preserves the baseline and gives clearer gains on challenging mismatch sets. It also remains stable under a broad single-view recipe, where a recent diffusion baseline becomes less reliable in controlled comparisons. These results suggest that effective label-free embedding enhancement in this setting does not require a highly structured formulation.
Figures & tables
Figure 1: Task overview and geometric intuition of the proposed vMF-PL objective. (a) An unaugmented utterance a and its acoustically degraded variants a~ are encoded by a frozen speaker embedding extractor g(⋅) , producing a clean embedding xc and variant embeddings xn,i . The blue marker denotes xc , and the remaining colors consistently identify the corresponding acoustic variants before and after enhancement. The lightweight enhancement mapping fθ(⋅) moves each degraded embedding toward xc without changing the downstream cosine-scoring pipeline. (b) On the unit hypersphere, faded and saturated markers of the same color denote variant embeddings before and after enhancement, respectively. Arrow thickness reflects the adaptive supervision strength induced by vMF-PL, so variants closer to the clean target receive stronger supervision and are associated with higher implied concentration. Shaded regions indicate relative concentration κi∝1/∥x^c,i−xc∥22 .
Backbone
Method
Vox1-O
Vox1-E
Vox1-H
VoxSRC23
CN-Celeb
VOiCES
VC-Mix
EER
minDCF
EER
minDCF
EER
minDCF
EER
minDCF
EER
minDCF
EER
minDCF
EER
minDCF
ECAPA
Baseline
0.91
0.065
1.17
0.081
2.39
0.152
5.64
0.362
18.28
0.626
6.50
0.366
2.96
0.262
+ SEED [ 23 ]
0.94
0.068
1.22
0.068 ↑
2.41
0.153
5.96
0.377
17.99 ↑
0.664
6.18 ↑
0.341 ↑
3.02
0.260 ↑
+ Ours (vMF-PL)
0.88 ↑
0.062 ↑
1.17 ↑
0.082
2.38 ↑
0.152 ↑
5.65
0.361 ↑
17.89 ↑
0.628
6.17 ↑
0.374
2.82 ↑
0.255 ↑
ResNet-34
Baseline
0.88
0.079
1.07
0.076
2.22
0.148
5.45
0.335
14.54
0.570
5.62
0.300
3.07
0.246
+ SEED [ 23 ]
0.94
0.080
1.10
0.079
2.24
0.149
5.37 ↑
0.328 ↑
13.92 ↑
0.569 ↑
5.63
0.294 ↑
2.68 ↑
0.224 ↑
Table 1: Main results on extensive benchmarks (EER (%) / minDCF@ Ptarget=0.05 ). Lower values are better. Bold values denote the best result for each backbone and metric. The symbol ↑ indicates performance equal to or better than the frozen baseline. All post-processing methods use the same 3-block MLP architecture. Each method is reported under the strongest stable configuration.
Method
Vox1-O
VoxSRC23
VOiCES
VC-Mix
Baseline
0.91
5.64
6.50
2.96
+ SEED (broad recipe)
2.49
10.65
10.85
5.54
+ Ours (vMF-PL)
0.88
5.65
6.17
2.82
Table 2: Controlled comparison under the same broad single-view recipe of Section 4.2 with ECAPA-TDNN. EER (%) is reported. Lower values are better.
Loss
Vox1-O
VoxSRC23
EER
minDCF
EER
minDCF
Standard MSE
1.05
0.084
6.38
0.413
vMF-PL
0.88
0.062
5.65
0.361
Table 3: Ablation of the objective on ECAPA-TDNN. The network, data, and augmentation pipeline are fixed. EER (%) and minDCF@ Ptarget=0.05 are reported. Lower values are better.
Using speaker embeddings as conditioning can strengthen speech enhancement, but most methods either require clean enrollment audio or rely on embeddings extracted from noisy speech, which are fragile under noise and domain shift. We propose G-MaP-SE, a guided enhancement framework that builds a clean-speech embedding prior with a Gaussian Mixture Model (GMM) and refines a noisy conditioning embedding by matching it to this prior. The matched prior embedding is then injected into a time-frequency enhancement backbone via a lightweight gated fusion module. Experiments on VoiceBank+DEMAND and DNS Challenge 2020 datasets show that the proposed prior matching consistently outperforms noisy conditioning and substantially narrows the gap to an oracle clean-conditioning upper bound, while requiring no enrollment audio at inference time. The code, audio samples, and checkpoint are available.
Deep neural network-based automatic speaker verification (ASV) systems achieve impressive performance but their embedding representations remain opaque, lacking a structured and perceptually verifiable explanation of the vocal characteristics they encode. Existing approaches either require annotation of speaker attributes or introduce alternative representations whose interpretability is unvalidated with listeners. We propose Listenable Interpretable Speaker Embeddings (LISE), a label-free framework that decomposes pretrained speaker embeddings into a small set of components. This decomposition yields a structured representation that supports the analysis of what information has been encoded by speaker embeddings. LISE preserves ASV performance with negligible EER degradation on x-vector and ECAPA-TDNN. Crucially, the interpretability of these components for human listeners is demonstrated through listening experiments, where participants distinguished speakers with 83.9% accuracy.
Xiaoliang Wu, Chongxin Gan, Ke Liu +2
University of Southampton, United Kingdom · The Hong Kong Polytechnic University, Hong Kong SAR, China · University of Edinburgh, United Kingdom
Automatic speaker verification must remain reliable across devices, rooms, and compression pipelines. We present ReDimNet2+, which scales training of the compact ReDimNet2 backbone across seven public corpora (63,934 speakers, about 8,675 hours). Analysis of a VoxBlink2 subset reveals a shift in predicted spectral coloration, motivating codec and waveform augmentation alongside this multi-corpus training, large-margin fine-tuning (LMFT), and graph-based retrieval reranking. With random 4-second evaluation windows for all models, ReDimNet2+ LMFT reduces pooled VoxCeleb1 EER from 2.42% to 0.82% and a 26-condition robustness stress-test EER from 7.21% to 1.99%. Under this shared local protocol, it reaches 0.35% EER on VoxCeleb1-O versus 0.787% for the best evaluated WeSpeaker checkpoint. On a VoxBlink2 retrieval subset, reranking improves the final model's Pr@k from 0.7413 to 0.7687.