Revisiting Label-Free Speaker Embedding Enhancement with vMF Profile Likelihood
Authors: Seunghwan Kim, Jinyong Kim, Sooyoung Yang, Youngjin Ko, Myungjoo Kang
Organizations: Interdisciplinary Program in Artificial Intelligence, Seoul National University, South Korea · Department of Mathematical Sciences, Seoul National University, South Korea · Research Institute of Mathematics, South Korea
Embedding enhancement improves speaker verification under acoustic mismatch without modifying a frozen backbone. Recent work has established a practical label-free setting for this task, but often adopts increasingly structured formulations. Here, the clean target is directly observed during training, making enhancement a matching problem on the unit hypersphere. We model the clean target with a von Mises--Fisher (vMF) likelihood and profile out a sample-wise concentration parameter, yielding a simple closed-form objective with adaptive weighting. Across VoxCeleb1, VoxSRC23, CN-Celeb, VOiCES, and VC-Mix, the proposed method largely preserves the baseline and gives clearer gains on challenging mismatch sets. It also remains stable under a broad single-view recipe, where a recent diffusion baseline becomes less reliable in controlled comparisons. These results suggest that effective label-free embedding enhancement in this setting does not require a highly structured formulation.
Figures & tables
Figure 1: Task overview and geometric intuition of the proposed vMF-PL objective. (a) An unaugmented utterance a and its acoustically degraded variants a~ are encoded by a frozen speaker embedding extractor g(⋅) , producing a clean embedding xc and variant embeddings xn,i . The blue marker denotes xc , and the remaining colors consistently identify the corresponding acoustic variants before and after enhancement. The lightweight enhancement mapping fθ(⋅) moves each degraded embedding toward xc without changing the downstream cosine-scoring pipeline. (b) On the unit hypersphere, faded and saturated markers of the same color denote variant embeddings before and after enhancement, respectively. Arrow thickness reflects the adaptive supervision strength induced by vMF-PL, so variants closer to the clean target receive stronger supervision and are associated with higher implied concentration. Shaded regions indicate relative concentration κi∝1/∥x^c,i−xc∥22 .
Backbone
Method
Vox1-O
Vox1-E
Vox1-H
VoxSRC23
CN-Celeb
VOiCES
VC-Mix
EER
minDCF
EER
minDCF
EER
minDCF
EER
minDCF
EER
minDCF
EER
minDCF
EER
minDCF
ECAPA
Baseline
0.91
0.065
1.17
0.081
2.39
0.152
5.64
0.362
18.28
0.626
6.50
0.366
2.96
0.262
+ SEED [ 23 ]
0.94
0.068
1.22
0.068 ↑
2.41
0.153
5.96
0.377
17.99 ↑
0.664
6.18 ↑
0.341 ↑
3.02
0.260 ↑
+ Ours (vMF-PL)
0.88 ↑
0.062 ↑
1.17 ↑
0.082
2.38 ↑
0.152 ↑
5.65
0.361 ↑
17.89 ↑
0.628
6.17 ↑
0.374
2.82 ↑
0.255 ↑
ResNet-34
Baseline
0.88
0.079
1.07
0.076
2.22
0.148
5.45
0.335
14.54
0.570
5.62
0.300
3.07
0.246
+ SEED [ 23 ]
0.94
0.080
1.10
0.079
2.24
0.149
5.37 ↑
0.328 ↑
13.92 ↑
0.569 ↑
5.63
0.294 ↑
2.68 ↑
0.224 ↑
Table 1: Main results on extensive benchmarks (EER (%) / minDCF@ Ptarget=0.05 ). Lower values are better. Bold values denote the best result for each backbone and metric. The symbol ↑ indicates performance equal to or better than the frozen baseline. All post-processing methods use the same 3-block MLP architecture. Each method is reported under the strongest stable configuration.
Method
Vox1-O
VoxSRC23
VOiCES
VC-Mix
Baseline
0.91
5.64
6.50
2.96
+ SEED (broad recipe)
2.49
10.65
10.85
5.54
+ Ours (vMF-PL)
0.88
5.65
6.17
2.82
Table 2: Controlled comparison under the same broad single-view recipe of Section 4.2 with ECAPA-TDNN. EER (%) is reported. Lower values are better.
Loss
Vox1-O
VoxSRC23
EER
minDCF
EER
minDCF
Standard MSE
1.05
0.084
6.38
0.413
vMF-PL
0.88
0.062
5.65
0.361
Table 3: Ablation of the objective on ECAPA-TDNN. The network, data, and augmentation pipeline are fixed. EER (%) and minDCF@ Ptarget=0.05 are reported. Lower values are better.