eess.ASSep 27, 2026

Rethinking Automated Voice Similarity by Shifting from EER to Embedding Geometry

Authors: Szu-Chi Chen, Jia-Kai Dong, Yi-Cheng Lin, Sung-Feng Huang, Hung-yi Lee

Organizations: National Taiwan University, Taipei, Taiwan · NVIDIA, Taiwan · Artificial Intelligence Center of Research Excellence (NTU AI-CoRE), NTU, Taiwan

Abstract

Speaker verification (SV) models are commonly assumed to better capture nuances among speaker characteristics as verification accuracy improves, leading to their widespread use as automated proxies for human voice similarity in speech generation tasks. However, by establishing a human perceptual alignment metric and conducting systematic analysis, we demonstrate that perceptual alignment is governed far more by how a model is trained (its learning objective) than by how well it performs (EER). Notably, standard margin-based classification losses (e.g., AAM-Softmax) yield substantially lower perceptual alignment than prototypical metric losses, while EER itself fails to track human judgment, directly challenging the community's implicit assumption. We trace this divergence to embedding geometry, where a model's effective dimensionality (deffd_{\mathrm{eff}}) tracks perceptual alignment with a −0.95-0.95 rank correlation, revealing that the dimensional spread favored by classification losses fundamentally clashes with the low-dimensional nature of human voice perception. Imposing a dimensionality bottleneck compresses deffd_{\mathrm{eff}} and raises perceptual alignment (ρalignρ_{\mathrm{align}}) from 0.08 to 0.74, establishing a principled geometric criterion for evaluating voice similarity.

Figures & tables

Explore similar work

CardsList
  1. Do speech foundation models perceive speaker similarity as humans do?

    Jun 4, 2026Minoru Kishi, Hayato Yagi, Shinnosuke Takamichi +1Speaker SimilaritySpeaker

  2. Simple Language Normalization Wins: Cross-Lingual Speaker Verification for the TidyVoice 2026 Challenge

    Jul 24, 2026Nina Hosseini-KivananiAutomatic Speaker VerificationCross-Lingual Consistency