cs.SDJun 19, 2026

LISE : Listenable Interpretable Speaker Embeddings

Authors: Xiaoliang WuChongxin GanKe LiuPeter BellJennifer Williams

Organizations: University of Southampton, United Kingdom · The Hong Kong Polytechnic University, Hong Kong SAR, China · University of Edinburgh, United Kingdom

Abstract

Deep neural network-based automatic speaker verification (ASV) systems achieve impressive performance but their embedding representations remain opaque, lacking a structured and perceptually verifiable explanation of the vocal characteristics they encode. Existing approaches either require annotation of speaker attributes or introduce alternative representations whose interpretability is unvalidated with listeners. We propose Listenable Interpretable Speaker Embeddings (LISE), a label-free framework that decomposes pretrained speaker embeddings into a small set of components. This decomposition yields a structured representation that supports the analysis of what information has been encoded by speaker embeddings. LISE preserves ASV performance with negligible EER degradation on x-vector and ECAPA-TDNN. Crucially, the interpretability of these components for human listeners is demonstrated through listening experiments, where participants distinguished speakers with 83.9% accuracy.

Explore similar work

CardsList