cs.SDMay 20, 2025

Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker Traits and Speech Properties

Authors: Tiantian Feng, Jihwan Lee, Anfeng Xu, Yoonjeong Lee, Thanathai Lertpetchpun, Xuan Shi, Helin Wang, Thomas Thebaud, +4 more

Abstract

We introduce Vox-Profile, a comprehensive benchmark to characterize rich speaker traits and speech properties using speech foundation models. Unlike prior work that targets a single dimension of speaker traits, Vox-Profile provides holistic and multi-dimensional profiles spanning both static speaker traits (e.g., age, sex, accent, voice quality) and dynamic speech traits (e.g., emotion, speech flow). This benchmark is grounded in speech science and linguistics, developed with domain experts to accurately index speaker and speech characteristics. We report benchmark experiments using over 15 publicly available speech datasets and several widely used speech foundation models that target various static and dynamic speaker and speech properties. In addition to benchmark experiments, we showcase several downstream applications supported by Vox-Profile. First, we show that Vox-Profile can augment existing speech recognition datasets to analyze ASR performance variability. Vox-Profile is also used as a tool to evaluate the performance of speech generation systems. Finally, we assess the quality of our automated profiles through comparison with human evaluation and show convergent validity. Vox-Profile is publicly available at: https://github.com/tiantiaf0627/vox-profile-release.

Explore similar work

CardsList
  1. ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions

    Date pendingThomas Thebaud, Junhyeok Lee, Laureano Moro-Velazquez +2Speech SynthesisSpeaker

  2. ChildVox: A Speech, Audio, and Large Audio-Language Model Benchmark in Understanding and Characterizing Sound across Childhood

    May 28, 2026Tiantian Feng, Anfeng Xu, Xuan Shi +10VocalizationsChildren