cs.SDSep 23, 2026

ASR ensembling for phoneme intelligibility evaluation of speech anonymizers

Authors: Victor Ménestrel, Sebastian Möller, Slim Ouni, Dorothea Kolossa

Organizations: Technische Universität Berlin, Faculty IV, Berlin, Germany · Université de Lorraine, CNRS, Inria, Loria, F-54000, Nancy, France

Abstract

We present the first phoneme-level intelligibility evaluation of speech anonymizers, assessing the performance of ASR-ensemble-based metrics against measured intelligibility from a crowdsourced listening test. Our results show that simple hard-voting ASR metric reaches correlations above 0.9 with human ratings when aggregated by feature, test-type, or condition, provided that multiple ASR models are combined; evaluating stimuli with and without a carrier sentence further improves the correlation at the stimulus level. However, posterior-probability-based confidence metrics bring no gain, which can be traced back to the insufficient calibration of the state-of-the-art open ASR models that were utilized here. All data, code, and evaluation tools are released as open source.

Figures & tables

Explore similar work

CardsList
  1. PHONOS: PHOnetic Neutralization for Online Streaming Applications

    Mar 27, 2026Waris Quamer, Mu-Ruei Tseng, Ghady Nasrallah +1Grapheme-To-PhonemeAnonymization

  2. A Large-Scale Per-Speaker Analysis of Re-identification Risk in Speech Anonymization

    Jun 5, 2026Orane Dufour, Paul Magron, Mickael Rouvier +1AnonymizationSpeaker