We present the first phoneme-level intelligibility evaluation of speech anonymizers, assessing the performance of ASR-ensemble-based metrics against measured intelligibility from a crowdsourced listening test. Our results show that simple hard-voting ASR metric reaches correlations above 0.9 with human ratings when aggregated by feature, test-type, or condition, provided that multiple ASR models are combined; evaluating stimuli with and without a carrier sentence further improves the correlation at the stimulus level. However, posterior-probability-based confidence metrics bring no gain, which can be traced back to the insufficient calibration of the state-of-the-art open ASR models that were utilized here. All data, code, and evaluation tools are released as open source.
Figures & tables
Condition
Intelligibility (%) ↑
WER (%) ↓
B2 (McAdams)
72.2
9.96
B3 (STTTS)
70.4
4.31
B4 (Neural Codec)
85.2
5.90
B5 (ASR-BN)
78.2
4.44
Qwen-TTS
90.2
N/A
Clean (reference)
91.8
1.84
Table 1: Intelligibility (human ratings) and reported WER by the VPC utility evaluation protocol, per condition. Note: Since Qwen-TTS is not an anonymizer, the WER from the VPC evaluation does not apply.
Aggregation
Carrier
Isoft↑
Ihard↑
ILLR↑
stimuli ( N=2006 )
w/o
0.4039
0.4100
0.3164
w/
0.3371
0.3484
0.3224
w/ + w/o
0.4424
0.4513
0.3852
C ( N=36 )
w/o
0.8159
0 .8208
0.7461
w/
0.7284
0.7251
0.7198
w/ + w/o
0.8085
0.8079
0.7708
Table 2: r2 between human intelligibility & ASR-based scores ( Isoft , Ihard , ILLR ), carrier condition (without / with / both), and aggregation level ( N = number of data points). A Steiger’s Z test confirms that rIsoft and rIhard differ significantly at the stimulus level ( p<0.05 ), but we cannot reject the null hypothesis at higher aggregation levels. rw/o and rw/+w/o differ significantly at the stimulus level ( p<0.05 ).
Model
w/o ↑
w/ ↑
w/ + w/o ↑
Cohere Transcribe
0.2392
0.1922
0.2874
wav2vec2-lv-60-espeak
0.1739
0.2163
0.2419
wav2vec2-xlsr-53-espeak
0.1714
0.1440
0.2041
Whisper-large-v3
0.1379
0.2069
0.2690
Whisper-large-v3-turbo
0.1185
0.1849
0.2424
Distil-Whisper-large-v3.5
0.1888
0.1943
0.2826
Table 3: r2 between listening-test intelligibility according to Eq. ( 1 ) and Isoft , per ASR model and carrier condition at the stimulus aggregation level. Model names are abbreviated. The last row reports the ensemble; the fourth column gives the score when stimuli with and without carrier are combined.
Speaker anonymization (SA) systems modify timbre while leaving regional or non-native accent cues intact, which is problematic because such cues can reveal a speaker's first-language or geographic background and narrow the anonymity set. To address this issue, we present PHONOS, a streaming module for real-time SA that performs accent neutralization in a privacy sense: reducing accent-origin cues by converting non-native segmental realizations toward a chosen target accent domain. Our approach pre-generates golden speaker utterances that preserve source timbre and rhythm but replace foreign segmentals with native ones using silence-aware DTW alignment and zero-shot voice conversion. These utterances supervise a causal accent translator that maps non-native content tokens to native equivalents with at most 40ms look-ahead, trained using joint cross-entropy and CTC losses. Our evaluations show an 81% reduction in non-native accent confidence, with listening-test accentedness ratings consistent with this shift. PHONOS also moves outputs away from the original speaker in embedding space, suggesting lower linkability under an embedding-based proxy, while running with ≤241ms end-to-end latency on a single GPU.
Waris Quamer, Mu-Ruei Tseng, Ghady Nasrallah +1
Department of Computer Science & Engineering, Texas A&M University, College Station, US
Popular ASR test sets adopt inconsistent conventions for numbers, disfluencies, entities, and casing, while standard normalizers erase the format distinctions users care about. Current benchmarks therefore cannot measure whether a model follows user preferences for output style. We introduce PreferenceASR, a test set evaluating ASR systems on their ability to follow natural-language preference instructions across four categories: normalization, entities, disfluencies, and case. Built from seven open-source corpora via a two-stage LLM-assisted pipeline with human verification, it is evaluated with a preference-aware normalizer that selectively skips steps matching the active instruction. Benchmarking four models shows rankings shift across preference types, exposing quality differences traditional evaluation obscures. We publicly release the dataset.
Nithin Rao Koluguri, Sasha Meister, Nikolay Karpov +4
Speech anonymization is commonly evaluated using averagecase metrics such as the equal error rate, which can hide large disparities in re-identification risks across individuals. In this paper, we conduct a large-scale per-speaker privacy analysis using a linkability-based metric under a worst-case scenario. Nearly 5,000 speakers are evaluated across multiple anonymization systems, attacker architectures, and conversation lengths. While linkability scores are highly polarized at the speaker level, the sets of easy to re-identify and hard to re-identify speakers vary substantially across configurations. We show that no single factor explains speaker vulnerability. Instead, the re-identification risk emerges from the interaction between the attacker, the anonymizer, and the amount of available speech. These results challenge the notion of intrinsic speaker-level privacy risks and emphasize the need for evaluation protocols that are explicitly conditioned on the attacker and anonymizer.
Orane Dufour, Paul Magron, Mickael Rouvier +1
Université de Lorraine, CNRS, Inria, LORIA, F-54000 Nancy, France · LIA, Avignon University, F-84911 Avignon, France