We present the first phoneme-level intelligibility evaluation of speech anonymizers, assessing the performance of ASR-ensemble-based metrics against measured intelligibility from a crowdsourced listening test. Our results show that simple hard-voting ASR metric reaches correlations above 0.9 with human ratings when aggregated by feature, test-type, or condition, provided that multiple ASR models are combined; evaluating stimuli with and without a carrier sentence further improves the correlation at the stimulus level. However, posterior-probability-based confidence metrics bring no gain, which can be traced back to the insufficient calibration of the state-of-the-art open ASR models that were utilized here. All data, code, and evaluation tools are released as open source.
Figures & tables
Condition
Intelligibility (%) ↑
WER (%) ↓
B2 (McAdams)
72.2
9.96
B3 (STTTS)
70.4
4.31
B4 (Neural Codec)
85.2
5.90
B5 (ASR-BN)
78.2
4.44
Qwen-TTS
90.2
N/A
Clean (reference)
91.8
1.84
Table 1: Intelligibility (human ratings) and reported WER by the VPC utility evaluation protocol, per condition. Note: Since Qwen-TTS is not an anonymizer, the WER from the VPC evaluation does not apply.
Aggregation
Carrier
Isoft↑
Ihard↑
ILLR↑
stimuli ( N=2006 )
w/o
0.4039
0.4100
0.3164
w/
0.3371
0.3484
0.3224
w/ + w/o
0.4424
0.4513
0.3852
C ( N=36 )
w/o
0.8159
0 .8208
0.7461
w/
0.7284
0.7251
0.7198
w/ + w/o
0.8085
0.8079
0.7708
Table 2: r2 between human intelligibility & ASR-based scores ( Isoft , Ihard , ILLR ), carrier condition (without / with / both), and aggregation level ( N = number of data points). A Steiger’s Z test confirms that rIsoft and rIhard differ significantly at the stimulus level ( p<0.05 ), but we cannot reject the null hypothesis at higher aggregation levels. rw/o and rw/+w/o differ significantly at the stimulus level ( p<0.05 ).
Model
w/o ↑
w/ ↑
w/ + w/o ↑
Cohere Transcribe
0.2392
0.1922
0.2874
wav2vec2-lv-60-espeak
0.1739
0.2163
0.2419
wav2vec2-xlsr-53-espeak
0.1714
0.1440
0.2041
Whisper-large-v3
0.1379
0.2069
0.2690
Whisper-large-v3-turbo
0.1185
0.1849
0.2424
Distil-Whisper-large-v3.5
0.1888
0.1943
0.2826
Table 3: r2 between listening-test intelligibility according to Eq. ( 1 ) and Isoft , per ASR model and carrier condition at the stimulus aggregation level. Model names are abbreviated. The last row reports the ensemble; the fourth column gives the score when stimuli with and without carrier are combined.