Exploiting Acoustic and Content-Oriented Speaker Verification Attacks Against Multilingual Voice Anonymization
Authors: Ridwan Arefeen, Ze Li, Rong Tong, Ming Li, Xiaoxiao Miao
Organizations: Singapore Institute of Technology, Singapore · Wuhan University, China · The Chinese University of Hong Kong, China · Duke Kunshan University, China
Attacker ASV systems for voice anonymization have been studied primarily in English, leaving their behavior in multilingual settings largely unexplored. Conventional ASV has shown that both acoustic and contextual information are important for multilingual speaker verification. Inspired by this, we investigate whether the same holds for attacker ASV on anonymized speech. We evaluate both acoustic- and content-oriented attackers on multilingual anonymized speech and construct a multilingual voice-converted dataset to improve cross-lingual generalization. Our results show that attacker effectiveness depends on the linguistic utility of the anonymized speech. Overall, acoustic-oriented attackers achieve better performance. However, when linguistic information is well preserved, the performance gap between content- and acoustic-oriented attackers narrows compared with conditions involving stronger speech distortion. The multilingual voice-converted dataset further improves performance and partially reduces the cross-lingual gap. These findings highlight the need for more comprehensive attacker modeling and evaluation protocols that consider both privacy and utility, rather than relying on a attacker strategy\footnote{Full code and pretrained models and MultiVC Dataset link are available at: https://github.com/monkeyDarefeen/DAST
Figures & tables
Fig. 1: Overview of the proposed evaluation.
Fig. 2: Average lazy-informed EER of the content-oriented (w2v-BERT-MFA) and acoustic-oriented (WavLM-ECAPA) attackers across six languages, under the lazy-informed threat model on the BM1 and BM2 baselines. Per-language ASR WER/CER is overlaid (right axis).
Language
BM1
BM2
English (WER)
5.4
16.0
German (WER)
9.7
24.2
Spanish (WER)
6.5
25.6
French (WER)
13.2
44.9
Japanese (CER)
9.7
45.1
Chinese (CER)
30.2
70.3
TABLE I: ASR utility metrics across BM1 and BM2 baselines. CER(%) is reported for Japanese and Chinese, WER(%) for the rest.
Fig. 3: Voice Converted Dataset Pipeline
Fig. 4: Language distribution of the evaluation corpora. (Left) Target corpus distribution (TidyVoiceX), where the top 5 languages comprise roughly 70% of the 321,709 utterances. (Right) Source corpus composition, where 70,113 utterances serve as the content reference pool for all seven voice conversion systems.
ID
Voice conversion system
Target fold
VC-1
CosyVoice2 [ 30 ]
Split 1
VC-2
Seed-VC [ 31 ]
Split 2
VC-3
YingMusic-SVC [ 32 ]
Split 3
VC-4
FreeVC [ 33 ]
Split 4
VC-5
KDVC [ 34 ]
Split 5
VC-6
LVC-VC [ 35 ]
Split 6
TABLE II: Voice conversion systems and their assigned target splits. Each system converts 34,022 target utterances using 70,113 source utterances, producing 102,066 converted samples per system.
Threat Model
ID
ASV Model
Training Data
En
De
Fr
Es
Cn
Ja
Avg.
Lazy-BM1
1
w2v-BERT-MFA
+SSTC
16.76
22.20
22.73
18.84
39.76
34.86
25.86
2
w2v-BERT-MFA
+SSTC + MultiVC
11.108
15.336
13.141
10.565
28.160
27.670
17.66
3
WavLM-ECAPA
+SSTC
5.52
11.57
11.93
8.23
31.25
26.13
15.77
4
WavLM-ECAPA
+SSTC + MultiVC
2.887
4.526
2.520
1.636
17.727
19.622
8.99
Lazy-BM2
5
w2v-BERT-MFA
+SSTC
35.80
34.90
39.82
34.46
46.86
45.85
39.61
6
w2v-BERT-MFA
+SSTC + MultiVC
35.708
34.507
38.948
33.972
46.069
44.602
38.97
TABLE III: EER (%) on BM1 and BM2 across six languages. Lower EER indicates a stronger attack. Best results are bolded independently for each block.