Exploiting Acoustic and Content-Oriented Speaker Verification Attacks Against Multilingual Voice Anonymization
Authors: Ridwan Arefeen, Ze Li, Rong Tong, Ming Li, Xiaoxiao Miao
Organizations: Singapore Institute of Technology, Singapore · Wuhan University, China · The Chinese University of Hong Kong, China · Duke Kunshan University, China
Attacker ASV systems for voice anonymization have been studied primarily in English, leaving their behavior in multilingual settings largely unexplored. Conventional ASV has shown that both acoustic and contextual information are important for multilingual speaker verification. Inspired by this, we investigate whether the same holds for attacker ASV on anonymized speech. We evaluate both acoustic- and content-oriented attackers on multilingual anonymized speech and construct a multilingual voice-converted dataset to improve cross-lingual generalization. Our results show that attacker effectiveness depends on the linguistic utility of the anonymized speech. Overall, acoustic-oriented attackers achieve better performance. However, when linguistic information is well preserved, the performance gap between content- and acoustic-oriented attackers narrows compared with conditions involving stronger speech distortion. The multilingual voice-converted dataset further improves performance and partially reduces the cross-lingual gap. These findings highlight the need for more comprehensive attacker modeling and evaluation protocols that consider both privacy and utility, rather than relying on a attacker strategy\footnote{Full code and pretrained models and MultiVC Dataset link are available at: https://github.com/monkeyDarefeen/DAST
Figures & tables
Fig. 1: Overview of the proposed evaluation.
Fig. 2: Average lazy-informed EER of the content-oriented (w2v-BERT-MFA) and acoustic-oriented (WavLM-ECAPA) attackers across six languages, under the lazy-informed threat model on the BM1 and BM2 baselines. Per-language ASR WER/CER is overlaid (right axis).
Language
BM1
BM2
English (WER)
5.4
16.0
German (WER)
9.7
24.2
Spanish (WER)
6.5
25.6
French (WER)
13.2
44.9
Japanese (CER)
9.7
45.1
Chinese (CER)
30.2
70.3
TABLE I: ASR utility metrics across BM1 and BM2 baselines. CER(%) is reported for Japanese and Chinese, WER(%) for the rest.
Fig. 3: Voice Converted Dataset Pipeline
Fig. 4: Language distribution of the evaluation corpora. (Left) Target corpus distribution (TidyVoiceX), where the top 5 languages comprise roughly 70% of the 321,709 utterances. (Right) Source corpus composition, where 70,113 utterances serve as the content reference pool for all seven voice conversion systems.
ID
Voice conversion system
Target fold
VC-1
CosyVoice2 [ 30 ]
Split 1
VC-2
Seed-VC [ 31 ]
Split 2
VC-3
YingMusic-SVC [ 32 ]
Split 3
VC-4
FreeVC [ 33 ]
Split 4
VC-5
KDVC [ 34 ]
Split 5
VC-6
LVC-VC [ 35 ]
Split 6
TABLE II: Voice conversion systems and their assigned target splits. Each system converts 34,022 target utterances using 70,113 source utterances, producing 102,066 converted samples per system.
Threat Model
ID
ASV Model
Training Data
En
De
Fr
Es
Cn
Ja
Avg.
Lazy-BM1
1
w2v-BERT-MFA
+SSTC
16.76
22.20
22.73
18.84
39.76
34.86
25.86
2
w2v-BERT-MFA
+SSTC + MultiVC
11.108
15.336
13.141
10.565
28.160
27.670
17.66
3
WavLM-ECAPA
+SSTC
5.52
11.57
11.93
8.23
31.25
26.13
15.77
4
WavLM-ECAPA
+SSTC + MultiVC
2.887
4.526
2.520
1.636
17.727
19.622
8.99
Lazy-BM2
5
w2v-BERT-MFA
+SSTC
35.80
34.90
39.82
34.46
46.86
45.85
39.61
6
w2v-BERT-MFA
+SSTC + MultiVC
35.708
34.507
38.948
33.972
46.069
44.602
38.97
TABLE III: EER (%) on BM1 and BM2 across six languages. Lower EER indicates a stronger attack. Best results are bolded independently for each block.
Cross-lingual mismatch remains a key source of overall degradation in modern speaker verification. The TidyVoice2026 Challenge targets this setting with text-independent verification, comprising 3,666 training and 808 development speakers in 40 languages and 2,200 evaluation speakers in 38 unseen languages, without language labels at test time. Starting from the official SimAM-ResNet34 baseline pretrained on VoxBlink2 and VoxCeleb2 and fine-tuned on TidyVoice, we revisit Nuisance Attribute Projection (NAP) as a simple language-normalization step in the embedding space. We estimate a compact language subspace from cross-language same-speaker differences and project embeddings onto its orthogonal complement before cosine scoring with Adaptive Symmetric score normalization. This reduces development EER from 2.97% with cosine and 2.70% with AS-Norm to 2.18% and yields a Codabench evaluation score of 8.40, showing that simple back-end language normalization can rival more complex systems.
Nina Hosseini-Kivanani
University of Luxembourg & Radio T´el´evisioun L¨etzebuerg (RTL), Luxembourg
Recent advances in text-to-speech and voice cloning make high-quality spoofing inexpensive and scalable, threatening voice authentication systems, especially automatic speaker verification (ASV). Existing defenses mainly address this threat through binary countermeasures (CMs) for deepfake detection or spoofing-aware speaker verification (SASV), where current systems are dominated by modular ASV-CM fusion and cascaded pipelines. Although large audio language models (LALMs) have shown promise on related audio tasks, including CM and ASV, their use for SASV remains unexplored, despite their capacity to produce natural-language rationales for auditing and robustness beyond discriminative predictions. This work systematically evaluates LALMs for SASV against conventional pipelines under zero-shot prompting, supervised adaptation, reasoning-oriented training, and reinforcement-learning-based optimization. Our results show that pretrained LALMs are near chance in the zero-shot setting, confirming that they are not natively suited to SASV, but that task-specific adaptation closes this gap. We further find that competitive SASV performance can be achieved through several distinct routes. These findings position LALMs as a promising and auditable foundation for unified SASV, while clarifying where conventional cascade systems still lead.
Voice privacy approaches that preserve the anonymity of speakers modify speech in an attempt to break the link with the true identity of the speaker. Current benchmarks measure speaker protection based on signal-to-signal comparisons. In this paper, we introduce an attribute-based perspective, where we measure privacy protection in terms of comparisons between sets of speaker attributes. First, we analyze privacy impact by calculating speaker uniqueness for ground truth attributes, attributes inferred on the original speech, and attributes inferred on speech protected with standard anonymization. Next, we examine a threat scenario involving only a single utterance per speaker and calculate attack error rates. Overall, we observe that inferred attributes still present a risk despite attribute inference errors. Our research points to the importance of considering both attribute-related threats and protection mechanisms in future voice privacy research.
Mehtab Ur Rahman, Martha Larson, Cristian Tejedor-Garcia
Centre for Language Studies · Institute for Computing and Information Sciences