Vietnamese speech research is constrained by resources that isolate automatic speech recognition from speaker, dialect, code-switching, and deepfake analysis. We introduce VietPrism, an open, multi-domain corpus that brings these dimensions together at scale: 993.4 hours and 403,941 bona fide utterances from 1,262 verified speakers across 8,388 real-world videos. To our knowledge, it is the first large-scale Vietnamese corpus to jointly provide transcripts, consistent speaker identities, five dialect groups, and naturally occurring Vietnamese--English code-switching, which constitutes nearly half of the corpus by duration. We further create over 3.1K hours of spoof speech with four open-source and commercial synthesis systems. Every spoof is conditioned on a verified speaker reference and paired with a transcript- and speaker-matched bona fide utterance, enabling unique controlled evaluation with reduced lexical and identity confounds. Zero-shot evaluation of five pretrained multilingual detectors reveals striking brittleness: EER greatly varies across detector--generator pairings, while recent multilingual detector DFA-1B degrades from 16.3% to 33.6% as speaker similarity increases. Dialect-stratified results expose further model-dependent disparities. By unifying natural linguistic diversity with controlled spoof generation, VietPrism provides a challenging foundation for Vietnamese speech modeling and trustworthy audio-deepfake detection.
Figures & tables
Dataset
CS
Dial.
Spk. B/S
Hrs. B/S
Avg. B/S
Tr
Utt. B/S
Speech corpora
VNSC [ 13 ]
-
3
50
100
-
-
-
VN-LVCSR [ 28 ]
-
-
-
25
-
-
-
VIVOS [ 15 ]
-
-
65
15
4.5
✓
12K
VDSPEC [ 10 ]
-
3
150
45.12
10
-
-
Viettel-CC [ 21 ]
-
-
-
85.8
-
-
-
Table 1: Comparison of VietPrism with existing Vietnamese speech and audio-deepfake datasets.
Figure 1: Overview of the VietPrism data generation pipeline. Human icons depict steps where we involve human screening.
Sub-category
Percentage
# Utt.
# Hours
Dialect
South
46.7%
175,187
463.6
North
50.0%
215,266
496.7
Central
2.0%
8,877
19.8
Southwest
0.7%
4,431
7.2
Overseas Vietnamese
0.6%
2,172
5.9
Topic
Real-estate
19.8%
65,093
196.3
Table 2: VietPrism distribution by dialect, topic, language, and gender. Percentages are duration-based except gender, which is speaker-based. Code-switched segments average 2.83 English words (8.12% of words).
Synthesizer
MOS ↑
Spk. Sim. ↑
Videos
Hours
#Spk.
#Seg.
Bona fide
2.88
—
8,388
993.4
1,262
403,941
MiniMax
2.93
0.55
8,388
1,032.9
1,262
403,941
OmniVoice
2.54
0.74
4,610
718.7
810
293,358
Higgs Audio v3
2.85
0.72
4,610
719.26
810
293,881
VoxCPM2
2.42
0.76
4,609
719.26
810
293,878
Overall (spoof)
2.71
0.68
8,388
3,190
1,262
1,290,883
Table 3: Naturalness (via MOS), speaker similarity (Spk. Sim.), and scale before similarity filtering. MOS uses a five-point scale for naturalness. Overall scores are segment-weighted, with deduplicated video, speaker counts (#Spk) and segment counts (#Seg.).
Generator
DFA
ADF
Mean
1B
500M
W2V2-L
XLS-R-2B
MMS-300M
MiniMax
13.14
16.60
47.12
56.30
52.13
37.06
OmniVoice
10.95
15.10
0.01
0.00
50.84
15.38
HIGGS
12.74
16.48
79.97
0.02
1.35
22.11
VoxCPM2
34.33
38.52
0.00
0.00
99.99
34.57
Mean
17.79
21.67
31.77
14.08
51.08
27.28
Table 4: Zero-shot detector performance in % EER ↓ .
Figure 2: Zero-shot EER by speaker similarity between synthesized speech and real reference, averaged across synthesizers.
Detector
South
Southwest
North
Central
Overseas Vietnamese
Macro
# Bona fide
173,195
4,431
215,266
8,877
2,172
–
ADF-XLS-R-2B
13.7
14.4
14.4
13.1
13.4
13.8
DFA-1B
16.6
28.7
18.3
17.5
26.9
21.6
DFA-500M
21.0
33.2
21.7
22.3
28.3
25.3
ADF-W2V2-L
31.2
35.5
32.0
33.3
14.2
29.2
ADF-MMS-300M
52.3
50.5
50.4
49.7
48.8
50.3
Table 5: Zero-shot EER (%, ↓ ) by dialect, averaged equally across generators. Macro is the unweighted dialect average. Detectors are sorted by Macro EER. Red and blue denotes the lowest and highest EER in respective column and row.
Figure 3: DFA’s EER (%) in zero-shot detection of MiniMax deepfake on code-switching (CS) utterances across dialects.
Speaker recognition has advanced rapidly with large-scale training datasets, yet Vietnamese remains under-resourced, with existing corpora limited in scale and acoustic diversity. Most large-scale datasets rely on facial cues to link speech with speaker identities, restricting data collection to recordings where speakers appear on camera. We propose a face-independent dataset construction pipeline and introduce VieSpeaker, a large-scale Vietnamese speaker recognition dataset. Our approach leverages textual metadata and large language model reasoning to infer speaker identities from transcripts and contextual information. VieSpeaker contains approximately 902 hours of speech from 4,715 speakers. Experiments show that models trained on VieSpeaker achieve improved robustness and generalization compared to existing Vietnamese datasets. This work demonstrates the feasibility of face-independent dataset construction and provides a new direction for building large-scale speech resources.
Viet Hoang Pham, Tran Trung Nguyen, Bao Thu Ho +2
Hanoi University of Science and Technology, Hanoi, Vietnam
Modern LLM-driven text-to-speech (TTS) and voice conversion (VC) systems produce synthetic speech that differs from the generators represented in many legacy spoofing benchmarks. This mismatch creates a temporal generalization gap that can overestimate detector robustness under real-world post-processing conditions. We bridge this gap by introducing VoxENES 2026, a bilingual (English and Spanish) benchmark of 53,628 audio samples generated using 10 contemporary speech synthesis methods and evaluated under 10 standardized post-processing conditions. Using VoxENES 2026, we benchmark eight pretrained detectors without fine-tuning and observe substantial performance degradation: the best model achieves 28.98% EER overall, while most perform near or below random chance across modern generators and perturbations. Our results highlight the reliance on brittle artifacts in current detectors and establish VoxENES 2026 as a practical testbed for developing robust audio spoofing countermeasures.
Recent advances in speech synthesis and voice conversion have greatly improved the naturalness and authenticity of generated audio. Meanwhile, evolving encoding, compression, and transmission mechanisms on social media platforms further obscure deepfake artifacts. These factors complicate reliable detection in real-world environments, underscoring the need for representative evaluation benchmarks. To this end, we introduce ML-ITW (Multilingual In-The-Wild), a multilingual dataset covering 14 languages, seven major platforms, and 180 public figures, totaling 28.39 hours of audio. We evaluate three detection paradigms: end-to-end neural models, self-supervised feature-based (SSL) methods, and audio large language models (Audio LLMs). Experimental results reveal significant performance degradation across diverse languages and real-world acoustic conditions, highlighting the limited generalization ability of existing detectors in practical scenarios. The ML-ITW dataset is publicly available.
Daixian Li, Jun Xue, Zhuolin Yi +4
School of Cyber Science and Engineering, Wuhan University