Vietnamese speech research is constrained by resources that isolate automatic speech recognition from speaker, dialect, code-switching, and deepfake analysis. We introduce VietPrism, an open, multi-domain corpus that brings these dimensions together at scale: 993.4 hours and 403,941 bona fide utterances from 1,262 verified speakers across 8,388 real-world videos. To our knowledge, it is the first large-scale Vietnamese corpus to jointly provide transcripts, consistent speaker identities, five dialect groups, and naturally occurring Vietnamese--English code-switching, which constitutes nearly half of the corpus by duration. We further create over 3.1K hours of spoof speech with four open-source and commercial synthesis systems. Every spoof is conditioned on a verified speaker reference and paired with a transcript- and speaker-matched bona fide utterance, enabling unique controlled evaluation with reduced lexical and identity confounds. Zero-shot evaluation of five pretrained multilingual detectors reveals striking brittleness: EER greatly varies across detector--generator pairings, while recent multilingual detector DFA-1B degrades from 16.3% to 33.6% as speaker similarity increases. Dialect-stratified results expose further model-dependent disparities. By unifying natural linguistic diversity with controlled spoof generation, VietPrism provides a challenging foundation for Vietnamese speech modeling and trustworthy audio-deepfake detection.
Figures & tables
Dataset
CS
Dial.
Spk. B/S
Hrs. B/S
Avg. B/S
Tr
Utt. B/S
Speech corpora
VNSC [ 13 ]
-
3
50
100
-
-
-
VN-LVCSR [ 28 ]
-
-
-
25
-
-
-
VIVOS [ 15 ]
-
-
65
15
4.5
✓
12K
VDSPEC [ 10 ]
-
3
150
45.12
10
-
-
Viettel-CC [ 21 ]
-
-
-
85.8
-
-
-
Table 1: Comparison of VietPrism with existing Vietnamese speech and audio-deepfake datasets.
Figure 1: Overview of the VietPrism data generation pipeline. Human icons depict steps where we involve human screening.
Sub-category
Percentage
# Utt.
# Hours
Dialect
South
46.7%
175,187
463.6
North
50.0%
215,266
496.7
Central
2.0%
8,877
19.8
Southwest
0.7%
4,431
7.2
Overseas Vietnamese
0.6%
2,172
5.9
Topic
Real-estate
19.8%
65,093
196.3
Table 2: VietPrism distribution by dialect, topic, language, and gender. Percentages are duration-based except gender, which is speaker-based. Code-switched segments average 2.83 English words (8.12% of words).
Synthesizer
MOS ↑
Spk. Sim. ↑
Videos
Hours
#Spk.
#Seg.
Bona fide
2.88
—
8,388
993.4
1,262
403,941
MiniMax
2.93
0.55
8,388
1,032.9
1,262
403,941
OmniVoice
2.54
0.74
4,610
718.7
810
293,358
Higgs Audio v3
2.85
0.72
4,610
719.26
810
293,881
VoxCPM2
2.42
0.76
4,609
719.26
810
293,878
Overall (spoof)
2.71
0.68
8,388
3,190
1,262
1,290,883
Table 3: Naturalness (via MOS), speaker similarity (Spk. Sim.), and scale before similarity filtering. MOS uses a five-point scale for naturalness. Overall scores are segment-weighted, with deduplicated video, speaker counts (#Spk) and segment counts (#Seg.).
Generator
DFA
ADF
Mean
1B
500M
W2V2-L
XLS-R-2B
MMS-300M
MiniMax
13.14
16.60
47.12
56.30
52.13
37.06
OmniVoice
10.95
15.10
0.01
0.00
50.84
15.38
HIGGS
12.74
16.48
79.97
0.02
1.35
22.11
VoxCPM2
34.33
38.52
0.00
0.00
99.99
34.57
Mean
17.79
21.67
31.77
14.08
51.08
27.28
Table 4: Zero-shot detector performance in % EER ↓ .
Figure 2: Zero-shot EER by speaker similarity between synthesized speech and real reference, averaged across synthesizers.
Detector
South
Southwest
North
Central
Overseas Vietnamese
Macro
# Bona fide
173,195
4,431
215,266
8,877
2,172
–
ADF-XLS-R-2B
13.7
14.4
14.4
13.1
13.4
13.8
DFA-1B
16.6
28.7
18.3
17.5
26.9
21.6
DFA-500M
21.0
33.2
21.7
22.3
28.3
25.3
ADF-W2V2-L
31.2
35.5
32.0
33.3
14.2
29.2
ADF-MMS-300M
52.3
50.5
50.4
49.7
48.8
50.3
Table 5: Zero-shot EER (%, ↓ ) by dialect, averaged equally across generators. Macro is the unweighted dialect average. Detectors are sorted by Macro EER. Red and blue denotes the lowest and highest EER in respective column and row.
Figure 3: DFA’s EER (%) in zero-shot detection of MiniMax deepfake on code-switching (CS) utterances across dialects.