Organizations: School of Electrical and Computer Engineering, Ben-Gurion University of the Negev, Israel · Department of Electrical Engineering, Afeka the Academic College of Engineering, Israel · Avignon University, LIA, France
Self-supervised learning (SSL) countermeasures (CMs) have shown strong performance in recent years. However, they often show degraded performance while facing unseen spoofing attacks and mismatched conditions. This study examines the Voxtral audio-language model (ALM) framework for spoofing detection, as a step toward combining CM capabilities within the ALM framework. We analyze how Voxtral captures spoofing cues through audio-text processing and propose an instruction-guided approach that uses label-sequence likelihoods to evaluate bonafide and spoofed speech. Experiments on the ASVspoof databases show that without task-specific adaptation, the LLM layers emphasize semantic representations, reducing the separability of spoof-discriminative acoustic cues compared to the Whisper-based audio encoder. Consequently, spoofing-related information becomes less separable after language-model processing. We also applied lightweight adaptation using weight-decomposed low-rank adaptation (DoRA) to the Voxtral model and propose the Spooftral model, achieving an equal error rate (EER) of 4.25% on the ASVspoof5 evaluation set.
Figures & tables
Fig. 1 : Overview of the Spooftral model based on Voxtral architecture, where audio and text inputs are jointly processed and spoofing decisions are obtained with likelihood scoring.
Database
Set
bonafide
Spoof
ASVspoof2019 LA
Train
2,580
22,800
Dev.
2,548
22,296
Eval.
7,355
63,882
ASVspoof2021 LA
Eval.
14,816
133,360
ASVspoof2021 DF
Eval.
14,869
519,059
ASVspoof5
Train
18,797
163,560
TABLE I : Number of bonafide and spoofed utterances in ASVspoof2019 LA, ASVspoof2021 LA, ASVspoof2021 DF, and ASVspoof5 databases.
Fig. 2 : PaCMAP projections of audio adapter (left) and LLM (right) embeddings across ASVspoof5 sets: (a) training set, (b) Dev. set, (c) Eval. set.
Database
LLM
Audio Adapter
Dev.
Eval.
Dev.
Eval.
ASVspoof2019 LA
6.60 (5.96,7.09)
7.17 (6.91,7.55)
2.79 (2.37,3.07)
5.48 (5.20,5.73)
ASVspoof2021 LA
9.26 (8.74,9.84)
17.34 (16.97,17.70)
2.79 (2.37,3.07)
13.57 (13.30,13.87)
ASVspoof2021 DF
7.02 (6.40,7.50)
19.20 (18.96,19.52)
2.79 (2.37,3.07)
9.64 (9.49,9.86)
ASVspoof5
20.73 (20.49,20.96)
19.09 (18.98,19.19)
11.58 (11.40,11.74)
9.75 (9.68,9.83)
TABLE II : Linear probing performance after the audio adapter and LLM layers on ASVspoof databases (EER %). CI is shown below each value.
Database
Spooftral
Spooftral-Enc
Dev.
Eval.
Dev.
Eval.
ASVspoof2019 LA
0.02 (0.00,0.11)
0.41 (0.35,0.47)
0.03 (0.00,0.11)
0.40 (0.34,0.50)
ASVspoof2021 LA
0.02 (0.00,0.11)
6.46 (6.25,6.69)
0.03 (0.00,0.12)
6.70 (6.43,6.89)
ASVspoof2021 DF
0.05 (0.00,0.09)
3.88 (3.78,3.97)
0.05 (0.00,0.11)
4.76 (4.61,4.91)
ASVspoof5
4.50 (4.41,4.61)
4.25 (4.20,4.31)
5.29 (5.18,5.41)
5.08 (5.03,5.14)
TABLE III : Performance comparison of Spooftral and Spooftral-Enc across the ASVspoof databases (EER %). CI is shown below each value.
Instruction Interface
Dev. EER
Eval. EER
Prompt: “Classify the audio as real or fake.” Labels: real / fake
5.67 (5.56,5.82)
4.47 (4.42,4.52)
Prompt: “Is this speech spoof or bonafide?” Labels: bonafide / spoof
4.77 (4.66,4.89)
4.46 (4.39,4.53)
TABLE IV : Instruction interface ablation for Spooftral on ASVspoof5 without margin regularization (EER %). CI below each value.
System
Dev.
Eval.
S10 Fusion † [ 19 ]
–
11.24
Fusion of WavLM-ResNet18-SA † ∗ [ 57 ]
0.64
7.01
SSL-IVSPT ∗ [ 58 ]
0.76
5.99
SLIM † ∗ [ 59 ]
–
5.50
Best open-condition submission † ∗ [ 60 ]
–
2.59
Spooftral (LLM Attention DoRA)
7.89 (7.75,8.03)
7.01 (6.94,7.08)
TABLE V : Comparison of Spooftral and Spooftral-Enc versus other SSL-based CMs on ASVspoof5 (EER %). CI below each value.
Fig. 3 : Per-attack spoof detection rate (%) on the ASVspoof5 development and evaluation sets using the EER threshold of each set.
Recent advances in text-to-speech and voice cloning make high-quality spoofing inexpensive and scalable, threatening voice authentication systems, especially automatic speaker verification (ASV). Existing defenses mainly address this threat through binary countermeasures (CMs) for deepfake detection or spoofing-aware speaker verification (SASV), where current systems are dominated by modular ASV-CM fusion and cascaded pipelines. Although large audio language models (LALMs) have shown promise on related audio tasks, including CM and ASV, their use for SASV remains unexplored, despite their capacity to produce natural-language rationales for auditing and robustness beyond discriminative predictions. This work systematically evaluates LALMs for SASV against conventional pipelines under zero-shot prompting, supervised adaptation, reasoning-oriented training, and reinforcement-learning-based optimization. Our results show that pretrained LALMs are near chance in the zero-shot setting, confirming that they are not natively suited to SASV, but that task-specific adaptation closes this gap. We further find that competitive SASV performance can be achieved through several distinct routes. These findings position LALMs as a promising and auditable foundation for unified SASV, while clarifying where conventional cascade systems still lead.
We introduce LRLspoof, a large-scale multilingual synthetic-speech corpus for cross-lingual spoof detection, comprising 2,732 hours of audio generated with 24 open-source TTS systems across 66 languages, including 45 low-resource languages under our operational definition. To evaluate robustness without requiring target-domain bonafide speech, we benchmark 11 publicly available countermeasures using threshold transfer: for each model we calibrate an EER operating point on pooled external benchmarks and apply the resulting threshold, reporting spoof rejection rate (SRR). Results show model-dependent cross-lingual disparity, with spoof rejection varying markedly across languages even under controlled conditions, highlighting language as an independent source of domain shift in spoof detection. The dataset is publicly available at \href{https://huggingface.co/datasets/lab260/LRLspoof}{\textbf{\underline{\textit{HuggingFace}}}} and \href{https://modelscope.cn/datasets/lab260/LRLspoof}{\textbf{\underline{\textit{ModelScope}}}}
Spoofed speech detection is increasingly challenged by realistic synthesis, voice conversion, and replay attacks, with cross-dataset generalization remaining a major limitation. This work we propose a Temporal Pyramid Adapter that utilize parallel temporal convolutions with varying receptive fields to capture multi-scale spoofing cues, ranging from local artifacts to global prosodic irregularities. We also integrated self-supervised XLS-R representations combined with front-end adapters, including Mel, Sinc, and a Temporal Pyramid design for multi-scale temporal modeling. The proposed model is evaluated cross multiple benchmark including ASVspoof 2017, ASVspoof 2021 (DF/LA), PartialSpoof, DiffSSD, and multilingual HQ-MPSD datasets. Experimental results demonstrate that Temporal Pyramid model obtained AUC of 99.24% and a EER of 3.87% on the PartialSpoof database, which is significantly outperforming the base model and several SOTA baseline such as LCNN-BLSTM (9.87% EER) and TRACE (8.08% EER). Additionally, multilingual evaluations confirm that while spoofing artifact are independent from language. While self-supervised representations improve robustness, performance degrades under domain and language shifts, highlighting the need for better adaptation and calibration strategies.
Mahtab Masoudi Nezhad, Nima Karimian
Lane Department of Computer Science and Electrical Engineering West Virginia University Morgantown, WV · Bellini College of Artificial Intelligence, Cybersecurity and Computing University of South Florida Tampa, FL