eess.ASSep 23, 2026

Spooftral: Can Voxtral Audio-Language Model Detect Speech Spoofing?

Authors: Avishai Weizman, Yehuda Ben-Shimol, Itshak Lapidot

Organizations: School of Electrical and Computer Engineering, Ben-Gurion University of the Negev, Israel · Department of Electrical Engineering, Afeka the Academic College of Engineering, Israel · Avignon University, LIA, France

Abstract

Self-supervised learning (SSL) countermeasures (CMs) have shown strong performance in recent years. However, they often show degraded performance while facing unseen spoofing attacks and mismatched conditions. This study examines the Voxtral audio-language model (ALM) framework for spoofing detection, as a step toward combining CM capabilities within the ALM framework. We analyze how Voxtral captures spoofing cues through audio-text processing and propose an instruction-guided approach that uses label-sequence likelihoods to evaluate bonafide and spoofed speech. Experiments on the ASVspoof databases show that without task-specific adaptation, the LLM layers emphasize semantic representations, reducing the separability of spoof-discriminative acoustic cues compared to the Whisper-based audio encoder. Consequently, spoofing-related information becomes less separable after language-model processing. We also applied lightweight adaptation using weight-decomposed low-rank adaptation (DoRA) to the Voxtral model and propose the Spooftral model, achieving an equal error rate (EER) of 4.25% on the ASVspoof5 evaluation set.

Figures & tables

Explore similar work

CardsList
  1. Large Audio Language Models for Spoofing-Aware Speaker Verification

    Jul 16, 2026Sofya Savelyeva, Mariia Perunova, Evgeny Kushnir +3Automatic Speaker VerificationLarge Audio Language Models

  2. When Spoof Detectors Travel: Evaluation Across 66 Languages in the Low-Resource Language Spoofing Corpus

    Mar 2, 2026Kirill Borodin, Vasiliy Kudryavtsev, Maxim Maslov +2SpoofingMultilingual Automatic Speech Recognition

  3. Robust Spoofed Speech Detection via Temporal Pyramid Modeling

    Jun 15, 2026Mahtab Masoudi Nezhad, Nima KarimianSpoofingSelf-Supervised Speech Models