cs.SDSep 30, 2026

SEAR: Spoofing Evidence-Grounded Audio Reasoning Benchmark for Audio Language Models

Authors: Rong Wan, Suliu Qin, Jiaxi Li, Wei Xie, Wenwu Wang, Xiaolong Han, Lu Yin, Xilu Wang

Organizations: University of Surrey · Singapore University of Technology and Design · Guangxi University · Shenzhen University of Advanced Technology

Abstract

Audio language models (ALMs) are increasingly used for audio deepfake detection (ADD), yet existing benchmarks assess their verdicts or rationale plausibility without verifying the underlying acoustic evidence. To address this issue, we first introduce spoofing evidence-grounded audio reasoning (SEAR), a four-task AQA benchmark to evaluate ALM-based ADD through acoustic evidence identification and quantification, deepfake detection, and forensic rationale generation. We further propose a bona-fide-based acoustic evidence agent (BAEA), which equips a frozen ALM with controlled acoustic tools under \textsc{fixed} or \textsc{adaptive} evidence-acquisition policies. Experiments with six ALMs reveal a clear gap between plausible rationales and verifiable acoustic evidence reasoning, while BAEA-\textsc{Fixed} improves final verdicts and forensic rationales on both evaluation partitions. Controlled interventions further show that misleading evidence degrades both detection and grounding performance.

Figures & tables

Explore similar work

CardsList
  1. SE-ADD: Self-Evolving Audio Deepfake Detection with Mistake-Driven Supervision

    Sep 30, 2026Rong Wan, Wei Xie, Jiaxi Li +4Environmental Sound Deepfake DetectionAudio Understanding

  2. FoeGlass: Simple In-Context Learning Is Enough for Red Teaming Audio Deepfake Detectors

    Jun 3, 2026Sepehr Dehdashtian, Jacob H Seidman, Vishnu N Boddeti +1Deepfake DetectionAudio Understanding

  3. MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection

    Aug 10, 2026Yanqiu Li, Yang Xiao, Jisheng Bai +3Environmental Sound Deepfake DetectionAudio Deepfake Detection