eess.ASOct 7, 2026

DuRe-ST: Dual-Relation Spectro-Temporal Modeling for Speech Deepfake Detection

Authors: Shaole Li, Siqing Qin, Youzhi Tu, Kong Aik Lee

Organizations: Department of Electrical and Electronic Engineering The Hong Kong Polytechnic University, Hong Kong SAR

Abstract

Previous speech deepfake detectors can adaptively capture spectro-temporal dependencies through graph attention, yet they largely overlook the co-variation between spectral and temporal representations. To address this gap, we construct a normalized affinity graph from their joint covariance and apply polynomial graph filtering to capture higher-order covariance-induced dependencies. We first develop Cov-ST to isolate the contribution of covariance-based relational modeling. Although it improves detection performance, its sensitivity to the polynomial order suggests limited robustness when covariance relations are modeled alone. We therefore propose DuRe-ST, which jointly exploits covariance-induced and graph-attention-induced relations to capture complementary second-order co-variation and adaptive spectro-temporal dependencies. Experiments show that DuRe-ST achieves an average relative EER reduction of 25.9% over XLSR-AASIST on the ASVspoof benchmarks and 28.4% across four cross-dataset benchmarks with only 4-8k additional trainable back-end parameters. It further outperforms the strongest publicly available comparison models by 2.2-13.2% in relative EER on four benchmarks, while remaining smaller than the publicly available models considered.

Figures & tables

Explore similar work

CardsList
  1. Time-Frequency Consistency Learning for Robust Speech Deepfake Detection

    Jul 20, 2026Jun Xue, Zhuolin Yi, Yanzhen Ren +6Audio Deepfake DetectionDistortion

  2. Deepfake Audio Detection Using Self-supervised Fusion Representations

    May 5, 2026Khalid Zaman, Qixuan Huang, Muhammad Uzair +1Audio UnderstandingSelf-Supervised Representations

  3. WST-Graph: Topology-Preserving Wavelet Scattering Front-End for Speech Deepfake Detection

    Sep 24, 2026Kwok-Ho Ng, Tingting Song, Bingwen Feng +1Audio Deepfake DetectionWav2Vec