eess.ASSep 24, 2026

WST-Graph: Topology-Preserving Wavelet Scattering Front-End for Speech Deepfake Detection

Authors: Kwok-Ho Ng, Tingting Song, Bingwen Feng, Zhihua Xia

Organizations: College of Cyber Security, Jinan University, Guangzhou, China

Abstract

The acoustic front-end determines which forensic cues a speech deepfake detector can exploit. The wavelet scattering transform (WST) provides stable multiscale coefficients with explicit coordinates, yet direct flattening obscures the parent relation between paths. We introduce WST-Graph, reconstructing these paths as a sparse modulation-carrier grid for an AASIST graph backend. Modulation-level normalization and length-aware adaptive local attention pooling produce fixed relative-time representations while retaining the acoustic axes before learned adaptation. This yields a waveform-to-graph interface with a fixed, parameter-free WST. Our configurations remain competitive with AASIST while using approximately 60% fewer trainable parameters and show clear gains on selected out-of-domain benchmarks. These results underscore the value of preserving parent-child relations within the carrier-modulation topology when constructing a compact, physically grounded interface for graph-based speech deepfake detection. Code will be released at https://github.com/saki-ciallo/wst-graph.

Figures & tables

Explore similar work

CardsList
  1. FlowFake: Liquid Networks for Audio Deepfake Detection

    Jun 17, 2026Shivaay Dhondiyal, Divyansh Sharma, Dinesh Kumar VishwakarmaNeural AudioSpeaker

  2. Why Do You Say It Like That? A Phoneme-level Framework for Explainable Speech Deepfake Detection

    Jul 9, 2026Anna Taylor, Michele Panariello, Massimiliano Todisco +3Audio Deepfake DetectionSpoofing

  3. Alethia: A Foundational Encoder for Voice Deepfakes

    Apr 30, 2026Yi Zhu, Brahmi Dwivedi, Jayaram Raghuram +1Deepfake DetectionOpencode