WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
Authors: Xi Xuan, Davide Carbone, Wenxin Zhang, Tomi H. Kinnunen
Organizations: Computational Speech Group at the University of Eastern Finland · Laboratoire de Physique de l’Ecole Normale Sup´erieure, Universit´e PSL, CNRS, Sorbonne Universit´e, Universit´e de Paris · School of Computer Science and Technology at the University of Chinese Academy of Sciences and the Department of Mathematics at the University of Toronto
Existing front-ends for speech deepfake detection are primarily categorized into two types. Hand-crafted filterbank features are transparent but limited in capturing higher-level information. SSL features, in turn, lack interpretability and may overlook fine-grained spectral anomalies. We propose WaveScat, a novel family of feature extractors that combines the best of both worlds via the wavelet scattering transform (WST), which cascades wavelet convolutions with modulus nonlinearities to produce deformation-stable, multi-scale features. Experiments on the recent Deepfake-Eval-2024 benchmark, together with cross-dataset evaluations on SpoofCeleb, In-the-Wild, and ASVspoof 5, show that WaveScat outperforms existing front-ends by a wide margin. Our analysis reveals that a small averaging scale combined with high-frequency and directional resolutions is critical for capturing subtle artifacts. This underscores the value of stable and translation-invariant features for speech deepfake detection. The code and supplementary materials are available at https://github.com/xxuan-acoustics/WaveScat.