The acoustic front-end determines which forensic cues a speech deepfake detector can exploit. The wavelet scattering transform (WST) provides stable multiscale coefficients with explicit coordinates, yet direct flattening obscures the parent relation between paths. We introduce WST-Graph, reconstructing these paths as a sparse modulation-carrier grid for an AASIST graph backend. Modulation-level normalization and length-aware adaptive local attention pooling produce fixed relative-time representations while retaining the acoustic axes before learned adaptation. This yields a waveform-to-graph interface with a fixed, parameter-free WST. Our configurations remain competitive with AASIST while using approximately 60% fewer trainable parameters and show clear gains on selected out-of-domain benchmarks. These results underscore the value of preserving parent-child relations within the carrier-modulation topology when constructing a compact, physically grounded interface for graph-based speech deepfake detection. Code will be released at https://github.com/saki-ciallo/wst-graph.
Figures & tables
Setting
Param
K=32
K=48
K=64
Uniform
86,024
21.81 (18.48)
19.25 (16.17)
22.12 (18.49)
ALAP
86,073
21.46 (19.91)
22.05 (19.38)
19.14 (14.91)
Table 1: Uniform bin averaging versus ALAP at different temporal resolutions. Results are reported as three-seed average EER (%) with the best result in brackets. Bold indicates best results.
N
Param
Type D
Param
Type G
Mixed
1
91k
8.38 (6.15)
95k
8.21 (8.10)
M1: 6.74 (5.68)
2
96k
6.52 (5.77)
104k
6.80 (6.32)
M2: 5.92 (4.22)
3
100k
4.98 (4.80)
113k
7.03 (5.88)
M3: 5.46 (5.01)
4
105k
5.74 (4.77)
122k
6.82 (5.64)
M4: 6.32 (4.88)
Table 2: Results are reported in EER (%), with configurations M1 (DG), M2 (GD), M3 (DDG), and M4 (DGG).
Mean \ Std.
path
order
modulation
path
4.16 (3.76)
4.72 (3.31)
4.04 (3.42)
order
5.17 (4.48)
4.26 (3.69)
3.96 (3.68)
modulation
4.20 (4.01)
4.59 (4.16)
3.80 (3.33)
Table 3: Rows and columns specify the groupings for mean centering and standard-deviation scaling, respectively, with diagonal entries corresponding to log_path , log_order , and log_modulation . Results are reported as “Avg. (best)”. Bold indicates the lowest EER.
Node
max
max+mean
gem+mean
mean+std
EER (%)
3.80 (3.33)
3.86 (3.49)
3.53 (3.00)
3.83 (3.31)
atten
max+atten
Asym-A
Asym-B
EER (%)
3.25 (3.06)
3.38 (2.54)
3.57 (2.64)
3.58 (3.26)
Table 4: Results are reported in average EER (%), with columns evaluating alternative graph-node pooling operators.
C\K
Param
K=32
K=48
K=64
32
83,195
4.51 (4.20)
3.81 (3.57)
3.39 (3.23)
48
99,739
4.28 (3.86)
3.61 (3.20)
3.47 (3.35)
64
119,867
4.47 (3.96)
3.97 (3.34)
3.25 (3.06)
J\Q1
Param
Q1=6
Q1=8
Q1=10
6
119,725
7.06 (5.88)
7.51 (6.71)
6.95 (6.45)
8
119,867
3.29 (2.92)
3.25 (3.06)
4.07 (2.97)
Table 5: Hyperparameter optimization results across varying grid capacities and acoustic resolutions. Results are reported in EER (%).
System
AASIST
AASIST-L
Graph-Q81
Graph-Q82
Param
297k
85k
119k
120k
ITW
45.41
43.07
46.07
44.46
ASV19LA
2.74
3.45
3.07
2.92
ASV21LA
14.82
15.48
13.92
8.53
ASV21DF
19.96
21.25
21.00
18.15
ASV5T1
37.94
34.07
33.84
35.55
Table 6: Out-of-domain evaluation of reproduced AASIST baselines and our proposed configurations, reporting single-seed results.
Audio deepfakes generated by neural text-to-speech and voice-cloning systems threaten speaker verification and public discourse at scale. The core challenge is cross-dataset generalization: detectors trained on one synthesis pipeline collapse on unseen forgeries. We argue that this failure is primarily because of structural synthetic speech artifacts which are multi-timescale trajectory anomalies. Though every existing detector aggregates a fixed-window frame statistics, this misaligns the architecture with the signal. We propose FlowFake, a Liquid Time-Constant (LTC) architecture whose hidden state evolves via a learned ODE, with per-neuron adaptive time constants simultaneously resolving spectral (10ms) and prosodic (2s) cues. At only 34K parameters FlowFake achieves formal BIBO stability and O(dt^4) integration error. On a four-dataset cross domain benchmark (ASVspoof2019-LA, FakeOrReal, InTheWild, MLAAD), FlowFake reaches 75.29% on ASVspoof2019 trained only on FakeOrReal and 79.97% trained only on MLAAD. It outperforms RawGAT-ST and Whisper-DF on every evaluated pair and matching SSL Wav2vec2 (300x larger) at 0.01% of its parameter count. The source code is available on : https://github.com/GhostRider2023/FlowFake
As the accuracy of speech deepfake detection improves with the use of self-supervised representations such as wav2vec 2.0 and HuBERT, understanding why the speech is classified as bona fide or deepfake remains an open challenge. In pursuit of more trustworthy and interpretable artificial intelligence, we introduce a phoneme-level analysis framework that connects model predictions to phonetic units. Our post-hoc explainability method is generally applicable to a variety of speech deepfake detection systems based on convolutional neural networks since it leverages Gradient-weighted Class Activation Mapping in conjunction with speech recognition to generate saliency maps aligned with phonemes and pauses. This pipeline reveals statistically significant attack- and speaker-dependent phonetic cues associated with spoofed speech in terms that humans can understand. Experiments using ASVspoof 5 show comparable detection performance to similar architectures while providing linguistic interpretations across speakers and spoofing conditions.
Anna Taylor, Michele Panariello, Massimiliano Todisco +3
EURECOM, Sophia Antipolis, France · Laboratoire Informatique d’Avignon, Avignon Université, France
Existing voice deepfake detection and localization models rely heavily on representations extracted from speech foundation models (SFMs). However, downstream finetuning has now reached a state of diminishing returns. In this paper, we shift the focus to pretraining and propose a novel recipe that combines bottleneck masked embedding prediction with flow-matching based spectrogram reconstruction. The outcome, Alethia, is the first foundational audio encoder for various voice deepfake detection and localization tasks. We evaluate on 5 different tasks with 56 benchmark datasets, and note Alethia significantly outperforms state-of-the-art SFMs with superior robustness to real-world perturbations and zero-shot generalization to unseen domains (e.g., singing deepfakes). We also demonstrate the limitation of discrete targets in masked token prediction, and show the importance of continuous embedding prediction and generative pretraining for capturing deepfake artifacts.