The acoustic front-end determines which forensic cues a speech deepfake detector can exploit. The wavelet scattering transform (WST) provides stable multiscale coefficients with explicit coordinates, yet direct flattening obscures the parent relation between paths. We introduce WST-Graph, reconstructing these paths as a sparse modulation-carrier grid for an AASIST graph backend. Modulation-level normalization and length-aware adaptive local attention pooling produce fixed relative-time representations while retaining the acoustic axes before learned adaptation. This yields a waveform-to-graph interface with a fixed, parameter-free WST. Our configurations remain competitive with AASIST while using approximately 60% fewer trainable parameters and show clear gains on selected out-of-domain benchmarks. These results underscore the value of preserving parent-child relations within the carrier-modulation topology when constructing a compact, physically grounded interface for graph-based speech deepfake detection. Code will be released at https://github.com/saki-ciallo/wst-graph.
Figures & tables
Setting
Param
K=32
K=48
K=64
Uniform
86,024
21.81 (18.48)
19.25 (16.17)
22.12 (18.49)
ALAP
86,073
21.46 (19.91)
22.05 (19.38)
19.14 (14.91)
Table 1: Uniform bin averaging versus ALAP at different temporal resolutions. Results are reported as three-seed average EER (%) with the best result in brackets. Bold indicates best results.
N
Param
Type D
Param
Type G
Mixed
1
91k
8.38 (6.15)
95k
8.21 (8.10)
M1: 6.74 (5.68)
2
96k
6.52 (5.77)
104k
6.80 (6.32)
M2: 5.92 (4.22)
3
100k
4.98 (4.80)
113k
7.03 (5.88)
M3: 5.46 (5.01)
4
105k
5.74 (4.77)
122k
6.82 (5.64)
M4: 6.32 (4.88)
Table 2: Results are reported in EER (%), with configurations M1 (DG), M2 (GD), M3 (DDG), and M4 (DGG).
Mean \ Std.
path
order
modulation
path
4.16 (3.76)
4.72 (3.31)
4.04 (3.42)
order
5.17 (4.48)
4.26 (3.69)
3.96 (3.68)
modulation
4.20 (4.01)
4.59 (4.16)
3.80 (3.33)
Table 3: Rows and columns specify the groupings for mean centering and standard-deviation scaling, respectively, with diagonal entries corresponding to log_path , log_order , and log_modulation . Results are reported as “Avg. (best)”. Bold indicates the lowest EER.
Node
max
max+mean
gem+mean
mean+std
EER (%)
3.80 (3.33)
3.86 (3.49)
3.53 (3.00)
3.83 (3.31)
atten
max+atten
Asym-A
Asym-B
EER (%)
3.25 (3.06)
3.38 (2.54)
3.57 (2.64)
3.58 (3.26)
Table 4: Results are reported in average EER (%), with columns evaluating alternative graph-node pooling operators.
C\K
Param
K=32
K=48
K=64
32
83,195
4.51 (4.20)
3.81 (3.57)
3.39 (3.23)
48
99,739
4.28 (3.86)
3.61 (3.20)
3.47 (3.35)
64
119,867
4.47 (3.96)
3.97 (3.34)
3.25 (3.06)
J\Q1
Param
Q1=6
Q1=8
Q1=10
6
119,725
7.06 (5.88)
7.51 (6.71)
6.95 (6.45)
8
119,867
3.29 (2.92)
3.25 (3.06)
4.07 (2.97)
Table 5: Hyperparameter optimization results across varying grid capacities and acoustic resolutions. Results are reported in EER (%).
System
AASIST
AASIST-L
Graph-Q81
Graph-Q82
Param
297k
85k
119k
120k
ITW
45.41
43.07
46.07
44.46
ASV19LA
2.74
3.45
3.07
2.92
ASV21LA
14.82
15.48
13.92
8.53
ASV21DF
19.96
21.25
21.00
18.15
ASV5T1
37.94
34.07
33.84
35.55
Table 6: Out-of-domain evaluation of reproduced AASIST baselines and our proposed configurations, reporting single-seed results.