Domain-Adaptive Dual-Gating Mixture of Experts for Generalizable Speech Deepfake Detection
Authors: Siqing Qin, Zhe Li, Kong Aik Lee, Man-Wai Mak
Organizations: Dept. of Electrical and Electronic Engineering, The Hong Kong Polytechnic University · Speech, Language, and Cognition Laboratory, The University of Hong Kong
Recent advances in speech deepfake detection (SDD) have leveraged the Mixture of Experts (MoE) to enhance generalization capacity. However, existing gating networks often overlook the acoustic and temporal cues of deepfakes. In this work, we propose a novel domain-adaptive dual-gating MoE (DADGMoE) framework for SDD under unseen attack types and acoustic conditions. Our innovative dual-gating mechanism leverages Sinc-layer-based filters to process both low-level acoustic signals (raw waveforms) and high-level speech representations from a large self-supervised learning (SSL) model. It further incorporates domain prototypes to guide expert routing based on implicit deepfake patterns. The lightweight affine experts process the routed inputs. Experiments show that our DADGMoE significantly outperforms the baseline, achieving up to a 40.8% relative EER reduction on challenging out-of-dataset benchmarks. This framework demonstrates superior generalization capabilities and efficient design.
Figures & tables
Figure 1 : Illustration of the proposed dual-gating network, highlighting the novel Sinc layer-based artifact filtering on both raw waveforms and SSL features, and the end-to-end learned domain prototypes for routing.
Figure 2 : Overview of the proposed DADGMoE framework.
Model
Params (M)
21DF
ITW
FoR
ADDR1
ADDR2
XLSR-AASIST
317.84
3.69
10.46
7.47*
27.74*
21.93*
DADGMoE
318.01
2.54
6.35
4.42
23.85
21.61
Table 1 : Overall SDD performance (EER%) and parameter efficiency of the proposed DADGMoE framework compared to the baseline Model. * denotes results reported by [ 29 ] .
Model
21DF
ITW
FoR
XLSR-AASIST
3.69
10.46
7.47*
DADGMoE
2.54
6.35
4.42
w/o Prototypes
2.54 (+0.00)
8.46 (+2.11)
5.31 (+0.89)
w/o SSL-Gating
2.66 (+0.12)
8.06 (+1.71)
6.67 (+2.25)
w/o Raw-Gating
3.89 (+1.35)
9.88 (+3.53)
12.67 (+8.25)
Table 2 : Ablation Study of DADGMoE Components (EER%). The full model is DADGMoE (5 experts; top-k=2). Values in parentheses indicate the absolute EER difference compared to the proposed method. * denotes results reported by [ 29 ]
Top- k Value
21DF
ITW
FoR
k=1
2.54
8.35
6.24
k=2
2.54
6.35
4.42
k=3
3.19
8.70
9.32
k=4
2.40
6.85
5.44
k=5
2.57
7.20
6.45
Table 3 : Impact of Top- k Routing Strategy for DADGMoE (5 experts) on Performance (EER%).
System
21DF
ITW
FoR
XLSR-MoE [ 9 ]
2.54
9.17
-
XLSR-SLS [ 30 ]
1.92
7.46
5.07*
XLSR-Nes2Net-X [ 31 ]
1.78
6.60
6.31*
XLSR-AASIST-SAM [ 8 ]
3.44
6.34
5.18
Wav2DF-TSL [ 13 ]
1.95
6.83
-
XLSR-Conformer-TCM [ 32 ]
2.06
7.79
10.68*
Table 4 : Comparison of DADGMoE (5 experts; top-k = 2) with state-of-the-art SDD Systems (EER%). * denotes results reported in [ 29 ] .
Figure 3 : Expert specialization in DADGMoE. Average gate weights for bonafide and spoof samples from the training set across experts, demonstrating clear roles (e.g., Expert 2 for bonafide, Expert 4 for spoof).
Figure 4 : The heatmap visualizes the gating probability for each expert across samples from 21DF, ITW, and FoR datasets. Clear vertical banding patterns indicate expert specialization towards specific deepfake domains.
Dept. of Electrical and Electronic Engineering, The Hong Kong Polytechnic University · College of Computing and Data Science, Nanyang Technological University