DGS-MLDG: Domain Gradient Surgery Guided Meta-Learning for Domain Generalization in Speech Deepfake Detection
Authors: Siqing Qin, Kong Aik Lee, Youzhi Tu, Eng Siong Chng, Man-Wai Mak
Organizations: Dept. of Electrical and Electronic Engineering, The Hong Kong Polytechnic University · College of Computing and Data Science, Nanyang Technological University
Speech deepfake detection faces significant challenges due to domain shifts. Domain generalization (DG), particularly meta-learning for domain generalization (MLDG), offers a promising solution by simulating and mitigating domain shifts. However, MLDG is often hindered by conflicting gradients between its meta-train and meta-test objectives, leading to suboptimal performance. To address this problem, we propose domain gradient surgery (DGS), a meta-learning method that resolves conflicts through an asymmetric projection strategy. DGS removes the destructive component from the meta-test gradient, ensuring a conflict-free optimization trajectory versus the meta-train gradient. Furthermore, we introduce layer-wise DGS (LW-DGS), an efficient variant of DGS that dynamically identifies and intervenes only conflict-prone layers. Extensive experiments on challenging benchmarks demonstrate that DGS-MLDG and LW-DGS-MLDG achieve an average relative EER reduction of 5.29% and 4.04%, respectively.
Figures & tables
Figure 1: Negative cosine similarity between the meta-train gradient and meta-test gradient of MLDG during training.
Figure 2: Overview of the proposed DGS-MLDG framework. (a) The training pipeline simulates domain shift via meta-splits and resolving conflicts via the selective mechanism (SM). (b) The backbone architecture used for deepfake detection. (c) Geometric comparison showing how DGS corrects the update direction by projecting conflicting gradients, avoiding the oscillation of Naive Gradient Aggregation (NGA).
Methods
ASV21 ↓
ASV5 ↓
CFAD ↓
ADD2023 R1 ↓
ADD2023 R2 ↓
In-the-wild ↓
CodecFake ↓
Avg. Rel. Imp. ( % ) ↑
ERM
0.92
8.57
24.33
19.16
23.23
6.71
7.66
-
MLDG
1.05
9.69
24.16
18.14
23.03
6.47
7.24
2.53
DGS-MLDG (ours)
1.00
10.24
23.86
17.76
23.19
5.23
6.76
5.29
LW-DGS-MLDG (ours)
1.02
9.55
23.83
17.88
22.68
6.15
7.27
4.04
Table 1: Comparison of EER (%) among main methods. The last column shows the cross-dataset average relative improvement w.r.t. the ERM baseline ( ↑ ). Bold and underline denote the lowest and the second-lowest EERs for each dataset.
Experiments
ASV21 ↓
ASV5 ↓
CFAD ↓
ADD2023 R1 ↓
ADD2023 R2 ↓
In-the-wild ↓
CodecFake ↓
Ablation ①
0.97
9.82
24.27
20.43
23.68
6.15
7.27
Ablation ②
1.16
9.57
23.93
18.03
22.75
6.58
7.27
Table 2: Ablation study: EER (%) under all evaluation settings.
System
In-the-wild
CodecFake
XLSR-Mamba
6.71
7.66
w/ MLDG training
6.47
7.24
+ DGS (ours)
5.23
6.76
+ LW-DGS (ours)
6.15
7.27
+ PCGrad
6.17
7.69
+ GradVac
6.34
7.40
Table 3: Comparison of EER (%) results across different gradient alignment methods.
Figure 3: Cosine similarity comparison across MLDG, DGS-MLDG, and LW-DGS-MLDG. Conflicts decrease via proposed DGS and LW-DGS.
Recent advances in speech deepfake detection (SDD) have leveraged the Mixture of Experts (MoE) to enhance generalization capacity. However, existing gating networks often overlook the acoustic and temporal cues of deepfakes. In this work, we propose a novel domain-adaptive dual-gating MoE (DADGMoE) framework for SDD under unseen attack types and acoustic conditions. Our innovative dual-gating mechanism leverages Sinc-layer-based filters to process both low-level acoustic signals (raw waveforms) and high-level speech representations from a large self-supervised learning (SSL) model. It further incorporates domain prototypes to guide expert routing based on implicit deepfake patterns. The lightweight affine experts process the routed inputs. Experiments show that our DADGMoE significantly outperforms the baseline, achieving up to a 40.8% relative EER reduction on challenging out-of-dataset benchmarks. This framework demonstrates superior generalization capabilities and efficient design.
Siqing Qin, Zhe Li, Kong Aik Lee +1
Dept. of Electrical and Electronic Engineering, The Hong Kong Polytechnic University · Speech, Language, and Cognition Laboratory, The University of Hong Kong
Recent advances in AI-based speech synthesis have enabled highly realistic speech, increasing the importance of speech deepfake detection (SDD) in preventing misuse. While mainstream Self-Supervised Learning (SSL)-based detectors achieve strong performance, they suffer from poor generalization to unseen domains and often overlook fine-grained signal artifacts due to a bias towards global semantic consistency. In this paper, we conduct the first detailed empirical and visual analysis to validate these limitations explicitly. Our investigation reveals two critical architectural vulnerabilities: (1) a systemic failure to capture localized spoofing traces, and (2) a severe lack of adaptability to domain-driven shifts in SSL layer importance, rendering static aggregation strategies prone to overfitting. To address these vulnerabilities, we propose the Global-Local Adaptive Detector (GLAD). Specifically, to capture localized forgeries, GLAD employs a Hierarchical Global-Local (HGL) backbone that explicitly bridges the granularity gap by fusing global linguistic and acoustic features with fine-grained local signal details. To counter layer importance shifts in out-of-distribution (OOD) scenarios, we introduce a Hierarchical Adaptive Gating (HAG) mechanism that dynamically recalibrates layer-wise focus in a sample-specific manner. Finally, to address shortcut learning induced by environmental biases, we introduce SaniBoost, a composite data augmentation strategy for robust signal standardization and noise sanitization. Extensive experiments demonstrate that GLAD significantly outperforms state-of-the-art methods, particularly on unseen domain cases.The code will be released upon publication.
Zelin Zhao, Guanjie Huang, Danny Hin Kwok Tsang +1
The Hong Kong University of Science and Technology (Guangzhou)
Recent neural audio codec-based speech generation (CodecFake) produces highly realistic audio, posing a challenge to existing deepfake countermeasure models. While using codec resynthesized speech (CoRS) as proxy data improves performance, it often suffers from limited generalization. We propose Domain-Shift Feature Augmentation (DSFA), which simulates "in-the-wild" variations by transforming deterministic feature statistics into stochastic distributions during fine-tuning. To evaluate generalization, we further introduce Codec-based Speech Generation Extension Evaluation (CoSG ExtEval) dataset, a more challenging extension of the CoSG Eval (from CodecFake+) dataset, featuring 40 unseen generative models and long-form audio. Experimental results demonstrate that combining a post-trained SSL backbone with DSFA effectively narrows the proxy-to-wild domain gap. This approach achieves state-of-the-art performance across diverse CodecFake attacks in both CoSG Eval and CoSG ExtEval.
Xuanjun Chen, Yun-Shing Wu, Wei-Chung Lu +4
Graduate Institute of Communication Engineering, National Taiwan University · Graduate Institute of Networking and Multimedia, National Taiwan University · Department of Information Management, National Taiwan University +2