Audio deepfake detectors remain vulnerable to adversarial perturbations that suppress the acoustic cues used for detection, allowing manipulated utterances to evade the detector. Although existing defenses can improve robustness, they require retraining the detector or introduce additional distortion. Diffusion-based purification instead leaves the pretrained detector unchanged, but existing methods use the same purification strength for all inputs, creating a trade-off between removing adversarial perturbations and preserving the subtle spoofing cues needed for detection. In this paper, we propose Detection-Guided Adaptive Purification (DGAP), a diffusion-based defense that adjusts purification strength per input. Building on the observation that a light purification perturbs the detector score of an adversarial input far more than that of a benign one, the framework uses the resulting score shift as a reference-free indicator of adversarial manipulation. Inputs with small shifts are passed unchanged, whereas flagged inputs undergo stronger purification before final detection. We evaluate the framework against three adversarial attack settings across three deepfake detectors, and compare it with nine existing defenses. Our results show that DGAP achieves the best defense performance across all detectors while leaving benign inputs nearly unaffected, and remains effective under the defense-aware adaptive attack.
Figures & tables
Fig. 1: Comparison of (a) an undefended baseline, (b) a fixed purification defense, and (c) the proposed DGAP defense under adversarial deepfake attacks.
Fig. 2: Illustration of the forward and reverse diffusion process.
Fig. 3: Overview of the proposed DGAP framework.
Detector
PGD- ℓ∞
PGD- ℓ2
C&W
AASIST
94.0 (20)
94.5 (10)
95.0
RawGAT-ST
73.2 (100)
91.2 (50)
97.9
Res-TSSDNet
82.5 (20)
97.8 (10)
100.0
TABLE I: White-box attack success rate (%) on the undefended detectors, with the number of PGD iterations in parentheses.
Fig. 4: Parameter selection on the development split. Each panel plots the objective y versus (a) the probe level tg , (b) the purification level tp , and (c) the gate coefficient γ . Panel (b) also shows the uniform baseline. Filled markers denote per-attack optima, the grey dashed curve is the objective averaged over the three attacks, and the grey marker denotes the single attack-agnostic configuration selected per detector.
Fig. 5: Score-gap distributions at the probe level tg for benign and adversarial inputs.
AASIST
RawGAT-ST
Res-TSSDNet
Method
EERc
ℓ∞
ℓ2
C&W
y
EERc
ℓ∞
ℓ2
C&W
y
EERc
ℓ∞
ℓ2
C&W
y
No defense
1.17
61.92
49.42
46.50
53.78
1.25
21.04
26.47
20.70
23.99
2.17
57.49
50.08
38.47
50.85
QT
4.08
4.08
5.42
7.00
9.58
2.17
9.16
19.26
15.35
16.76
8.33
17.43
12.71
54.29
36.48
DS
2.67
12.75
13.08
22.25
18.70
2.67
20.36
26.47
33.51
29.45
3.00
29.71
34.18
30.64
34.51
LPF
2.08
12.17
14.83
32.83
22.02
1.83
17.56
22.31
23.66
23.01
4.50
26.94
34.18
36.36
36.99
BPF
2.58
8.75
9.58
19.50
15.19
1.83
11.71
17.56
16.20
16.99
3.08
30.22
32.24
28.11
33.27
TABLE II: EERc and per-attack EERadv (%) of the defense methods, and the objective y with the adversarial term averaged over the three attacks. Best in bold , second best underlined .
Detector
Configuration
t
EERc
ℓ∞
ℓ2
C&W
AASIST
Uniform, optimal t
160
1.90
20.80
17.60
9.00
Uniform at tp
280
6.30
24.40
20.00
14.60
Gated at tp (Ours)
280
2.30
9.30
10.30
5.20
RawGAT-ST
Uniform, optimal t
160
1.30
23.65
26.45
8.02
Uniform at tp
300
4.00
27.76
30.16
13.53
Gated at tp (Ours)
300
1.30
11.72
14.53
2.91
TABLE III: Ablation of the gate on the development split: EERc and per-attack EERadv (%). Uniform purification at the per-detector optimal level (equivalent to DiffPure), uniform purification at the stronger level tp , and the gated method at the same tp . Best in bold .
AASIST
RawGAT-ST
Method
EERc
EERadv
y
EERc
EERadv
y
No defense
0.00
64.17
64.17
0.83
23.33
24.17
QT
3.33
16.67
20.00
0.83
20.00
20.83
DS
0.83
25.00
25.83
2.50
21.67
24.17
LPF
1.67
30.00
31.67
1.67
19.17
20.83
BPF
1.67
23.33
25.00
1.67
14.17
15.83
TABLE IV: EERc , EERadv , and the objective y (%) under the adaptive BPDA+EOT attack. Best in bold , second best underlined .
Fig. 6: Score-gap distributions under the adaptive BPDA+EOT attack.
We present Proteus, a framework developed at Resemble AI for automated robustness testing of our audio deepfake detection system. Given a detector, Proteus systematically searches over sequences of everyday audio transformations (codec transcoding, additive noise, reverberation, dynamic-range compression, and VoIP simulation) to find combinations that fool the detector while preserving speech quality. We propose two complementary search strategies: (1) a breadth-first search that exhaustively maps augmentation effectiveness across the parameter space, and (2) a Q-learning agent designed to efficiently discover deeper attack chains by exploiting structural patterns in the BFS data. We report findings from continuous deployment of Proteus against our production detector, showing that specific augmentation chains can reliably flip detection verdicts while preserving speech intelligibility and speaker identity. We discuss how these findings are used to harden the detector through targeted retraining.
Nicolas M. Müller, Aditya Tirumala Bukkapatnam, Zohaib Ahmed
The rapid advancement of speech synthesis and voice conversion technologies has made audio deepfakes increasingly realistic, posing serious security risks in practical applications. While existing detection methods achieve strong performance under controlled conditions, they often fail to generalize under real-world perturbations and corruptions. In this paper, we propose ROGUE, a framework that dynamically constructs robust detection workflows by orchestrating multiple detection tools. ROGUE formulates workflow generation as a sequential decision-making problem and introduces a dual-agent paradigm, where a perturbation agent generates audio perturbations and a policy agent learns to select and execute detection tools under perturbed conditions. Through adversarial learning, ROGUE enables perturbation-aware tool selection, adaptive execution strategies, and improved robustness to distribution shifts. Extensive experiments across multiple datasets and real-world corruptions demonstrate that ROGUE consistently outperforms strong baselines in both robustness and generalization. Our results highlight the effectiveness of adversarially optimized workflow generation for building reliable audio deepfake detection systems in real-world deployment settings.
Xiang Li, Pin-Yu Chen, Wenqi Wei
Department of Computer Science and Information, Fordham University, NY, USA · IBM Research, Yorktown Heights, NY, USA
Generalizing to unseen attacks remains challenging for audio deepfake detectors, and collecting training data covering all potential attacks is impractical. We explore recurrent refinement in an already-trained SSL-based detector without additional data or changes to its original parameters. However, directly recycling encoder outputs as inputs degrades detection in our diagnostic. We propose CoReLoop, which makes this reuse effective by adapting recurrent inputs to the frozen encoder, controlling state updates, and aligning refined outputs with the frozen classifier. By training only lightweight refinement modules and loop-specific low-rank adapters on the original data, CoReLoop enables additional refinement while preserving the detector's original first-pass prediction. On 14 cross-domain test sets, the 24-layer model reduces pooled equal error rate (EER) from 4.85% to 3.74% with two passes, with approximately 10M trainable parameters out of 598M. To selectively apply this refinement, an optional halting head chooses the depth for each utterance, achieving 3.73% pooled EER with an average of 1.18 passes.
Kunyu Feng, Yuxiang Wang, Li Wang +2
The Chinese University of Hong Kong, Shenzhen · Amphion Technology Co., Ltd.