Audio deepfake detectors remain vulnerable to adversarial perturbations that suppress the acoustic cues used for detection, allowing manipulated utterances to evade the detector. Although existing defenses can improve robustness, they require retraining the detector or introduce additional distortion. Diffusion-based purification instead leaves the pretrained detector unchanged, but existing methods use the same purification strength for all inputs, creating a trade-off between removing adversarial perturbations and preserving the subtle spoofing cues needed for detection. In this paper, we propose Detection-Guided Adaptive Purification (DGAP), a diffusion-based defense that adjusts purification strength per input. Building on the observation that a light purification perturbs the detector score of an adversarial input far more than that of a benign one, the framework uses the resulting score shift as a reference-free indicator of adversarial manipulation. Inputs with small shifts are passed unchanged, whereas flagged inputs undergo stronger purification before final detection. We evaluate the framework against three adversarial attack settings across three deepfake detectors, and compare it with nine existing defenses. Our results show that DGAP achieves the best defense performance across all detectors while leaving benign inputs nearly unaffected, and remains effective under the defense-aware adaptive attack.
Figures & tables
Fig. 1: Comparison of (a) an undefended baseline, (b) a fixed purification defense, and (c) the proposed DGAP defense under adversarial deepfake attacks.
Fig. 2: Illustration of the forward and reverse diffusion process.
Fig. 3: Overview of the proposed DGAP framework.
Detector
PGD- ℓ∞
PGD- ℓ2
C&W
AASIST
94.0 (20)
94.5 (10)
95.0
RawGAT-ST
73.2 (100)
91.2 (50)
97.9
Res-TSSDNet
82.5 (20)
97.8 (10)
100.0
TABLE I: White-box attack success rate (%) on the undefended detectors, with the number of PGD iterations in parentheses.
Fig. 4: Parameter selection on the development split. Each panel plots the objective y versus (a) the probe level tg , (b) the purification level tp , and (c) the gate coefficient γ . Panel (b) also shows the uniform baseline. Filled markers denote per-attack optima, the grey dashed curve is the objective averaged over the three attacks, and the grey marker denotes the single attack-agnostic configuration selected per detector.
Fig. 5: Score-gap distributions at the probe level tg for benign and adversarial inputs.
AASIST
RawGAT-ST
Res-TSSDNet
Method
EERc
ℓ∞
ℓ2
C&W
y
EERc
ℓ∞
ℓ2
C&W
y
EERc
ℓ∞
ℓ2
C&W
y
No defense
1.17
61.92
49.42
46.50
53.78
1.25
21.04
26.47
20.70
23.99
2.17
57.49
50.08
38.47
50.85
QT
4.08
4.08
5.42
7.00
9.58
2.17
9.16
19.26
15.35
16.76
8.33
17.43
12.71
54.29
36.48
DS
2.67
12.75
13.08
22.25
18.70
2.67
20.36
26.47
33.51
29.45
3.00
29.71
34.18
30.64
34.51
LPF
2.08
12.17
14.83
32.83
22.02
1.83
17.56
22.31
23.66
23.01
4.50
26.94
34.18
36.36
36.99
BPF
2.58
8.75
9.58
19.50
15.19
1.83
11.71
17.56
16.20
16.99
3.08
30.22
32.24
28.11
33.27
TABLE II: EERc and per-attack EERadv (%) of the defense methods, and the objective y with the adversarial term averaged over the three attacks. Best in bold , second best underlined .
Detector
Configuration
t
EERc
ℓ∞
ℓ2
C&W
AASIST
Uniform, optimal t
160
1.90
20.80
17.60
9.00
Uniform at tp
280
6.30
24.40
20.00
14.60
Gated at tp (Ours)
280
2.30
9.30
10.30
5.20
RawGAT-ST
Uniform, optimal t
160
1.30
23.65
26.45
8.02
Uniform at tp
300
4.00
27.76
30.16
13.53
Gated at tp (Ours)
300
1.30
11.72
14.53
2.91
TABLE III: Ablation of the gate on the development split: EERc and per-attack EERadv (%). Uniform purification at the per-detector optimal level (equivalent to DiffPure), uniform purification at the stronger level tp , and the gated method at the same tp . Best in bold .
AASIST
RawGAT-ST
Method
EERc
EERadv
y
EERc
EERadv
y
No defense
0.00
64.17
64.17
0.83
23.33
24.17
QT
3.33
16.67
20.00
0.83
20.00
20.83
DS
0.83
25.00
25.83
2.50
21.67
24.17
LPF
1.67
30.00
31.67
1.67
19.17
20.83
BPF
1.67
23.33
25.00
1.67
14.17
15.83
TABLE IV: EERc , EERadv , and the objective y (%) under the adaptive BPDA+EOT attack. Best in bold , second best underlined .
Fig. 6: Score-gap distributions under the adaptive BPDA+EOT attack.