Backdoor as Probe: Test-Time Adversarial Defense for CLIP
Authors: Zhongqi Wang, Jie Zhang, Nie Sen, Zhiyu Chen, Shiguang Shan, Xilin Chen
Organizations: Key Laboratory of AI Safety of CAS, Institute of Computing Technology, Chinese Academy of Sciences (CAS), Beijing, China · University of Chinese Academy of Sciences, Beijing, China · Xuzhou University of Technology, China
Test-time adversarial defense improves the robustness of vision-language foundation models such as CLIP without retraining. However, adversarial activation shifts are typically treated as distortions to suppress, rather than signals to exploit. We turn these shifts into defense signals by repurposing the trigger-to-target mechanism of backdoors. The key is to implant a defender-controlled backdoor as a probe that is weakly activated by clean inputs but strongly activated by adversarial shifts. Based on this insight, we propose \emph{Backdoor as Probe} (BaP), a test-time adversarial defense for CLIP. BaP constructs the probe through a closed-form model edit to a selected MLP layer. It projects the average adversarial activation shift and a defender-specified semantic direction onto the layer's low-energy input and output activation subspaces to obtain the trigger and target directions, respectively. At inference time, adversarial inputs produce measurable responses along the target direction for detection. BaP then selectively rectifies detected inputs by optimizing a small perturbation that steers their representations away from adversarial shifts and toward the clean subspace. Experiments across 16 benchmarks show that BaP improves average robust accuracy from 1.0% to 52.3% while retaining clean accuracy, achieving performance comparable to state-of-the-art methods with up to a 5.7× inference speedup. BaP further shows the generalization to adversarial attacks on large vision-language models. Project page: https://robin-wzq.github.io/Backdoor-as-Probe/
Figures & tables
Figure 1: Comparison of backdoor attacks, adversarial attacks, and BaP. Backdoor attacks implant triggers to induce attacker-specified behavior, whereas adversarial attacks perturb inputs to corrupt model predictions. BaP instead implants a defender-controlled probe for test-time defense.
Method
Adversarial detection
Adversarial rectification
Open vocabulary
Trapdoor
✓
✗
✗
AI-Shielder
✗
✓
✗
BaP
✓
✓
✓
Table 1: Comparison of backdoor-based adversarial defenses.
Figure 2: Overview of BaP. (a) Defender-controlled probe implantation. BaP constructs an attack-sensitive direction in a low-energy clean subspace and implants it through an edit. (b) Probe-guided detection and rectification. The implanted probe detects suspicious inputs and selectively guides their representations toward the clean subspace.
Figure 3: Visualization of the probe response p(x) before and after the edit. It contains 1,000 clean and 1,000 adversarial images from STL-10 using CLIP ViT-B/16. The edited probe amplifies adversarial activation shifts into a separable response space.
Dataset
Original
Test-Time Defense
Δ
CLIP
R-TPT
LPF
HD
Anti-Adv
TTE
TTC
ET3
BaP (Ours)
Type
Name
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
ImageNet
63.9
0.0
66.7
51.0
58.1
30.5
59.7
4.1
61.5
23.9
66.2
23.2
40.9
27.8
58.6
10.2
60.3
39.6
-3.6
+39.6
CIFAR10
88.1
0.5
81.6
69.2
89.0
40.4
84.1
11.8
82.8
63.5
85.5
29.8
90.0
28.2
77.7
29.2
86.8
59.1
-1.3
+58.6
CIFAR100
59.6
0.1
51.8
36.4
63.4
19.6
57.6
7.9
51.5
34.9
60.4
14.2
63.1
11.1
50.0
13.9
57.5
34.1
-2.1
+34.0
STL10
97.5
4.8
96.8
92.7
96.9
77.6
96.8
34.0
97.2
83.1
97.6
75.6
96.4
51.1
92.8
52.1
96.7
93.4
-0.8
+88.6
Table 2: Top-1 zero-shot accuracy (%) under 10-step PGD with ℓ∞=1/255 . “Clean” and “Rob.” denote accuracies on clean and adversarial samples, respectively. The final two columns report the performance of our BaP compared to the original CLIP.
Table 3: Efficiency analysis on an NVIDIA GeForce RTX 4090 GPU.
Method
Clean
Cross-modal
Targeted
Label-free
PGD
AA
PGD
DLR
AA
PGD
AA
CLIP
63.9
0.0
0.0
0.0
0.0
0.0
0.1
0.0
TTC
40.9
2.7
0.3
25.4
6.8
6.2
10.2
1.6
ET3
58.6
3.3
0.9
24.5
16.5
2.3
15.9
8.5
BaP
60.3 ↓ 3.6
34.6 ↑ 34.6
35.1 ↑ 35.1
36.1 ↑ 36.1
39.3 ↑ 39.3
36.9 ↑ 36.9
42.4 ↑ 42.3
40.0 ↑ 40.0
Table 4: ImageNet top-1 accuracy (%) across attack objectives at ℓ∞=4/255 with 50 steps. ‘AA’ denotes ‘AutoAttack’.
Method
General
Fine-Grained
Scene
Domain
Clean
PGD
AutoAttack
Clean
PGD
AutoAttack
Clean
PGD
AutoAttack
Clean
PGD
AutoAttack
CLIP
79.4
0.0
0.0
76.2
0.0
0.0
40.3
0.0
0.0
39.2
0.0
0.0
R-TPT
78.5
36.2
28.4
76.7
46.4
42.8
41.7
26.7
26.4
38.1
27.5
23.2
TTE
80.7
13.4
11.5
73.6
1.7
0.4
40.6
1.2
1.4
38.0
13.4
11.5
LPF
78.6
37.4
30.8
68.2
11.3
8.2
37.0
8.0
6.9
37.1
18.6
16.7
HD
76.9
0.5
0.1
72.8
0.0
0.0
37.4
0.0
0.0
37.3
0.1
0.0
Table 5: Comparison of zero-shot classification accuracy under stronger attacks at ℓ∞=4/255 .
Method
CLIP-B/32
CLIP-L/14
General
FG
Scene
Domain
General
FG
Scene
Domain
clean
rob
clean
rob
clean
rob
clean
rob
clean
rob
clean
rob
clean
rob
clean
rob
CLIP
76.7
4.0
71.8
0.3
38.8
0.4
35.9
6.6
83.7
4.0
84.2
0.3
45.3
0.2
46.8
0.3
R-TPT
72.9
41.9
71.4
45.4
38.7
29.6
35.7
27.2
84.2
76.8
83.5
70.3
47.1
37.5
42.4
37.3
LPF
74.8
38.0
62.0
18.9
36.2
11.3
33.7
17.4
84.0
66.9
78.8
51.7
44.2
26.2
44.5
28.5
HD
76.0
15.9
68.6
4.9
35.1
3.1
35.1
13.6
82.7
38.3
79.7
9.3
43.0
6.2
44.2
15.9
Table 6: Comparison of zero-shot classification accuracy on CLIP-B/32 and CLIP-L/14 under 10-step PGD at ℓ∞=1/255 . More detailed results are provided in Appendix F .
Table 10
Figure 4: Qualitative results of BaP against M-Attack and FOA-Attack on image captioning.
Method
M-Attack
FOA-Attack
Clean
Rob.
Clean
Rob.
Origin Model
100.0
21.1
100.0
19.1
+TTC defense
84.4
23.8
84.4
20.0
+ET3 defense
87.2
19.4
87.2
16.7
+BaP defense
99.9 ↓ 0.1
24.6 ↑ 3.5
99.9 ↓ 0.1
23.6 ↑ 4.5
Table 9: Defense performance against M-Attack and FOA-Attack.
Figure 5: The sensitivity to the edited layer l , probe norm ν , and response scale γ . More ablation results are in Appendix D .
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: The sensitivity to the threshold τ .
Target prompt
Clean
Robustness
a white teapot
61.1
52.3
a red sports car
61.1
52.3
a wooden chair
61.0
52.3
a cute cat
61.1
52.1
Appendix
Table 10: Target-prompt sensitivity.
Figure 7: The system prompt for computing GPTScore.
Dataset
Original
Test-Time Defense
Δ
CLIP
R-TPT
LPF
HD
Anti-Adv
TTE
TTC
ET3
BaP (Ours)
Type
Name
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
ImageNet
57.8
0.2
57.4
30.2
52.0
17.4
55.2
3.5
55.7
10.8
61.3
24.1
44.0
24.6
51.4
12.4
53.4
43.6
-4.4
+43.4
CIFAR10
86.1
0.6
76.7
35.1
84.9
27.2
86.5
4.2
84.4
41.8
86.0
33.7
87.6
43.3
74.1
30.4
87.1
63.1
+1.0
+62.5
CIFAR100
57.2
0.3
41.0
14.7
56.0
10.3
61.6
4.2
53.5
20.9
58.8
15.5
58.8
19.1
43.5
14.8
55.5
35.5
-1.7
+35.2
STL10
96.2
12.3
96.5
78.5
96.0
67.2
95.2
32.6
95.3
67.8
97.4
84.8
96.5
72.5
90.7
61.1
95.6
92.5
-0.6
+80.2
Appendix
Table 11: Performance comparison across different dataset types and method categories on CLIP-B/32 . Δ reports the difference between BaP and the original CLIP.
Dataset
Original
Test-Time Defense
Δ
CLIP
R-TPT
LPF
HD
Anti-Adv
TTE
TTC
ET3
BaP (Ours)
Type
Name
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
ImageNet
68.1
0.6
71.2
60.0
64.2
43.2
66.4
10.1
68.0
37.1
71.5
32.4
54.7
35.8
66.3
16.1
65.9
60.7
-2.2
+60.1
CIFAR10
93.4
1.1
90.2
82.1
94.3
71.0
92.1
40.5
89.3
78.5
92.3
47.0
94.1
18.4
86.7
45.9
93.2
53.7
-0.2
+52.6
CIFAR100
65.0
0.1
65.7
58.3
72.8
40.3
67.9
26.7
64.7
51.8
72.9
37.9
71.0
7.0
62.7
27.7
62.3
28.6
-2.7
+28.5
STL10
99.4
12.5
98.8
93.4
99.2
91.7
98.6
67.8
99.0
93.3
98.6
88.9
99.4
53.2
98.2
72.4
99.1
92.9
-0.3
+80.4
Appendix
Table 12: Performance comparison across different dataset types and method categories on CLIP-L/14 . Δ reports the difference between BaP and the original CLIP.
Figure 8: ROC curves of CLIP-B/16 for adversarial sample detection on 16 datasets.
Figure 9: ROC curves of CLIP-B/32 for adversarial sample detection on 16 datasets.
Figure 10: ROC curves of CLIP-L/14 for adversarial sample detection on 16 datasets.
Figure 11: ROC curves of CLIP-L/14@336 for LVLM attacks.