Backdoor as Probe: Test-Time Adversarial Defense for CLIP
Authors: Zhongqi Wang, Jie Zhang, Nie Sen, Zhiyu Chen, Shiguang Shan, Xilin Chen
Organizations: Key Laboratory of AI Safety of CAS, Institute of Computing Technology, Chinese Academy of Sciences (CAS), Beijing, China · University of Chinese Academy of Sciences, Beijing, China · Xuzhou University of Technology, China
Test-time adversarial defense improves the robustness of vision-language foundation models such as CLIP without retraining. However, adversarial activation shifts are typically treated as distortions to suppress, rather than signals to exploit. We turn these shifts into defense signals by repurposing the trigger-to-target mechanism of backdoors. The key is to implant a defender-controlled backdoor as a probe that is weakly activated by clean inputs but strongly activated by adversarial shifts. Based on this insight, we propose \emph{Backdoor as Probe} (BaP), a test-time adversarial defense for CLIP. BaP constructs the probe through a closed-form model edit to a selected MLP layer. It projects the average adversarial activation shift and a defender-specified semantic direction onto the layer's low-energy input and output activation subspaces to obtain the trigger and target directions, respectively. At inference time, adversarial inputs produce measurable responses along the target direction for detection. BaP then selectively rectifies detected inputs by optimizing a small perturbation that steers their representations away from adversarial shifts and toward the clean subspace. Experiments across 16 benchmarks show that BaP improves average robust accuracy from 1.0% to 52.3% while retaining clean accuracy, achieving performance comparable to state-of-the-art methods with up to a 5.7× inference speedup. BaP further shows the generalization to adversarial attacks on large vision-language models. Project page: https://robin-wzq.github.io/Backdoor-as-Probe/
Figures & tables
Figure 1: Comparison of backdoor attacks, adversarial attacks, and BaP. Backdoor attacks implant triggers to induce attacker-specified behavior, whereas adversarial attacks perturb inputs to corrupt model predictions. BaP instead implants a defender-controlled probe for test-time defense.
Method
Adversarial detection
Adversarial rectification
Open vocabulary
Trapdoor
✓
✗
✗
AI-Shielder
✗
✓
✗
BaP
✓
✓
✓
Table 1: Comparison of backdoor-based adversarial defenses.
Figure 2: Overview of BaP. (a) Defender-controlled probe implantation. BaP constructs an attack-sensitive direction in a low-energy clean subspace and implants it through an edit. (b) Probe-guided detection and rectification. The implanted probe detects suspicious inputs and selectively guides their representations toward the clean subspace.
Figure 3: Visualization of the probe response p(x) before and after the edit. It contains 1,000 clean and 1,000 adversarial images from STL-10 using CLIP ViT-B/16. The edited probe amplifies adversarial activation shifts into a separable response space.
Dataset
Original
Test-Time Defense
Δ
CLIP
R-TPT
LPF
HD
Anti-Adv
TTE
TTC
ET3
BaP (Ours)
Type
Name
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
ImageNet
63.9
0.0
66.7
51.0
58.1
30.5
59.7
4.1
61.5
23.9
66.2
23.2
40.9
27.8
58.6
10.2
60.3
39.6
-3.6
+39.6
CIFAR10
88.1
0.5
81.6
69.2
89.0
40.4
84.1
11.8
82.8
63.5
85.5
29.8
90.0
28.2
77.7
29.2
86.8
59.1
-1.3
+58.6
CIFAR100
59.6
0.1
51.8
36.4
63.4
19.6
57.6
7.9
51.5
34.9
60.4
14.2
63.1
11.1
50.0
13.9
57.5
34.1
-2.1
+34.0
STL10
97.5
4.8
96.8
92.7
96.9
77.6
96.8
34.0
97.2
83.1
97.6
75.6
96.4
51.1
92.8
52.1
96.7
93.4
-0.8
+88.6
Table 2: Top-1 zero-shot accuracy (%) under 10-step PGD with ℓ∞=1/255 . “Clean” and “Rob.” denote accuracies on clean and adversarial samples, respectively. The final two columns report the performance of our BaP compared to the original CLIP.
Table 3: Efficiency analysis on an NVIDIA GeForce RTX 4090 GPU.
Method
Clean
Cross-modal
Targeted
Label-free
PGD
AA
PGD
DLR
AA
PGD
AA
CLIP
63.9
0.0
0.0
0.0
0.0
0.0
0.1
0.0
TTC
40.9
2.7
0.3
25.4
6.8
6.2
10.2
1.6
ET3
58.6
3.3
0.9
24.5
16.5
2.3
15.9
8.5
BaP
60.3 ↓ 3.6
34.6 ↑ 34.6
35.1 ↑ 35.1
36.1 ↑ 36.1
39.3 ↑ 39.3
36.9 ↑ 36.9
42.4 ↑ 42.3
40.0 ↑ 40.0
Table 4: ImageNet top-1 accuracy (%) across attack objectives at ℓ∞=4/255 with 50 steps. ‘AA’ denotes ‘AutoAttack’.
Method
General
Fine-Grained
Scene
Domain
Clean
PGD
AutoAttack
Clean
PGD
AutoAttack
Clean
PGD
AutoAttack
Clean
PGD
AutoAttack
CLIP
79.4
0.0
0.0
76.2
0.0
0.0
40.3
0.0
0.0
39.2
0.0
0.0
R-TPT
78.5
36.2
28.4
76.7
46.4
42.8
41.7
26.7
26.4
38.1
27.5
23.2
TTE
80.7
13.4
11.5
73.6
1.7
0.4
40.6
1.2
1.4
38.0
13.4
11.5
LPF
78.6
37.4
30.8
68.2
11.3
8.2
37.0
8.0
6.9
37.1
18.6
16.7
HD
76.9
0.5
0.1
72.8
0.0
0.0
37.4
0.0
0.0
37.3
0.1
0.0
Table 5: Comparison of zero-shot classification accuracy under stronger attacks at ℓ∞=4/255 .
Method
CLIP-B/32
CLIP-L/14
General
FG
Scene
Domain
General
FG
Scene
Domain
clean
rob
clean
rob
clean
rob
clean
rob
clean
rob
clean
rob
clean
rob
clean
rob
CLIP
76.7
4.0
71.8
0.3
38.8
0.4
35.9
6.6
83.7
4.0
84.2
0.3
45.3
0.2
46.8
0.3
R-TPT
72.9
41.9
71.4
45.4
38.7
29.6
35.7
27.2
84.2
76.8
83.5
70.3
47.1
37.5
42.4
37.3
LPF
74.8
38.0
62.0
18.9
36.2
11.3
33.7
17.4
84.0
66.9
78.8
51.7
44.2
26.2
44.5
28.5
HD
76.0
15.9
68.6
4.9
35.1
3.1
35.1
13.6
82.7
38.3
79.7
9.3
43.0
6.2
44.2
15.9
Table 6: Comparison of zero-shot classification accuracy on CLIP-B/32 and CLIP-L/14 under 10-step PGD at ℓ∞=1/255 . More detailed results are provided in Appendix F .
Table 10
Figure 4: Qualitative results of BaP against M-Attack and FOA-Attack on image captioning.
Method
M-Attack
FOA-Attack
Clean
Rob.
Clean
Rob.
Origin Model
100.0
21.1
100.0
19.1
+TTC defense
84.4
23.8
84.4
20.0
+ET3 defense
87.2
19.4
87.2
16.7
+BaP defense
99.9 ↓ 0.1
24.6 ↑ 3.5
99.9 ↓ 0.1
23.6 ↑ 4.5
Table 9: Defense performance against M-Attack and FOA-Attack.
Figure 5: The sensitivity to the edited layer l , probe norm ν , and response scale γ . More ablation results are in Appendix D .
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: The sensitivity to the threshold τ .
Target prompt
Clean
Robustness
a white teapot
61.1
52.3
a red sports car
61.1
52.3
a wooden chair
61.0
52.3
a cute cat
61.1
52.1
Appendix
Table 10: Target-prompt sensitivity.
Figure 7: The system prompt for computing GPTScore.
Dataset
Original
Test-Time Defense
Δ
CLIP
R-TPT
LPF
HD
Anti-Adv
TTE
TTC
ET3
BaP (Ours)
Type
Name
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
ImageNet
57.8
0.2
57.4
30.2
52.0
17.4
55.2
3.5
55.7
10.8
61.3
24.1
44.0
24.6
51.4
12.4
53.4
43.6
-4.4
+43.4
CIFAR10
86.1
0.6
76.7
35.1
84.9
27.2
86.5
4.2
84.4
41.8
86.0
33.7
87.6
43.3
74.1
30.4
87.1
63.1
+1.0
+62.5
CIFAR100
57.2
0.3
41.0
14.7
56.0
10.3
61.6
4.2
53.5
20.9
58.8
15.5
58.8
19.1
43.5
14.8
55.5
35.5
-1.7
+35.2
STL10
96.2
12.3
96.5
78.5
96.0
67.2
95.2
32.6
95.3
67.8
97.4
84.8
96.5
72.5
90.7
61.1
95.6
92.5
-0.6
+80.2
Appendix
Table 11: Performance comparison across different dataset types and method categories on CLIP-B/32 . Δ reports the difference between BaP and the original CLIP.
Dataset
Original
Test-Time Defense
Δ
CLIP
R-TPT
LPF
HD
Anti-Adv
TTE
TTC
ET3
BaP (Ours)
Type
Name
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
Clean
Rob.
ImageNet
68.1
0.6
71.2
60.0
64.2
43.2
66.4
10.1
68.0
37.1
71.5
32.4
54.7
35.8
66.3
16.1
65.9
60.7
-2.2
+60.1
CIFAR10
93.4
1.1
90.2
82.1
94.3
71.0
92.1
40.5
89.3
78.5
92.3
47.0
94.1
18.4
86.7
45.9
93.2
53.7
-0.2
+52.6
CIFAR100
65.0
0.1
65.7
58.3
72.8
40.3
67.9
26.7
64.7
51.8
72.9
37.9
71.0
7.0
62.7
27.7
62.3
28.6
-2.7
+28.5
STL10
99.4
12.5
98.8
93.4
99.2
91.7
98.6
67.8
99.0
93.3
98.6
88.9
99.4
53.2
98.2
72.4
99.1
92.9
-0.3
+80.4
Appendix
Table 12: Performance comparison across different dataset types and method categories on CLIP-L/14 . Δ reports the difference between BaP and the original CLIP.
Figure 8: ROC curves of CLIP-B/16 for adversarial sample detection on 16 datasets.
Figure 9: ROC curves of CLIP-B/32 for adversarial sample detection on 16 datasets.
Figure 10: ROC curves of CLIP-L/14 for adversarial sample detection on 16 datasets.
Figure 11: ROC curves of CLIP-L/14@336 for LVLM attacks.
Training-free test-time defenses offer a practical way to improve the adversarial robustness of CLIP-style vision--language models without modifying the pretrained model. However, their correction strength is typically fixed for a narrow range of attack budgets, even though the attack budget is unknown at inference and the required correction varies across samples. We show that this mismatch causes existing defenses to degrade sharply as attacks strengthen. We introduce ReACT-CLIP, a response-conditioned test-time defense that separately determines how strongly each input should be corrected and whether defensive intervention is necessary. Our key observation is that the relative increase in CLIP visual-feature drift between low- and high-noise probes provides a graded, sample-specific proxy for correction demand. ReACT-CLIP maps this relative cross-noise drift to the Gaussian noise scale used to construct a stable, noise-averaged feature anchor, enabling the corrective reach to adapt to each input. To determine whether intervention is necessary, we further observe that clean inputs retain stable class-probability distributions under weak spatial augmentations, whereas adversarial inputs exhibit greater variation. ReACT-CLIP quantifies this variation using a prediction-instability score computed by Jensen--Shannon divergence and combines it with relative cross-noise drift to form the defensive intervention score. ReACT-CLIP requires no model or prompt training, and its correction-strength mapping is calibrated once and fixed across datasets and attack budgets. Across 12 downstream datasets, as well as ImageNet and its distribution-shifted variants, ReACT-CLIP delivers substantial robustness gains across diverse attack types and strengths while largely preserving clean accuracy.
Vision-language models (VLMs) such as CLIP show strong zero-shot generalization but remain highly vulnerable to adversarial attacks. Adversarial training improves robustness but is computationally expensive, motivating test-time defenses. Recent approaches exploit how CLIP's visual representations respond to stochastic perturbations: aggregating predictions across noisy views, constructing Gaussian noise-averaged anchors and interpolating features toward them, or applying counter-perturbations. These strategies improve robustness but often degrade clean accuracy, yielding an unfavorable clean-robust trade-off. We revisit stochastic test-time defenses and identify an underexplored noise-regime transition in CLIP's representation space. Prior work explored perturbations mainly in the weak-noise regime, where adversarial examples can appear unusually stable (false stability). Our analysis shows this reverses as perturbation strength grows: beyond the weak-noise regime, adversarial representations become markedly more unstable than clean ones, giving a clearer separation signal. The transition is consistent across uniform and Gaussian noise, photometric and geometric transforms, datasets, and diverse attacks. It largely disappears in adversarially trained models, suggesting it is tied to the fragile local-basin geometry of adversarial representations in non-robust CLIP. We propose a training-free, plug-in drift-gated mechanism that uses high-noise feature drift as a lightweight gating signal to trigger existing test-time defenses only when adversarial-like instability is detected. Across 13 datasets it consistently improves the clean-robust trade-off. On eight fine-grained datasets, mean clean+adversarial accuracy rises from 65.7% to 71.4% for counterattack defenses and 68.4% to 73.2% for noise-anchoring; on ImageNet and four shifted variants, from 56.1% to 66.2% and 62.1% to 67.6%.
Vision-Language Models (VLMs), such as CLIP, have shown strong zero-shot generalization but remain highly vulnerable to adversarial perturbations, posing serious risks in real-world applications. Test-time defenses for VLMs have recently emerged as a promising and efficient approach to defend against adversarial attacks without requiring costly large-scale retraining. In this work, we uncover a surprising phenomenon: under diverse input transformations, adversarial images in CLIP's feature space consistently shift along a dominant direction, in contrast to the dispersed patterns of clean images. We hypothesize that this dominant shift, termed the Defense Direction, opposes the adversarial shift, pointing features back toward their correct class centers. Building on this insight, we propose Directional Bias-guided Defense (DBD), a test-time framework that estimates the Defense Direction and employs a DB-score-based two-stream reconstruction strategy to recover robust representations. Experiments on 15 datasets demonstrate that DBD not only achieves SOTA adversarial robustness while preserving clean accuracy, but also reveals the counterintuitive result that adversarial accuracy can even surpass clean accuracy. This demonstrates that adversarial perturbations inherently encode directional priors about the true decision boundary.
Liangsheng Liu, Si Chen, Jiamin Wu +5
University of Science and Technology of China · The Chinese University of Hong Kong · Zhejiang University +2