Large vision-language models (LVLMs) are increasingly deployed in safety-critical applications, yet they remain vulnerable to backdoor attacks. Defending against such attacks remains costly, as existing methods require either extensive retraining on clean data or per-query intervention at inference time. To address this limitation, we propose OrthoPurify, a more efficient method to purify backdoored model weights via one-step orthogonal projection. Specifically, through structural analysis of backdoor weight updates, we find that the backdoor is encoded by diverting a small number of weight update directions from task adaptation to backdoor shortcut encoding, a phenomenon we term direction hijacking. However, identifying these hijacked directions requires a benign reference model, which is typically inaccessible to the defender. We show that a pseudo-benign model, obtained by fine-tuning the pretrained weights on only a small set of clean samples, provides a sufficient approximation, as the dominant update directions stabilize within the first few gradient steps. OrthoPurify uses this pseudo-benign reference to isolate the hijacked directions and removes them through a single projection on the weight update. Extensive experiments show that OrthoPurify reduces the attack success rate to near zero while preserving the original performance across diverse benchmarks, without retraining the backdoored model or introducing inference-time overhead. Our code is publicly available at https://github.com/womeimingzi/OrthoPurify.
Figures & tables
Figure 1: Overview of OrthoPurify. SVD analysis of the backdoored weight update reveals direction hijacking : most principal directions are shared with a benign reference ( θ<30∘ ), while a few are nearly orthogonal ( θ>70∘ ) and encode the backdoor exclusively. A pseudo-benign reference , obtained by fine-tuning pretrained weights on a small clean set for a few gradient steps, provides a sufficient approximation. OrthoPurify identifies the hijacked directions via principal angle thresholding and removes them through a single orthogonal projection.
Figure 2: Direction hijacking in backdoor weights. (a) Principal angles between the backdoor and benign weight update subspaces ( k=5 ) across six attack types on LLaVA-1.5-7B. Most directions are well aligned ( <30∘ ), while one direction consistently exceeds 70∘ (red), indicating direction hijacking. The dashed line marks the 50∘ threshold used in our method. (b) Hijacked energy retention ratio after applying each defense method. Values below 1.0 indicate removal of hijacked energy; values above 1.0 indicate amplification. All existing defenses retain or amplify the hijacked energy, while OrthoPurify removes most of it.
Figure 3: Verification that hijacked directions encode the backdoor. ASR under three conditions on LLaVA-1.5-7B (COCO): no defense (original backdoored model), removing hijacked directions from the weight update, and retaining only hijacked directions. Removing the hijacked directions eliminates the backdoor (ASR → 0%). Retaining only the hijacked directions yields a low ASR, showing that the backdoor relies on co-opting directions within the task update rather than encoding an independent shortcut.
Figure 4: Supporting evidence for direction hijacking and pseudo-benign approximation. (a) Singular value spectrum of the backdoor weight updates (colored) and the benign weight update (black) on LLaVA-1.5-7B. The spectra are nearly identical across all attacks, confirming that the backdoor does not alter the energy distribution of the update but only redirects a few directions. (b) Cosine similarity between the hijacked directions identified using the pseudo-benign model and those identified using the true benign model, as a function of fine-tuning steps (batch size 8). The similarity exceeds 0.97 within 2 steps across all attack types.
BadNet
Blended
WaNet
ISSBA
TrojVLM
VLOOD
Model
Method
ASR
CU
ASR
CU
ASR
CU
ASR
CU
ASR
CU
ASR
CU
LLaVA-1.5
No defense
99.41
121.48
97.85
129.98
98.63
121.48
98.44
126.74
100.00
107.66
81.64
117.36
Clean FT
98.44
134.64
93.16
137.22
96.29
129.83
83.59
135.85
96.09
128.21
28.32
127.60
Fine-Pruning
0.59
129.09
86.91
128.90
52.34
126.89
16.99
129.75
3.71
110.11
0.78
128.41
ANP
74.20
123.82
86.00
121.95
81.60
124.54
8.80
125.77
44.80
120.53
25.20
125.02
CLP
59.77
115.82
51.76
122.16
26.17
120.42
2.93
116.11
50.59
106.36
0.98
120.10
Table 1: Results on image captioning (COCO). ASR ( ↓ %) and CU (CIDEr on clean inputs, ↑ ) of each defense against six backdoor attacks on LLaVA-1.5-7B and Qwen3-VL-8B.
Figure 5: Ablation study. Top row: ASR ( ↓ ); bottom row: CIDEr ( ↑ ). (a) Effect of total clean sample count (batch size =n/8 , 1 epoch). (b) Effect of fine-tuning steps (batch size 8 fixed). (c) Effect of subspace dimension k ( θ=50∘ fixed). (d) Effect of angle threshold θ ( k=10 fixed). All experiments use LLaVA-1.5-7B on COCO under four attack types.
BadNet
TrojVLM
Scope
Model
Method
ASR
CU
ASR
CU
No defense
100
124.37
100
124.49
LLaVA-7B (LoRA)
OrthoPurify
0
123.52
1.56
124.36
No defense
100
129.47
100
127.36
LLaVA-13B (Adapter)
OrthoPurify
0
131.19
0
134.59
No defense
100
62.57
99.80
60.60
Table 3: Results under different fine-tuning scopes. ASR ( ↓ %) and CU (CIDEr on clean inputs, ↑ ) with and without OrthoPurify, where the backdoor is injected by fine-tuning the adapter only (Adapter), the vision encoder and adapter (Mixed), or the full model (Full).
Figure 6: Adaptive attack via alignment regularization . Solid: ASR; dashed: CU. The attacker faces an unavoidable trade-off at λ≈0.28 .
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Visual comparison of the six trigger types. Top row: BadNet, Blended, ISSBA. Bottom row: WaNet, TrojVLM, VLOOD.
Model
Attack
Dataset
Poison Rate
LLaVA-1.5-7B
Blended
VQAv2
0.3
ISSBA
COCO
0.15
ISSBA
VQAv2
0.2
VLOOD
VQAv2
0.2
Qwen3-VL-8B
VLOOD
COCO
0.2
VLOOD
VQAv2
0.2
Appendix
Table 4: Non-default poison rates. All unlisted combinations use the default rate of 0.1.
Method
Samples
Retrain-free
ASR (%)
GPU Time (s)
Clean FT
1000
×
98.44
1521
Fine-Pruning
1000
×
0.59
1766
ANP
500
×
74.2
9367
CLP
0
✓
59.77
—
OrthoPurify
64
✓
0
46
Appendix
Table 5: Efficiency comparison between defense methods against BadNet on LLaVA-1.5-7B (COCO). CLP operates on CPU without GPU computation.
Figure 8: Resistance to backdoor reactivation. ASR after re-poisoning defended models on LLaVA-1.5-7B. We re-poison the defended models with varying numbers of poisoned samples and report the reactivated ASR (%).
Figure 9: Qualitative examples on LLaVA-1.5-7B (COCO). Each row shows a different attack type. For each image, we show the output of the backdoored model and the purified model. On clean inputs, both models produce similar captions. On triggered inputs, the backdoored model outputs the target phrase while the purified model generates a normal caption.
Attack
Condition
m
ρ
ASR (%)
CIDEr
BadNet
No defense
–
–
99.41
121.48
Hijacked (ours)
3
0.034
0.0
125.85
Random
3
0.035 ± 0.006
70.7 ± 39.2
124.29 ± 3.24
Most-aligned
3
0.154
83.79
112.67
Blended
No defense
–
–
97.85
129.98
Hijacked (ours)
5
0.058
0.0
131.79
Appendix
Table 6: Direction specificity control. m : number of directions removed (matched across conditions). ρ : fraction of adaptation energy removed ( ∥ΔWD∥F2/∥ΔW∥F2 ). Random reports mean ± std over 5 seeds. All experiments use k=10 , θ=50∘ on LLaVA-1.5-7B (COCO).
Method
BadNet ASR
TrojVLM ASR
Latency (ms/query)
Overhead
No Defense
99.41
100.0
1469
—
PurMM
0.0
50.0
4692
+219%
CleanSight
0.0
11.7
2202
+50%
OrthoPurify
0.0
0.0
1469
0%
Appendix
Table 7: Comparison with test-time defenses. ASR (% ↓ ), per-query latency, and inference overhead on LLaVA-1.5-7B (COCO).
Model
Task
CU (before)
CU (after)
LLaVA-1.5-7B
COCO (CIDEr ↑ )
125.23
133.07
Qwen3-VL
COCO (CIDEr ↑ )
67.61
70.07
LLaVA-1.5-7B
VQAv2 (Score ↑ )
71.03
66.93
Qwen3-VL
VQAv2 (Score ↑ )
84.24
85.87
Appendix
Table 8: Benign model safety. CU before and after OrthoPurify on clean fine-tuned models (poison rate =0 ).
No defense
θ=50∘
θ=30∘
θ=20∘
λ
ASR
CU
ASR
CU
ASR
CU
ASR
CU
0.0
100.0
112.7
0.0
105.8
0.0
109.1
0.0
112.2
0.1
97.7
122.4
95.3
122.9
0.8
119.3
0.0
116.2
0.2
97.7
117.9
96.1
122.0
15.6
119.7
0.0
123.3
0.25
99.2
122.0
95.3
120.1
25.0
117.8
0.0
120.3
0.26
96.1
125.4
95.3
119.4
21.1
121.2
0.0
116.9
Appendix
Table 9: Adaptive attack via alignment regularization (Strategy 1). ASR (% ↓ ) and CU (CIDEr ↑ ) across regularization strengths λ and defense thresholds θ .
No defense
OrthoPurify
λ
σ1
ASR
CIDEr
ASR
CIDEr
Hijacked dirs
0
3.94
100.0
123.14
0.0
126.08
7
0.01
3.68
100.0
127.27
0.0
127.39
7
0.1
2.07
100.0
129.53
7.6
123.62
16
0.5
1.47
99.4
123.57
0.0
116.87
16
1.0
1.51
99.8
129.87
0.0
120.52
16
Appendix
Table 10: Adaptive attack via energy dispersal (Strategy 2). σ1 denotes the largest singular value of the Layer 2 weight update.
BadNet
TrojVLM
Pseudo-benign source
ASR
CIDEr
ASR
CIDEr
COCO (same dist.)
0.0
121.48 → 125.85
0.0
107.66 → 130.08
Flickr30k (diff. images)
0.0
121.48 → 111.80
0.0
107.66 → 122.37
Compact COCO (diff. style)
0.0
121.48 → 112.03
0.0
107.66 → 121.21
VQAv2 (diff. task)
0.0
121.48 → 119.29
0.0
107.66 → 108.63
Appendix
Table 11: Clean data mismatch. ASR (% ↓ ) and CIDEr ( ↑ ) under different pseudo-benign data sources.
Setting
ASR (%)
CIDEr
Backdoored (surrogate attack)
100.0
127.55
After OrthoPurify
0.0
126.72
Clean baseline
0.0
125.23
Appendix
Table 12: Black-box transferability. BadNet attack trained on a surrogate model and transferred to the real pretrained weights.
Target
ASR (before)
ASR (after)
“you have been hacked lol”
20.5
0.0
“a beautiful sunset over the ocean …”
40.6
0.0
“access denied please contact administrator”
38.8
0.0
Overall
99.8
0.0
Appendix
Table 13: Multi-target attack. Per-target ASR (%) before and after OrthoPurify.
pr
ASR (before)
ASR (after)
CIDEr (before)
CIDEr (after)
0.0
0.0
0.0
125.23
133.07
0.1
99.41
0.0
121.48
125.85
0.2
100.0
0.0
126.05
126.12
0.3
100.0
0.0
127.75
127.35
0.4
100.0
0.0
124.13
123.38
0.5
100.0
0.0
122.50
120.18
Appendix
Table 14: Poison rate sweep. ASR (% ↓ ) and CIDEr ( ↑ ) across poison rates on BadNet (LLaVA-1.5-7B, COCO).