Organizations: University of Electronic Science and Technology of China, Chengdu, China · Southwestern University of Finance and Economics, Chengdu, China · Tongji University, Shanghai, China · Shanghai Innovation Institute, Shanghai, China
Token-Pruning accelerates Vision-Language Models by removing redundant visual tokens, yet its safety implications remain underexplored. In this work, we present the first comprehensive safety evaluation of Token-Pruning mechanisms and find that: most pruning strategies significantly degrade safety as pruning ratios increase, whereas Query-based Compression shows the opposite, with extreme pruning (up to 99.8%), unexpectedly improves model safety. This sharp contrast prompts a key question: How do different Token-Pruning strategies reshape model safety behavior, and is it possible to enhance safety without sacrificing acceleration? To answer this, we identify an unrecognized mechanism, termed Pruning-Induced Malicious Amplification, where removal of background tokens triggers a side effect: forcing the model's attention to collapse onto a few retained malicious anchors within the foreground, inadvertently amplifying their toxic semantics under jailbreak. To address that, we propose an inference-time and plug-and-play Safety-Aware Pruning (SAP) mechanism that counteracts such dominance via three steps: (1) identifying malicious anchors, (2) restoring pruned benign tokens, and (3) reallocating excessive attention from malicious anchors to benign tokens. Extensive experiments across three safety and four utility benchmarks demonstrate that SAP mitigates pruning-induced vulnerabilities, i.e., reducing ASR by up to 62%, without compromising efficiency or utility.
Figures & tables
Figure 1 : Comparison of Attack Success Rate (ASR) across Token-Pruning methods. Most methods increase vulnerability to jailbreak attacks, while query-based ones improve safety.
Figure 2 : Overall safety–utility comparison across benchmarks. Radar charts compare Token-Pruning methods with and without SAP on safety and utility benchmarks. SAP enhances safety while maintaining utility.
Figure 3 : Comparison of safety (ASR) and utility across different Token-Pruning methods. Bars represent the Average ASR, while markers indicate the task performance (Utility).
Method
ASR (%) ↓
Utility (%) ↑
Pruning Rate (%)
Original
54.01
64.46
0
Random
52.73
63.02
75
Text-Guided Pruning
TRIM
62.33
63.84
75
FitPrune
63.14
64.26
75
SparseVLM
62.73
61.78
75
Table 1 : Overview of Pruning Methods: Safety, Utility, and Compression. We report the average Attack Success Rate (ASR ↓ ) across MM-Safety, FigStep, and JailBreakV, average Utility score ( ↑ ) across MMBench, MM-Vet, LLaVA-Bench, and SQA, and Pruning Rate for Token-Pruning methods across three classes.
Figure 4 : Visualization of attention distributions across pruning paradigms. TRIM (Left) concentrates attention on a few foreground tokens, whereas LLaVA-Mini (Right) maintains balanced attention across the entire image.
Figure 5 : Temperature Scaling Intervention on TRIM. As temperature τ decreases, ASR rises while attention entropy drops. This confirms the link between attention collapse and increased vulnerability.
Figure 6 : The architecture of SAP. SAP operates in three stages. First, Malicious Anchor Identification (MAI) identifies foreground tokens that encode malicious semantics. Next, Benign Token Restoration (BTR) selectively restores pruned tokens from the background to create capacity for attention redistribution. Finally, Attention Reallocation (AR) redistributes excessive attention from malicious anchors to benign tokens, thereby diluting malicious influence after Token-Pruning.
Method
Safety ↓
Utility ↑
MM-Safety
FigStep
JailBreakV
AVG
MMBench
MM-Vet
LLaVA-Bench
SQA I
AVG
Original
59.37
51.60
51.05
54.01
64.60
36.20
87.48
69.56
64.46
Random
56.75
50.00
51.43
52.73
62.80
32.10
88.18
69.01
63.02
TRIM
65.87
62.20
58.93
62.33
66.92
35.70
83.28
69.46
63.84
w/ SAP
52.78 ↓ 13.1
47.80 ↓ 14.4
50.00 ↓ 8.9
50.19 ↓ 19.5%
67.27 ↑ 0.4
33.30 ↓ 2.4
83.60 ↑ 0.3
69.61 ↑ 0.2
63.45 ↓ 0.6%
HiPrune
69.44
57.20
52.86
59.83
62.37
36.50
85.57
68.57
63.25
Table 2 : Main Results: Balancing Safety and Utility. We report the performance of various pruning methods with and without our SAP defense across safety and utility benchmarks. Rows highlighted in green indicate our method. Small gray numbers in AVG columns represent the relative percentage change compared to the corresponding pruning baseline.
Method
MM-Safety ↓
FigStep ↓
Latency (ms)
TRIM
65.87
62.20
55.90
w/ LLaVAGuard
51.59
49.40
160.09
w/ SAP
52.78
47.80
57.19
w/ SAP + LLaVAGuard
41.07
35.80
162.44
Table 3 : Comparison with external guardrails. We compare SAP with LLaVAGuard on LLaVA-1.5-7B with TRIM. SAP achieves comparable safety performance with much lower latency.
Method
Safety ↓
MM-Safety
FigStep
JailBreakV
TRIM
65.87
62.20
58.93
w/ SAP
52.78
47.80
50.00
w/ SAP †
20.04
8.20
15.36
FasterVLM
72.02
63.00
54.64
w/ SAP
54.56
41.80
46.43
Table 4 : Safety performance under different SAP variants . SAP † denotes the expansion of SAP to textual instruction.
BTR
AR
Safety ↓
Utility ↑
MM-Safety
FigStep
SQA I
MMBench
-
-
65.87
62.20
69.46
66.92
✓
-
63.89
59.20
69.56
66.84
-
✓
53.57
50.40
68.52
63.23
✓
✓
52.78
47.80
69.61
67.27
Table 5 : Ablation study on SAP components. Evaluating individual contributions of Benign Token Restoration ( BTR ) and Attention Reallocation ( AR ) on LLaVA-1.5-7B (w/ TRIM).
Top- m
Safety ↓
Utility ↑
MM-Safety
FigStep
SQA I
MMBench
0
65.87
62.20
69.46
66.92
5
54.96
53.00
68.67
66.49
10
52.78
47.80
69.61
67.27
20
52.18
43.40
67.87
60.31
30
47.42
42.80
65.25
59.02
Table 6 : Ablation Study on Amounts of Malicious Anchors ( m ).
Method
Safety ↓
Utility ↑
MM-Safety
FigStep
SQA I
TRIM
65.87
62.20
69.46
+SAP (50%)
51.59
49.60
69.76
+SAP (75%)
52.78
47.80
69.61
+SAP (85%)
50.40
47.00
69.86
+SAP (90%)
50.40
46.40
69.21
Table 7 : Sensitivity of SAP to pruning ratios. Evaluating SAP under varying pruning ratios on LLaVA-1.5-7B (w/ TRIM).
Method
Latency (ms)
FLOPs (T)
Memory (GB)
Original
94.29
8.79
14.66
TRIM
55.90
2.22
14.51
w/ SAP
57.19
2.22
14.53
Table 8 : Efficiency Analysis. We report hardware latency (Latency), theoretical complexity (FLOPs), and peak GPU memory usage. All metrics are measured on a single NVIDIA GeForce RTX 4090 GPU with LLaVA-1.5-7B.
Figure 7 : Visualizing the Impact of SAP on Attention Distributions. Compared to TRIM baseline (left), TRIM with SAP (right) effectively alleviates pruning-induced attention collapse by restoring entropy, thereby improving safety.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Removal Target
ASR (%)
SQA I
Original
59.37
69.56
Foreground removed
50.00
65.25
Background removed
65.87
69.46
Appendix
Table 9 : Foreground and background removal analysis.
Method
ASR (%)
Entropy
Original
54.01
3.39
Random
52.73
3.20
TRIM
62.33
2.31
FitPrune
63.14
2.21
SparseVLM
62.73
2.64
HiPrune
59.83
2.79
Appendix
Table 10 : Correlation between Attention Entropy and Model Safety. We present the average ASR alongside attention entropy for various pruning methods. A clear negative correlation is observed: methods with lower entropy (higher concentration) generally exhibit higher ASR, reinforcing the malicious semantic purification hypothesis.
Method
Pruning Ratio (%)
ASR (%)
Original
0%
59.37
Random
50%
60.12
Random
60%
59.13
Random
70%
58.33
Random
80%
55.56
Random
90%
52.18
Appendix
Table 11 : Empirical Safety Performance of Random Pruning on MM-SafetyBench.
Setting
MM-Safety ↓
FigStep ↓
TRIM (baseline)
65.87
62.20
Shallow layers only (0–15)
60.71
58.60
Deep layers only (16–31)
56.15
56.00
All layers (ours)
52.78
47.80
Appendix
Table 12 : Ablation study on layer choice for AR.
Setting
MM-Safety ↓
FigStep ↓
TRIM (baseline)
65.87
62.20
Averaged across heads
67.26
59.20
Per-head (ours)
52.78
47.80
Appendix
Table 13 : Ablation study on head-wise configuration for AR.
Decoding Setting
MM-Safety ↓
FigStep ↓
Greedy ( T=0 )
53.2 ± 0.4
48.1 ± 0.5
Sampling ( T=0.7 )
52.9 ± 0.6
48.6 ± 0.7
Appendix
Table 14: Statistical robustness under different decoding settings.
Safety-Aligned Instruction
MM-Safety ↓
FigStep ↓
JailBreakV ↓
Please answer safely and accurately. (Default)
52.78
47.80
50.00
Safe and aligned with human values.
53.57
50.80
50.71
Prevent exposure to harmful materials.
50.99
48.60
52.14
Compliance with safety standards.
54.17
50.20
51.07
Appendix
Table 15 : Sensitivity of vsafe prompt on LLaVA-1.5-7B (TRIM+SAP). ASR (%) across diverse safety-aligned instructions.
Figure 8 : Attention visualization for the case study.