Organizations: School of Intelligence Science and Technology, Nanjing University · School of Electronic Science and Engineering, Nanjing University · Geely Automobile Research Institute (Ningbo) Co., Ltd., 315000 · Alpha Labs, Goertek
Vision-Language Models (VLMs) excel at visual understanding and reasoning but often incur substantial inference costs due to the large number of visual tokens. Recent visual token pruning methods increasingly follow a two-stage paradigm: they first remove visually redundant tokens after the vision encoder and then discard tokens irrelevant to the textual query within the Large Language Model (LLM). However, since the first stage typically relies solely on vision-encoder saliency, it may prematurely eliminate query-relevant tokens, depriving the subsequent text-guided stage of critical visual evidence. Our empirical analysis shows that incorporating query guidance into first-stage pruning better preserves task-relevant evidence and consistently improves performance over vision-only saliency-based pruning. We further find that high-variance attention heads are more sensitive to the textual query and yield more discriminative text-to-vision attention signals for second-stage pruning. Motivated by these findings, we propose TReVS, a training-free framework that combines textual relevance with vision-encoder saliency for pre-LLM pruning and leverages high-variance attention heads to remove task-irrelevant tokens at shallow-to-intermediate layers of the LLM. On LLaVA-1.5-7B, TReVS retains 92.8% of the unpruned baseline performance while pruning 94.4% of visual tokens, outperforming prior state-of-the-art methods.
Figures & tables
Figure 1: (a–c) A two-stage baseline guided only by [CLS] attention during pre-LLM pruning discards visual tokens representing the queried grass, leaving incomplete evidence for subsequent query-guided pruning. TReVS uses textual relevance in pre-LLM pruning, preserving this evidence for the in-LLM stage. (d) TReVS achieves the best performance across six image-understanding benchmarks.
Figure 2: Effect of textual guidance on pre-LLM pruning. (a) Vision-encoder saliency focuses on the dominant fire hydrant, whereas textual relevance identifies the queried car. Their integration preserves both. (b) Performance relative to dense inference under different textual ratios. (c) Mean Spearman correlation between the two signals within the top- K saliency-selected tokens.
Figure 3: Head sensitivity and positional attention patterns. (a) Mean textual sensitivity of cumulative head subsets ranked by text-to-vision attention variance. (b) Average text-to-vision attention across visual-token positions at selected LLM layers.
Figure 4: The architecture of TReVS. Stage 1 performs query-aware pre-LLM pruning by preserving visually salient, query-relevant, and diverse tokens. Stage 2 uses high-variance attention heads to guide in-LLM visual token pruning.
Method
GQA
SQA I
VQA T
POPE
MME
MMB
MMB CN
RelAcc.
Upper Bound, 576 Tokens (100%)
Vanilla
61.9
69.5
58.2
85.9
1862
64.7
58.3
100.0%
Retain Averaged 128 Tokens ( ↓ 77.8% )
FastV (ECCV 2024)
49.6
60.2
50.6
59.6
1490
56.1
51.4
82.6%
SparseVLM (ICML 2025)
56.0
67.1
54.9
80.5
1696
60.0
51.1
92.4%
DivPrune (CVPR 2025)
59.3
69.0
56.1
86.7
1718
62.0
54.8
96.4%
Table 1: Performance comparison on LLaVA-1.5-7B under matched average token budgets. RelAcc. averages benchmark-wise performance relative to the unpruned model. Bold and underlined entries denote the best and second-best results.
Method
GQA
VQA T
MME
MMB
RelAcc.
Upper Bound, 2,880 Tokens (100%)
Vanilla
64.2
61.3
1842
67.9
100.0%
Retain Averaged 320 Tokens ( ↓ 88.9% )
SparseVLM
57.7
55.9
1694
64.3
91.9%
DivPrune
61.1
56.2
1724
63.9
93.6%
VisionZip
59.3
58.9
1702
63.1
93.4%
Table 2: Performance comparison on LLaVA-NeXT-7B under matched average token budgets.
Method
TGIF
MSVD
MSRVTT
RelAcc.
Upper Bound, 2,048 Tokens (100%)
Video-LLaVA
48.7
70.1
57.4
100.0%
Retain Averaged 136 Tokens ( ↓ 93.4% )
FastV
30.4
44.3
39.4
64.8%
SparseVLM
44.7
68.2
31.0
81.0%
VisionZip
42.4
63.5
52.1
89.5%
Table 3: Performance comparison on Video-LLaVA-7B with 136 average retained tokens.
Stage 1
Stage 2
RelAcc. (%)
[CLS] Attn
All Heads
91.8
TReVS
All Heads
92.7
High-Variance Heads
93.2
Table 4: Ablation study of the two-stage token-pruning strategies on LLaVA-1.5-7B under the 32-token preset. RelAcc. denotes the average relative accuracy over TextVQA, MMBench, GQA, and POPE.