Organizations: School of Intelligence Science and Technology, Nanjing University · School of Electronic Science and Engineering, Nanjing University · Geely Automobile Research Institute (Ningbo) Co., Ltd., 315000 · Alpha Labs, Goertek
Vision-Language Models (VLMs) excel at visual understanding and reasoning but often incur substantial inference costs due to the large number of visual tokens. Recent visual token pruning methods increasingly follow a two-stage paradigm: they first remove visually redundant tokens after the vision encoder and then discard tokens irrelevant to the textual query within the Large Language Model (LLM). However, since the first stage typically relies solely on vision-encoder saliency, it may prematurely eliminate query-relevant tokens, depriving the subsequent text-guided stage of critical visual evidence. Our empirical analysis shows that incorporating query guidance into first-stage pruning better preserves task-relevant evidence and consistently improves performance over vision-only saliency-based pruning. We further find that high-variance attention heads are more sensitive to the textual query and yield more discriminative text-to-vision attention signals for second-stage pruning. Motivated by these findings, we propose TReVS, a training-free framework that combines textual relevance with vision-encoder saliency for pre-LLM pruning and leverages high-variance attention heads to remove task-irrelevant tokens at shallow-to-intermediate layers of the LLM. On LLaVA-1.5-7B, TReVS retains 92.8% of the unpruned baseline performance while pruning 94.4% of visual tokens, outperforming prior state-of-the-art methods.
Figures & tables
Figure 1: (a–c) A two-stage baseline guided only by [CLS] attention during pre-LLM pruning discards visual tokens representing the queried grass, leaving incomplete evidence for subsequent query-guided pruning. TReVS uses textual relevance in pre-LLM pruning, preserving this evidence for the in-LLM stage. (d) TReVS achieves the best performance across six image-understanding benchmarks.
Figure 2: Effect of textual guidance on pre-LLM pruning. (a) Vision-encoder saliency focuses on the dominant fire hydrant, whereas textual relevance identifies the queried car. Their integration preserves both. (b) Performance relative to dense inference under different textual ratios. (c) Mean Spearman correlation between the two signals within the top- K saliency-selected tokens.
Figure 3: Head sensitivity and positional attention patterns. (a) Mean textual sensitivity of cumulative head subsets ranked by text-to-vision attention variance. (b) Average text-to-vision attention across visual-token positions at selected LLM layers.
Figure 4: The architecture of TReVS. Stage 1 performs query-aware pre-LLM pruning by preserving visually salient, query-relevant, and diverse tokens. Stage 2 uses high-variance attention heads to guide in-LLM visual token pruning.
Method
GQA
SQA I
VQA T
POPE
MME
MMB
MMB CN
RelAcc.
Upper Bound, 576 Tokens (100%)
Vanilla
61.9
69.5
58.2
85.9
1862
64.7
58.3
100.0%
Retain Averaged 128 Tokens ( ↓ 77.8% )
FastV (ECCV 2024)
49.6
60.2
50.6
59.6
1490
56.1
51.4
82.6%
SparseVLM (ICML 2025)
56.0
67.1
54.9
80.5
1696
60.0
51.1
92.4%
DivPrune (CVPR 2025)
59.3
69.0
56.1
86.7
1718
62.0
54.8
96.4%
Table 1: Performance comparison on LLaVA-1.5-7B under matched average token budgets. RelAcc. averages benchmark-wise performance relative to the unpruned model. Bold and underlined entries denote the best and second-best results.
Method
GQA
VQA T
MME
MMB
RelAcc.
Upper Bound, 2,880 Tokens (100%)
Vanilla
64.2
61.3
1842
67.9
100.0%
Retain Averaged 320 Tokens ( ↓ 88.9% )
SparseVLM
57.7
55.9
1694
64.3
91.9%
DivPrune
61.1
56.2
1724
63.9
93.6%
VisionZip
59.3
58.9
1702
63.1
93.4%
Table 2: Performance comparison on LLaVA-NeXT-7B under matched average token budgets.
Method
TGIF
MSVD
MSRVTT
RelAcc.
Upper Bound, 2,048 Tokens (100%)
Video-LLaVA
48.7
70.1
57.4
100.0%
Retain Averaged 136 Tokens ( ↓ 93.4% )
FastV
30.4
44.3
39.4
64.8%
SparseVLM
44.7
68.2
31.0
81.0%
VisionZip
42.4
63.5
52.1
89.5%
Table 3: Performance comparison on Video-LLaVA-7B with 136 average retained tokens.
Stage 1
Stage 2
RelAcc. (%)
[CLS] Attn
All Heads
91.8
TReVS
All Heads
92.7
High-Variance Heads
93.2
Table 4: Ablation study of the two-stage token-pruning strategies on LLaVA-1.5-7B under the 32-token preset. RelAcc. denotes the average relative accuracy over TextVQA, MMBench, GQA, and POPE.
Visual token pruning has been widely studied as a practical approach to reducing the computational cost of large vision-language models. However, it struggles to preserve essential visual information, which can lead to substantial performance degradation. In particular, image-based token selection can overlook task-relevant details, while text-guided token selection may fail to capture the text--visual relationships needed for complex reasoning. We find that applying textual guidance too early can limit its ability to identify answer-relevant visual regions, whereas text-to-visual attention becomes more informative at intermediate decoder depths. This finding motivates our training-free method, which separates early vision-guided pruning from deferred text-guided reselection. We first prune visual tokens using vision-encoder attention, retain additional candidates until the decoder midpoint, and then use text-to-visual attention to determine the final visual-token set. Across eight benchmarks and three models, our method outperforms the best-performing baselines by an average of 11.10 and 16.84 percentage points in performance recovery at 80% and 90% pruning, respectively, with comparable or lower LLM-prefill latency than most baselines. The source code is publicly available at https://github.com/kmc3661/DeFT
Minchan Kang, Kyeonghye Park, Seoyoung Cho +2
Korea Advanced Institute of Science and Technology (KAIST) · Hanbat National University
Vision-Language Models (VLMs) have recently demonstrated remarkable capabilities in visual understanding and reasoning, but they also impose significant computational burdens due to long visual sequence inputs. Recent works address this issue by pruning unimportant visual tokens, achieving substantial computational reduction while maintaining model performance. The core of token pruning lies in determining token importance, with current approaches primarily relying on attention scores from vision encoders or Large Language Models (LLMs). In this paper, we analyze the effectiveness of attention mechanisms in both vision encoders and LLMs. We find that vision encoders suffer from attention sink, leading to poor focus on informative foreground regions, while in LLMs, although prior studies have identified attention bias toward token positions, text-to-vision attention demonstrates resistance to this bias and enables effective pruning guidance in middle layers. Based on these observations, we propose LearnPruner, a two-stage token pruning framework that first removes redundant vision tokens via a learnable pruning module after the vision encoder, then retains only task-relevant tokens in the LLM's middle layer. Experimental results show that our LearnPruner can preserve approximately 95% of the original performance while using only 5.5% of vision tokens, and achieve 3.2× inference acceleration, demonstrating a superior accuracy-efficiency trade-off.
Vision-language models (VLMs) rely on long visual token sequences for visual understanding, making the prefill stage expensive in both computation and memory. Most existing pruning methods follow an absolute-ranking paradigm, assigning importance scores to visual tokens and retaining a fixed top-K subset. In this work, we argue that this paradigm is fundamentally brittle: attention sinks distort token importance rankings, while image redundancy and query-dependent visual evidence make fixed token budgets unreliable across inputs. We propose OccamToken, a training-free framework that replaces absolute token ranking with register-anchored relative evidence testing. Instead of asking which tokens are globally important, OccamToken evaluates whether a visual token provides information beyond a register-based reference. Our key insight is that register tokens naturally absorb low-information attention patterns, making them a stable reference for identifying genuinely informative visual evidence. Based on this principle, OccamToken performs both image-adaptive redundancy pruning and query-adaptive relevance pruning through dynamic thresholds derived from register attention. Across LLaVA-NeXT, LLaVA-v1.5, and Qwen3-VL, OccamToken consistently improves the accuracy-efficiency trade-off without additional training. Notably, on LLaVA-NeXT, it reduces 2,880 visual tokens to approximately 40 while preserving over 93% of the original accuracy, enabling stable visual token compression even in the extreme 1.4% retention regime.