Large Vision-Language Models (LVLMs) face significant computational inefficiencies caused by the large number of visual tokens. Existing visual token pruning methods mainly focus on either retaining individually important tokens or selecting mutually diverse ones. In this work, we revisit visual token pruning from a coverage perspective and formulate it as a biased attention coverage maximization problem. The key idea is to select a compact token subset whose encoder-side outgoing attention can jointly cover the image while assigning higher coverage priority to more informative regions. From this perspective, we propose ACPruner, a training-free visual token pruning framework for efficient LVLM inference. ACPruner first estimates token importance by combining intra-modal saliency and inter-modal relevance, then derives token-wise coverage from attention patterns within the vision encoder, and finally performs greedy selection to maximize the proposed coverage objective. Extensive experiments across multiple LVLM backbones, including LLaVA-1.5-7B/13B, LLaVA-NeXT-7B/13B, Qwen2.5-VL-7B, and LLaVA-OneVision-7B, show that ACPruner consistently achieves strong performance retention while delivering substantial end-to-end inference speedups.
Figures & tables
Figure 1: The left part illustrates the core idea of ACPruner. The right part shows how ACPruner is integrated with LVLM inference pipelines and provides a workflow visualization of the method.
Method
Venue
MME
GQA
SQA
POPE
TextVQA
VizWiz
VQA-v2
MMB
MM-Vet
Perf. Ret.
Upper Bound, 576 visual tokens (pruning ratio = 0%)
LLaVA-1.5-7B
CVPR’24
1862
61.9
69.5
85.9
58.2
50.0
78.4
64.7
31.3
100.0%
Retain 192 visual tokens (pruning ratio = 66.7%)
FastV
ECCV’24
1612
52.7
67.3
64.8
52.5
50.8
67.1
61.2
27.7
89.4%
PruMerge
ICCV’25
1632
54.3
67.9
71.3
54.3
50.1
70.6
59.6
-
91.5%
FasterVLM
arXiv’24
1780
59.3
70.0
85.3
57.3
50.1
75.2
63.5
31.8
98.4%
Table 1: Performance comparison on LLaVA-1.5-7B under different pruning ratios. “Perf. Ret.” denotes the average performance retention rate relative to the full-token baseline.
Method
Venue
MME
GQA
SQA
POPE
TextVQA
VizWiz
VQA-v2
MMB
Perf. Ret.
Upper Bound, up to 2880 visual tokens (pruning ratio = 0%)
LLaVA-NeXT-7B
CVPR’24
1842
64.3
70.2
86.5
61.3
55.2
81.3
67.9
100.0%
Retain up to 640 visual tokens (pruning ratio = 77.8%)
FastV
ECCV’24
1807
58.9
67.4
79.5
58.1
53.9
77.0
63.1
94.7%
PruMerge
ICCV’25
1790
60.8
67.8
85.3
54.9
57.9
78.2
64.6
96.6%
DART
EMNLP’25
1793
61.3
68.2
85.0
59.5
57.0
78.3
64.9
97.5%
Table 2: Performance comparison on LLaVA-NeXT-7B under different pruning ratios. “Perf. Ret.” denotes the average performance retention rate relative to the full-token baseline.
VideoMME
Method
Venue
MVBench
LongVideo
NextQA
Overall
Short
Medium
Long
AVG.
Perf. Ret.
Pruning Ratio = 0%
LLaVA-OV-7B
TMLR’25
56.9
56.4
79.0
58.6
70.3
56.6
48.8
62.7
100.0%
Pruning Ratio = 75%
FastV
ECCV’24
54.7
55.5
77.5
56.2
68.0
54.6
46.0
61.0
97.1%
DivPrune
CVPR’25
54.6
55.7
77.7
57.1
69.0
54.6
47.9
61.3
97.6%
Table 3: Performance comparison on LLaVA-OneVision-7B for video understanding tasks. “Perf. Ret.” denotes the average performance retention rate relative to the full-token baseline.
Method
MME
POPE
TextVQA
MMB
GQA
SQA
VQA-v2
Perf. Ret.
Pruning ratio = 0%
Qwen2.5-VL-7B
2310
86.3
84.8
83.9
60.9
88.9
82.9
100.0%
Pruning ratio = 66.7%
FastV
2072
82.2
77.9
75.7
58.0
78.5
80.4
92.5%
DivPrune
2198
85.6
80.1
81.6
59.0
86.6
80.9
96.8%
VisionZIP
2317
85.8
80.4
78.9
56.6
80.5
80.7
95.6%
Table 4: Performance comparison on Qwen2.5-VL-7B under different pruning ratios. “Perf. Ret.” denotes the average performance retention rate relative to the full-token baseline.
Sample-level Inference Time Breakdown (MilliSeconds)
Method
# Tokens
FLOPs (Tera)
KV-cache (MB)
Eval. Time (Seconds)
Score
Visual Encoding
Token Pruning
LLM Inference
Total
LLaVA-1.5-7B
576
8.5
318.6
511.8
1862
14.6
-
201.4
215.7
+ Ours
64 ( ↓88.9% )
1.6 ( ↓81.2% )
62.6 ( ↓80.4% )
299.1 ( ↓41.6% )
1728
14.6
10.2
101.3
126.1
LLaVA-1.5-13B
576
16.5
497.9
742.9
1818
14.6
-
298.5
313.1
+ Ours
64 ( ↓88.9% )
3.2 ( ↓80.6% )
97.9 ( ↓80.4% )
380.3 ( ↓48.8% )
1820
14.6
10.5
135.2
160.3
LLaVA-NeXT-7B
2880
30.6
1084.7
1184.4
1842
35.9
-
463.2
499.1
Table 5: Practical efficiency analysis of ACPruner on MME. The left section shows benchmark-level efficiency statistics while the right section shows sample-level runtime breakdowns.
Figure 2: Ablation of α and β on LLaVA-1.5-7B and Qwen2.5-VL-7B under a 77.8% pruning ratio.
MME
GQA
SQA
POPE
TextVQA
Perf. Ret.
Full method
1812
59.8
69.4
87.2
57.6
98.9%
Setting-1
1768
59.3
68.6
86.0
57.4
97.6%
Setting-2
1715
59.4
69.1
87.0
54.6
96.5%
Setting-3
1700
59.3
68.7
86.3
50.9
94.8%
Setting-4
1780
57.9
68.3
84.8
53.3
95.5%
Setting-5
1749
59.0
68.6
86.8
53.1
96.0%
Table 6: Ablation of ACPruner’s main components on LLaVA-1.5-7B under a 77.8% pruning ratio.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Δ(vi∣S)=∑j=1Twˉjmax(cij−mj,0)
Appendix
Algorithm 1 ACPruner
Method
Venue
MME
GQA
SQA
POPE
TextVQA
VizWiz
VQA-v2
MMB
MM-Vet
Perf. Ret.
Upper Bound, 576 visual tokens (pruning ratio = 0%)
LLaVA-1.5-13B
CVPR’24
1818
63.2
72.8
85.9
61.3
56.6
80.0
67.7
35.3
100.0%
Retain 192 visual tokens (pruning ratio = 66.7%)
VisionZIP
CVPR’25
1754
59.1
73.5
85.1
59.5
54.9
78.0
66.9
37.5
98.5%
SCOPE
NeurIPS’25
1775
59.7
73.8
86.7
60.0
55.0
78.1
67.6
39.4
99.8%
PruneSID
ICLR’26
1770
59.6
72.8
86.4
58.6
56.0
78.0
65.9
38.0
98.8%
Appendix
Table 7: Performance comparison on LLaVA-1.5-13B under different pruning ratios. “Perf. Ret.” denotes the average performance retention rate relative to the full-token baseline.
Method
Venue
MME
GQA
SQA
POPE
TextVQA
VizWiz
VQA-v2
MMB
Perf. Ret.
Upper Bound, 2880 visual tokens (pruning ratio = 0%)
LLaVA-NeXT-13B
CVPR’24
1901
65.4
73.5
86.2
64.3
64.0
81.8
70.0
100.0%
Retain up to 640 visual tokens (pruning ratio = 77.8%)
VisionZIP
CVPR’25
1871
63.0
71.2
85.7
62.2
56.0
79.7
68.6
96.3%
SCOPE
NeurIPS’25
1897
63.6
72.5
86.4
62.4
60.0
79.5
69.3
97.9%
PruneSID
ICLR’26
1817
62.4
70.1
85.6
60.2
60.2
79.1
67.0
96.1%
Appendix
Table 8: Performance comparison on LLaVA-NeXT-13B under different pruning ratios. “Perf. Ret.” denotes the average performance retention rate relative to the full-token baseline.
MME
GQA
SQA
POPE
TextVQA
Perf. Ret.
CLS attention
1812
59.8
69.4
87.2
57.6
98.9%
Patch incoming attention
1780
59.2
68.9
87.1
57.1
98.0%
L2 norm
1707
59.7
68.5
87.0
53.9
96.1%
Similarity with global mean
1665
59.1
68.6
87.2
48.8
93.8%
Appendix
Table 9: Ablation of different intra-modality importance variants on LLaVA-1.5-7B under a 77.8% pruning ratio. The blue row corresponds to ACPruner’s design choice.
MME
GQA
SQA
POPE
Perf. Ret.
Noun
1812
59.8
69.4
87.2
98.8%
Noun + Adj
1804
59.8
69.2
87.2
98.6%
Noun + Adj + Verb
1800
59.7
69.1
87.2
98.5%
Full text
1782
59.5
68.9
86.9
98.0%
Appendix
Table 10: Ablation of noun extraction on LLaVA-1.5-7B under a 77.8% pruning ratio. The blue row corresponds to ACPruner’s design choice.
MME
GQA
SQA
POPE
TextVQA
Perf. Ret.
Token-level Max
1813
59.7
69.2
86.9
57.5
98.7%
Token-level Mean
1806
59.6
68.7
86.8
56.9
98.2%
Word-level Max
1812
59.8
69.4
87.2
57.6
98.9%
Word-level Mean
1768
59.3
69.2
86.0
57.4
97.8%
Appendix
Table 11: Ablation of different inter-modality importance variants on LLaVA-1.5-7B under a 77.8% pruning ratio. The blue row corresponds to ACPruner’s design choice.
Layer Index
MME
GQA
SQA
POPE
TextVQA
Perf. Ret.
Initial Layers
0
1797
59.5
68.8
87.1
55.1
97.5%
1
1746
59.7
69.2
87.2
53.6
96.7%
Middle Layers
9
1753
59.4
68.3
87.0
55.6
97.0%
10
1756
59.5
68.1
86.8
55.7
97.0%
Appendix
Table 12: Ablation of CLS attention layer on LLaVA-1.5-7B under a 77.8% pruning ratio. The blue row corresponds to ACPruner’s design choice.
Layer Index
MME
GQA
SQA
POPE
TextVQA
Perf. Ret.
Initial Layers
0
1738
58.8
68.5
85.2
52.2
89.7%
1
1745
59.4
68.4
86.3
53.7
92.3%
Middle Layers
9
1721
60.0
68.4
86.5
53.3
91.6%
10
1784
59.5
68.1
86.6
56.3
96.7%
Appendix
Table 13: Ablation of patch attention layer on LLaVA-1.5-7B under a 77.8% pruning ratio. The blue row corresponds to ACPruner’s design choice.
Figure 3: Step-wise runtime distribution of ACPruner on different LVLM backbones.
Figure 4: Visualization of ACPruner’s pruning results on LLaVA-1.5-7B under different token budgets. The masked patches correspond to visual tokens pruned by ACPruner. For each image, we provide two different text queries to demonstrate the query-relevant nature of ACPruner.
Visual token pruning reduces the inference cost of vision-language models (VLMs), but most methods only ask which tokens to keep. This retained-token view can keep redundant high-scoring tokens while leaving discarded evidence without a close representative. We propose CoverPruner, a training-free pruner that asks the complementary demand-side question: after a token is removed, which surviving original token represents it for the target VLM? CoverPruner formulates pruning as Representational Coverage Maximization (RCM), covering the full projected visual-token set with query-weighted demand. It instantiates RCM with projector-space coverage and a lightweight first-layer attention probe. Across multiple VLM architectures and compression rates, CoverPruner achieves the best average accuracy among all compared methods, with the largest gains usually appearing under aggressive compression.
Qingchan Zhu, Weihang You, Hanqi Jiang +3
School of Computing, University of Georgia · College of Engineering, Northeastern University
Vision-Language Models (VLMs) have recently demonstrated remarkable capabilities in visual understanding and reasoning, but they also impose significant computational burdens due to long visual sequence inputs. Recent works address this issue by pruning unimportant visual tokens, achieving substantial computational reduction while maintaining model performance. The core of token pruning lies in determining token importance, with current approaches primarily relying on attention scores from vision encoders or Large Language Models (LLMs). In this paper, we analyze the effectiveness of attention mechanisms in both vision encoders and LLMs. We find that vision encoders suffer from attention sink, leading to poor focus on informative foreground regions, while in LLMs, although prior studies have identified attention bias toward token positions, text-to-vision attention demonstrates resistance to this bias and enables effective pruning guidance in middle layers. Based on these observations, we propose LearnPruner, a two-stage token pruning framework that first removes redundant vision tokens via a learnable pruning module after the vision encoder, then retains only task-relevant tokens in the LLM's middle layer. Experimental results show that our LearnPruner can preserve approximately 95% of the original performance while using only 5.5% of vision tokens, and achieve 3.2× inference acceleration, demonstrating a superior accuracy-efficiency trade-off.
Vision-language models (VLMs) rely on long visual token sequences for visual understanding, making the prefill stage expensive in both computation and memory. Most existing pruning methods follow an absolute-ranking paradigm, assigning importance scores to visual tokens and retaining a fixed top-K subset. In this work, we argue that this paradigm is fundamentally brittle: attention sinks distort token importance rankings, while image redundancy and query-dependent visual evidence make fixed token budgets unreliable across inputs. We propose OccamToken, a training-free framework that replaces absolute token ranking with register-anchored relative evidence testing. Instead of asking which tokens are globally important, OccamToken evaluates whether a visual token provides information beyond a register-based reference. Our key insight is that register tokens naturally absorb low-information attention patterns, making them a stable reference for identifying genuinely informative visual evidence. Based on this principle, OccamToken performs both image-adaptive redundancy pruning and query-adaptive relevance pruning through dynamic thresholds derived from register attention. Across LLaVA-NeXT, LLaVA-v1.5, and Qwen3-VL, OccamToken consistently improves the accuracy-efficiency trade-off without additional training. Notably, on LLaVA-NeXT, it reduces 2,880 visual tokens to approximately 40 while preserving over 93% of the original accuracy, enabling stable visual token compression even in the extreme 1.4% retention regime.