Large Vision-Language Models (LVLMs) face significant computational inefficiencies caused by the large number of visual tokens. Existing visual token pruning methods mainly focus on either retaining individually important tokens or selecting mutually diverse ones. In this work, we revisit visual token pruning from a coverage perspective and formulate it as a biased attention coverage maximization problem. The key idea is to select a compact token subset whose encoder-side outgoing attention can jointly cover the image while assigning higher coverage priority to more informative regions. From this perspective, we propose ACPruner, a training-free visual token pruning framework for efficient LVLM inference. ACPruner first estimates token importance by combining intra-modal saliency and inter-modal relevance, then derives token-wise coverage from attention patterns within the vision encoder, and finally performs greedy selection to maximize the proposed coverage objective. Extensive experiments across multiple LVLM backbones, including LLaVA-1.5-7B/13B, LLaVA-NeXT-7B/13B, Qwen2.5-VL-7B, and LLaVA-OneVision-7B, show that ACPruner consistently achieves strong performance retention while delivering substantial end-to-end inference speedups.
Figures & tables
Figure 1: The left part illustrates the core idea of ACPruner. The right part shows how ACPruner is integrated with LVLM inference pipelines and provides a workflow visualization of the method.
Method
Venue
MME
GQA
SQA
POPE
TextVQA
VizWiz
VQA-v2
MMB
MM-Vet
Perf. Ret.
Upper Bound, 576 visual tokens (pruning ratio = 0%)
LLaVA-1.5-7B
CVPR’24
1862
61.9
69.5
85.9
58.2
50.0
78.4
64.7
31.3
100.0%
Retain 192 visual tokens (pruning ratio = 66.7%)
FastV
ECCV’24
1612
52.7
67.3
64.8
52.5
50.8
67.1
61.2
27.7
89.4%
PruMerge
ICCV’25
1632
54.3
67.9
71.3
54.3
50.1
70.6
59.6
-
91.5%
FasterVLM
arXiv’24
1780
59.3
70.0
85.3
57.3
50.1
75.2
63.5
31.8
98.4%
Table 1: Performance comparison on LLaVA-1.5-7B under different pruning ratios. “Perf. Ret.” denotes the average performance retention rate relative to the full-token baseline.
Method
Venue
MME
GQA
SQA
POPE
TextVQA
VizWiz
VQA-v2
MMB
Perf. Ret.
Upper Bound, up to 2880 visual tokens (pruning ratio = 0%)
LLaVA-NeXT-7B
CVPR’24
1842
64.3
70.2
86.5
61.3
55.2
81.3
67.9
100.0%
Retain up to 640 visual tokens (pruning ratio = 77.8%)
FastV
ECCV’24
1807
58.9
67.4
79.5
58.1
53.9
77.0
63.1
94.7%
PruMerge
ICCV’25
1790
60.8
67.8
85.3
54.9
57.9
78.2
64.6
96.6%
DART
EMNLP’25
1793
61.3
68.2
85.0
59.5
57.0
78.3
64.9
97.5%
Table 2: Performance comparison on LLaVA-NeXT-7B under different pruning ratios. “Perf. Ret.” denotes the average performance retention rate relative to the full-token baseline.
VideoMME
Method
Venue
MVBench
LongVideo
NextQA
Overall
Short
Medium
Long
AVG.
Perf. Ret.
Pruning Ratio = 0%
LLaVA-OV-7B
TMLR’25
56.9
56.4
79.0
58.6
70.3
56.6
48.8
62.7
100.0%
Pruning Ratio = 75%
FastV
ECCV’24
54.7
55.5
77.5
56.2
68.0
54.6
46.0
61.0
97.1%
DivPrune
CVPR’25
54.6
55.7
77.7
57.1
69.0
54.6
47.9
61.3
97.6%
Table 3: Performance comparison on LLaVA-OneVision-7B for video understanding tasks. “Perf. Ret.” denotes the average performance retention rate relative to the full-token baseline.
Method
MME
POPE
TextVQA
MMB
GQA
SQA
VQA-v2
Perf. Ret.
Pruning ratio = 0%
Qwen2.5-VL-7B
2310
86.3
84.8
83.9
60.9
88.9
82.9
100.0%
Pruning ratio = 66.7%
FastV
2072
82.2
77.9
75.7
58.0
78.5
80.4
92.5%
DivPrune
2198
85.6
80.1
81.6
59.0
86.6
80.9
96.8%
VisionZIP
2317
85.8
80.4
78.9
56.6
80.5
80.7
95.6%
Table 4: Performance comparison on Qwen2.5-VL-7B under different pruning ratios. “Perf. Ret.” denotes the average performance retention rate relative to the full-token baseline.
Sample-level Inference Time Breakdown (MilliSeconds)
Method
# Tokens
FLOPs (Tera)
KV-cache (MB)
Eval. Time (Seconds)
Score
Visual Encoding
Token Pruning
LLM Inference
Total
LLaVA-1.5-7B
576
8.5
318.6
511.8
1862
14.6
-
201.4
215.7
+ Ours
64 ( ↓88.9% )
1.6 ( ↓81.2% )
62.6 ( ↓80.4% )
299.1 ( ↓41.6% )
1728
14.6
10.2
101.3
126.1
LLaVA-1.5-13B
576
16.5
497.9
742.9
1818
14.6
-
298.5
313.1
+ Ours
64 ( ↓88.9% )
3.2 ( ↓80.6% )
97.9 ( ↓80.4% )
380.3 ( ↓48.8% )
1820
14.6
10.5
135.2
160.3
LLaVA-NeXT-7B
2880
30.6
1084.7
1184.4
1842
35.9
-
463.2
499.1
Table 5: Practical efficiency analysis of ACPruner on MME. The left section shows benchmark-level efficiency statistics while the right section shows sample-level runtime breakdowns.
Figure 2: Ablation of α and β on LLaVA-1.5-7B and Qwen2.5-VL-7B under a 77.8% pruning ratio.
MME
GQA
SQA
POPE
TextVQA
Perf. Ret.
Full method
1812
59.8
69.4
87.2
57.6
98.9%
Setting-1
1768
59.3
68.6
86.0
57.4
97.6%
Setting-2
1715
59.4
69.1
87.0
54.6
96.5%
Setting-3
1700
59.3
68.7
86.3
50.9
94.8%
Setting-4
1780
57.9
68.3
84.8
53.3
95.5%
Setting-5
1749
59.0
68.6
86.8
53.1
96.0%
Table 6: Ablation of ACPruner’s main components on LLaVA-1.5-7B under a 77.8% pruning ratio.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Δ(vi∣S)=∑j=1Twˉjmax(cij−mj,0)
Appendix
Algorithm 1 ACPruner
Method
Venue
MME
GQA
SQA
POPE
TextVQA
VizWiz
VQA-v2
MMB
MM-Vet
Perf. Ret.
Upper Bound, 576 visual tokens (pruning ratio = 0%)
LLaVA-1.5-13B
CVPR’24
1818
63.2
72.8
85.9
61.3
56.6
80.0
67.7
35.3
100.0%
Retain 192 visual tokens (pruning ratio = 66.7%)
VisionZIP
CVPR’25
1754
59.1
73.5
85.1
59.5
54.9
78.0
66.9
37.5
98.5%
SCOPE
NeurIPS’25
1775
59.7
73.8
86.7
60.0
55.0
78.1
67.6
39.4
99.8%
PruneSID
ICLR’26
1770
59.6
72.8
86.4
58.6
56.0
78.0
65.9
38.0
98.8%
Appendix
Table 7: Performance comparison on LLaVA-1.5-13B under different pruning ratios. “Perf. Ret.” denotes the average performance retention rate relative to the full-token baseline.
Method
Venue
MME
GQA
SQA
POPE
TextVQA
VizWiz
VQA-v2
MMB
Perf. Ret.
Upper Bound, 2880 visual tokens (pruning ratio = 0%)
LLaVA-NeXT-13B
CVPR’24
1901
65.4
73.5
86.2
64.3
64.0
81.8
70.0
100.0%
Retain up to 640 visual tokens (pruning ratio = 77.8%)
VisionZIP
CVPR’25
1871
63.0
71.2
85.7
62.2
56.0
79.7
68.6
96.3%
SCOPE
NeurIPS’25
1897
63.6
72.5
86.4
62.4
60.0
79.5
69.3
97.9%
PruneSID
ICLR’26
1817
62.4
70.1
85.6
60.2
60.2
79.1
67.0
96.1%
Appendix
Table 8: Performance comparison on LLaVA-NeXT-13B under different pruning ratios. “Perf. Ret.” denotes the average performance retention rate relative to the full-token baseline.
MME
GQA
SQA
POPE
TextVQA
Perf. Ret.
CLS attention
1812
59.8
69.4
87.2
57.6
98.9%
Patch incoming attention
1780
59.2
68.9
87.1
57.1
98.0%
L2 norm
1707
59.7
68.5
87.0
53.9
96.1%
Similarity with global mean
1665
59.1
68.6
87.2
48.8
93.8%
Appendix
Table 9: Ablation of different intra-modality importance variants on LLaVA-1.5-7B under a 77.8% pruning ratio. The blue row corresponds to ACPruner’s design choice.
MME
GQA
SQA
POPE
Perf. Ret.
Noun
1812
59.8
69.4
87.2
98.8%
Noun + Adj
1804
59.8
69.2
87.2
98.6%
Noun + Adj + Verb
1800
59.7
69.1
87.2
98.5%
Full text
1782
59.5
68.9
86.9
98.0%
Appendix
Table 10: Ablation of noun extraction on LLaVA-1.5-7B under a 77.8% pruning ratio. The blue row corresponds to ACPruner’s design choice.
MME
GQA
SQA
POPE
TextVQA
Perf. Ret.
Token-level Max
1813
59.7
69.2
86.9
57.5
98.7%
Token-level Mean
1806
59.6
68.7
86.8
56.9
98.2%
Word-level Max
1812
59.8
69.4
87.2
57.6
98.9%
Word-level Mean
1768
59.3
69.2
86.0
57.4
97.8%
Appendix
Table 11: Ablation of different inter-modality importance variants on LLaVA-1.5-7B under a 77.8% pruning ratio. The blue row corresponds to ACPruner’s design choice.
Layer Index
MME
GQA
SQA
POPE
TextVQA
Perf. Ret.
Initial Layers
0
1797
59.5
68.8
87.1
55.1
97.5%
1
1746
59.7
69.2
87.2
53.6
96.7%
Middle Layers
9
1753
59.4
68.3
87.0
55.6
97.0%
10
1756
59.5
68.1
86.8
55.7
97.0%
Appendix
Table 12: Ablation of CLS attention layer on LLaVA-1.5-7B under a 77.8% pruning ratio. The blue row corresponds to ACPruner’s design choice.
Layer Index
MME
GQA
SQA
POPE
TextVQA
Perf. Ret.
Initial Layers
0
1738
58.8
68.5
85.2
52.2
89.7%
1
1745
59.4
68.4
86.3
53.7
92.3%
Middle Layers
9
1721
60.0
68.4
86.5
53.3
91.6%
10
1784
59.5
68.1
86.6
56.3
96.7%
Appendix
Table 13: Ablation of patch attention layer on LLaVA-1.5-7B under a 77.8% pruning ratio. The blue row corresponds to ACPruner’s design choice.
Figure 3: Step-wise runtime distribution of ACPruner on different LVLM backbones.
Figure 4: Visualization of ACPruner’s pruning results on LLaVA-1.5-7B under different token budgets. The masked patches correspond to visual tokens pruned by ACPruner. For each image, we provide two different text queries to demonstrate the query-relevant nature of ACPruner.