Visual token pruning has been widely studied as a practical approach to reducing the computational cost of large vision-language models. However, it struggles to preserve essential visual information, which can lead to substantial performance degradation. In particular, image-based token selection can overlook task-relevant details, while text-guided token selection may fail to capture the text--visual relationships needed for complex reasoning. We find that applying textual guidance too early can limit its ability to identify answer-relevant visual regions, whereas text-to-visual attention becomes more informative at intermediate decoder depths. This finding motivates our training-free method, which separates early vision-guided pruning from deferred text-guided reselection. We first prune visual tokens using vision-encoder attention, retain additional candidates until the decoder midpoint, and then use text-to-visual attention to determine the final visual-token set. Across eight benchmarks and three models, our method outperforms the best-performing baselines by an average of 11.10 and 16.84 percentage points in performance recovery at 80% and 90% pruning, respectively, with comparable or lower LLM-prefill latency than most baselines. The source code is publicly available at https://github.com/kmc3661/DeFT
Figures & tables
Figure 1: Failure cases on Qwen3-VL-8B at 80% pruning. Image-based visual token pruning discards the visual region containing the requested text, while text-guided visual token pruning preserves related support but loses the corresponding value. Gray: discarded regions; orange: merged support. Percentages denote full-benchmark performance losses relative to Dense.
Figure 2: Text-guided evidence identification across decoder depth. Top: correlation with masking-based region importance. Bottom: matching-versus-swapped question gain. Results average four QA benchmarks equally, each with 300 images and two questions per image.
Boundary
Qwen3-VL-4B
Qwen3-VL-8B
LLaVA-OV-1.5-8B
Mean
D/4 (9 blocks)
82.119
86.551
85.898
84.856
2D/4 (18 blocks)
90.860
94.787
92.769
92.806
Observed peak (22 blocks)
89.185
94.014
92.299
91.833
3D/4 (27 blocks)
87.508
93.431
91.162
90.701
Table 1: Selection-depth ablation at 80% pruning under matched decoder visual-token processing budgets. Observed peak denotes block 22, where the mean question gain in Figure 2 is highest. Values are eight-task Dense-relative recovery (%), with the best results in bold.
Figure 3: Overview of our two-stage pruning. Before the LLM, vision-encoder attention guides visual-token pruning while preserving an expanded candidate set. At the decoder midpoint, text-to-visual attention selects the final visual-token set. The reserve fraction α∈[0.1,0.3] controls the number of additional candidates; we use α=0.2 by default.
Method
TextVQA
MMMU
AI2D
MMStar
ChartQA
NoCaps
TextCaps
InfoVQA
Avg. Rel.
Dense
81.30
46.22
81.87
56.40
82.92
66.51
81.86
77.43
100.00
Prune 70% of vision tokens
FastV (ECCV2024)
74.89
43.33
72.22
46.53
45.64
63.16
84.84
44.51
83.46
SparseVLM (ICML2025)
73.21
45.44
73.87
48.80
50.72
65.68
74.94
45.78
84.46
VisPruner (ICCV2025)
64.76
44.33
72.31
47.07
65.00
66.45
68.25
46.48
83.63
ZOO-Prune (CVPR2026)
65.64
45.11
76.75
51.20
72.84
66.12
58.97
53.92
86.48
Table 2: Qwen3-VL-4B results across eight benchmarks. Avg. Rel.: mean performance relative to Dense (%); higher is better. Bold/underline: best/second-best pruned results at each ratio.
Method
TextVQA
MMMU
AI2D
MMStar
ChartQA
NoCaps
TextCaps
InfoVQA
Avg. Rel.
Dense
82.95
51.67
83.68
62.73
83.48
63.54
81.62
81.19
100.00
Prune 70% of vision tokens
FastV (ECCV2024)
74.54
49.44
72.73
50.27
43.60
63.11
84.87
42.53
82.56
SparseVLM (ICML2025)
76.12
51.33
76.68
54.93
47.52
63.54
78.13
48.95
85.41
VisPruner (ICCV2025)
75.85
51.56
78.47
52.87
66.68
63.13
79.61
52.60
88.85
ZOO-Prune (CVPR2026)
69.31
51.00
78.14
56.27
70.80
63.39
63.42
54.70
86.87
Table 3: Qwen3-VL-8B results across eight benchmarks. Avg. Rel.: mean performance relative to Dense (%); higher is better. Bold/underline: best/second-best pruned results at each ratio.
Method
TextVQA
MMMU
AI2D
MMStar
ChartQA
NoCaps
TextCaps
InfoVQA
Avg. Rel.
Dense
79.71
55.78
84.62
67.33
86.64
110.40
123.13
77.41
100.00
Prune 70% of vision tokens
FastV (ECCV2024)
66.67
54.33
75.84
51.80
48.20
103.94
114.19
43.81
80.84
SparseVLM (ICML2025)
74.68
55.00
78.21
57.33
62.76
108.44
113.21
48.83
86.94
VisPruner (ICCV2025)
50.50
52.89
74.29
54.20
53.40
109.30
75.89
37.70
74.68
ZOO-Prune (CVPR2026)
71.09
55.22
81.83
62.87
77.40
110.07
107.16
54.16
90.54
Table 4: LLaVA-OneVision-1.5-8B results across eight benchmarks. Avg. Rel.: mean performance relative to Dense (%); higher is better. Bold/underline: best/second-best pruned results at each ratio.
Figure 4: TextVQA and ChartQA examples on Qwen3-VL-8B at 80% pruning. Purple: final tokens outside the image-only top- K ; gray: discarded regions; orange: merged support.
Figure 5: Candidate preservation and text-guided reselection at 80% pruning. A: immediate image-only top- K selection; B: preserve additional candidates until the midpoint, then retain the same initial top- K as A; C: use the same candidates and token-count trajectory as B, but reselect the final K tokens using text-to-visual attention (Ours). Values are Dense-relative recovery (%); the dashed line denotes Dense (100%).
Figure 6: Pruning-score ablations at 80% pruning. (a) Encoder-based candidate selection before the LLM versus after block 1, with or without text attention. (b) Midpoint scoring with fixed candidates and token budgets. Results average Dense-relative recovery across InfoVQA, ChartQA, TextCaps, and MMStar.
Figure 7: Performance–latency trade-off as the reserve fraction α varies at 80% pruning. Recovery is averaged equally across eight benchmarks relative to Dense. Latency includes token-selection overhead and is averaged over three timed repetitions on one RTX A6000 GPU, using 100 fixed InfoVQA inputs stratified by image size and aspect ratio.
80% pruning
90% pruning
Method
LLM prefill
End-to-end
LLM prefill
End-to-end
Dense
226.66
364.13
226.66
364.13
FastV
116.46
256.15
106.97
247.22
SparseVLM
134.79
276.36
122.16
263.63
VisPruner
131.25
292.39
123.80
284.79
ZOO-Prune
137.28
401.81
111.69
375.04
Table 5: Qwen3-VL-8B inference latency (ms). Measurements average three timed repetitions on one RTX A6000 GPU using 100 fixed InfoVQA inputs stratified by image size and aspect ratio. Prefill includes token-selection overhead. End-to-end spans GPU-ready inputs through the first output token, excluding CPU preprocessing and input transfer.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Reserve-size trade-offs at 80% pruning across the three backbones. Recovery is averaged equally across eight benchmarks relative to Dense. Top: analytical decoder-prefill FLOPs, excluding the vision frontend and token-selection overhead. Bottom: measured GPU-input-to-first-token latency, including the vision frontend and token-selection overhead but excluding CPU preprocessing and input transfer. Labels indicate α∈{0.10,0.15,0.20,0.25,0.30} , the fraction of otherwise pruned tokens preserved as additional candidates until midpoint reselection.
Qwen3-VL-4B
Qwen3-VL-8B
LLaVA-OV-1.5-8B
Task
Ret. (%)
Δ (pp)
Ret. (%)
Δ (pp)
Ret. (%)
Δ (pp)
InfoVQA
42.09
+18.37
42.96
+14.09
44.28
+9.28
ChartQA
39.03
+13.27
42.00
+6.04
41.18
+4.57
TextCaps
40.65
+11.15
41.47
+6.47
44.74
+5.75
MMStar
43.84
+3.78
41.55
+0.11
44.38
+3.27
Appendix
Table 6: Effect of candidate preservation and midpoint reselection at 80% pruning. Ret. is the input-averaged percentage of final tokens selected from outside the initial vision-guided top- K . Δ is the Dense-relative recovery gain (percentage points) from candidate preservation to text-guided reselection under the same token-count trajectory.
Method
TextVQA
MMMU
AI2D
MMStar
ChartQA
NoCaps
TextCaps
InfoVQA
Avg. Rel.
Dense
82.95
51.67
83.68
62.73
83.48
63.54
81.62
81.19
100.00
Matched to Ours at 80% pruning
PyramidDrop
61.77
48.22
70.76
47.33
45.80
62.41
66.92
33.58
75.53
FlowCut
77.23
50.44
75.71
51.07
50.40
62.47
83.83
47.29
85.28
Ours
80.48
51.33
81.96
57.47
75.20
63.49
83.95
64.57
94.79
Matched to Ours at 90% pruning
Appendix
Table 7: Comparison with two additional progressive pruning methods on Qwen3-VL-8B under matched per-input visual token–block budgets. The 80% and 90% settings denote Ours’ final pruning ratios; baseline final token counts may differ under the matched workloads. Avg. Rel. is the eight-task mean performance relative to Dense. Bold/underline indicate the best/second-best pruned results.
Figure 9: TextVQA examples across three backbones at 80% pruning. Columns compare the original image, ZOO-Prune, SparseVLM, and Ours.
Figure 10: ChartQA examples across three backbones at 80% pruning. Columns compare the original image, ZOO-Prune, SparseVLM, and Ours.
Figure 11: AI2D examples across three backbones at 80% pruning. Columns compare the original image, ZOO-Prune, SparseVLM, and Ours.
Figure 12: MMMU examples across three backbones at 80% pruning. Columns compare the original image, ZOO-Prune, SparseVLM, and Ours.
Figure 13: InfoVQA examples across three backbones at 80% pruning. Columns compare the original image, ZOO-Prune, SparseVLM, and Ours.
Figure 14: MMStar examples across three backbones at 80% pruning. Columns compare the original image, ZOO-Prune, SparseVLM, and Ours.
Figure 15: TextCaps examples across three backbones at 80% pruning. Columns compare the original image, ZOO-Prune, SparseVLM, and Ours.
Figure 16: NoCaps examples across three backbones at 80% pruning. Columns compare the original image, ZOO-Prune, SparseVLM, and Ours.
Vision-Language Models (VLMs) excel at visual understanding and reasoning but often incur substantial inference costs due to the large number of visual tokens. Recent visual token pruning methods increasingly follow a two-stage paradigm: they first remove visually redundant tokens after the vision encoder and then discard tokens irrelevant to the textual query within the Large Language Model (LLM). However, since the first stage typically relies solely on vision-encoder saliency, it may prematurely eliminate query-relevant tokens, depriving the subsequent text-guided stage of critical visual evidence. Our empirical analysis shows that incorporating query guidance into first-stage pruning better preserves task-relevant evidence and consistently improves performance over vision-only saliency-based pruning. We further find that high-variance attention heads are more sensitive to the textual query and yield more discriminative text-to-vision attention signals for second-stage pruning. Motivated by these findings, we propose TReVS, a training-free framework that combines textual relevance with vision-encoder saliency for pre-LLM pruning and leverages high-variance attention heads to remove task-irrelevant tokens at shallow-to-intermediate layers of the LLM. On LLaVA-1.5-7B, TReVS retains 92.8% of the unpruned baseline performance while pruning 94.4% of visual tokens, outperforming prior state-of-the-art methods.
Jing Wang, Zhiping Wu, Dongdong Ren +3
School of Intelligence Science and Technology, Nanjing University · School of Electronic Science and Engineering, Nanjing University · Geely Automobile Research Institute (Ningbo) Co., Ltd., 315000 +1
Vision-Language Models (VLMs) have recently demonstrated remarkable capabilities in visual understanding and reasoning, but they also impose significant computational burdens due to long visual sequence inputs. Recent works address this issue by pruning unimportant visual tokens, achieving substantial computational reduction while maintaining model performance. The core of token pruning lies in determining token importance, with current approaches primarily relying on attention scores from vision encoders or Large Language Models (LLMs). In this paper, we analyze the effectiveness of attention mechanisms in both vision encoders and LLMs. We find that vision encoders suffer from attention sink, leading to poor focus on informative foreground regions, while in LLMs, although prior studies have identified attention bias toward token positions, text-to-vision attention demonstrates resistance to this bias and enables effective pruning guidance in middle layers. Based on these observations, we propose LearnPruner, a two-stage token pruning framework that first removes redundant vision tokens via a learnable pruning module after the vision encoder, then retains only task-relevant tokens in the LLM's middle layer. Experimental results show that our LearnPruner can preserve approximately 95% of the original performance while using only 5.5% of vision tokens, and achieve 3.2× inference acceleration, demonstrating a superior accuracy-efficiency trade-off.
Visual token pruning reduces the computational cost of Vision-Language Models (VLMs) by removing redundant visual tokens. The key is to learn a score that measures whether a token is useful. Existing methods typically rely on Gumbel-Softmax to approximate discrete selection during training. Such selectors make the score depend on the behavior of a relaxed pruning operator, not directly on the consequence of information loss. In this paper, we propose DiffPrune, which gives token scores a direct meaning. During training, DiffPrune keeps all tokens and weakens each token's information according to its score. If weakening a token hurts the task, the scorer is pushed to protect it; if not, the token can receive a lower score. Because the loss is differentiated through this actual information-throttling path, the scorer avoids the unstable surrogate path of relaxed token selection. DiffPrune implements this idea with an Information Throttler, which injects variance-preserving noise into visual tokens, where high-score tokens remain close to their original representations, while low-score tokens carry less original information. At inference, the throttler is removed, and hard top-K pruning is applied using the learned scores. Across ten VLM benchmarks, DiffPrune retains 96.5% of full-model accuracy while accelerating LLM prefill by 2.85x, with only 0.69 ms inference overhead. Code will be publicly available.
Landi He, Mingde Yao, Shawn Young +1
Shenzhen University of Advanced Technology · CUHK MMLab, CPII under InnoHK