Visual token pruning has been widely studied as a practical approach to reducing the computational cost of large vision-language models. However, it struggles to preserve essential visual information, which can lead to substantial performance degradation. In particular, image-based token selection can overlook task-relevant details, while text-guided token selection may fail to capture the text--visual relationships needed for complex reasoning. We find that applying textual guidance too early can limit its ability to identify answer-relevant visual regions, whereas text-to-visual attention becomes more informative at intermediate decoder depths. This finding motivates our training-free method, which separates early vision-guided pruning from deferred text-guided reselection. We first prune visual tokens using vision-encoder attention, retain additional candidates until the decoder midpoint, and then use text-to-visual attention to determine the final visual-token set. Across eight benchmarks and three models, our method outperforms the best-performing baselines by an average of 11.10 and 16.84 percentage points in performance recovery at 80% and 90% pruning, respectively, with comparable or lower LLM-prefill latency than most baselines. The source code is publicly available at https://github.com/kmc3661/DeFT
Figures & tables
Figure 1: Failure cases on Qwen3-VL-8B at 80% pruning. Image-based visual token pruning discards the visual region containing the requested text, while text-guided visual token pruning preserves related support but loses the corresponding value. Gray: discarded regions; orange: merged support. Percentages denote full-benchmark performance losses relative to Dense.
Figure 2: Text-guided evidence identification across decoder depth. Top: correlation with masking-based region importance. Bottom: matching-versus-swapped question gain. Results average four QA benchmarks equally, each with 300 images and two questions per image.
Boundary
Qwen3-VL-4B
Qwen3-VL-8B
LLaVA-OV-1.5-8B
Mean
D/4 (9 blocks)
82.119
86.551
85.898
84.856
2D/4 (18 blocks)
90.860
94.787
92.769
92.806
Observed peak (22 blocks)
89.185
94.014
92.299
91.833
3D/4 (27 blocks)
87.508
93.431
91.162
90.701
Table 1: Selection-depth ablation at 80% pruning under matched decoder visual-token processing budgets. Observed peak denotes block 22, where the mean question gain in Figure 2 is highest. Values are eight-task Dense-relative recovery (%), with the best results in bold.
Figure 3: Overview of our two-stage pruning. Before the LLM, vision-encoder attention guides visual-token pruning while preserving an expanded candidate set. At the decoder midpoint, text-to-visual attention selects the final visual-token set. The reserve fraction α∈[0.1,0.3] controls the number of additional candidates; we use α=0.2 by default.
Method
TextVQA
MMMU
AI2D
MMStar
ChartQA
NoCaps
TextCaps
InfoVQA
Avg. Rel.
Dense
81.30
46.22
81.87
56.40
82.92
66.51
81.86
77.43
100.00
Prune 70% of vision tokens
FastV (ECCV2024)
74.89
43.33
72.22
46.53
45.64
63.16
84.84
44.51
83.46
SparseVLM (ICML2025)
73.21
45.44
73.87
48.80
50.72
65.68
74.94
45.78
84.46
VisPruner (ICCV2025)
64.76
44.33
72.31
47.07
65.00
66.45
68.25
46.48
83.63
ZOO-Prune (CVPR2026)
65.64
45.11
76.75
51.20
72.84
66.12
58.97
53.92
86.48
Table 2: Qwen3-VL-4B results across eight benchmarks. Avg. Rel.: mean performance relative to Dense (%); higher is better. Bold/underline: best/second-best pruned results at each ratio.
Method
TextVQA
MMMU
AI2D
MMStar
ChartQA
NoCaps
TextCaps
InfoVQA
Avg. Rel.
Dense
82.95
51.67
83.68
62.73
83.48
63.54
81.62
81.19
100.00
Prune 70% of vision tokens
FastV (ECCV2024)
74.54
49.44
72.73
50.27
43.60
63.11
84.87
42.53
82.56
SparseVLM (ICML2025)
76.12
51.33
76.68
54.93
47.52
63.54
78.13
48.95
85.41
VisPruner (ICCV2025)
75.85
51.56
78.47
52.87
66.68
63.13
79.61
52.60
88.85
ZOO-Prune (CVPR2026)
69.31
51.00
78.14
56.27
70.80
63.39
63.42
54.70
86.87
Table 3: Qwen3-VL-8B results across eight benchmarks. Avg. Rel.: mean performance relative to Dense (%); higher is better. Bold/underline: best/second-best pruned results at each ratio.
Method
TextVQA
MMMU
AI2D
MMStar
ChartQA
NoCaps
TextCaps
InfoVQA
Avg. Rel.
Dense
79.71
55.78
84.62
67.33
86.64
110.40
123.13
77.41
100.00
Prune 70% of vision tokens
FastV (ECCV2024)
66.67
54.33
75.84
51.80
48.20
103.94
114.19
43.81
80.84
SparseVLM (ICML2025)
74.68
55.00
78.21
57.33
62.76
108.44
113.21
48.83
86.94
VisPruner (ICCV2025)
50.50
52.89
74.29
54.20
53.40
109.30
75.89
37.70
74.68
ZOO-Prune (CVPR2026)
71.09
55.22
81.83
62.87
77.40
110.07
107.16
54.16
90.54
Table 4: LLaVA-OneVision-1.5-8B results across eight benchmarks. Avg. Rel.: mean performance relative to Dense (%); higher is better. Bold/underline: best/second-best pruned results at each ratio.
Figure 4: TextVQA and ChartQA examples on Qwen3-VL-8B at 80% pruning. Purple: final tokens outside the image-only top- K ; gray: discarded regions; orange: merged support.
Figure 5: Candidate preservation and text-guided reselection at 80% pruning. A: immediate image-only top- K selection; B: preserve additional candidates until the midpoint, then retain the same initial top- K as A; C: use the same candidates and token-count trajectory as B, but reselect the final K tokens using text-to-visual attention (Ours). Values are Dense-relative recovery (%); the dashed line denotes Dense (100%).
Figure 6: Pruning-score ablations at 80% pruning. (a) Encoder-based candidate selection before the LLM versus after block 1, with or without text attention. (b) Midpoint scoring with fixed candidates and token budgets. Results average Dense-relative recovery across InfoVQA, ChartQA, TextCaps, and MMStar.
Figure 7: Performance–latency trade-off as the reserve fraction α varies at 80% pruning. Recovery is averaged equally across eight benchmarks relative to Dense. Latency includes token-selection overhead and is averaged over three timed repetitions on one RTX A6000 GPU, using 100 fixed InfoVQA inputs stratified by image size and aspect ratio.
80% pruning
90% pruning
Method
LLM prefill
End-to-end
LLM prefill
End-to-end
Dense
226.66
364.13
226.66
364.13
FastV
116.46
256.15
106.97
247.22
SparseVLM
134.79
276.36
122.16
263.63
VisPruner
131.25
292.39
123.80
284.79
ZOO-Prune
137.28
401.81
111.69
375.04
Table 5: Qwen3-VL-8B inference latency (ms). Measurements average three timed repetitions on one RTX A6000 GPU using 100 fixed InfoVQA inputs stratified by image size and aspect ratio. Prefill includes token-selection overhead. End-to-end spans GPU-ready inputs through the first output token, excluding CPU preprocessing and input transfer.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Reserve-size trade-offs at 80% pruning across the three backbones. Recovery is averaged equally across eight benchmarks relative to Dense. Top: analytical decoder-prefill FLOPs, excluding the vision frontend and token-selection overhead. Bottom: measured GPU-input-to-first-token latency, including the vision frontend and token-selection overhead but excluding CPU preprocessing and input transfer. Labels indicate α∈{0.10,0.15,0.20,0.25,0.30} , the fraction of otherwise pruned tokens preserved as additional candidates until midpoint reselection.
Qwen3-VL-4B
Qwen3-VL-8B
LLaVA-OV-1.5-8B
Task
Ret. (%)
Δ (pp)
Ret. (%)
Δ (pp)
Ret. (%)
Δ (pp)
InfoVQA
42.09
+18.37
42.96
+14.09
44.28
+9.28
ChartQA
39.03
+13.27
42.00
+6.04
41.18
+4.57
TextCaps
40.65
+11.15
41.47
+6.47
44.74
+5.75
MMStar
43.84
+3.78
41.55
+0.11
44.38
+3.27
Appendix
Table 6: Effect of candidate preservation and midpoint reselection at 80% pruning. Ret. is the input-averaged percentage of final tokens selected from outside the initial vision-guided top- K . Δ is the Dense-relative recovery gain (percentage points) from candidate preservation to text-guided reselection under the same token-count trajectory.
Method
TextVQA
MMMU
AI2D
MMStar
ChartQA
NoCaps
TextCaps
InfoVQA
Avg. Rel.
Dense
82.95
51.67
83.68
62.73
83.48
63.54
81.62
81.19
100.00
Matched to Ours at 80% pruning
PyramidDrop
61.77
48.22
70.76
47.33
45.80
62.41
66.92
33.58
75.53
FlowCut
77.23
50.44
75.71
51.07
50.40
62.47
83.83
47.29
85.28
Ours
80.48
51.33
81.96
57.47
75.20
63.49
83.95
64.57
94.79
Matched to Ours at 90% pruning
Appendix
Table 7: Comparison with two additional progressive pruning methods on Qwen3-VL-8B under matched per-input visual token–block budgets. The 80% and 90% settings denote Ours’ final pruning ratios; baseline final token counts may differ under the matched workloads. Avg. Rel. is the eight-task mean performance relative to Dense. Bold/underline indicate the best/second-best pruned results.
Figure 9: TextVQA examples across three backbones at 80% pruning. Columns compare the original image, ZOO-Prune, SparseVLM, and Ours.
Figure 10: ChartQA examples across three backbones at 80% pruning. Columns compare the original image, ZOO-Prune, SparseVLM, and Ours.
Figure 11: AI2D examples across three backbones at 80% pruning. Columns compare the original image, ZOO-Prune, SparseVLM, and Ours.
Figure 12: MMMU examples across three backbones at 80% pruning. Columns compare the original image, ZOO-Prune, SparseVLM, and Ours.
Figure 13: InfoVQA examples across three backbones at 80% pruning. Columns compare the original image, ZOO-Prune, SparseVLM, and Ours.
Figure 14: MMStar examples across three backbones at 80% pruning. Columns compare the original image, ZOO-Prune, SparseVLM, and Ours.
Figure 15: TextCaps examples across three backbones at 80% pruning. Columns compare the original image, ZOO-Prune, SparseVLM, and Ours.
Figure 16: NoCaps examples across three backbones at 80% pruning. Columns compare the original image, ZOO-Prune, SparseVLM, and Ours.
School of Intelligence Science and Technology, Nanjing University · School of Electronic Science and Engineering, Nanjing University · Geely Automobile Research Institute (Ningbo) Co., Ltd., 315000 +1