Adaptive Visual Token Reduction for Accelerated Image Understanding
Authors: Seyoung Jeong, Jong Pil Yun, Sang Jun Lee
Organizations: Jeonbuk National University, Jeonju, Republic of Korea · Korea Institute of Industrial Technology (KITECH), Incheon, Republic of Korea · Chung-Ang University, Seoul, Republic of Korea
Large Vision-Language Models achieve strong VQA performance, but processing high-resolution, information-rich images requires substantial computation, motivating visual token reduction. However, existing methods often prune individual tokens or rely on fixed-size cropping, limiting their ability to preserve spatially structured information such as horizontally or vertically elongated text. To address this limitation, we propose ReFIT, an instruction-guided visual token reduction framework for efficient LVLM inference. ReFIT consists of Relevance-Guided Window Reshaping (RWR) and Instruction-Guided Token Refinement (ITR), where RWR captures instruction-relevant regions by adapting to their spatial characteristics, while ITR further removes unnecessary visual tokens. Experiments on four VQA benchmarks demonstrate that ReFIT improves answer accuracy while reducing computational cost, and qualitative results demonstrate its effectiveness in localizing relevant regions and removing unnecessary visual information.
Figures & tables
Figure 1: Overall architecture of ReFIT. Given an input image and instruction, RWR adaptively reshapes the window configuration to localize instruction-relevant regions, which are then cropped and re-encoded for fine-grained visual representation. ITR further removes unnecessary visual tokens within the cropped regions using an instruction-guided relevance map, and the remaining tokens are forwarded to the LLM for answer generation.
InfoVQA
SPDocVQA
MPDocVQA
GQA
LLM
Method
ANLS↑
FLOPs (T)↓
ANLS↑
FLOPs (T) ↓
ANLS↑
FLOPs (T)↓
Acc↑
FLOPs (T)↓
Vanilla [ 10 ]
0.2552
38.98
0.6628
51.68
0.3758
50.01
0.7598
30.39
Random sampling
0.2387
24.62
0.4888
27.70
0.2988
26.66
0.7484
19.26
ToMe [ 3 ]
0.1975
26.67
0.3215
39.41
0.2135
38.32
0.7293
19.10
FastV [ 4 ]
0.2306
26.22
0.6099
28.10
0.3523
29.31
0.7478
19.23
Pdrop [ 18 ]
0.2335
26.00
0.5507
26.93
0.3637
30.82
0.7436
20.33
Table 1: Quantitative comparison with state-of-the-art methods on VQA benchmarks: InfoVQA, SPDocVQA, MPDocVQA, and GQA
Figure 2: Qualitative comparison of instruction-relevant region cropping and answer generation between existing and proposed methods. Red boxes indicate the ground-truth regions, while green boxes indicate the regions cropped by each model. The proposed method better localizes horizontally extended text and further removes unnecessary visual tokens within the cropped regions.
RWR
ITR
InfoVQA
SPDocVQA
MPDocVQA
GQA
ANLS ↑
ANLS ↑
ANLS ↑
Acc ↑
0.3024
0.6472
0.3866
0.7608
✓
0.3058
0.6529
0.3948
0.7642
✓
0.3077
0.6552
0.3907
0.7612
✓
✓
0.3189
0.6638
0.4125
0.7669
Table 2: Ablation study on the effectiveness of RWR and ITR.
Region selection
InfoVQA
SPDocVQA
MPDocVQA
GQA
ANLS ↑
Acc ↑
ANLS ↑
Acc ↑
ANLS ↑
Acc ↑
Acc ↑
Single best window
0.3077
0.2535
0.6514
0.5431
0.3994
0.3019
0.7532
Random multi-window
0.3179
0.2667
0.6643
0.5524
0.4077
0.3071
0.7526
Fixed Top- K IoU
0.3186
0.2596
0.6552
0.5487
0.4060
0.3044
0.7569
RWR (Ours)
0.3189
0.2645
0.6638
0.5562
0.4125
0.3112
0.7669
Table 3: Ablation study on region selection strategies.
Token selection
InfoVQA
SPDocVQA
MPDocVQA
GQA
ANLS ↑
Acc ↑
ANLS ↑
Acc ↑
ANLS ↑
Acc ↑
Acc ↑
Random
0.3180
0.2638
0.6634
0.5554
0.4119
0.3104
0.7567
Spatially Connected
0.3185
0.2642
0.6638
0.5562
0.4120
0.3106
0.7568
ITR (Ours)
0.3189
0.2645
0.6638
0.5562
0.4125
0.3112
0.7669
Table 4: Ablation study on token selection methods.
Department of Electrical and Computer Engineering, Sungkyunkwan University, Republic of Korea · Department of Semiconductor Systems Engineering, Sungkyunkwan University, Republic of Korea