The inference efficiency of Multimodal Large Language Models (MLLMs) is severely constrained by massive visual token sequences induced by high-resolution inputs, with computational cost scaling quadratically. Existing approaches primarily focus on downstream token compression, while overlooking a fundamental upstream inefficiency: input resolution is treated as a static, task-agnostic hyperparameter. We propose Task-Conditioned Resolution Routing (TCRR), which formulates visual compression as a task-conditioned decision and employs a lightweight cross-modal router that conditions backbone visual representations on textual semantics via feature-wise modulation and cross-attention to predict the minimal sufficient compression level per query. To support this, we curate a dataset of 500k samples across 12 task categories, labeled via a teacher-oracle pipeline to approximate Pareto-optimal compression scales. Extensive experiments across diverse architectures show that TCRR achieves a superior efficiency frontier, specifically reducing visual FLOPs by 40.9% and latency by 53.7% on Qwen3-VL-8B while preserving competitive performance. Further analysis of scaling behavior confirms that dynamically routing visual compression enables optimal resource allocation without modifying the MLLM backbone.
Figures & tables
Figure 1: Task-Conditioned Resolution Routing. By adapting resolution to textual intent, TCRR slashes computational overhead without sacrificing accuracy.
Task Category
Proportion
Resolution Demand & Key Factor
General Vision-Language
VQA
22%
Mixed (Context-dependent complexity)
Captioning
10%
Low (Global semantic gist)
Grounding
5%
Medium (Bounding box localization)
Counting
5%
High (Object separation)
Document & Text
Table 1: Task distribution and resolution in Res-500k . The dataset is explicitly balanced to cover diverse resolution priors, ranging from semantic gist (Low) to pixel-level precision (Very High).
Figure 2: The architecture of TCRR. The lightweight router ( bottom ) fuses visual features and the text prompt via cross-attention to predict an optimal continuous scale factor, dynamically resizing the raw image before it enters the unmodified MLLM backbone ( top ).
Model
MMB
MMS
MMMU
MathV
Hallu.
OCR
MMVet
TVQA
ChQA
DVQA
IVQA
MRWC
Avg.
FLOPs (T)
Lat. (ms)
2B
76.8
54.7
46.3
51.1
47.1
84.0
40.7
79.6
78.0
92.5
72.2
55.2
64.9
8.5
524.1
+TCRR
76.8
54.7
45.7
51.4
46.4
84.0
40.5
79.7
77.5
92.2
69.4
53.3
64.3
4.8 (-43%)
236.9 (-55%)
4B
82.9
62.3
55.0
64.2
55.1
84.8
49.7
81.6
81.8
94.6
79.4
62.7
71.2
14.7
672.0
+TCRR
82.6
62.7
54.9
64.5
54.2
84.9
49.9
81.6
81.3
94.8
77.9
60.9
70.8
8.5 (-42%)
319.8 (-52%)
8B
85.0
65.3
57.4
65.6
58.0
87.0
52.9
83.3
82.8
95.6
83.1
64.5
73.4
23.7
1037.7
+TCRR
84.9
65.4
54.0
65.7
57.5
87.1
53.2
83.4
82.7
95.4
81.3
62.6
72.8
14.0 (-41%)
480.8 (-54%)
Table 2: Main Results on Qwen3-VL Family. TCRR achieves substantial reductions in FLOPs and latency while preserving average accuracy across 12 multimodal benchmarks. Abbreviations: MMB (MMBench), MMS (MMStar), MathV (MathVista), Hallu. (HallusionBench), TVQA (TextVQA), ChQA (ChartQA), DVQA (DocVQA), IVQA (InfoVQA), MRWC (MME-RealWorld-CN).
Table 5Table 6Table 7Table 8
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Latency (ms)
Share (%)
Visual Encoder (MobileNetV4)
12.4
68.5%
Text Encoder (BERT-Tiny)
2.1
11.6%
Interaction Head (FiLM+Attn)
0.8
4.4%
Data Move / Overhead
2.8
15.5%
Total End-to-End
18.1 ms
100%
Appendix
Table 10: TCRR Latency Profiling (NVIDIA H20) .
Parameter
Value
Architecture
Visual Backbone
MobileNetV4-Medium (e250_r384)
Text Backbone
BERT-Tiny (prajjwal1/bert-tiny)
Input Resolution
384×384
Max Text Length
192
MLP Hidden Dim
512
Appendix
Table 11: TCRR Training Hyperparameters.
Architecture
Acc.
MAE
FLOPs
Lat. (ms)
R. Lat. (ms)
Token Red.
Multi-stage Fusion
72.7
1.16
14.0
499.4
22.8
45.6%
TCRR (Single-stage)
72.8
1.14
14.0
480.8
18.1
45.8%
Appendix
Table 12: Router Architecture Comparison. Comparison of fusion paradigms. The proposed shallow, single-stage fusion achieves the optimal balance of accuracy and computational overhead compared to a heavier multi-stage alternative. “R. Lat.” denotes the latency specifically incurred by the router module.
Benchmark
R1
R2
R3
R4
R5
R6
R7
R8
ChartQA_TEST
8
144
48
2130
147
23
0
0
DocVQA_VAL
0
0
12
45
196
4792
257
47
HallusionBench
14
93
153
262
230
182
11
6
InfoVQA_VAL
0
0
12
223
368
1478
630
90
MMBench_DEV_EN_V11
919
974
2787
12
162
22
0
0
MME-RealWorld-CN
0
0
0
20
141
3577
1930
249
Appendix
Table 13: Resolution Distribution Across Benchmarks. We classify sample quantities into eight intervals: R1:<2562 , R2:2562–3842 , R3:3842–5122 , R4:5122–7682 , R5:7682–10242 , R6:10242–20482 , R7:20482–40962 , and R8:>40962 (pixels 2 ). Using indexed headers allows for a cleaner layout and improved readability.
Dataset
0.1
0.2
0.3
0.4
0.5
0.6
0.7
0.8
0.9
1.0
ChartQA_TEST
314
734
626
239
114
57
61
25
24
134
DocVQA_VAL
4356
599
169
84
37
20
19
11
9
33
HallusionBench
632
53
33
11
5
5
1
2
1
5
InfoVQA_VAL
1325
505
337
202
116
55
56
36
27
123
MathVista_MINI
303
79
41
26
8
12
7
5
4
11
MMBench_DEV_EN_V11
930
75
37
15
8
11
3
5
10
14
Appendix
Table 14: Distribution of selected compression ratios across benchmarks.
Task Type
Prompt
TCRR Routing
Visual Input
Analysis
Global Semantics
“Is this picture taken during the day or at night?”
Area Scale: 10% (Eq. Res: ∼1214×798 )
While pruning 90% of visual tokens, the model retains sufficient global illumination features to correctly output “Daytime” . This is the optimal computational sweet spot for high-confidence predictions.
Fine Details
“What is the text inside the red rectangular area on the building in the bottom center?”
Area Scale: 100% (Orig. Res: 3840×2523 )
The model dynamically determines that pixel-level precision is required. It bypasses down-sampling to accurately read: “MUZIUM TEKSTIL NEGARA / National Textile Museum” .
Appendix
Table 14
Task Type
Prompt
TCRR Routing
Visual Input
Analysis
Global Semantics
“What is the functional scene of this room?”
Area Scale: 10% (Eq. Res: ∼648×648 )
After aggressively removing 90% of local redundancy, the model relies on core structural features to correctly output: “Computer lab or library reading area” .
Fine Details
“What brands are the computers on the desk?”
Area Scale: 90% (Eq. Res: ∼1943×1943 )
TCRR intelligently assigns a high resolution to extract micro-features (the tiny logo on the back of the distant Dell monitor would blur entirely at low resolutions). Output: “Apple and Dell” .
Appendix
Table 15
Task Type
Prompt
TCRR Routing
Visual Input
Analysis
Information Extraction
“Which airport was the most used?”
Area Scale: 80% (Eq. Res: ∼716×748 )
Recognizing dense axis texts and fine-grained data bars, TCRR adopts an extremely conservative down-sampling strategy. It marginally reduces resolution to save 20% compute while preserving impeccable legibility, correctly outputting: “Paris-Charles-de-Gaulle” .
Appendix
Table 16
Task Type
Prompt
TCRR Routing
Visual Input
Analysis
Document Parsing
“What is the % of raw material imported in the current year?”
Area Scale: 90% (Eq. Res: ∼1593×2220 )
To extract a specific percentage from a financially dense table, TCRR assigns near-maximum resolution (90%). This strictly guarantees that microscopic digits do not suffer from aliasing before entering the ViT, enabling the model to precisely extract: “79.23%” .
Appendix
Table 17
Task Type
Prompt
TCRR Routing
Visual Input
Analysis
Mood Perception
“What is the overall mood or atmosphere conveyed by this landscape?”
Area Scale: 20% (Eq. Res: ∼259×170 )
Pixel-level textures (individual leaves, tiny ripples) are redundant for mood perception. TCRR aggressively scales the image down to 20% to capture global color blocks and contours, accurately predicting: “Serene and natural” at ultra-low latency.
Appendix
Table 18
Task Type
Prompt
TCRR Routing
Visual Input
Analysis
Basic Recognition
“What is the main animal lying right in the center?”
Area Scale: 10% (Eq. Res: ∼215×144 )
Iconic global features of a Shiba Inu (reddish-gold coat, triangular ears, silhouette) remain distinctly recognizable even when uniformly compressed to 10% of the original area. TCRR assigns the lowest resolution routing, maximizing efficiency while fully preserving semantic accuracy: “Shiba Inu” .
Long-context reasoning in large language models incurs substantial computation and memory costs. Visual text compression (VTC) reduces input length by rendering text as images, but fixed-resolution rendering creates a compression-performance trade-off: low DPI saves tokens at the expense of legibility, whereas high DPI spends tokens on irrelevant content. We introduce FocusVTC, which breaks this trade-off through adaptive resolution while preserving general multimodal capabilities. It combines compressed low-DPI global views with selective region enhancement, integrating enhanced views into ongoing reasoning. We construct 29.4K high-quality Reasoning-Evidence Localization (REL) chain-of-thought examples (REL-CoT) that link reasoning traces to page indices and bounding boxes. Multi-resolution REL supervised fine-tuning (REL-SFT) teaches the model to localize relevant regions, and Group Relative Policy Optimization learns when to enhance resolution and how to use the resulting observations, without a separate continual-pretraining stage. At 72 DPI on RULER v1, FocusVTC scores 87.4 at 2.9× input compression, including tool observations, versus 57.5 for Glyph at 3.0× input compression. It surpasses its text-input backbone on LongBench (56.40 versus 55.86), improves the MRCR macro-average by 13.91 points, and achieves a 51.19 macro-average on VTCBench. The MRCR latency evaluation also shows a 2.79× online end-to-end speedup over Text. General multimodal capabilities are preserved, with MMMU increasing from 65.12 to 66.73 and MME from 2424.02 to 2457.62.
FangZhi Zhong, Xuerui Qiu, Yuqi Pan +4
Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences · Zhongguancun Academy +1
Recent Multimodal Large Language Models (MLLMs) have demonstrated strong performance on vision-language understanding tasks, yet their inference efficiency is often hampered by the large number of visual tokens, particularly in high-resolution or multi-image scenarios. To address this issue, we propose EvoComp, a visual token compression framework that significantly reduces token count while preserving task accuracy. EvoComp introduces a lightweight encoder-only transformer-based compressor that selects the most informative and non-redundant visual tokens by jointly considering visual and textual contexts. A core challenge lies in providing effective supervision for training the compressor. To this end, we design an evolutionary labeling strategy that searches for token subsets minimizing the MLLM's output loss, while enforcing semantic diversity through vocabulary-based token grouping. We further train the compressor using a tailored loss function combining the GHM loss to mitigate class and difficulty imbalance, and a cosine similarity regularization to encourage semantic separation between retained and discarded tokens. Extensive experiments across multiple vision-language benchmarks show that EvoComp outperforms existing methods based on attention or similarity heuristics. Notably, it retains 99.3% of the original accuracy under 3x token compression and delivers up to 1.6x speedup on mobile devices.
Visual encoding constitutes a major computational bottleneck in Multimodal Large Language Models (MLLMs), especially for high-resolution image inputs. The prevailing practice typically adopts global encoding followed by post-ViT compression. Global encoding produces massive token sequences, while post-ViT compression incurs the full quadratic attention cost of the ViT before any token reduction takes place. In this work, we revisit this convention along two dimensions: the encoding strategy and visual token compression. First, controlled experiments show that slice-based encoding outperforms global encoding across benchmarks, suggesting that preserving local details through sliced views can be more beneficial than applying global attention for fine-grained perception. Second, we introduce intra-ViT early compression, which reduces tokens in shallow ViT layers and substantially lowers visual-encoding FLOPs while preserving downstream performance. By integrating intra-ViT compression into the slice-based encoding framework, we present LLaVA-UHD v4, an efficient and compute-controllable visual encoding scheme tailored for high-resolution inputs. Across a diverse set of benchmarks covering document understanding, OCR, and general VQA, LLaVA-UHD v4 reduces visual-encoding FLOPs by 55.8% while matching or even surpassing baseline performance. These results suggest that visual-encoding efficiency can be substantially improved without sacrificing downstream performance, providing a practical design direction for efficient high-resolution MLLMs. All model weights and code will be publicly released to support further research.