The inference efficiency of Multimodal Large Language Models (MLLMs) is severely constrained by massive visual token sequences induced by high-resolution inputs, with computational cost scaling quadratically. Existing approaches primarily focus on downstream token compression, while overlooking a fundamental upstream inefficiency: input resolution is treated as a static, task-agnostic hyperparameter. We propose Task-Conditioned Resolution Routing (TCRR), which formulates visual compression as a task-conditioned decision and employs a lightweight cross-modal router that conditions backbone visual representations on textual semantics via feature-wise modulation and cross-attention to predict the minimal sufficient compression level per query. To support this, we curate a dataset of 500k samples across 12 task categories, labeled via a teacher-oracle pipeline to approximate Pareto-optimal compression scales. Extensive experiments across diverse architectures show that TCRR achieves a superior efficiency frontier, specifically reducing visual FLOPs by 40.9% and latency by 53.7% on Qwen3-VL-8B while preserving competitive performance. Further analysis of scaling behavior confirms that dynamically routing visual compression enables optimal resource allocation without modifying the MLLM backbone.
Figures & tables
Figure 1: Task-Conditioned Resolution Routing. By adapting resolution to textual intent, TCRR slashes computational overhead without sacrificing accuracy.
Task Category
Proportion
Resolution Demand & Key Factor
General Vision-Language
VQA
22%
Mixed (Context-dependent complexity)
Captioning
10%
Low (Global semantic gist)
Grounding
5%
Medium (Bounding box localization)
Counting
5%
High (Object separation)
Document & Text
Table 1: Task distribution and resolution in Res-500k . The dataset is explicitly balanced to cover diverse resolution priors, ranging from semantic gist (Low) to pixel-level precision (Very High).
Figure 2: The architecture of TCRR. The lightweight router ( bottom ) fuses visual features and the text prompt via cross-attention to predict an optimal continuous scale factor, dynamically resizing the raw image before it enters the unmodified MLLM backbone ( top ).
Model
MMB
MMS
MMMU
MathV
Hallu.
OCR
MMVet
TVQA
ChQA
DVQA
IVQA
MRWC
Avg.
FLOPs (T)
Lat. (ms)
2B
76.8
54.7
46.3
51.1
47.1
84.0
40.7
79.6
78.0
92.5
72.2
55.2
64.9
8.5
524.1
+TCRR
76.8
54.7
45.7
51.4
46.4
84.0
40.5
79.7
77.5
92.2
69.4
53.3
64.3
4.8 (-43%)
236.9 (-55%)
4B
82.9
62.3
55.0
64.2
55.1
84.8
49.7
81.6
81.8
94.6
79.4
62.7
71.2
14.7
672.0
+TCRR
82.6
62.7
54.9
64.5
54.2
84.9
49.9
81.6
81.3
94.8
77.9
60.9
70.8
8.5 (-42%)
319.8 (-52%)
8B
85.0
65.3
57.4
65.6
58.0
87.0
52.9
83.3
82.8
95.6
83.1
64.5
73.4
23.7
1037.7
+TCRR
84.9
65.4
54.0
65.7
57.5
87.1
53.2
83.4
82.7
95.4
81.3
62.6
72.8
14.0 (-41%)
480.8 (-54%)
Table 2: Main Results on Qwen3-VL Family. TCRR achieves substantial reductions in FLOPs and latency while preserving average accuracy across 12 multimodal benchmarks. Abbreviations: MMB (MMBench), MMS (MMStar), MathV (MathVista), Hallu. (HallusionBench), TVQA (TextVQA), ChQA (ChartQA), DVQA (DocVQA), IVQA (InfoVQA), MRWC (MME-RealWorld-CN).
Table 5Table 6Table 7Table 8
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Latency (ms)
Share (%)
Visual Encoder (MobileNetV4)
12.4
68.5%
Text Encoder (BERT-Tiny)
2.1
11.6%
Interaction Head (FiLM+Attn)
0.8
4.4%
Data Move / Overhead
2.8
15.5%
Total End-to-End
18.1 ms
100%
Appendix
Table 10: TCRR Latency Profiling (NVIDIA H20) .
Parameter
Value
Architecture
Visual Backbone
MobileNetV4-Medium (e250_r384)
Text Backbone
BERT-Tiny (prajjwal1/bert-tiny)
Input Resolution
384×384
Max Text Length
192
MLP Hidden Dim
512
Appendix
Table 11: TCRR Training Hyperparameters.
Architecture
Acc.
MAE
FLOPs
Lat. (ms)
R. Lat. (ms)
Token Red.
Multi-stage Fusion
72.7
1.16
14.0
499.4
22.8
45.6%
TCRR (Single-stage)
72.8
1.14
14.0
480.8
18.1
45.8%
Appendix
Table 12: Router Architecture Comparison. Comparison of fusion paradigms. The proposed shallow, single-stage fusion achieves the optimal balance of accuracy and computational overhead compared to a heavier multi-stage alternative. “R. Lat.” denotes the latency specifically incurred by the router module.
Benchmark
R1
R2
R3
R4
R5
R6
R7
R8
ChartQA_TEST
8
144
48
2130
147
23
0
0
DocVQA_VAL
0
0
12
45
196
4792
257
47
HallusionBench
14
93
153
262
230
182
11
6
InfoVQA_VAL
0
0
12
223
368
1478
630
90
MMBench_DEV_EN_V11
919
974
2787
12
162
22
0
0
MME-RealWorld-CN
0
0
0
20
141
3577
1930
249
Appendix
Table 13: Resolution Distribution Across Benchmarks. We classify sample quantities into eight intervals: R1:<2562 , R2:2562–3842 , R3:3842–5122 , R4:5122–7682 , R5:7682–10242 , R6:10242–20482 , R7:20482–40962 , and R8:>40962 (pixels 2 ). Using indexed headers allows for a cleaner layout and improved readability.
Dataset
0.1
0.2
0.3
0.4
0.5
0.6
0.7
0.8
0.9
1.0
ChartQA_TEST
314
734
626
239
114
57
61
25
24
134
DocVQA_VAL
4356
599
169
84
37
20
19
11
9
33
HallusionBench
632
53
33
11
5
5
1
2
1
5
InfoVQA_VAL
1325
505
337
202
116
55
56
36
27
123
MathVista_MINI
303
79
41
26
8
12
7
5
4
11
MMBench_DEV_EN_V11
930
75
37
15
8
11
3
5
10
14
Appendix
Table 14: Distribution of selected compression ratios across benchmarks.
Task Type
Prompt
TCRR Routing
Visual Input
Analysis
Global Semantics
“Is this picture taken during the day or at night?”
Area Scale: 10% (Eq. Res: ∼1214×798 )
While pruning 90% of visual tokens, the model retains sufficient global illumination features to correctly output “Daytime” . This is the optimal computational sweet spot for high-confidence predictions.
Fine Details
“What is the text inside the red rectangular area on the building in the bottom center?”
Area Scale: 100% (Orig. Res: 3840×2523 )
The model dynamically determines that pixel-level precision is required. It bypasses down-sampling to accurately read: “MUZIUM TEKSTIL NEGARA / National Textile Museum” .
Appendix
Table 14
Task Type
Prompt
TCRR Routing
Visual Input
Analysis
Global Semantics
“What is the functional scene of this room?”
Area Scale: 10% (Eq. Res: ∼648×648 )
After aggressively removing 90% of local redundancy, the model relies on core structural features to correctly output: “Computer lab or library reading area” .
Fine Details
“What brands are the computers on the desk?”
Area Scale: 90% (Eq. Res: ∼1943×1943 )
TCRR intelligently assigns a high resolution to extract micro-features (the tiny logo on the back of the distant Dell monitor would blur entirely at low resolutions). Output: “Apple and Dell” .
Appendix
Table 15
Task Type
Prompt
TCRR Routing
Visual Input
Analysis
Information Extraction
“Which airport was the most used?”
Area Scale: 80% (Eq. Res: ∼716×748 )
Recognizing dense axis texts and fine-grained data bars, TCRR adopts an extremely conservative down-sampling strategy. It marginally reduces resolution to save 20% compute while preserving impeccable legibility, correctly outputting: “Paris-Charles-de-Gaulle” .
Appendix
Table 16
Task Type
Prompt
TCRR Routing
Visual Input
Analysis
Document Parsing
“What is the % of raw material imported in the current year?”
Area Scale: 90% (Eq. Res: ∼1593×2220 )
To extract a specific percentage from a financially dense table, TCRR assigns near-maximum resolution (90%). This strictly guarantees that microscopic digits do not suffer from aliasing before entering the ViT, enabling the model to precisely extract: “79.23%” .
Appendix
Table 17
Task Type
Prompt
TCRR Routing
Visual Input
Analysis
Mood Perception
“What is the overall mood or atmosphere conveyed by this landscape?”
Area Scale: 20% (Eq. Res: ∼259×170 )
Pixel-level textures (individual leaves, tiny ripples) are redundant for mood perception. TCRR aggressively scales the image down to 20% to capture global color blocks and contours, accurately predicting: “Serene and natural” at ultra-low latency.
Appendix
Table 18
Task Type
Prompt
TCRR Routing
Visual Input
Analysis
Basic Recognition
“What is the main animal lying right in the center?”
Area Scale: 10% (Eq. Res: ∼215×144 )
Iconic global features of a Shiba Inu (reddish-gold coat, triangular ears, silhouette) remain distinctly recognizable even when uniformly compressed to 10% of the original area. TCRR assigns the lowest resolution routing, maximizing efficiency while fully preserving semantic accuracy: “Shiba Inu” .
Sep 29, 2026·FangZhi Zhong, Xuerui Qiu, Yuqi Pan +4Text-Driven
Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences · Zhongguancun Academy +1