Pre-trained vision encoders contain layer-wise visual representations that differ in spatial granularity, semantic abstraction, and sensitivity to local details. However, most Multimodal Large Language Models (MLLMs) rely on only the final or penultimate vision encoder representations or fixed aggregation rules, making visual abstraction largely query-agnostic and limiting access to fine-grained cues such as small objects, spatial details, text, and subtle visual attributes. In this work, we propose Mixture of Layers (MoL), an instruction-conditioned layer routing approach at the visual patch level that dynamically aggregates query-relevant latent representations from intermediate vision encoder layers. Given a text query, MoL predicts routing probabilities over vision encoder layers and performs a top-k sparse aggregation over selected hidden states at either the image level, patch level, or through a hybrid routing mechanism. In doing so, MoL enables query-adaptive access to layer-specific visual features for fine-grained visual reasoning. Our experiments across 7 fine-grained visual reasoning tasks demonstrate substantial performance improvements, especially across fine-grained visual grounding and understanding tasks such as +18.9% improvement on V* in overall accuracy, +4.5% on HRBench4K, and +16.3% on CharXiv compared to the baseline MLLMs, without resorting to multi-resolution inputs, simple interleaving of multiple vision encoders, or increasing the number of patch tokens. We study vision encoders' receptive field scales across different layers and their sampling behaviors to provide an in-depth analysis of why layer-wise sampling is helpful, demonstrating that conditional visual representations are a key step towards better visual perception and reasoning in MLLMs. Our project page is available at https://wjdghks950.github.io/mol.github.io/.
Figures & tables
Figure 1 : Overview of Mixture of Layers . This figure illustrates a fine-grained visual reasoning scenario where the MLLM should ground its reasoning based on visual representations that accurately capture the query-relevant visual features. MoL can be divided into three complementary variants: (i) MoL layer that generates a query-guided ( tˉ ) layer-level probability distribution over the input image; (ii) MoL patch , which generates a query-guided distribution over layers for each visual patch; (iii) MoL hybrid that measures the disagreement between the MoL layer and MoL patch and assigns gating weights to the reserve layer (§ 3.1 ) representations, which subsequently are pushed through weighted linear combination to form the final patch embedding sequence. We leave out some arrows (e.g., text query to LLM backbone) for brevity.
Method
V ∗
MMStar
HRBench8K
HRBench4K
RealWorldQA
NaturalBench
CharXiv
all
GPT4V-hard
OCR
direct attr.
rel. pos.
LLaVA-v1.5-13B (Vicuna-13B + CLIP) ( Liu et al., 2024 ; Radford et al., 2021 )
LLaVA-v1.5-13B
71.43
52.94
50.00
65.22
93.42
30.00
36.88
43.75
51.64
68.93
23.84
MoL layer
76.89
58.82
53.33
75.65
92.11
29.53
38.00
40.62
52.35
69.71
26.25
MoL patch
84.87
58.82
60.00
86.09
98.68
30.13
38.50
42.25
51.82
69.41
29.17
MoL hybrid
84.03
64.71
56.67
85.22
97.37
31.07
38.62
43.88
53.14
70.36
25.56
Table 1 : Performance across backbones and router types on fine-grained visual reasoning benchmarks. Metrics reported are accuracy. V ∗ benchmark is broken into its constituent sub-categories for detailed analysis. Best Δ denotes the maximum improvement (in absolute accuracy points) achieved by any MoL variant over the corresponding backbone baseline for each metric. We include MoL layer for the Vicuna-13B settings with CLIP and DINOv2 w/ Txt as global-routing comparisons. The multi-encoder results are reported separately in Table 3 .
Method
V ∗
MMStar
HRBench8K
HRBench4K
RealWorldQA
NaturalBench
CharXiv
all
GPT4V-hard
OCR
direct attr.
rel. pos.
LLaVA-v1.5-13B
71.43
52.94
50.00
65.22
93.42
30.00
36.88
43.75
51.64
68.93
23.84
TokenPacker ( Li et al., 2025b )
73.53
47.06
63.33
66.09
93.74
29.07
38.00
43.38
54.12
67.63
22.76
MLVF ( Lin et al., 2025 )
78.92
45.88
54.00
79.13
96.05
29.84
36.88
42.50
51.90
69.80
24.72
DeepStack-L ( Meng et al., 2024 )
65.13
64.71
46.67
56.52
85.53
29.26
37.63
43.75
57.77
69.10
21.31
Dense Connector ( Yao et al., 2024 )
81.35
47.05
56.67
83.47
98.68
30.06
37.12
43.75
52.41
71.41
25.23
Table 2 : Performance comparison against existing state-of-the-art methods on multi-layer visual feature integration. With all the existing baselines providing LLaVA-baseline checkpoints, we use LLaVA-v1.5-13B ( Liu et al., 2024 ) as our baseline in this experiment.
Method
Visual Input
V ∗
MMStar
HRBench8K
HRBench4K
RealWorldQA
NaturalBench
CharXiv
all
GPT4V-hard
OCR
direct attr.
rel. pos.
Interleaved-MoF ( Tong et al., 2024b )
CLIP + DINOv2
82.35
58.82
53.33
82.61
98.68
31.87
35.75
44.12
57.77
67.26
26.73
MoL hybrid
CLIP
84.03
64.71
56.67
85.22
97.37
31.07
38.62
43.88
53.14
70.36
25.56
MoL hybrid
DINOv2 †
84.45
70.59
50.00
86.09
98.68
30.13
36.50
39.87
52.29
67.89
26.92
MoL hybrid - Dual Enc.
CLIP + DINOv2 †
90.34
82.35
66.67
92.17
98.68
32.00
37.62
43.50
52.55
69.72
31.13
Best Δ
+7.99
+23.53
+13.34
+9.56
+0.00
+0.13
+2.87
-0.24
-4.63
+3.10
+4.40
Table 3 : Comparison of single- and multi-encoder routing with a Vicuna-13B backbone. Interleaved-MoF interleaves CLIP and DINOv2 patch inputs. MoL routes within one or both encoders; the dual-encoder variant uses concatenation/upsampling followed by a learned projector. Best Δ reports the largest improvement among the three MoL configurations over Interleaved-MoF, computed from the displayed scores in percentage points. † denotes DINOv2 with an aligned text encoder.
Figure 2 : Layer selection behaviors of Mixture of Layers . Average normalized routing probability across vision-encoder layers for LLaVA-v1.5-13B on V* Bench, MMStar, and GQA categories. Background intensity shows category-normalized receptive field (RF) ( Park et al., 2023 ) ; lighter regions indicate more localized layers. Markers denote selected top- k layers ( k=4 ). DINOv2 variants are shown in Appendix Figure 4 .
Figure 3 : Self-attention heatmap visualization of the selected layers aggregated across all the queries of the layers on the selected layers of CLIP-ViT-L14. For additional qualitative analysis on the selected layers of the DINOv2 encoder, please refer to Figure 5 .
Model
V*
GQA
TextVQA
MMMU
MoL image
77.31 (+5.88)
62.94 (-0.32)
58.17 (-3.08)
36.00 (-0.89)
MoL hybrid
84.03 (+12.60)
63.75 (+0.49)
61.11 (-0.14)
37.00 (+0.11)
Table 5 : Ablation on vision encoder-aligned text encoder. MoL image uses image-only routing, where the pooled embeddings of each vision encoder layer is converted to logits to form a distribution over the layers, while MoL hybrid uses text-conditioned routing. We denote the improvement / degradation compared to the baseline model in green / red . We use a Vicuna-13B + CLIP backbone.
Table 6 : Ablations on routing design. (a) Removing the reserve layer degrades performance, indicating its importance as a fallback when routing is uncertain. In tandem, having an anchor ( reserve ) layer with complementary k layers helps MLLMs perform grounded fine-grained visual reasoning. (b) Varying k shows that selecting a small set of layers suffices for strong fine-grained reasoning.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Backbone
Stage
Data
ZeRO
Epochs
Global Batch
Projector LR
Router LR
Vicuna-13B + CLIP
Pretrain
BLIP/LAION-CC-SBU 558K
2
1
256
1×10−3
1×10−5 – 5×10−5
Vicuna-13B + CLIP
Finetune
LLaVA-v1.5 mix 665K
3
1
128
2×10−5
2×10−5
Llama-3-8B + SigLIP
Pretrain
Bunny pretrain 2M
2
1
128
1×10−3
1×10−4
Llama-3-8B + SigLIP
Finetune
Bunny instruction 695K
3
1
128
2×10−5 – 2×10−4
2×10−5 – 2×10−4
Phi-1.5 + SigLIP
Pretrain
Bunny pretrain 2M
2
1
256
5×10−4
5×10−4
Phi-1.5 + SigLIP
Finetune
Bunny instruction 695K
3
1
128
5×10−5 – 2×10−4
5×10−5 – 2×10−4
Appendix
Table 7: Training setup. The Vicuna-13B + CLIP setting follows the LLaVA-v1.5 two-stage training recipe. The Llama-3-8B + SigLIP and Phi-1.5 + SigLIP settings follow the Bunny two-stage recipe and use LoRA during instruction tuning.
Variant
Implementation key
Routing granularity
k
Reserve layer
Reserve mode
Text conditioning
MoL layer
attn-token
image-level layer set
1–4
−2
local or scaled gate
attention
MoL patch
patch_layer
patch-specific layer set
1–4
−2
local gate
attention
MoL hybrid
hybrid
global-local fused routing
1–4
−2
disagreement gate
attention
Appendix
Table 8: Router configuration summary. The reserve layer is the penultimate vision-encoder layer used as an anchor representation. Appendix tables use normalized method names rather than implementation keys.
Method
MMBench
MMMU
MMStar
MMVet
MMVP
CV-Bench
NaturalBench
CharXiv
CountBenchQA
GQA
POPE
TextVQA
LLaVA-v1.5-13B
76.49
36.89
30.00
41.11
65.33
62.60
68.93
25.20
50.71
63.26
85.60
61.25
MoL layer
75.81
36.56
29.53
–
64.67
57.88
69.71
27.85
48.68
63.43
86.53
60.64
MoL patch
76.40
38.22
31.93
–
65.00
63.28
69.41
30.89
48.68
63.73
86.23
61.09
MoL hybrid
75.94
37.00
31.07
41.88
65.67
62.46
70.45
26.97
51.93
63.75
86.47
61.11
Appendix
Table 9 : Full benchmark matrix for the primary LLaVA-v1.5 family, part 1. MoL layer is the per-metric best of the attn-layer and attn-token layer-level implementations. MoL patch and MoL hybrid are the best available scores within their normalized method families.
Method
DocVQA
MME-RW
HallusionBench
RealWorldQA
ChartQA
HRBench4K
HRBench8K
V ∗ Bench
DTD
LLaVA-v1.5-13B
23.98
32.33
51.64
55.29
20.28
43.75
36.88
71.43
35.41
MoL layer
24.68
30.39
55.09
55.82
20.56
41.38
38.00
76.89
31.94
MoL patch
25.28
33.18
52.97
55.42
21.08
43.75
39.38
84.87
–
MoL hybrid
25.21
32.34
53.23
56.60
22.44
44.00
38.62
84.03
–
Appendix
Table 10 : Full benchmark matrix for the primary LLaVA-v1.5 family, part 2. MME-RW denotes MME-RealWorld. Empty entries indicate unavailable scores in the generated benchmark table.
Benchmark
MoL patch
MoL hybrid
k=1
k=2
k=4
Best
k=1
k=2
k=4
Best
V ∗ Bench
76.89
73.95
79.41
79.41
84.03
76.89
67.65
84.03
HRBench8K
39.38
36.50
37.25
39.38
38.62
38.25
34.88
38.62
HRBench4K
43.12
42.75
43.75
43.75
43.88
44.00
43.50
44.00
MMStar
30.13
31.93
29.93
31.93
30.47
30.00
29.73
30.47
NaturalBench
69.41
68.07
68.61
69.41
69.96
68.91
70.45
70.45
Appendix
Table 11 : Fine-grained visual reasoning under different k values. A small selected set of layers is usually sufficient; the best k varies by task and routing family.
Benchmark
MoL hybrid
MoL hybrid No Reserve
V ∗ Bench
82.35
78.57 (-3.78)
HRBench8K
38.62
35.50 (-3.12)
HRBench4K
43.88
38.75 (-5.13)
POPE
86.50
82.10 (-4.40)
MMMU
37.00
34.00 (-3.00)
Appendix
Table 12: Reserve-layer ablation. Removing the reserve layer consistently hurts performance, supporting its role as a stable anchor representation.
Figure 4 : DINOv2 layer sampling with receptive-field background on V ∗ Bench, MMStar, and GQA. Curves show the normalized routing probability assigned to each DINOv2 vision-encoder layer for MoL patch and MoL hybrid across question categories. Background intensity denotes the category-normalized receptive field (RF), where lighter regions indicate more localized layers and darker regions indicate broader spatial context. Type-wise markers indicate the selected top- k layers.
Figure 5 : Self-attention heatmap visualization of the selected layers aggregated across all the queries of the layers for DINOv2 encoder for the selected layers.
Latent visual reasoning equips vision--language models with continuous intermediate states that can process visual evidence without explicit textual reasoning traces or repeated image operations. However, existing methods often allow multiple latent tokens to access the same visual evidence through shared value projections, providing no mechanism for them to extract complementary visual information; simply increasing the latent budget can therefore yield redundant latent representations. We argue that effective latent reasoning should encourage different latent tokens to extract complementary visual information, and thereby act as specialized visual experts. Based on this insight, we propose MoLE, a Mixture of Latent Experts framework that controls both what visual evidence each latent visual expert observes and how it transforms that evidence. MoLE isolates latent visual experts during evidence extraction and uses dedicated latent summary experts to aggregate the complementary representations of latent visual experts. A two-stage training pipeline first forces visual evidence through this latent pathway and then restores direct visual access, requiring neither predefined expert roles nor intermediate visual targets. Across five visual reasoning benchmarks, MoLE achieves an average score of 78.6, outperforming data-matched supervised fine-tuning by 4.9 and the strongest evaluated latent visual reasoning baseline at the same latent budget by 3.6. Representation analyses show lower latent-state similarity and more diverse visual attention, while masking the latent pathway reduces average performance by 9.2. These results demonstrate that specializing latent computation is more effective than merely increasing the number of latent tokens.
Yingcheng Liu, Tianyi Jiang, Yujuan Ding +5
Tongji University, China · Hong Kong Polytechnic University, Hong Kong · Alibaba Group, China +1
Vision-Language Models often struggle with complex visual reasoning due to the visual information loss in textual CoT. Existing methods either add the cost of tool calls or rely on localized patch-based embeddings that are insufficient to extract semantics in multi-step reasoning. We propose "Decompose, Look, and Reason" (DLR), a reinforced latent reasoning framework that dynamically decomposes queries into textual premises, extracts premise-conditioned continuous visual latents, and deduces answers through grounded rationales. We introduce a three-stage training pipeline and propose a novel Spherical Gaussian Latent Policy, to enable effective exploration in the latent space. Extensive experiments on vision-centric benchmarks show that DLR consistently outperforms strong baselines, including text-only, interleaved multimodal CoT, and latent reasoning methods, while providing superior stepwise interpretability.
Mengdan Zhu, Senhao Cheng, Liang Zhao
Emory University · University of Michigan, Ann Arbor
Recent multimodal large language models (MLLMs) achieve strong performance on visual reasoning benchmarks, yet it remains unclear to what extent such performance reflects reasoning directly grounded in visual evidence. We introduce VisReason, a benchmark for vision-centric reasoning in everyday scenarios where perception and inference are tightly coupled. VisReason contains 1,505 questions across 10 categories spanning perceptual, structural, and conceptual reasoning. Our evaluation shows that VisReason poses a qualitatively different challenge from existing benchmarks, exposing substantial gaps between humans and current MLLMs and revealing limited benefits from test-time reasoning strategies. VisReason offers a focused diagnostic for evaluating vision-centric reasoning beyond language.
Longteng Guo, Yifan Wang, Pengkang Huo +4
Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences