MoLE: Mixture of Latent Experts for Complementary Visual Reasoning
Authors: Yingcheng Liu, Tianyi Jiang, Yujuan Ding, jiangbo Ai, Xun Jiang, Guoqing Wang, Wei Ye, Yi Bin
Organizations: Tongji University, China · Hong Kong Polytechnic University, Hong Kong · Alibaba Group, China · University of Electronic Science and Technology of China, China
Latent visual reasoning equips vision--language models with continuous intermediate states that can process visual evidence without explicit textual reasoning traces or repeated image operations. However, existing methods often allow multiple latent tokens to access the same visual evidence through shared value projections, providing no mechanism for them to extract complementary visual information; simply increasing the latent budget can therefore yield redundant latent representations. We argue that effective latent reasoning should encourage different latent tokens to extract complementary visual information, and thereby act as specialized visual experts. Based on this insight, we propose MoLE, a Mixture of Latent Experts framework that controls both what visual evidence each latent visual expert observes and how it transforms that evidence. MoLE isolates latent visual experts during evidence extraction and uses dedicated latent summary experts to aggregate the complementary representations of latent visual experts. A two-stage training pipeline first forces visual evidence through this latent pathway and then restores direct visual access, requiring neither predefined expert roles nor intermediate visual targets. Across five visual reasoning benchmarks, MoLE achieves an average score of 78.6, outperforming data-matched supervised fine-tuning by 4.9 and the strongest evaluated latent visual reasoning baseline at the same latent budget by 3.6. Representation analyses show lower latent-state similarity and more diverse visual attention, while masking the latent pathway reduces average performance by 9.2. These results demonstrate that specializing latent computation is more effective than merely increasing the number of latent tokens.
Figures & tables
Figure 1: Two paths to enhancing latent visual reasoning. Simply increasing latent count while sharing visual evidence and transformations can produce redundant representations. Explicitly specializing what each latent observes and how it transforms that evidence encourages complementary representations and stronger reasoning without increasing the latent budget.
Figure 2: Overview of MoLE . (a) A layer-wise router sparsely assigns visual tokens to isolated latent visual experts; latent summary experts aggregate their representations, and answer tokens attend to both latent sets. (b) An Expert-Specific Value Adapter (ESVA) is used within the attention block when a latent visual expert attends to the routed visual tokens. Importantly, the ESVA is activated only for expert–visual-token connections selected by the router; all other attention interactions use the standard attention computation without an ESVA.
CVBench
Method
2D
3D
BLINK
RWQA
V*
POPE
Avg.
Proprietary models
GPT-4o †
74.6
83.9
63.0
69.7
42.9
85.6
70.0
Claude-4-Sonnet †
73.5
79.1
39.6
63.7
15.2
84.6
59.3
Base and data-matched models
Qwen2.5-VL-7B
75.4
73.3
55.6
68.9
76.6
85.0
72.5
Table 1: Overall performance across five visual reasoning benchmarks.
Figure 3: Complementarity, component ablations, and latent-state interventions. (a) Lower latent-state similarity and higher attention-pattern diversity indicate more complementary latent representations. (b) Performance after independently removing each component; the dashed line marks the full MoLE result. (c) Performance under inference-time interventions on the learned latent states.
Budget
E
S
Avg.
8
4
4
78.6
6
2
77.2
7
1
77.4
8
0
75.7
16
8
8
77.8
12
4
76.8
Table 2: Latent-allocation, training-stage, and training-objective ablations. (a) Allocation of visual ( E ) and summary ( S ) experts. (b) Training stages and objectives.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Attention pathways in the two-stage training pipeline. Dashed arrows denote permitted information flow between token groups.
Task family
Source group
Samples
Visual math reasoning
MathV360K identity group ( Shi et al., 2024 ) ; IconQA ( Lu et al., 2021 )
3,150+3,150
Inferential VQA
ALLaVA-LAION; ALLaVA-VFLAN ( Chen et al., 2024 )
3,150+3,150
Knowledge VQA
A-OKVQA ( Schwenk et al., 2022 )
6,300
Compositional reasoning
CLEVR ( Johnson et al., 2017 )
6,300
Safety classification
Hateful Memes ( Kiela et al., 2020 )
6,300
Detailed captioning
Image textualization; numeric-ID group; sa ; ShareGPT-4o ( Cui et al., 2024 )
4×1,575
Appendix
Table 3: Composition of the task-balanced 63K training set. Counts separated by “+” correspond to the listed sources in the same order.
Benchmark
Tokens/image
Reported metric
POPE
64–512
Unweighted mean F1 over the adversarial, popular, and random subsets
BLINK
64–512
Unweighted mean accuracy over the 14 tasks
V*Bench
256–4,096
Accuracy over all test examples
RealWorldQA
256–4,096
Accuracy over all test examples
CV-Bench
256–2,048
2D: mean of ADE20K and COCO accuracies; 3D: Omni3D accuracy
Appendix
Table 4: Per-image visual-token budgets and reported metrics for the five evaluation benchmarks.
Model
CVBench
Count
Depth
Dist.
GPT-4o
79.2
65.6
86.7
81.0
Claude-4-Sonnet
76.3
62.2
77.7
80.5
Appendix
Table 5: CVBench-related scores for GPT-4o and Claude-4-Sonnet as reported by CoVT.
CVBench
Method
2D
3D
BLINK
RWQA
V*
POPE
Avg.
Qwen2.5-VL-7B
75.4
73.3
55.6
68.9
76.6
85.0
72.5
Parameter-matched adapter
74.9
75.4
56.5
66.7
76.3
86.8
72.8
MoLE
78.2
87.8
63.0
71.5
80.6
90.2
78.6
Appendix
Table 6: Comparison with a parameter-matched shared-adapter baseline.
CVBench
Model
2D
3D
BLINK
RWQA
V*
POPE
Avg.
Qwen2.5-VL-7B
75.4
73.3
55.6
68.9
76.6
85.0
72.5
MoLE (seed 42)
78.2
87.8
63.0
71.5
80.6
90.2
78.6
MoLE (seed 45)
77.5
87.1
61.3
72.0
79.4
89.9
77.9
MoLE (seed 48)
78.3
86.8
62.1
71.2
79.8
89.9
78.0
Mean ± std.
78.0±0.4
87.2±0.5
62.1±0.9
71.6±0.4
79.9±0.6
90.0±0.2
78.2±0.4
Appendix
Table 7: Performance of the default MoLE model across three random seeds. Mean and sample standard deviation are computed over the three runs.
CVBench
Model
2D
3D
BLINK
RWQA
V*
POPE
Avg.
Qwen3-VL
72.8
86.6
45.7
65.6
75.4
89.3
72.6
MoLE (42)
74.7
88.3
55.5
66.3
78.0
90.2
75.5
MoLE (45)
73.4
87.9
55.5
66.0
77.4
89.9
75.0
MoLE (48)
73.9
88.4
55.7
66.7
77.8
90.0
75.4
Mean ± std.
74.0±0.7
88.2±0.3
55.6±0.1
66.3±0.4
77.7±0.3
90.0±0.2
75.3±0.3
Appendix
Table 8: Backbone and scale extension on Qwen3-VL. Mean and sample standard deviation are computed over three MoLE runs at each model scale.
Router Top- K
Avg.
K=1
77.8
K=2
78.6
K=3
77.4
K=4
77.9
Appendix
Table 9: Ablations of (a) routing density and (b) ESVA hidden dimension. Only average performance is reported.
Continuous latent-space reasoning offers a compact alternative to textual chain-of-thought for multimodal models, enabling high-dimensional visual evidence to be integrated without explicit reasoning tokens. However, we identify a previously overlooked optimization pathology in existing latent visual reasoning methods: although visual latents become semantically enriched during training, their contribution to final answer prediction is systematically suppressed. Within the shared parameter space, the autoregressive objective favors shortcut reliance on direct visual input, driving latent tokens toward transition-like states rather than informative reasoning content. We term this phenomenon Silenced Visual Latents. To address it, we disentangle the two conflicting objectives by directly optimizing the latent reasoning at inference time, keeping backbone parameters frozen. In Stage I, visual latents are warmed up via query-guided contrastive latent--visual alignment, improving semantic quality while preventing latent collapse. In Stage II, the latent reasoning is further optimized via a confidence-progression reward, which incentivizes predicted token distributions along the latent span to become progressively more concentrated, routing predictions through the latent reasoning rather than bypassing it. Experiments across eight benchmarks and four model backbones show that inference-time latent optimization, without any parameter updates, effectively unleashes the suppressed reasoning capacity of visual latents.
Xin Zhang, Qiqi Tao, Jiawei Du +2
Centre for Frontier AI Research, Agency for Science, Technology and Research, Singapore · Institute of High Performance Computing, Agency for Science, Technology and Research, Singapore · Singapore University of Technology and Design +1
Latent visual reasoning (LVR) enables multimodal large language models (MLLMs) to perform intermediate computation in continuous latent tokens rather than expressing every reasoning step in words. However, unlike textual CoT, latent reasoning is not directly observable, making it difficult to supervise what latent tokens learn. In this work, we first conduct a thorough analysis of latent-token behavior and identify a latent evidence-credit gap: latent tokens respond only weakly to image perturbations that alter the correct answer. We hypothesize that this issue stems from the lack of explicit supervision during GRPO training. These findings suggest that a final-answer reward provides too little guidance on what visual evidence to preserve or how credit should be assigned across latent tokens. To bridge this gap, we propose ReaLVR, which brings visual-evidence supervision to the model's own free-running latent trajectories. ReaLVR contrasts correct and model-generated wrong answers to determine where stronger supervision is needed, and relevant and mismatched visual evidence to specify what to preserve. Across three model families, ReaLVR consistently outperforms evaluated LVR baselines, achieving the highest five-task average of 63.7% on Qwen2.5-VL-7B. Crucially, we are the first to scale visual reasoning in latent space, showing that our framework continues to deliver robust improvements at frontier model scales up to 235B. Further analyses show more question-sensitive latent-token positions, stronger alignment with relevant visual regions, and greater fixed-context dependence on the most attended latent tokens.
Xi Xiao, Tianchen Zhao, Youngeun Kim +10
University of Alabama at Birmingham · Amazon AGI · Work done during an internship at Amazon AGI.
Recent latent visual reasoning methods achieve substantial gains by inserting continuous latent tokens into multimodal language models. These gains are commonly attributed to the tokens encoding visual evidence; recent analyses, however, reveal a paradox: the tokens are loosely tied to the image and contribute little to the answer. Critically, these analyses treat latent tokens as a single unit, obscuring the true source of the gains. We therefore decompose latent tokens into three testable components: latent slots, boundary markers, and format, and develop a state-of-the-art method as a probe under favorable conditions. Across six method-stage settings and four perception-heavy benchmarks, latent slots fail every prediction of the visual-memory account. Strikingly, retaining only the boundary markers preserves 78 to 100% of the gain in several settings, while the model attends to the image more narrowly at latent positions than at answer positions. The gain therefore comes from boundary markers, format, and this attention pattern, not from latent slots. How each method engages this mechanism depends on its training supervision: at matched accuracy, mechanisms can still differ markedly. Latent visual reasoning thus needs evaluation not only by accuracy but by what the model actually relies on.
Garvin Guo, Yu Chen, Xiang Wang +4
Amap, Alibaba Group · Shanghai Innovation Institute