MoLE: Mixture of Latent Experts for Complementary Visual Reasoning
Organizations: Tongji University, China · Hong Kong Polytechnic University, Hong Kong · Alibaba Group, China · University of Electronic Science and Technology of China, China
Abstract
Latent visual reasoning equips vision--language models with continuous intermediate states that can process visual evidence without explicit textual reasoning traces or repeated image operations. However, existing methods often allow multiple latent tokens to access the same visual evidence through shared value projections, providing no mechanism for them to extract complementary visual information; simply increasing the latent budget can therefore yield redundant latent representations. We argue that effective latent reasoning should encourage different latent tokens to extract complementary visual information, and thereby act as specialized visual experts. Based on this insight, we propose MoLE, a Mixture of Latent Experts framework that controls both what visual evidence each latent visual expert observes and how it transforms that evidence. MoLE isolates latent visual experts during evidence extraction and uses dedicated latent summary experts to aggregate the complementary representations of latent visual experts. A two-stage training pipeline first forces visual evidence through this latent pathway and then restores direct visual access, requiring neither predefined expert roles nor intermediate visual targets. Across five visual reasoning benchmarks, MoLE achieves an average score of 78.6, outperforming data-matched supervised fine-tuning by 4.9 and the strongest evaluated latent visual reasoning baseline at the same latent budget by 3.6. Representation analyses show lower latent-state similarity and more diverse visual attention, while masking the latent pathway reduces average performance by 9.2. These results demonstrate that specializing latent computation is more effective than merely increasing the number of latent tokens.
Figures & tables
| CVBench | |||||||
| Method | 2D | 3D | BLINK | RWQA | V* | POPE | Avg. |
| Proprietary models | |||||||
| GPT-4o † | 74.6 | 83.9 | 63.0 | 69.7 | 42.9 | 85.6 | 70.0 |
| Claude-4-Sonnet † | 73.5 | 79.1 | 39.6 | 63.7 | 15.2 | 84.6 | 59.3 |
| Base and data-matched models | |||||||
| Qwen2.5-VL-7B | 75.4 | 73.3 | 55.6 | 68.9 | 76.6 | 85.0 | 72.5 |
| Budget | Avg. | ||
| 8 | 4 | 4 | 78.6 |
| 6 | 2 | 77.2 | |
| 7 | 1 | 77.4 | |
| 8 | 0 | 75.7 | |
| 16 | 8 | 8 | 77.8 |
| 12 | 4 | 76.8 |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Task family | Source group | Samples |
|---|---|---|
| Visual math reasoning | MathV360K identity group ( Shi et al., 2024 ) ; IconQA ( Lu et al., 2021 ) | |
| Inferential VQA | ALLaVA-LAION; ALLaVA-VFLAN ( Chen et al., 2024 ) | |
| Knowledge VQA | A-OKVQA ( Schwenk et al., 2022 ) | |
| Compositional reasoning | CLEVR ( Johnson et al., 2017 ) | |
| Safety classification | Hateful Memes ( Kiela et al., 2020 ) | |
| Detailed captioning | Image textualization; numeric-ID group; sa ; ShareGPT-4o ( Cui et al., 2024 ) |
| Benchmark | Tokens/image | Reported metric |
|---|---|---|
| POPE | 64–512 | Unweighted mean F1 over the adversarial, popular, and random subsets |
| BLINK | 64–512 | Unweighted mean accuracy over the 14 tasks |
| V*Bench | 256–4,096 | Accuracy over all test examples |
| RealWorldQA | 256–4,096 | Accuracy over all test examples |
| CV-Bench | 256–2,048 | 2D: mean of ADE20K and COCO accuracies; 3D: Omni3D accuracy |
| Model | CVBench | Count | Depth | Dist. |
|---|---|---|---|---|
| GPT-4o | 79.2 | 65.6 | 86.7 | 81.0 |
| Claude-4-Sonnet | 76.3 | 62.2 | 77.7 | 80.5 |
| CVBench | |||||||
|---|---|---|---|---|---|---|---|
| Method | 2D | 3D | BLINK | RWQA | V* | POPE | Avg. |
| Qwen2.5-VL-7B | 75.4 | 73.3 | 55.6 | 68.9 | 76.6 | 85.0 | 72.5 |
| Parameter-matched adapter | 74.9 | 75.4 | 56.5 | 66.7 | 76.3 | 86.8 | 72.8 |
| MoLE | 78.2 | 87.8 | 63.0 | 71.5 | 80.6 | 90.2 | 78.6 |
| CVBench | |||||||
| Model | 2D | 3D | BLINK | RWQA | V* | POPE | Avg. |
| Qwen2.5-VL-7B | 75.4 | 73.3 | 55.6 | 68.9 | 76.6 | 85.0 | 72.5 |
| MoLE (seed 42) | 78.2 | 87.8 | 63.0 | 71.5 | 80.6 | 90.2 | 78.6 |
| MoLE (seed 45) | 77.5 | 87.1 | 61.3 | 72.0 | 79.4 | 89.9 | 77.9 |
| MoLE (seed 48) | 78.3 | 86.8 | 62.1 | 71.2 | 79.8 | 89.9 | 78.0 |
| Mean std. | |||||||
| CVBench | |||||||
| Model | 2D | 3D | BLINK | RWQA | V* | POPE | Avg. |
| Qwen3-VL | 72.8 | 86.6 | 45.7 | 65.6 | 75.4 | 89.3 | 72.6 |
| MoLE (42) | 74.7 | 88.3 | 55.5 | 66.3 | 78.0 | 90.2 | 75.5 |
| MoLE (45) | 73.4 | 87.9 | 55.5 | 66.0 | 77.4 | 89.9 | 75.0 |
| MoLE (48) | 73.9 | 88.4 | 55.7 | 66.7 | 77.8 | 90.0 | 75.4 |
| Mean std. | |||||||
| Router Top- | Avg. |
|---|---|
| 77.8 | |
| 78.6 | |
| 77.4 | |
| 77.9 |
| Model | Time (s) | FLOPs | Peak memory (GiB) |
|---|---|---|---|
| Qwen2.5-VL-7B | 0.374478 | 15.836116 | |
| MoLE | 0.460499 | 16.081371 | |
| Relative overhead |