cs.CVSep 28, 2026

Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence

Authors: Xi Xiao, Tianchen Zhao, Youngeun Kim, Zhuowei Li, Linghan Xu, Jiaye Wu, Zheng Zhang, Xiang Xu, +5 more

Organizations: University of Alabama at Birmingham · Amazon AGI · Work done during an internship at Amazon AGI.

Abstract

Latent visual reasoning (LVR) enables multimodal large language models (MLLMs) to perform intermediate computation in continuous latent tokens rather than expressing every reasoning step in words. However, unlike textual CoT, latent reasoning is not directly observable, making it difficult to supervise what latent tokens learn. In this work, we first conduct a thorough analysis of latent-token behavior and identify a latent evidence-credit gap: latent tokens respond only weakly to image perturbations that alter the correct answer. We hypothesize that this issue stems from the lack of explicit supervision during GRPO training. These findings suggest that a final-answer reward provides too little guidance on what visual evidence to preserve or how credit should be assigned across latent tokens. To bridge this gap, we propose ReaLVR, which brings visual-evidence supervision to the model's own free-running latent trajectories. ReaLVR contrasts correct and model-generated wrong answers to determine where stronger supervision is needed, and relevant and mismatched visual evidence to specify what to preserve. Across three model families, ReaLVR consistently outperforms evaluated LVR baselines, achieving the highest five-task average of 63.7% on Qwen2.5-VL-7B. Crucially, we are the first to scale visual reasoning in latent space, showing that our framework continues to deliver robust improvements at frontier model scales up to 235B. Further analyses show more question-sensitive latent-token positions, stronger alignment with relevant visual regions, and greater fixed-context dependence on the most attended latent tokens.

Figures & tables

Appendix figures & tables22 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Leveraging Latent Visual Reasoning in Silence

    May 18, 2026Dongyao Zhu, Zhen Wang, Xi Xiao +7Latent Visual ReasoningVisual Reasoning

  2. LUT: Latent Utility Training for Visual Reasoning

    Aug 1, 2026Jiaxuan Kang, Siyu Chen, Mingda Li +6Latent Visual ReasoningVisual Reasoning

  3. Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs

    May 4, 2026Xin Zhang, Qiqi Tao, Jiawei Du +2Latent Visual ReasoningEfficient Latent Reasoning