Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence
Organizations: University of Alabama at Birmingham · Amazon AGI · Work done during an internship at Amazon AGI.
Abstract
Latent visual reasoning (LVR) enables multimodal large language models (MLLMs) to perform intermediate computation in continuous latent tokens rather than expressing every reasoning step in words. However, unlike textual CoT, latent reasoning is not directly observable, making it difficult to supervise what latent tokens learn. In this work, we first conduct a thorough analysis of latent-token behavior and identify a latent evidence-credit gap: latent tokens respond only weakly to image perturbations that alter the correct answer. We hypothesize that this issue stems from the lack of explicit supervision during GRPO training. These findings suggest that a final-answer reward provides too little guidance on what visual evidence to preserve or how credit should be assigned across latent tokens. To bridge this gap, we propose ReaLVR, which brings visual-evidence supervision to the model's own free-running latent trajectories. ReaLVR contrasts correct and model-generated wrong answers to determine where stronger supervision is needed, and relevant and mismatched visual evidence to specify what to preserve. Across three model families, ReaLVR consistently outperforms evaluated LVR baselines, achieving the highest five-task average of 63.7% on Qwen2.5-VL-7B. Crucially, we are the first to scale visual reasoning in latent space, showing that our framework continues to deliver robust improvements at frontier model scales up to 235B. Further analyses show more question-sensitive latent-token positions, stronger alignment with relevant visual regions, and greater fixed-context dependence on the most attended latent tokens.
Figures & tables
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
| Item | Value |
|---|---|
| Accelerator | AMD Instinct MI250X (2 GCDs per accelerator) |
| Per node | 4 MI250X 8 GPUs (GCDs), 64 GB HBM2e per GPU |
| Accelerators, 235B | 200 nodes 4 800 MI250X (1,600 GCDs) |
| Accelerators, 30B | 64 nodes 4 256 MI250X (512 GCDs) |
| Parallelism | DeepSpeed ZeRO-3 (parameters, gradients, optimizer states) |
| Offload | None |
| Hyperparameter | Value |
|---|---|
| Per-GPU batch grad. accumulation | |
| Global batch (sequences) | 1,600 (235B) / 512 ( 30B) |
| Optimizer | AdamW |
| Peak learning rate | |
| LR schedule / warmup ratio / weight decay | cosine / 0.03 / 0.1 |
| Reconstruction loss | MSE, weight |
| Hyperparameter | Value |
|---|---|
| Objective | GRPO evidence loss |
| Group size (rollouts per prompt) | 8 |
| Sampling temperature / top- / top- | 0.6 / 1.0 / off |
| KL weight | 0 |
| Per-GPU batch grad. accumulation | |
| Prompts per step | 1,600 (235B) / 512 ( 30B) |
| Benchmark | Latent-token budget (difference from ) | ||||||
|---|---|---|---|---|---|---|---|
| HR-8K | |||||||
| MMVP | |||||||
| BLINK | |||||||
| MME-RealWorld | |||||||
| Mean | |||||||
| Variant | MMVP | BLINK | HR-4K | HR-8K | MME-RW | Avg. |
|---|---|---|---|---|---|---|
| LVR-RL (no evidence loss) | 64.2 | 53.6 | 69.6 | 64.4 | 50.1 | 60.4 |
| Where to supervise (answer contrast) | ||||||
| Uniform routing ( ) | 69.5 | 54.6 | 71.0 | 65.6 | 51.3 | 62.4 |
| Raw attention ( ) | 70.6 | 55.1 | 71.3 | 66.0 | 51.7 | 62.9 |
| Selective only ( ) | 71.2 | 55.3 | 71.4 | 66.2 | 51.9 | 63.2 |
| What to preserve (visual contrast) | ||||||
| Example | Latent tokens | Off-diag. cos. | Adjacent cos. | Max cos. |
|---|---|---|---|---|
| VISCOT 18 | 126 | 0.7342 | 0.9417 | 0.9951 |
| VISCOT 23 | 63 | 0.8408 | 0.9215 | 0.9958 |
| Family | Condition | Mean RCS |
|---|---|---|
| Reference | No augmentation | 0.99999997 |
| Image augmentation | Brightness/contrast/rotation/crop | 0.99917778 |
| Mask sweep | Strength 0.05 | 0.99987373 |
| Mask sweep | Strength 0.30 | 0.99974626 |
| Blur+jitter | Strength 0.10 | 0.99994875 |
| Blur+jitter | Strength 0.50 | 0.99988020 |
| BLINK | VISCOT | |||
|---|---|---|---|---|
| Generation schedule | MSE | Cosine | MSE | Cosine |
| Free-running | 1.4218 | 0.2145 | 1.7700 | 0.2863 |
| Force | – | 0.4221 | – | 0.4898 |
| Force | – | 0.6145 | – | 0.6619 |
| Force all | – | 1.0000 | – | 1.0000 |
| Diagnostic | Evaluation | Reported values |
|---|---|---|
| Answer-to-latent-token attention | MMVP | Acc. ; answer-to-LVR mass ; top-1 attention . |
| BLINK | Acc. ; mass ; top-1 attention . | |
| HR-4K / HR-8K | Acc. / ; correct-answer mass and top-1 mass are higher, with smaller deltas at 8K. | |
| Residual injection | generated rollout records | Best toward counterfactual prediction: to ; logit-margin deltas up to . |
| Representation | ARI | Bootstrap ARI | Linear accuracy | MLP accuracy |
|---|---|---|---|---|
| 0.9489 | 0.9492 | 0.9994 | 0.9960 | |
| LVR mean | 0.3885 | 0.3905 | 0.9991 | 0.9822 |
| 0.1298 | 0.1430 | 0.9966 | 0.9684 | |
| Final hidden | 0.2714 | 0.2558 | 0.9954 | 0.9641 |
| Method | Positions | Donor-answer lift (pp) | Donor-margin shift (nats) | |
|---|---|---|---|---|
| LVR-RL | Contrast | 1 | +1.2 | +0.02 |
| Random | 1 | +0.7 | +0.01 | |
| Contrast | 2 | +2.4 | +0.04 | |
| Random | 2 | +1.4 | +0.02 | |
| Contrast | 4 | +4.1 | +0.07 | |
| Random | 4 | +2.9 | +0.04 |