Latent visual reasoning (LVR) enables multimodal large language models (MLLMs) to perform intermediate computation in continuous latent tokens rather than expressing every reasoning step in words. However, unlike textual CoT, latent reasoning is not directly observable, making it difficult to supervise what latent tokens learn. In this work, we first conduct a thorough analysis of latent-token behavior and identify a latent evidence-credit gap: latent tokens respond only weakly to image perturbations that alter the correct answer. We hypothesize that this issue stems from the lack of explicit supervision during GRPO training. These findings suggest that a final-answer reward provides too little guidance on what visual evidence to preserve or how credit should be assigned across latent tokens. To bridge this gap, we propose ReaLVR, which brings visual-evidence supervision to the model's own free-running latent trajectories. ReaLVR contrasts correct and model-generated wrong answers to determine where stronger supervision is needed, and relevant and mismatched visual evidence to specify what to preserve. Across three model families, ReaLVR consistently outperforms evaluated LVR baselines, achieving the highest five-task average of 63.7% on Qwen2.5-VL-7B. Crucially, we are the first to scale visual reasoning in latent space, showing that our framework continues to deliver robust improvements at frontier model scales up to 235B. Further analyses show more question-sensitive latent-token positions, stronger alignment with relevant visual regions, and greater fixed-context dependence on the most attended latent tokens.
Figures & tables
Figure 1: ReaLVR connects visual evidence to latent reasoning. (a) ReaLVR identifies the stroller missed by LVR-7B in a complex scene; dashed boxes mark image details. (b) Visual and answer contrast determine what to preserve and where to supervise , during training only. (c) Replacing the top-8 tokens ranked by answer-to-token attention, with other states fixed, decreases correct-answer probability by 4 and 11 percentage points. The larger ReaLVR drop indicates greater local answer dependence.
Figure 2: ReaLVR learns visual grounding on its own latent trajectory. The model first generates continuous latent states from the image and question. Relevant and mismatched visual prototypes specify what these states should preserve. Correct and wrong answers are separately teacher-forced after the shared latent span; their attention contrast determines the supervision weights. The weights are detached, and the visual loss trains the latent-generation process. Both supervision branches are used only during training; inference follows the original LVR procedure.
Figure 3: Visual evidence and latent-token readout. Attention of LVR (a) and ReaLVR (b) on Qwen2.5-VL-7B , averaged over all layers and heads and 200 HR-Bench-4K examples, with image, question, and answer tokens averaged into 8/4/4 bins and the K=8 latent positions shown individually. The gray upper triangle is the causal mask: each query can attend only to its own and earlier positions ( Vaswani et al., 2017 ) . Each row sums to one over allowed keys. Green tick labels mark the same visual-evidence keys ( 4 , 6 , and 7 ), the image bins overlapping the annotated ROI. The narrow frames show latent queries reading these visual keys; the bottom frames show answer queries reading latent states. Panel (a) shows weak readout along both links; panel (b) shows stronger readout along both. Please see Appendix I for further attention analyses.
Figure 4: Latent changes and answer updates after image edits. (a) Four edits of the same reference image. (b) Cosine distance between the mean-pooled latent trajectories for the original and edited inputs; zero denotes no change. Each point is the average over 512 original–edited pairs per edit type with the question held fixed, measured for both LVR and ReaLVR on Qwen2.5-VL-7B . The broken vertical axis enlarges the LVR range to make its small fluctuations visible; values are unchanged. (c) Each user question is followed by the two model replies, read from the original image to the edited image. See Appendix H for more analysis.
Figure 5: Image and token views of the same three visual questions. ReaLVR and LVR on the same fence, bus, and icy-ground questions: image saliency overlays (top) and token-pair maps (bottom). Each pair is labeled with its question and answer.
Figure 6: A real visual answer and its latent-state text readout. ReaLVR ( Gemma-3-12B ) answers “27B” from an HR-Bench image (a). The original vocabulary head reads each of its eight latent states as < , as it does all 376 states from 47 questions (b), yet predicts 7 after 2 , B after 7 , and </ after B during answer generation (c). The latent readouts do not form an intermediate explanation.
Table A.1: Hardware and distributed training configuration. Each MI250X accelerator contributes two ROCm-visible 64 GB GCDs.
Hyperparameter
Value
Per-GPU batch × grad. accumulation
1×1
Global batch (sequences)
1,600 (235B) / 512 ( ≤ 30B)
Optimizer
AdamW
Peak learning rate
1×10−5
LR schedule / warmup ratio / weight decay
cosine / 0.03 / 0.1
Reconstruction loss
MSE, weight 0.1
Appendix
Table A.2: Stage-1 hyperparameters.
Hyperparameter
Value
Objective
GRPO + evidence loss
Group size G (rollouts per prompt)
8
Sampling temperature / top- p / top- k
0.6 / 1.0 / off
KL weight β
0
Per-GPU batch × grad. accumulation
1×1
Prompts per step
1,600 (235B) / 512 ( ≤ 30B)
Appendix
Table A.3: Stage-2 hyperparameters (GRPO with ReaLVR evidence supervision).
Figure B.1: From final reward to latent evidence credit. The output reward indicates whether the answer is correct. ReaLVR uses relevant and mismatched visual evidence to determine what the latent tokens should preserve, then contrasts ground-truth and wrong-answer readouts to determine where supervision should be stronger. Intervention separately evaluates whether the answer depends on the latent tokens and is not part of the training objective.
Benchmark
Latent-token budget K (difference from K=8 )
0
2
4
8†
12
16
20
HR-8K
−3.0
−2.0
−1.5
+0.0
−1.0
−0.5
+0.0
MMVP
−0.7
−1.0
−0.3
+0.0
−1.0
−1.0
+0.0
BLINK
−4.0
+0.0
−1.1
+0.0
+1.2
+2.3
+0.9
MME-RealWorld
−4.4
+1.9
+2.1
+0.0
−0.2
−1.4
−2.1
Mean
−3.0
−0.3
−0.2
+0.0
−0.2
−0.1
−0.3
Appendix
Table F.1: Sensitivity to the inference-time latent budget. The ReaLVR ( Qwen2.5-VL-7B ) checkpoint is held fixed while the latent-token budget K is varied at inference; K=0 decodes the answer without a latent span. K=8 is the budget used during training and serves as the reference ( † ): entries are differences in accuracy (percentage points) from that column, which is therefore 0.0 by construction. The sweep is run on a fixed evaluation subset so that all budgets are scored identically; its absolute scores are consequently not on the same scale as Table 1 , and the differences reported here are meaningful only within this sweep. Under the full evaluation protocol of Table 1 , the K=8 setting scores 72.0 on MMVP , 55.8 on BLINK , 66.6 on HR-8K and 52.2 on MME-RealWorld . Bold marks each row’s largest value, including ties. The mean is the unweighted average across the four listed benchmarks.
Variant
MMVP
BLINK
HR-4K
HR-8K
MME-RW
Avg.
LVR-RL (no evidence loss)
64.2
53.6
69.6
64.4
50.1
60.4
Where to supervise (answer contrast)
Uniform routing ( η=1 )
69.5
54.6
71.0
65.6
51.3
62.4
Raw attention ( γt=rt+ )
70.6
55.1
71.3
66.0
51.7
62.9
Selective only ( η=0 )
71.2
55.3
71.4
66.2
51.9
63.2
What to preserve (visual contrast)
Appendix
Table F.2: Component ablation on Qwen2.5-VL-7B . Each row removes or replaces one ingredient of ReaLVR; all other settings follow Appendix A . “Uniform routing” sets η=1 so every latent position receives weight 1/K and the answer contrast is unused. “Selective only” sets η=0 . “No negatives” replaces the margin with a plain cosine alignment to p+ . “Raw attention” routes with rt+ instead of [rt+−rt−]+ . “Undetached” removes the stop-gradient on wt . “Off-policy” applies the evidence loss to the saved rollout latents instead of a regenerated trajectory. LVR-RL is the λev=0 endpoint.
Figure G.1: Fixed-context latent-token dependence. Each cell gives the correct-answer probability after replacing the top- k answer-read latent tokens, with all remaining latent states held fixed. Columns increase k from 0 to 8 . ReaLVR’s probability falls from 0.70 to 0.59 , the largest endpoint drop among the four methods.
Figure G.2: Target-region attention enrichment. Layer-wise answer attention on the annotated region relative to matched background windows. Values above 1 indicate target-region enrichment; green and red curves correspond to correct and incorrect generations, respectively. Their overlap shows that alignment alone does not determine answer correctness.
Example
Latent tokens
Off-diag. cos.
Adjacent cos.
Max cos.
VISCOT 18
126
0.7342
0.9417
0.9951
VISCOT 23
63
0.8408
0.9215
0.9958
Appendix
Table G.1: Latent-token similarity diagnostics. (a) Raw within-trajectory cosine similarity on two VISCOT examples. (b) Context-normalized similarity on BLINK ( N=697 ), with same-prefix dummy and ordinary continuation tokens as references.
Family
Condition
Mean RCS
Reference
No augmentation
0.99999997
Image augmentation
Brightness/contrast/rotation/crop
0.99917778
Mask sweep
Strength 0.05
0.99987373
Mask sweep
Strength 0.30
0.99974626
Blur+jitter
Strength 0.10
0.99994875
Blur+jitter
Strength 0.50
0.99988020
Appendix
Table G.2: Representation stability and effective rank. (a) Representation consistency (RCS) is the mean pairwise cosine between repeated mean LVR trajectories under each condition. (b) Rank90, Rank95, and Rank99 count the principal directions needed to explain 90%, 95%, and 99% of the variance; PR is the participation ratio.
BLINK
VISCOT
Generation schedule
MSE
Cosine
MSE
Cosine
Free-running
1.4218
0.2145
1.7700
0.2863
Force k=4
–
0.4221
–
0.4898
Force k=8
–
0.6145
–
0.6619
Force all k=16
–
1.0000
–
1.0000
Appendix
Table G.3: Alignment with visual targets under target forcing. The first k latent positions receive the visual target vectors; later positions are generated autoregressively. MSE and cosine similarity compare the resulting trajectory with the visual targets.
Acc. 49.28% ; mass Δ=+0.0029 ; top-1 attention Δ=+0.0217 .
HR-4K / HR-8K
Acc. 56.88% / 48.62% ; correct-answer mass and top-1 mass are higher, with smaller deltas at 8K.
Residual injection
47,012 generated rollout records
Best Δ toward counterfactual prediction: +0.0068 to +0.0137 ; logit-margin deltas up to +0.0358 .
Appendix
Table G.4: Answer readout and residual-injection diagnostics. Summary statistics from generated rollouts. Attention deltas compare correct and incorrect predictions.
Representation
ARI
Bootstrap ARI
Linear accuracy
MLP accuracy
hpre
0.9489
0.9492
0.9994
0.9960
LVR mean
0.3885
0.3905
0.9991
0.9822
hlvr,last
0.1298
0.1430
0.9966
0.9684
Final hidden
0.2714
0.2558
0.9954
0.9641
Appendix
Table G.5: Task decodability and sensitivity to answer-changing edits. (a) Task-label structure on BLINK ( N=697 ). (b) Sensitivity to synthetic image edits ( N=2,048 ; 512 pairs per edit). LVR distance uses the mean-pooled latent trajectory; final distance uses the final hidden representation.
Figure H.1: Attention and saliency cases (01–03). Top to bottom: the cap on a table, the bus beside a person, and the chair beside a tennis player.
Figure H.2: Attention and saliency cases (04–06). Top to bottom: the object behind the blender, the furniture left of the heater, and a furniture-material comparison.
Figure H.3: Attention and saliency cases (07–09). Top to bottom: pans left of the pots, the teddy-bear answer comparison, and a bear in water. In the middle example (Case 08), ReaLVR answers “teddy bear” and LVR-7B answers “blanket” to “What is touching the bed?”
Figure H.4: Attention and saliency cases (10–12). Top to bottom: the object in front of an airplane’s front wheel, the chair-and-bird scene, and the vehicle containing the pilot.
Figure H.5: One question from each of the five benchmarks. Images, questions, and answers come from the official datasets ( Tong et al., 2024 ; Fu et al., 2024 ; Wang et al., 2025 ; Zhang et al., 2025 ) . Robot bubbles show annotated answers, not recorded ReaLVR predictions; the chart calculation is added for clarity. Insets enlarge source-image details, and the chart is cropped to the relevant panel. Point and value labels are enlarged for readability. Multiple-choice options are omitted, and the chart question is shortened.
Figure H.6: Position-wise latent variation. (a) Cross-example variation. (b) Variation across questions about the same image. Heatmap columns index latent positions 1 – 16 . (c) Reported top-token variation gap: 0.07 for ReaLVR, 0.02 for Monet, and 0.01 for each LVR variant. ReaLVR concentrates more of the measured variation at particular latent positions.
Figure I.1: Visual evidence and answer readout. (a) Target-region attention relative to same-area background windows, shown by decoder layer and latent step; 1 means no preference. Both heatmaps share a color scale. (b) Raw attention mass assigned to all image tokens before the answer. (c) Raw attention to each latent state from the query used to predict the first answer token, before that token is supplied. The panels illustrate aggregate post-softmax quantities.
Figure I.2: Question-dependent evidence selection. The photograph is HRBench-8K example 798 ( Wang et al., 2025 ) . Q1 adapts its spatial-relation question; Q2 and both paraphrases are author-written. Dashed boxes mark the manually specified regions for the boat-and-buildings question and the animal question. All maps share a density scale normalized to a full-image mean of 1 . The chart reports mean density within each question’s region and compares that question with its paraphrase ( Q1′ or Q2′ ). This uniform baseline differs from the matched-background baseline in Figure I.1 .
Method
Positions
k
Donor-answer lift (pp)
Donor-margin shift (nats)
LVR-RL
Contrast
1
+1.2
+0.02
Random
1
+0.7
+0.01
Contrast
2
+2.4
+0.04
Random
2
+1.4
+0.02
Contrast
4
+4.1
+0.07
Random
4
+2.9
+0.04
Appendix
Table J.2: Fixed-context latent exchange. The recipient image and question remain fixed. Donor-answer lift and donor-margin shift are differences from the self-swap control. Contrast and random rows at the same k use the same donors. The full-span row does not test position selection. Donor-answer lift is a percentage-point change in donor-answer frequency.
Centre for Frontier AI Research, Agency for Science, Technology and Research, Singapore · Institute of High Performance Computing, Agency for Science, Technology and Research, Singapore · Singapore University of Technology and Design +1