Aligning Thoughts with Answers: Probability Rewards to Tame Thinking Drift
Authors: Pengzhan Sun, Shiu-hong Kao, Shijie Li, Yongyi Su, Junbin Xiao, Arjun Reddy Akula, Angela Yao
Organizations: National University of Singapore · A*STAR Institute of Advanced Intelligence and Computing, Singapore · South China University of Technology · University of Science and Technology of China · Google DeepMind
This paper studies \textbf{thinking--answer consistency} in vision-language models. We focus on Visual Intention Grounding, where a model infers a target object based on a human intention query and predicts a bounding box. We reveal that previous IoU-based reinforcement learning (RL) frameworks suffer from ``thinking drift'', where the model produces a correct bounding box, despite having an incorrect reasoning process pointing to a different target object. Thus, we propose \textbf{Rita} (\textit{ReInforcing Thinking--Answer consistency}) as a novel RL paradigm to tame the drift. Specifically, Rita introduces two reasoning-label-free RL rewards, constructed from the conditional probability of reference answers: a \textbf{thinking reward} and a \textbf{consistency reward}. It also adopts a difficulty-aware \textbf{data filtering} strategy that selects informative easy-to-medium samples for RL using rollout error rate and reward variance. Extensive experiments on EgoIntention and the new RefEgo-Int benchmarks show that Rita performs consistently superior to the supervised finetuning approaches and vanilla RL-finetuned frameworks.
Figures & tables
Figure 1: Base RFT with IoU reward alone fails to enforce semantic consistency. In the Base RFT ( red ), which relies solely on the IoU reward, the model predicts a reasonable bounding box but provides irrelevant reasoning, leading to a label error. In contrast, in our Rita ( green ), the integration of thinking and consistency rewards encourages consistency between the reasoning process and the final prediction, correcting the label while maintaining spatial accuracy.
Figure 2: Overview of the Rita framework. Given an input image and human intention query, the VLM samples multiple think-then-answer rollouts (Q,Tk,Ak) , from which we derive a grounding reward via IoU with the ground-truth box and a format reward. We then overwrite Ak with the reference answer A⋆ and run the VLM in teacher-forced mode to obtain logits for the answer tokens, which define the thinking reward Rthink . The consistency reward Rcons scales this score by the predicted box’s IoU minus a margin δ . During the GRPO-based optimization, we conduct difficulty-and-variance-aware data sampling computed from IoU rollouts to improve training efficiency.
Method
Context
Uncommon
Overall
P@0.5
mIoU
P@0.5
mIoU
P@0.5
mIoU
Qwen-VL ( Bai et al., 2023 )
32.1
0.350
26.1
0.298
29.1
0.324
MiniGPT-v2 ( Chen et al., 2023c )
46.0
0.429
40.9
0.373
43.4
0.401
Reason-to-Ground ( Sun et al., 2025 )
49.9
0.445
44.7
0.396
47.3
0.421
Qwen2.5-VL-3B-Instruct ( Bai et al., 2025a )
64.8
0.581
56.1
0.483
60.4
0.532
Qwen2.5-VL-7B-Instruct ( Bai et al., 2025a )
67.2
0.594
61.8
0.514
64.5
0.554
Table 1: EgoIntention results split by Context / Uncommon and Overall. Metrics are P@0.5 ( ↑ ) and mIoU ( ↑ ). Numbers in parentheses give the P@0.5 gain of Rita over the Qwen2.5-VL-Instruct baseline of the same size.
RFT Rewards
Context Split (%)
Uncommon Split (%)
RIoU
Rthink
Rcons
Acc think
Acc align
Acc think
Acc align
✓
83.7
61.4
71.5
54.4
✓
✓
83.1
61.2
70.9
54.0
✓
✓
85.1
62.3
73.7
55.6
✓
✓
✓
86.6
63.7
75.0
56.5
Table 2: Effectiveness of reward components on reasoning performance. We evaluate Acc think and Acc align (%) across Context and Uncommon splits. RIoU , Rthink , and Rcons represent grounding, thinking, and consistency rewards used during RFT. Best results are in bold .
Figure 3: Grounding accuracy (Precision@0.5) on the EgoIntention dataset for five data filtering strategies across training update steps.
Figure 4: Qualitative comparison of the IoU-reward baseline (pink) and Rita (green), with key reasoning phrases underlined . White boxes show overlapping predictions in the first three examples; colored boxes show differing predictions in the last example.
Figure 5: Zero-shot generalization across source RL training checkpoints on RefEgo-Int. We evaluate (a) visual grounding precision (P@0.5) and (b) intention reasoning accuracy using model checkpoints saved at different steps during RL training on the EgoIntention source dataset.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Screenshot of the VGG Image Annotator interface used to verify intention sentences and correct bounding boxes.
Camera
Referred objects
Target visible
RefEgo
RefEgo-Int
Moved
Multiple
No
1264 (2.00%)
27 (2.05%)
Moved
Multiple
Yes
3796 (6.01%)
85 (6.45%)
Moved
Unique
No
1290 (2.04%)
26 (1.97%)
Moved
Unique
Yes
4470 (7.08%)
79 (6.00%)
Stationary
Multiple
No
6858 (10.87%)
154 (11.69%)
Stationary
Multiple
Yes
20762 (32.89%)
489 (37.13%)
Appendix
Table 3: Challenge distribution of RefEgo (annotation level) and RefEgo-Int (one frame per clip) over camera motion, number of referred objects, and target visibility.
Bin
#Samples
E[ncorrect]
E[μIoU]
E[σIoU2]
E[Frac(IoU=0)]
Easy
9757
7.583
0.851
0.027
0.023
Medium
1988
4.118
0.479
0.133
0.240
Hard
1536
1.411
0.213
0.078
0.452
XHard
2346
0.000
0.054
0.004
0.676
Appendix
Table 4: IoU rollout statistics across difficulty bins. We bin samples by the number of correct rollouts ( ncorrect out of 8; IoU ≥0.5 ). For each bin, we report the number of samples, the mean ncorrect , the mean of per-sample mean IoU, the mean of per-sample IoU variance, and the mean fraction of zero-IoU rollouts (averaged per sample).
Before RL (CoT cold start)
After IoU-only RL (step 500)
P@0.5 (%), Context / Uncommon / Overall
64.90 / 57.82 / 61.36
65.86 / 59.60 / 62.73
Drift (% of samples), Context / Uncommon / Overall
3.54 / 4.48 / 4.01
5.04 / 4.92 / 4.98
Drift among correct boxes (%), Overall
6.54
7.94
Appendix
Table 5: Thinking drift of UniVG-R1 before and after IoU-only RL on EgoIntention (5,000 test samples per split). Drift is the percentage of samples whose box is correct (IoU ≥0.5 ) but whose predicted category is wrong; the last row normalizes drift by the number of correct boxes.
Step
Context
Uncommon
δ=0.1
δ=0.5
δ=0.1
δ=0.5
100
66.3 / 83.8 / 61.2
66.0 / 84.6 / 61.5
60.0 / 71.1 / 53.1
60.9 / 73.3 / 54.9
200
66.6 / 84.1 / 61.6
66.6 / 84.4 / 61.7
61.2 / 73.3 / 55.2
60.5 / 72.7 / 54.3
300
66.8 / 84.0 / 61.7
67.4 / 86.2 / 62.8
61.2 / 72.7 / 54.8
62.2 / 74.2 / 55.7
400
67.6 / 86.7 / 63.6
67.4 / 85.7 / 62.5
62.4 / 75.2 / 56.8
62.2 / 74.1 / 55.9
500
67.9 / 86.6 / 63.7
67.4 / 85.6 / 62.6
62.4 / 75.0 / 56.5
62.4 / 74.5 / 56.0
Appendix
Table 6: Ablation on the IoU margin δ of the consistency reward. We report P@0.5 / Acc think / Acc align (%) on the EgoIntention Context and Uncommon splits at different RL update steps. Both runs use Qwen2.5-VL-3B-Instruct and the full Rita reward; only δ differs. The main paper uses δ=0.1 .
Method
100
200
300
400
500
Base RFT ( RIoU+Rfmt )
82.2
81.0
81.9
83.0
83.7
+ Rthink w/o reference-answer overwrite
83.7
83.1
83.7
82.2
83.1
+ Rthink w/ reference-answer overwrite
83.8
83.6
83.7
84.2
85.1
Appendix
Table 7: Ablation on the effect of reference-answer overwrite for thinking reward. We report Acc think (%) on the EgoIntention Context split at different RL update steps. Base RFT uses only the IoU and format rewards. Without overwrite, the thinking reward does not improve over Base RFT.
Reinforcement Learning (RL) is an important paradigm for improving the reasoning capabilities of Vision-Language Models (VLMs). However, directly applying RL to rollout multimodal reasoning can lead to instability, due to the exploitation of language priors, the neglect of visual evidence, and the generation of reasoning traces that are fluent yet not visually grounded. The question arises: Can initially steer the policy toward visually faithful reasoning regime before applying reinforcement learning? To this end, we propose a Faithful Warm-Start (FWS) strategy that first curates samples with explicit vision-language causal relationships from six general VQA benchmarks to construct the FaithfulQA dataset, where each of the image-question pairs gains a certain degree of visual observations, question requirements, commonsense knowledge, domain knowledge, and the final answer. Subsequently, a VLM-based judge is employed to further purify the dataset, ensuring strong causal consistency and visual faithfulness. This warm-start stage equips the model with the capability to understand causally grounded vision-language patterns before subsequent RL optimization under sparse answer-level rewards. Experimental results show that such faithful supervision improves answer accuracy, stabilizes RL training, and reduces visually unsupported reasoning.
Peng, Lee, Yin Zhang +9
GMLab · Hong Kong Polytechnic University · XPENG Robotics
Vision-Language Models (VLMs) frequently suffer from visual perception errors and hallucinations that compromise answer accuracy in complex reasoning tasks. Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising solution by optimizing policies using answer correctness signals. Despite their effectiveness, prevailing RLVR methods face two critical limitations. First, much of the sampling budget is wasted on trajectories doomed to fail due to early visual description errors. Second, sparse rewards cannot distinguish whether failures stem from visual perception or reasoning stages. We introduce MIRL, a decoupled framework that addresses both limitations by leveraging mutual information (MI) between generated descriptions and visual inputs as a cheap pre-screening signal. This enables intelligent budget allocation toward high-potential trajectories via forking, while decoupled training provides independent MI-based rewards for visual perception optimization, resolving reward blindness. Experiments on six vision-language reasoning benchmarks demonstrate that MIRL achieves 70.22% average accuracy and successfully surpasses the performance of sampling 16 complete trajectories using only 10 pre-samples with top-6 selection (25% fewer complete trajectories). Our code is available at: https://anonymous.4open.science/r/mirl-main/.
Yin Zhang, Jiaxuan Zhao, Zonghan Wu +5
School of Mathematics, Tianjin University, Tianjin, China · Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China · Shanghai Advanced Institute of Finance (SAIFS), East China Normal University, Shanghai, China +4
Reinforcement learning with verifiable rewards (RLVR) has successfully elicited the reasoning capabilities of large language models, motivating its extension to multimodal scenarios. Existing methods primarily focus on improving the visual coverage of reasoning traces and mitigating visual hallucinations, but underestimate the semantic inconsistency between the reasoning process and the final answer. In this paper, we delve into thinking-answer inconsistency in RLVR for large vision-language models (LVLMs), showing thorough analyses of rollouts collected throughout Group Relative Policy Optimization (GRPO) training process and post-RLVR evaluation outputs that this issue persists during training and remains present during inference. Motivated by the analysis, we propose Consistency-Oriented Reasoning Alignment (CORA), which introduces thinking-answer semantic consistency into RLVR through a lightweight plug-and-play consistency reward model, and further incorporates Hybrid Reward Advantage Splitting (HRAS) to stably coordinate task and consistency optimization. Extensive experiments across representative multimodal reasoning benchmarks and mainstream LVLMs show that CORA improves task performance while effectively mitigating thinking-answer inconsistency, leading to more faithful reasoning traces.
Jiayue Cao, Zhicong Lu, Xuehan Sun +6
University of Chinese Academy of Sciences · 2Wuhan University · 3Tsinghua University +1