Aligning Thoughts with Answers: Probability Rewards to Tame Thinking Drift
Authors: Pengzhan Sun, Shiu-hong Kao, Shijie Li, Yongyi Su, Junbin Xiao, Arjun Reddy Akula, Angela Yao
Organizations: National University of Singapore · A*STAR Institute of Advanced Intelligence and Computing, Singapore · South China University of Technology · University of Science and Technology of China · Google DeepMind
This paper studies \textbf{thinking--answer consistency} in vision-language models. We focus on Visual Intention Grounding, where a model infers a target object based on a human intention query and predicts a bounding box. We reveal that previous IoU-based reinforcement learning (RL) frameworks suffer from ``thinking drift'', where the model produces a correct bounding box, despite having an incorrect reasoning process pointing to a different target object. Thus, we propose \textbf{Rita} (\textit{ReInforcing Thinking--Answer consistency}) as a novel RL paradigm to tame the drift. Specifically, Rita introduces two reasoning-label-free RL rewards, constructed from the conditional probability of reference answers: a \textbf{thinking reward} and a \textbf{consistency reward}. It also adopts a difficulty-aware \textbf{data filtering} strategy that selects informative easy-to-medium samples for RL using rollout error rate and reward variance. Extensive experiments on EgoIntention and the new RefEgo-Int benchmarks show that Rita performs consistently superior to the supervised finetuning approaches and vanilla RL-finetuned frameworks.
Figures & tables
Figure 1: Base RFT with IoU reward alone fails to enforce semantic consistency. In the Base RFT ( red ), which relies solely on the IoU reward, the model predicts a reasonable bounding box but provides irrelevant reasoning, leading to a label error. In contrast, in our Rita ( green ), the integration of thinking and consistency rewards encourages consistency between the reasoning process and the final prediction, correcting the label while maintaining spatial accuracy.
Figure 2: Overview of the Rita framework. Given an input image and human intention query, the VLM samples multiple think-then-answer rollouts (Q,Tk,Ak) , from which we derive a grounding reward via IoU with the ground-truth box and a format reward. We then overwrite Ak with the reference answer A⋆ and run the VLM in teacher-forced mode to obtain logits for the answer tokens, which define the thinking reward Rthink . The consistency reward Rcons scales this score by the predicted box’s IoU minus a margin δ . During the GRPO-based optimization, we conduct difficulty-and-variance-aware data sampling computed from IoU rollouts to improve training efficiency.
Method
Context
Uncommon
Overall
P@0.5
mIoU
P@0.5
mIoU
P@0.5
mIoU
Qwen-VL ( Bai et al., 2023 )
32.1
0.350
26.1
0.298
29.1
0.324
MiniGPT-v2 ( Chen et al., 2023c )
46.0
0.429
40.9
0.373
43.4
0.401
Reason-to-Ground ( Sun et al., 2025 )
49.9
0.445
44.7
0.396
47.3
0.421
Qwen2.5-VL-3B-Instruct ( Bai et al., 2025a )
64.8
0.581
56.1
0.483
60.4
0.532
Qwen2.5-VL-7B-Instruct ( Bai et al., 2025a )
67.2
0.594
61.8
0.514
64.5
0.554
Table 1: EgoIntention results split by Context / Uncommon and Overall. Metrics are P@0.5 ( ↑ ) and mIoU ( ↑ ). Numbers in parentheses give the P@0.5 gain of Rita over the Qwen2.5-VL-Instruct baseline of the same size.
RFT Rewards
Context Split (%)
Uncommon Split (%)
RIoU
Rthink
Rcons
Acc think
Acc align
Acc think
Acc align
✓
83.7
61.4
71.5
54.4
✓
✓
83.1
61.2
70.9
54.0
✓
✓
85.1
62.3
73.7
55.6
✓
✓
✓
86.6
63.7
75.0
56.5
Table 2: Effectiveness of reward components on reasoning performance. We evaluate Acc think and Acc align (%) across Context and Uncommon splits. RIoU , Rthink , and Rcons represent grounding, thinking, and consistency rewards used during RFT. Best results are in bold .
Figure 3: Grounding accuracy (Precision@0.5) on the EgoIntention dataset for five data filtering strategies across training update steps.
Figure 4: Qualitative comparison of the IoU-reward baseline (pink) and Rita (green), with key reasoning phrases underlined . White boxes show overlapping predictions in the first three examples; colored boxes show differing predictions in the last example.
Figure 5: Zero-shot generalization across source RL training checkpoints on RefEgo-Int. We evaluate (a) visual grounding precision (P@0.5) and (b) intention reasoning accuracy using model checkpoints saved at different steps during RL training on the EgoIntention source dataset.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Screenshot of the VGG Image Annotator interface used to verify intention sentences and correct bounding boxes.
Camera
Referred objects
Target visible
RefEgo
RefEgo-Int
Moved
Multiple
No
1264 (2.00%)
27 (2.05%)
Moved
Multiple
Yes
3796 (6.01%)
85 (6.45%)
Moved
Unique
No
1290 (2.04%)
26 (1.97%)
Moved
Unique
Yes
4470 (7.08%)
79 (6.00%)
Stationary
Multiple
No
6858 (10.87%)
154 (11.69%)
Stationary
Multiple
Yes
20762 (32.89%)
489 (37.13%)
Appendix
Table 3: Challenge distribution of RefEgo (annotation level) and RefEgo-Int (one frame per clip) over camera motion, number of referred objects, and target visibility.
Bin
#Samples
E[ncorrect]
E[μIoU]
E[σIoU2]
E[Frac(IoU=0)]
Easy
9757
7.583
0.851
0.027
0.023
Medium
1988
4.118
0.479
0.133
0.240
Hard
1536
1.411
0.213
0.078
0.452
XHard
2346
0.000
0.054
0.004
0.676
Appendix
Table 4: IoU rollout statistics across difficulty bins. We bin samples by the number of correct rollouts ( ncorrect out of 8; IoU ≥0.5 ). For each bin, we report the number of samples, the mean ncorrect , the mean of per-sample mean IoU, the mean of per-sample IoU variance, and the mean fraction of zero-IoU rollouts (averaged per sample).
Before RL (CoT cold start)
After IoU-only RL (step 500)
P@0.5 (%), Context / Uncommon / Overall
64.90 / 57.82 / 61.36
65.86 / 59.60 / 62.73
Drift (% of samples), Context / Uncommon / Overall
3.54 / 4.48 / 4.01
5.04 / 4.92 / 4.98
Drift among correct boxes (%), Overall
6.54
7.94
Appendix
Table 5: Thinking drift of UniVG-R1 before and after IoU-only RL on EgoIntention (5,000 test samples per split). Drift is the percentage of samples whose box is correct (IoU ≥0.5 ) but whose predicted category is wrong; the last row normalizes drift by the number of correct boxes.
Step
Context
Uncommon
δ=0.1
δ=0.5
δ=0.1
δ=0.5
100
66.3 / 83.8 / 61.2
66.0 / 84.6 / 61.5
60.0 / 71.1 / 53.1
60.9 / 73.3 / 54.9
200
66.6 / 84.1 / 61.6
66.6 / 84.4 / 61.7
61.2 / 73.3 / 55.2
60.5 / 72.7 / 54.3
300
66.8 / 84.0 / 61.7
67.4 / 86.2 / 62.8
61.2 / 72.7 / 54.8
62.2 / 74.2 / 55.7
400
67.6 / 86.7 / 63.6
67.4 / 85.7 / 62.5
62.4 / 75.2 / 56.8
62.2 / 74.1 / 55.9
500
67.9 / 86.6 / 63.7
67.4 / 85.6 / 62.6
62.4 / 75.0 / 56.5
62.4 / 74.5 / 56.0
Appendix
Table 6: Ablation on the IoU margin δ of the consistency reward. We report P@0.5 / Acc think / Acc align (%) on the EgoIntention Context and Uncommon splits at different RL update steps. Both runs use Qwen2.5-VL-3B-Instruct and the full Rita reward; only δ differs. The main paper uses δ=0.1 .
Method
100
200
300
400
500
Base RFT ( RIoU+Rfmt )
82.2
81.0
81.9
83.0
83.7
+ Rthink w/o reference-answer overwrite
83.7
83.1
83.7
82.2
83.1
+ Rthink w/ reference-answer overwrite
83.8
83.6
83.7
84.2
85.1
Appendix
Table 7: Ablation on the effect of reference-answer overwrite for thinking reward. We report Acc think (%) on the EgoIntention Context split at different RL update steps. Base RFT uses only the IoU and format rewards. Without overwrite, the thinking reward does not improve over Base RFT.
School of Mathematics, Tianjin University, Tianjin, China · Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China · Shanghai Advanced Institute of Finance (SAIFS), East China Normal University, Shanghai, China +4