cs.CVSep 30, 2026

Aligning Thoughts with Answers: Probability Rewards to Tame Thinking Drift

Authors: Pengzhan Sun, Shiu-hong Kao, Shijie Li, Yongyi Su, Junbin Xiao, Arjun Reddy Akula, Angela Yao

Organizations: National University of Singapore · A*STAR Institute of Advanced Intelligence and Computing, Singapore · South China University of Technology · University of Science and Technology of China · Google DeepMind

Abstract

This paper studies \textbf{thinking--answer consistency} in vision-language models. We focus on Visual Intention Grounding, where a model infers a target object based on a human intention query and predicts a bounding box. We reveal that previous IoU-based reinforcement learning (RL) frameworks suffer from ``thinking drift'', where the model produces a correct bounding box, despite having an incorrect reasoning process pointing to a different target object. Thus, we propose \textbf{Rita} (\textit{ReInforcing Thinking--Answer consistency}) as a novel RL paradigm to tame the drift. Specifically, Rita introduces two reasoning-label-free RL rewards, constructed from the conditional probability of reference answers: a \textbf{thinking reward} and a \textbf{consistency reward}. It also adopts a difficulty-aware \textbf{data filtering} strategy that selects informative easy-to-medium samples for RL using rollout error rate and reward variance. Extensive experiments on EgoIntention and the new RefEgo-Int benchmarks show that Rita performs consistently superior to the supervised finetuning approaches and vanilla RL-finetuned frameworks.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MIRL: Mutual Information-Guided Reinforcement Learning for Vision-Language Models

    May 2, 2026Yin Zhang, Jiaxuan Zhao, Zonghan Wu +5Recent Vision-Language ModelsReinforcement Learning With Verifiable Reward

  2. CORA: Analyzing and bridging thinking-answer gap in Multimodal RLVR via Consistency-Oriented Reasoning Alignment

    Jun 12, 2026Jiayue Cao, Zhicong Lu, Xuehan Sun +6Multimodal Reasoning BenchmarksReinforcement Learning With Verifiable Reward