cs.LGOct 1, 2026

Same Reward, Different Skills: When Multimodal RL Learns to Look

Authors: Haocun Ye, Xinlong Jiang, Qile Chen, Bingyu Wang, Teng Zhang, Shubai Chen, Tingyu Wu, Zhenkun Zheng, +1 more

Organizations: University of the Chinese Academy of Sciences · Institute of Computing Technology, Chinese Academy of Sciences · Independent Researcher

Abstract

Reinforcement learning with verifiable rewards (RLVR) improves vision-language benchmark scores even without visual information during training. With images at test, blind-trained models recover roughly half of the real-image gain at 3B and nearly four fifths at 7B. Prolonged real-image training can erode grounding while benchmark gains persist. Both findings expose the same gap: an image in the prompt is not an image in the learning signal. Our design rule, visual resolvability, asks that visual evidence be necessary for a correct answer and that the task remain learnable. We test it on counterfactual coordinate scenes in which the question stays fixed and the target is never named, so a correct answer requires finding the target in the image. With standard GRPO and correctness-and-format rewards, a 7B model raises its accuracy at finding the target (discovery) from 0.425 to 0.875 on held-out scenes denser than any it trained on, and it improves on question types it never trained on. Two controls locate the source of the gain. Replacing test images with gray canvases drops discovery to zero; training on gray canvases instead, at matched step 30 and in each of four seeds, yields essentially none of the gain even when the model is then tested with real images. The learned skill carries over to grounding tasks built independently of the training corpus. A caption that answers the training question, added to the same images, reward and budget, cuts the gain by nearly two thirds. Changing what reward requires changes what RL learns.

Figures & tables

Appendix figures & tables33 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Attend to Evidence: Evidence-Anchored Spatial Attention Supervision for Multimodal RLVR

    May 29, 2026Ruina Hu, Chen Wang, Lai Wei +5Multimodal Reasoning BenchmarksSpatial Supervision

  2. Improving Generalization Robustness of Multimodal RLVR

    Aug 9, 2026Pengfei Zhou, Zhiwei Tang, Xiaopeng Peng +11Reinforcement Learning With Verifiable RewardVerifiable Rewards

  3. MIRL: Mutual Information-Guided Reinforcement Learning for Vision-Language Models

    May 2, 2026Yin Zhang, Jiaxuan Zhao, Zonghan Wu +5Recent Vision-Language ModelsReinforcement Learning With Verifiable Reward