cs.AISep 29, 2026

Visual sensitivity is not claim retractability: persistence-aware credit assignment for multimodal reinforcement learning

Authors: Zhongan Bi, Kepeng Lin, Xuanang Gao, Yuhan Sun, Lianrun Zhang

Organizations: Zhejiang University · Huazhong University of Science and Technology · Shanghai Jiao Tong University

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) has been extended to Large Vision-Language Models (LVLMs), and perception-aware methods further encourage policies to rely on visual evidence. Yet relying on the image does not guarantee that visual claims are supported by it. Before RL training, 27.81% of the correctly answered responses of Qwen2.5-VL-7B on four multimodal reasoning benchmarks contain at least one direct visual claim that the image does not support. Since outcome-level RL rewards each response as a whole, these claims inherit the positive credit of the correct answer. We introduce a fixed-rollout counterfactual diagnostic that re-scores the same response under an intervened image to separate Evidence-Function Sensitivity (EFS), how strongly the model's predictions change, from claim persistence, whether the model keeps supporting the same claim rather than retracting it. The diagnostic reveals Sensitivity-Persistence Decoupling (SPD): under DAPO and VPPO, EFS increases and claims become more retractable overall, yet unsupported claims become significantly more persistent, whereas GRPO raises EFS without this deterioration. We therefore propose Persistence-Aware Credit Gating (PACG), which attenuates positive credit for unusually persistent visual claims and leaves all other credit unchanged. It requires no supported/unsupported labels and adds no inference cost. On Qwen2.5-VL-7B, PACG raises the nine-benchmark average over three seeds from 58.1% to 59.9% with DAPO and from 59.8% to 60.9% with VPPO, while making unsupported claims more retractable. The gains extend to a larger model, a newer backbone, and the accuracy of HallusionBench also improves consistently. These results suggest that visual sensitivity and claim retractability are complementary dimensions of multimodal credit assignment.

Figures & tables

Appendix figures & tables23 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Same Reward, Different Skills: When Multimodal RL Learns to Look

    Oct 1, 2026Haocun Ye, Xinlong Jiang, Qile Chen +6Multimodal Reasoning BenchmarksVisual Evidence

  2. Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence

    Sep 28, 2026Xi Xiao, Tianchen Zhao, Youngeun Kim +10Latent Visual ReasoningVisual Reasoning

  3. Attend to Evidence: Evidence-Anchored Spatial Attention Supervision for Multimodal RLVR

    May 29, 2026Ruina Hu, Chen Wang, Lai Wei +5Multimodal Reasoning BenchmarksSpatial Supervision