Reinforcement Learning with Verifiable Rewards (RLVR) has improved the reasoning capabilities of Multimodal Large Language Models (MLLMs), yet existing frameworks rely on coarse, sequence-level reward signals that lack the fine-grained supervision over the visually-grounded steps within a multimodal reasoning chain. We investigate this gap through the lens of two token-level metrics: visual dependency (i.e. how much a token's prediction relies on the input image features) and predictive entropy. Our empirical analysis reveals two key findings: (1) correct reasoning chains exhibit a markedly sharper entropy reduction as visual grounding intensifies, compared to incorrect ones; (2) pivotal tokens, those whose misprediction triggers reasoning collapse, are statistical outliers in the joint distribution of visual dependency and predictive entropy derived from correct chains. Motivated by these findings, we propose token-level perception-grounded advantage estimation (TPAE), which estimates token-level advantages by measuring each token's statistical consistency with the vision-entropy patterns of correct rollouts. TPAE leverages this granular score to modulate the sequence-level advantage, producing a fine-grained supervision signal that can be integrated into various RLVR frameworks. Extensive experiments on seven benchmarks show that TPAE consistently outperforms leading strong baselines, yielding more stable and efficient optimization for multimodal reasoning. The code is publicly available at https://github.com/Zhihan72/TPAE.
Figures & tables
Figure 1. Token-level joint distribution of visual dependency and predictive entropy, where outlier tokens from incorrect trajectories emerge as the triggers of reasoning collapse.
Figure 2. Empirical analysis of reasoning trajectories using Qwen2.5-VL-7B. (a) Distribution of average entropy and log-frequency across visual dependency levels (left). (b) Density of token-level vision-entropy misalignment relative to the reference distribution center (upper right). (c) Error analysis of divergent tokens through human annotation (lower right).
Figure 3. Overview of the TPAE algorithm. TPAE utilizes correct rollouts to establish reference distributions of vision-entropy states of unique tokens. From a hypothesis testing perspective, TPAE then applies a trust boundary τ to quantify token-level trustworthiness based on distributional alignment. This metric is then integrated with the rollout-level advantage from GRPO.
Models
Mathematical & Geometric
Logical
General
Avg.
MathVerse
We-Math
MathVision
DynaMath
Geo3k
LogicVista
MMMU-Pro
ThinkLite-VL-7B
42.50
65.17
27.34
47.44
37.77
39.15
28.53
41.13
OpenVLThinker-7B
41.39
65.63
25.90
51.41
38.60
43.85
32.33
42.73
NoisyRollout-7B
39.06
63.76
24.15
53.86
42.80
46.25
35.13
43.57
MM-Eureka-7B
44.56
64.18
27.89
49.80
39.62
47.32
32.11
43.64
Perception-R1-7B
42.99
69.05
25.42
52.84
45.19
44.32
34.36
44.88
Table 1. Main results (avg 8 acc %) across seven multimodal reasoning benchmarks. All evaluations utilize exact-match scoring on verifiable instances to ensure objective results, avoiding any LLM-as-a-judge. All comparative baselines are instantiated from the Qwen2.5-VL-7B backbone. Best performance of 7B models is marked with bold, second best with underline.
Models
MathVerse
We-Math
MathVision
DynaMath
Geo3k
LogicVista
MMMU-Pro
Avg.
Qwen2.5-VL-7B
28.53
46.13
18.51
45.20
36.23
42.51
25.64
34.68
+ GRPO w/o TPAE
42.42
65.88
27.27
51.12
42.12
43.71
34.53
43.86
+ GRPO w/ TPAE
45.50
69.07
30.02
56.59
46.73
47.20
38.27
47.63
+ DAPO w/o TPAE
44.22
66.18
27.14
53.12
43.11
45.86
35.19
44.97
+ DAPO w/ TPAE
47.15
71.70
30.56
57.39
48.14
48.55
37.45
48.71
Qwen3-VL-8B
40.98
65.58
27.15
60.40
53.83
51.48
33.86
47.61
Table 2. Performance (avg 8 acc %) comparison of multimodal reinforcement learning with ( w/ ) and without ( w/o ) TPAE.
Figure 4. Comparison of RLVR training dynamics on accuracy rewards. Solid lines denote running averages with a window size of 20. TPAE achieves consistently faster learning on both GRPO and DAPO across two model backbones.
Configuration
MathVerse
We-Math
MathVision
DynaMath
Geo3k
LogicVista
MMMU-Pro
Avg.
τ=2.45
45.15
67.99
29.30
55.72
45.74
46.76
38.09
46.96
τ=3.26
45.42
68.84
29.80
55.83
46.80
46.84
37.87
47.34
τ=3.03
45.50
69.07
30.02
56.59
46.73
47.20
38.27
47.63
Table 3. Ablation study on the trust boundary τ based on TPAE-G-Qwen2.5-7B.
Figure 5. Qualitative analysis of token-level trustworthiness in a failed reasoning trajectory. Dark red highlights successfully localize the pivotal triggers of multimodal reasoning collapse, demonstrating the effectiveness of our TPAE algorithm.
Reinforcement Learning with Verifiable Rewards (RLVR) has become an effective paradigm for improving the reasoning capability of Large Vision-Language Models (LVLMs). However, existing RLVR methods primarily rely on trajectory-level outcome rewards, which assign identical learning signals across all generated tokens. This coarse-grained credit assignment is fundamentally mismatched to multimodal reasoning, where only a sparse subset of tokens is causally grounded in visual evidence. Consequently, these pivotal perceptual tokens receive weak supervision and are often overwhelmed by language priors or reasoning-template tokens. To address this limitation, we propose Perception-Reinforced Policy Optimization (PRPO), a token-level reinforcement learning framework that explicitly identifies and reinforces pivotal perceptual tokens within long-horizon multimodal reasoning trajectories. PRPO introduces Robust Visual Dependency (RVD), a principled metric that identifies tokens whose predictions are both visually grounded and perturbation-stable, filtering out brittle or noisy visual tokens. Based on RVD, we further propose Perceptual Advantage Reshaping (PAR), a token-level credit assignment technique that amplifies perceptually informative tokens while preserving stable gradients for non-perceptual tokens. Extensive experiments on seven multimodal reasoning benchmarks demonstrate that PRPO consistently outperforms strong LVLM baselines across both 3B and 7B model scales, achieving average gains of 23.3% and 21.1%, respectively. PRPO achieves state-of-the-art performance with improved training efficiency and stronger cross-task generalization. Our findings highlight the importance of fine-grained credit assignment for scalable multimodal reinforcement learning.
Reinforcement learning from verifiable rewards (RLVR), especially with Group Relative Policy Optimization (GRPO), has shown strong potential for improving the reasoning capabilities of large vision-language models (LVLMs). However, in multimodal reasoning, final-answer rewards are typically assigned at the sequence level and do not distinguish the functional roles of different tokens, making it difficult to determine whether a correct answer is supported by task-relevant visual evidence. In this paper, we revisit multimodal RLVR from the perspective of role-aware token-level credit assignment, where structured responses are decomposed into perception tokens for extracting visual evidence and reasoning tokens for deriving answers from that evidence. Based on this perspective, we propose Structured Role-aware Policy Optimization (SRPO), which refines the sequence-level GRPO advantage into role-aware token-level advantages without changing the reward function. Specifically, SRPO assigns role-specific credit by using self-distilled on-policy contrasts: perception tokens are emphasized according to their visual dependency under original versus corrupted visual inputs, while reasoning tokens are emphasized according to their consistency with the generated perception. These role-specific signals are further unified through a shared trajectory-level baseline, yielding positive token weights that adjust relative update magnitudes while preserving the original GRPO reward and optimization direction, without requiring external reward models or separate teachers. Experiments across diverse multimodal reasoning benchmarks show that SRPO improves evidence-grounded reasoning, highlighting the importance of moving beyond uniform sequence-level credit toward role-aware optimization for reliable multimodal reasoning.
Bingqing Jiang, Difan Zou
School of Computing & Data Science, The University of Hong Kong · School of Computing & Data Science and Institute of Data Science, The University of Hong Kong
While token-level entropy is commonly recognized as effective for credit assignment in text-only reinforcement learning with verifiable rewards (RLVR), it remains unclear whether this mechanism still holds in visual reasoning. Our controlled study shows that this mechanism collapses in visual reasoning due to the omission of vision-sensitive tokens with naturally low entropy. Although existing multimodal RL methods increasingly acknowledge the importance of visual perception, they struggle to satisfy the inherent demand for interleaving precise perceptual grounding with semantic reasoning, either lacking systematic visual measurements or overlooking that token entropy primarily drives semantic exploration. To address this, we introduce VEPO (Vision-Entropy token-selection for Policy Optimization), an effective RL framework explicitly integrating visual sensitivity with token entropy via a principled multiplicative coupling, where VEPO redirects gradient credit toward tokens which are simultaneously visually grounded and highly informative. Extensive experiments demonstrate VEPO's leading performance, significantly outperforming the entropy-only baseline by 2.28 points at 7B-scale and 3.15 points at 3B-scale. Ablations further substantiate the soundness of our method.
Senjie Jin, Peixin Wang, Boyang Liu +8
College of Computer Science and Artificial Intelligence, Fudan University