Reinforcement Learning with Verifiable Rewards (RLVR) has been extended to Large Vision-Language Models (LVLMs), and perception-aware methods further encourage policies to rely on visual evidence. Yet relying on the image does not guarantee that visual claims are supported by it. Before RL training, 27.81% of the correctly answered responses of Qwen2.5-VL-7B on four multimodal reasoning benchmarks contain at least one direct visual claim that the image does not support. Since outcome-level RL rewards each response as a whole, these claims inherit the positive credit of the correct answer. We introduce a fixed-rollout counterfactual diagnostic that re-scores the same response under an intervened image to separate Evidence-Function Sensitivity (EFS), how strongly the model's predictions change, from claim persistence, whether the model keeps supporting the same claim rather than retracting it. The diagnostic reveals Sensitivity-Persistence Decoupling (SPD): under DAPO and VPPO, EFS increases and claims become more retractable overall, yet unsupported claims become significantly more persistent, whereas GRPO raises EFS without this deterioration. We therefore propose Persistence-Aware Credit Gating (PACG), which attenuates positive credit for unusually persistent visual claims and leaves all other credit unchanged. It requires no supported/unsupported labels and adds no inference cost. On Qwen2.5-VL-7B, PACG raises the nine-benchmark average over three seeds from 58.1% to 59.9% with DAPO and from 59.8% to 60.9% with VPPO, while making unsupported claims more retractable. The gains extend to a larger model, a newer backbone, and the accuracy of HallusionBench also improves consistently. These results suggest that visual sensitivity and claim retractability are complementary dimensions of multimodal credit assignment.
Figures & tables
Figure 1: A correct pre-RL answer can conceal an unsupported visual assertion. We fix this rewarded trace and compare its teacher-forced behavior across checkpoints, asking whether outcome RL makes the claim more sensitive to image intervention without making it more retractable.
Figure 2: Overview of PACG. A frozen router identifies direct visual claims in responses scored under the original and degraded images. The residual-corrected signed shift Δ(s) measures claim persistence and determines a soft gate on positive token credit. Negative advantages and non-routed tokens are unchanged, so the method only modulates positive credit for routed visual claims.
Backbone
Method
General Mathematical & Geometric Reasoning
Vision-Dependent Reasoning
Avg.
MMK12
Math Verse
Dyna Math
Math Vision
Geometry 3K
We- Math
Logic Vista
MMMU- Pro
Clever- Count
Qwen2.5-VL 7B
ThinkLite-VL
62.5
63.8
62.0
31.5
35.8
66.4
40.0
27.6
77.8
51.9
VL-Rethinker
69.3
68.8
65.7
29.5
40.7
68.5
45.8
39.7
82.0
56.7
NoisyRollout ∗
50.0
67.8
62.1
22.1
46.9
71.0
45.3
34.5
85.7
53.9
R1-ShareVL
70.9
68.2
63.9
27.5
41.2
69.9
45.4
35.1
81.2
55.9
MM-Eureka
67.5
65.4
64.8
27.1
40.8
65.5
45.8
35.3
77.8
54.4
Table 1: Main results across multimodal reasoning benchmarks and backbone families. All entries are accuracy (%). The first six datasets assess general mathematical and geometric reasoning, whereas the last three emphasize vision-dependent multimodal reasoning. Public 7B baselines are listed first, followed by our matched Qwen2.5-VL-7B runs. Bold and underlined values denote the best and second-best completed results within each backbone block. Additional blocks show transfer to a larger model scale and a newer backbone family.
Model
aAcc ↑
fAcc ↑
qAcc ↑
Base
65.19
38.15
33.41
GRPO
66.90
44.35
35.10
DAPO
67.17
43.77
35.32
VPPO
67.70
45.06
35.18
DAPO+PACG
67.64
44.64
35.54
VPPO+PACG
68.11
45.51
36.04
Table 2: HallusionBench (7B; %; single checkpoint per model).
Figure 3: Online accuracy reward. Dashed: parents; solid: PACG compositions; curves are trailing 10-step means for one representative seed per method.
Variant
MMK12
Math Verse
Dyna Math
Math Vision
Geometry 3K
We- Math
Logic Vista
MMMU- Pro
Clever- Count
Avg.
Unsup. Pers. ↓
DAPO
81.0
66.6
64.6
30.3
42.4
67.9
46.2
39.2
84.6
58.1
−0.1241
w/o Residual Corr.
80.9
69.0
65.6
30.9
45.4
69.1
48.9
40.1
81.3
59.0
−0.1278
w/o Direct Routing
46.3
40.8
36.7
16.2
10.4
44.0
31.7
30.3
85.2
38.0
−0.0374
w/o Positive-Only
80.8
70.3
66.1
31.3
44.5
70.6
47.7
39.6
82.6
59.3
−0.1301
PACG (full)
81.2
70.6
65.9
30.6
44.3
70.1
47.1
40.7
88.7
59.9
−0.1392
Table 3: Component ablations with DAPO as parent. w/o Residual Corr. : omit same-rollout residual calibration; w/o Direct Routing : gate all positive tokens instead of direct claims; w/o Positive-Only : also attenuate negative advantages. Accuracy and its nine-benchmark mean (Avg.) are percentages. Unsup. Pers. is the unsupported corrected persistence defined in Section 5.4 . Entries are point estimates; bold marks the best column value.
Figure 4: Aggregate improvement masks unsupported-claim deterioration. (a) Direct EFS ( ↑ ); (b) overall and (c) unsupported corrected persistence ( ↓ ). DAPO and VPPO improve aggregate retractability while worsening unsupported persistence; PACG mitigates this deterioration.
Model
Unsup. Pers. ↓
Img-vs-Text DID ↓
Acc. (%) ↑
Base
−0.1892
0.000
38.9
GRPO
−0.1963
−0.167
59.0
DAPO
−0.1241†
−0.022
62.2
VPPO
−0.1061†
−0.042
63.4
DAPO+PACG
−0.1392†
−0.016
63.0
VPPO+PACG
−0.1880†
−0.066
64.3
Table 5: Claim-level mechanism results. Accuracy (%) is evaluated on the CrossBench4-2000 questions using seed-0 checkpoints. DID is the Base-relative image-versus-text corrected-persistence contrast. † marks an unsupported-persistence change whose paired 95% question-clustered bootstrap CI excludes zero: versus Base for parent optimizers, and versus the corresponding parent for PACG. Full intervals are in Appendix E.2 .
Figure 5: Sensitivity–retractability phase trajectory. Direct EFS versus negative unsupported persistence on the fixed 500-question subset. Right: more sensitive, up: more retractable. Lines connect steps 0,50,100,150,200 ; dotted lines mark Base.
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Qwen2.5-VL-7B
Qwen2.5-VL-32B
Qwen3-VL-2B Thinking
Training corpus
ViRL39K
ViRL39K
ViRL39K
Training budget
2 epochs
2 epochs
200 steps
Optimizer
AdamW
AdamW
AdamW
Learning rate
10−6
10−6
10−6
Adam betas
(0.9,0.999)
(0.9,0.999)
(0.9,0.999)
Weight decay
0.01
0.01
0.01
Appendix
Table 6: Training configurations. All three backbones use an entropy penalty of 0.1; training budgets and response-length limits are backbone-specific.
Task
Precision
Recall
F1
κ
Direct visual claim
0.90
0.87
0.88
0.81
Direct vs. derived
0.85
0.83
0.84
0.76
Supported vs. unsupported
0.82
0.78
0.80
0.71
Human–human Cohen’s κ (same task order)
Direct visual claim
—
—
—
0.84
Direct vs. derived
—
—
—
0.79
Appendix
Table 7: Router and human audit on 500 CrossBench4 questions. Precision, recall, F1, and the first three κ values use human consensus as the reference.
Router
Precision
Recall
F1
Span Coverage
Timing (s)
Qwen3-4B
0.68
0.63
0.65
76%
108
Qwen3.8-27B
0.87
0.85
0.86
94%
425
Qwen3.6-35B-A3B
0.85
0.83
0.84
92%
165
Appendix
Table 8: Span-extractor comparison on 200 expert-annotated questions. Span coverage is the fraction of questions with at least one matched expert span. Timing is wall-clock seconds per batch on eight A100 GPUs; shading identifies the selected router.
Model
Direct Claim Rate (%)
Direct Claims / Response
Direct Token Ratio (%)
Response Length
Unsup. Claim Rate (%) ↓
Unsup. Resp. Rate (%) ↓
Accuracy (%)
Uncert./ Abst. (%)
Base
77.46 [76.08, 78.81]
2.463 [2.387, 2.537]
11.18 [10.83, 11.54]
506.62 [498.61, 514.84]
24.19 [23.14, 25.27]
32.19 [30.88, 33.48]
43.02 [41.56, 44.51]
10.2
GRPO
78.05 [76.60, 79.48]
2.382 [2.306, 2.459]
11.14 [10.75, 11.52]
526.64 [517.18, 535.97]
19.61 [18.53, 20.73]
26.61 [25.22, 27.99]
73.41 [71.76, 75.07]
9.5–10.5
DAPO
76.12 [74.66, 77.57]
2.148 [1.912, 2.314]
10.30 [9.94, 10.66]
516.82 [508.45, 525.13]
18.84 [17.79, 19.88]
25.83 [24.48, 27.17]
81.60 [79.87, 83.28]
9.5–10.5
DAPO + PACG
74.52 [73.07, 75.96]
2.211 [2.139, 2.282]
9.87 [9.54, 10.21]
528.46 [520.12, 536.96]
16.16 [14.11, 18.25]
25.62 [24.55, 27.10]
83.16 [81.56, 84.75]
9.0–11.0
VPPO
79.41 [78.08, 80.75]
2.645 [2.562, 2.733]
9.95 [9.62, 10.28]
681.46 [670.84, 692.14]
17.44 [16.51, 18.41]
25.61 [24.33, 26.87]
81.73 [80.16, 83.29]
9.5–10.5
VPPO + PACG
77.90 [76.46, 79.28]
2.463 [2.383, 2.543]
9.31 [8.99, 9.62]
697.47 [686.76, 708.45]
16.61 [15.69, 17.55]
24.12 [22.90, 25.37]
82.63 [81.06, 84.14]
9.0–11.0
Appendix
Table 9: Free-generation visual-claim audit on MMK12. Question-macro means with 95% question-cluster bootstrap intervals, using successfully annotated responses. Rates and accuracy are percentages; lengths are tokens. Uncertainty/abstention entries are point values or ranges, not confidence intervals.
Model
Median length (tokens)
Annotated responses
Direct claims
Base
462.0
16,000/16,000
39,401
GRPO
474.0
15,996/16,000
38,098
DAPO
480.0
15,999/16,000
34,360
DAPO + PACG
489.0
15,998/16,000
35,370
VPPO
635.0
15,990/16,000
42,295
VPPO + PACG
645.0
15,992/16,000
39,375
Appendix
Table 10: Response-length medians and annotation coverage. Coverage is out of 16,000 expected responses per model. Claim counts are exact high-confidence direct claims.
Model
Supported Persistence ↓
Base
−0.137
GRPO
−0.318
DAPO
−0.159
VPPO
−0.184
DAPO+PACG
−0.165
VPPO+PACG
−0.213
Appendix
Table 11: Supported-claim corrected persistence. Lower is better. Bold and underlining mark the best and second-best point estimates; blue shading identifies PACG variants.
Comparison
Metric
Effect
95% CI
DAPO − Base
Unsupported persistence
+0.0651
[0.022,0.108]
VPPO − Base
Unsupported persistence
+0.0831
[0.036,0.130]
DAPO+PACG − DAPO
Unsupported persistence
−0.0151
[−0.022,−0.008]
VPPO+PACG − VPPO
Unsupported persistence
−0.0819
[−0.115,−0.049]
DAPO − Base
Direct EFS
+0.0715
[0.025,0.118]
VPPO − Base
Direct EFS
+0.2330
[0.160,0.306]
Appendix
Table 12: Paired mechanism effects with question-clustered 95% bootstrap intervals. Each effect is the first model minus the second. Accuracy effects are on the [0,1] scale; other effects use their metric’s native scale.
Metric
GRPO
DAPO
VPPO
Positive routed-token rate, Ppos (%)
71.8
66.3
78.9
Effective positive update mass, Meff
0.35
0.93
1.00
Appendix
Table 13: Online positive-credit exposure over steps 181–200. Results concern routed direct visual claims in the parent training runs. Values are point estimates.
Figure 6: Sensitivity and unsupported-claim persistence diverge during training. DAPO and DAPO+PACG are evaluated at matched steps 0,50,100,150, and 200 . Direct EFS increases for both methods, whereas DAPO’s unsupported persistence becomes progressively less negative. PACG prevents the same drift and ends with greater unsupported retractability than DAPO. The temporal series uses 500 fixed questions sampled from CrossBench4-2000, with four Base rollouts each; subset averages can differ from full-set averages. Lines connect observed checkpoints rather than fitted curves.
Intervention
EFS ↑
Overall Pers. ↓
Unsup. Pers. ↓
PACG Effect ↓
Valid Rate
Random mask
0.31
−0.2558
−0.139
−0.015
96%
Coarse ink
0.28
−0.148
−0.125
−0.013
94%
Foreground/salient
0.33
−0.172
−0.151
−0.018
93%
Appendix
Table 14: Robustness to counterfactual corruption. Absolute DAPO+PACG scores and the matched DAPO+PACG minus DAPO change in unsupported persistence (PACG Effect). Entries are point estimates.
Figure 7: Three image interventions on the same example. Illustrative outputs of the evaluation operators using their default parameters on the original panel recovered from a saved MMK12 audit image. Random mask removes grid cells; coarse ink expands dark strokes; foreground masking removes their padded bounding rectangle. These are image-level controls, not verified claim-localized evidence removals. The illustration does not add observations to Table 14 .
Method
MMK12
Math Verse
Dyna Math
Math Vision
Geometry 3K
We-Math
Logic Vista
MMMU- Pro
Clever- Count
Avg.
GRPO
72.2
67.8
65.1
30.1
42.0
67.8
47.1
38.3
81.0
56.8
GRPO+PACG
73.5
68.1
65.0
30.2
42.3
68.8
47.3
39.0
81.7
57.3
Appendix
Table 15: PACG composed with plain GRPO. Accuracy (%); GRPO+PACG uses mean accuracy@8. Bold marks the higher displayed value in each column. Entries are point estimates.
Variant
Direct EFS ↑
Corr. Pers. ↓
Unsup. Pers. ↓
DAPO
0.2806
−0.2095
−0.1241
w/o Residual Corr.
0.4099
−0.2410
−0.1278
w/o Direct Routing
3.0366†
−0.1766
−0.0374
w/o Positive-Only
0.2046
−0.2224
−0.1301
PACG (full)
0.3096
−0.2558
−0.1392
Appendix
Table 16: Additional component mechanism metrics. Point estimates; lower persistence indicates stronger retraction. † denotes a degenerate EFS value associated with training collapse, not improved visual grounding.
Model
Avg. (%)
Δ vs. parent
p
DAPO
58.1±0.40
–
–
DAPO+PACG
59.9±0.52
+1.8
0.010
VPPO
59.8±0.40
–
–
VPPO+PACG
60.9±0.47
+1.1
0.038
Appendix
Table 17: Seed variance (Qwen2.5-VL-7B; nine-benchmark Avg., %; mean ± sample standard deviation over three training seeds). p : two-sided Welch’s t -test between each PACG composition and its parent ( n=3 per group); Δ is in percentage points.
m
α
gmin
MMK12 Accuracy (%) ↑
EFS ↑
Corrected Persistence ↓
Unsupported Persistence ↓
Positive-credit retention (%)
0.00
10
0.30
81.0
0.34
−0.10
−0.08
95
0.05
10
0.30
81.2
0.31
−0.23
−0.14
81
0.10
10
0.30
81.0
0.28
−0.28
−0.20
65
Appendix
Table 18: Sensitivity to the persistence margin. Gate strength α=10 and floor gmin=0.30 are fixed. Accuracy and retention are percentages. Shading marks the selected operating point; bold marks the best accuracy, EFS, and persistence values. Entries are point estimates.
Figure 8: PACG gate dynamics. (a) Mean span gate; (b) mean multiplier over positive direct-claim tokens; (c) fraction of these tokens downweighted; (d) fraction at the gate floor. Faint curves show observations and solid curves trailing 10-step means.
Figure 9: Online training-reward dynamics. Paired parent/PACG comparisons of reward/accuracy through step 200. Faint curves show logged values and solid curves trailing 10-step means.
Figure 10: Online VPPO visual-selector scores. Mean low-variance KL scores for (a) selected and (b) unselected tokens. The panels use different vertical scales.
Figure 11: Overlap between VPPO selection and PACG attenuation. (a) Fraction of VPPO-selected tokens downweighted by PACG; (b) fraction of PACG-downweighted tokens selected by VPPO. The two panels use different denominators.
Figure 12: Training dynamics across six runs. Accuracy reward, entropy, response length, clipping rate, gradient norm, and approximate policy-update KL. Dashed curves denote parents and solid curves their PACG compositions. Policy-update KL is distinct from reference-model KL and image-intervention EFS.
Figure 13: Online persistence and routing coverage. (a)–(c) Mean corrected delta, its 90th percentile, and margin-violation rate over eligible spans. (d)–(f) Routed-rollout rate, direct-token fraction, and intervention-valid rate. The 90th percentile describes the span distribution, not a confidence interval.
Stage
DAPO
PACG
Overhead vs. DAPO
Rollout
148.8
174.6
–
Original-image teacher forcing
46.7
49.2
–
Reference log-probability (KL)
44.9
–
–
Augmented scoring
–
51.2
–
Counterfactual teacher forcing
–
56.6
–
Blocking annotation wait
–
159.4
–
Appendix
Table 19: Training efficiency on Qwen2.5-VL-7B using 16 PPU-810E accelerators. Mean seconds per step over steps 1–200. Total time includes validation, checkpointing, and other work.
Figure 14: Training-time breakdown for PACG ( m=0.05 ). Mean seconds per step over steps 1–200. TF denotes teacher forcing. Augmented scoring corresponds to timing_s/aug .
Reinforcement learning with verifiable rewards (RLVR) improves vision-language benchmark scores even without visual information during training. With images at test, blind-trained models recover roughly half of the real-image gain at 3B and nearly four fifths at 7B. Prolonged real-image training can erode grounding while benchmark gains persist. Both findings expose the same gap: an image in the prompt is not an image in the learning signal. Our design rule, visual resolvability, asks that visual evidence be necessary for a correct answer and that the task remain learnable. We test it on counterfactual coordinate scenes in which the question stays fixed and the target is never named, so a correct answer requires finding the target in the image. With standard GRPO and correctness-and-format rewards, a 7B model raises its accuracy at finding the target (discovery) from 0.425 to 0.875 on held-out scenes denser than any it trained on, and it improves on question types it never trained on. Two controls locate the source of the gain. Replacing test images with gray canvases drops discovery to zero; training on gray canvases instead, at matched step 30 and in each of four seeds, yields essentially none of the gain even when the model is then tested with real images. The learned skill carries over to grounding tasks built independently of the training corpus. A caption that answers the training question, added to the same images, reward and budget, cuts the gain by nearly two thirds. Changing what reward requires changes what RL learns.
Haocun Ye, Xinlong Jiang, Qile Chen +6
University of the Chinese Academy of Sciences · Institute of Computing Technology, Chinese Academy of Sciences · Independent Researcher
Latent visual reasoning (LVR) enables multimodal large language models (MLLMs) to perform intermediate computation in continuous latent tokens rather than expressing every reasoning step in words. However, unlike textual CoT, latent reasoning is not directly observable, making it difficult to supervise what latent tokens learn. In this work, we first conduct a thorough analysis of latent-token behavior and identify a latent evidence-credit gap: latent tokens respond only weakly to image perturbations that alter the correct answer. We hypothesize that this issue stems from the lack of explicit supervision during GRPO training. These findings suggest that a final-answer reward provides too little guidance on what visual evidence to preserve or how credit should be assigned across latent tokens. To bridge this gap, we propose ReaLVR, which brings visual-evidence supervision to the model's own free-running latent trajectories. ReaLVR contrasts correct and model-generated wrong answers to determine where stronger supervision is needed, and relevant and mismatched visual evidence to specify what to preserve. Across three model families, ReaLVR consistently outperforms evaluated LVR baselines, achieving the highest five-task average of 63.7% on Qwen2.5-VL-7B. Crucially, we are the first to scale visual reasoning in latent space, showing that our framework continues to deliver robust improvements at frontier model scales up to 235B. Further analyses show more question-sensitive latent-token positions, stronger alignment with relevant visual regions, and greater fixed-context dependence on the most attended latent tokens.
Xi Xiao, Tianchen Zhao, Youngeun Kim +10
University of Alabama at Birmingham · Amazon AGI · Work done during an internship at Amazon AGI.
Reinforcement learning with verifiable rewards (RLVR) improves vision-language models (VLMs) by optimizing outcome rewards derived from final answers. However, such outcome-only rewards do not tell the model which image regions justify an answer. For questions that require visual grounding, these rewards cannot distinguish responses supported by relevant visual evidence from those produced by language-prior shortcuts or lucky guesses. We introduce EASE (Evidence-Anchored Spatial Attention), which augments multimodal RLVR with visual-evidence process supervision. EASE converts annotated evidence regions into a smoothed visual-token target and uses it to guide response-to-image attention during RL training, but only on high-reward trajectories. The annotations are used solely as privileged training labels, while inference requires only the original image and question. Across Qwen2.5-VL-7B, Qwen3-VL-4B, and Qwen3-VL-8B, EASE raises average scores over DAPO by 2.5 to 3.1 points on perception, hallucination, visual math, and multimodal reasoning benchmarks. Diagnostics and ablations show that EASE better aligns visual attention with annotated evidence regions.
Ruina Hu, Chen Wang, Lai Wei +5
1Harbin Institute of Technology · 2Zhongguancun Academy · 4Nankai University +4