Reinforcement Learning with Verifiable Rewards (RLVR) has been extended to Large Vision-Language Models (LVLMs), and perception-aware methods further encourage policies to rely on visual evidence. Yet relying on the image does not guarantee that visual claims are supported by it. Before RL training, 27.81% of the correctly answered responses of Qwen2.5-VL-7B on four multimodal reasoning benchmarks contain at least one direct visual claim that the image does not support. Since outcome-level RL rewards each response as a whole, these claims inherit the positive credit of the correct answer. We introduce a fixed-rollout counterfactual diagnostic that re-scores the same response under an intervened image to separate Evidence-Function Sensitivity (EFS), how strongly the model's predictions change, from claim persistence, whether the model keeps supporting the same claim rather than retracting it. The diagnostic reveals Sensitivity-Persistence Decoupling (SPD): under DAPO and VPPO, EFS increases and claims become more retractable overall, yet unsupported claims become significantly more persistent, whereas GRPO raises EFS without this deterioration. We therefore propose Persistence-Aware Credit Gating (PACG), which attenuates positive credit for unusually persistent visual claims and leaves all other credit unchanged. It requires no supported/unsupported labels and adds no inference cost. On Qwen2.5-VL-7B, PACG raises the nine-benchmark average over three seeds from 58.1% to 59.9% with DAPO and from 59.8% to 60.9% with VPPO, while making unsupported claims more retractable. The gains extend to a larger model, a newer backbone, and the accuracy of HallusionBench also improves consistently. These results suggest that visual sensitivity and claim retractability are complementary dimensions of multimodal credit assignment.
Figures & tables
Figure 1: A correct pre-RL answer can conceal an unsupported visual assertion. We fix this rewarded trace and compare its teacher-forced behavior across checkpoints, asking whether outcome RL makes the claim more sensitive to image intervention without making it more retractable.
Figure 2: Overview of PACG. A frozen router identifies direct visual claims in responses scored under the original and degraded images. The residual-corrected signed shift Δ(s) measures claim persistence and determines a soft gate on positive token credit. Negative advantages and non-routed tokens are unchanged, so the method only modulates positive credit for routed visual claims.
Backbone
Method
General Mathematical & Geometric Reasoning
Vision-Dependent Reasoning
Avg.
MMK12
Math Verse
Dyna Math
Math Vision
Geometry 3K
We- Math
Logic Vista
MMMU- Pro
Clever- Count
Qwen2.5-VL 7B
ThinkLite-VL
62.5
63.8
62.0
31.5
35.8
66.4
40.0
27.6
77.8
51.9
VL-Rethinker
69.3
68.8
65.7
29.5
40.7
68.5
45.8
39.7
82.0
56.7
NoisyRollout ∗
50.0
67.8
62.1
22.1
46.9
71.0
45.3
34.5
85.7
53.9
R1-ShareVL
70.9
68.2
63.9
27.5
41.2
69.9
45.4
35.1
81.2
55.9
MM-Eureka
67.5
65.4
64.8
27.1
40.8
65.5
45.8
35.3
77.8
54.4
Table 1: Main results across multimodal reasoning benchmarks and backbone families. All entries are accuracy (%). The first six datasets assess general mathematical and geometric reasoning, whereas the last three emphasize vision-dependent multimodal reasoning. Public 7B baselines are listed first, followed by our matched Qwen2.5-VL-7B runs. Bold and underlined values denote the best and second-best completed results within each backbone block. Additional blocks show transfer to a larger model scale and a newer backbone family.
Model
aAcc ↑
fAcc ↑
qAcc ↑
Base
65.19
38.15
33.41
GRPO
66.90
44.35
35.10
DAPO
67.17
43.77
35.32
VPPO
67.70
45.06
35.18
DAPO+PACG
67.64
44.64
35.54
VPPO+PACG
68.11
45.51
36.04
Table 2: HallusionBench (7B; %; single checkpoint per model).
Figure 3: Online accuracy reward. Dashed: parents; solid: PACG compositions; curves are trailing 10-step means for one representative seed per method.
Variant
MMK12
Math Verse
Dyna Math
Math Vision
Geometry 3K
We- Math
Logic Vista
MMMU- Pro
Clever- Count
Avg.
Unsup. Pers. ↓
DAPO
81.0
66.6
64.6
30.3
42.4
67.9
46.2
39.2
84.6
58.1
−0.1241
w/o Residual Corr.
80.9
69.0
65.6
30.9
45.4
69.1
48.9
40.1
81.3
59.0
−0.1278
w/o Direct Routing
46.3
40.8
36.7
16.2
10.4
44.0
31.7
30.3
85.2
38.0
−0.0374
w/o Positive-Only
80.8
70.3
66.1
31.3
44.5
70.6
47.7
39.6
82.6
59.3
−0.1301
PACG (full)
81.2
70.6
65.9
30.6
44.3
70.1
47.1
40.7
88.7
59.9
−0.1392
Table 3: Component ablations with DAPO as parent. w/o Residual Corr. : omit same-rollout residual calibration; w/o Direct Routing : gate all positive tokens instead of direct claims; w/o Positive-Only : also attenuate negative advantages. Accuracy and its nine-benchmark mean (Avg.) are percentages. Unsup. Pers. is the unsupported corrected persistence defined in Section 5.4 . Entries are point estimates; bold marks the best column value.
Figure 4: Aggregate improvement masks unsupported-claim deterioration. (a) Direct EFS ( ↑ ); (b) overall and (c) unsupported corrected persistence ( ↓ ). DAPO and VPPO improve aggregate retractability while worsening unsupported persistence; PACG mitigates this deterioration.
Model
Unsup. Pers. ↓
Img-vs-Text DID ↓
Acc. (%) ↑
Base
−0.1892
0.000
38.9
GRPO
−0.1963
−0.167
59.0
DAPO
−0.1241†
−0.022
62.2
VPPO
−0.1061†
−0.042
63.4
DAPO+PACG
−0.1392†
−0.016
63.0
VPPO+PACG
−0.1880†
−0.066
64.3
Table 5: Claim-level mechanism results. Accuracy (%) is evaluated on the CrossBench4-2000 questions using seed-0 checkpoints. DID is the Base-relative image-versus-text corrected-persistence contrast. † marks an unsupported-persistence change whose paired 95% question-clustered bootstrap CI excludes zero: versus Base for parent optimizers, and versus the corresponding parent for PACG. Full intervals are in Appendix E.2 .
Figure 5: Sensitivity–retractability phase trajectory. Direct EFS versus negative unsupported persistence on the fixed 500-question subset. Right: more sensitive, up: more retractable. Lines connect steps 0,50,100,150,200 ; dotted lines mark Base.
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Qwen2.5-VL-7B
Qwen2.5-VL-32B
Qwen3-VL-2B Thinking
Training corpus
ViRL39K
ViRL39K
ViRL39K
Training budget
2 epochs
2 epochs
200 steps
Optimizer
AdamW
AdamW
AdamW
Learning rate
10−6
10−6
10−6
Adam betas
(0.9,0.999)
(0.9,0.999)
(0.9,0.999)
Weight decay
0.01
0.01
0.01
Appendix
Table 6: Training configurations. All three backbones use an entropy penalty of 0.1; training budgets and response-length limits are backbone-specific.
Task
Precision
Recall
F1
κ
Direct visual claim
0.90
0.87
0.88
0.81
Direct vs. derived
0.85
0.83
0.84
0.76
Supported vs. unsupported
0.82
0.78
0.80
0.71
Human–human Cohen’s κ (same task order)
Direct visual claim
—
—
—
0.84
Direct vs. derived
—
—
—
0.79
Appendix
Table 7: Router and human audit on 500 CrossBench4 questions. Precision, recall, F1, and the first three κ values use human consensus as the reference.
Router
Precision
Recall
F1
Span Coverage
Timing (s)
Qwen3-4B
0.68
0.63
0.65
76%
108
Qwen3.8-27B
0.87
0.85
0.86
94%
425
Qwen3.6-35B-A3B
0.85
0.83
0.84
92%
165
Appendix
Table 8: Span-extractor comparison on 200 expert-annotated questions. Span coverage is the fraction of questions with at least one matched expert span. Timing is wall-clock seconds per batch on eight A100 GPUs; shading identifies the selected router.
Model
Direct Claim Rate (%)
Direct Claims / Response
Direct Token Ratio (%)
Response Length
Unsup. Claim Rate (%) ↓
Unsup. Resp. Rate (%) ↓
Accuracy (%)
Uncert./ Abst. (%)
Base
77.46 [76.08, 78.81]
2.463 [2.387, 2.537]
11.18 [10.83, 11.54]
506.62 [498.61, 514.84]
24.19 [23.14, 25.27]
32.19 [30.88, 33.48]
43.02 [41.56, 44.51]
10.2
GRPO
78.05 [76.60, 79.48]
2.382 [2.306, 2.459]
11.14 [10.75, 11.52]
526.64 [517.18, 535.97]
19.61 [18.53, 20.73]
26.61 [25.22, 27.99]
73.41 [71.76, 75.07]
9.5–10.5
DAPO
76.12 [74.66, 77.57]
2.148 [1.912, 2.314]
10.30 [9.94, 10.66]
516.82 [508.45, 525.13]
18.84 [17.79, 19.88]
25.83 [24.48, 27.17]
81.60 [79.87, 83.28]
9.5–10.5
DAPO + PACG
74.52 [73.07, 75.96]
2.211 [2.139, 2.282]
9.87 [9.54, 10.21]
528.46 [520.12, 536.96]
16.16 [14.11, 18.25]
25.62 [24.55, 27.10]
83.16 [81.56, 84.75]
9.0–11.0
VPPO
79.41 [78.08, 80.75]
2.645 [2.562, 2.733]
9.95 [9.62, 10.28]
681.46 [670.84, 692.14]
17.44 [16.51, 18.41]
25.61 [24.33, 26.87]
81.73 [80.16, 83.29]
9.5–10.5
VPPO + PACG
77.90 [76.46, 79.28]
2.463 [2.383, 2.543]
9.31 [8.99, 9.62]
697.47 [686.76, 708.45]
16.61 [15.69, 17.55]
24.12 [22.90, 25.37]
82.63 [81.06, 84.14]
9.0–11.0
Appendix
Table 9: Free-generation visual-claim audit on MMK12. Question-macro means with 95% question-cluster bootstrap intervals, using successfully annotated responses. Rates and accuracy are percentages; lengths are tokens. Uncertainty/abstention entries are point values or ranges, not confidence intervals.
Model
Median length (tokens)
Annotated responses
Direct claims
Base
462.0
16,000/16,000
39,401
GRPO
474.0
15,996/16,000
38,098
DAPO
480.0
15,999/16,000
34,360
DAPO + PACG
489.0
15,998/16,000
35,370
VPPO
635.0
15,990/16,000
42,295
VPPO + PACG
645.0
15,992/16,000
39,375
Appendix
Table 10: Response-length medians and annotation coverage. Coverage is out of 16,000 expected responses per model. Claim counts are exact high-confidence direct claims.
Model
Supported Persistence ↓
Base
−0.137
GRPO
−0.318
DAPO
−0.159
VPPO
−0.184
DAPO+PACG
−0.165
VPPO+PACG
−0.213
Appendix
Table 11: Supported-claim corrected persistence. Lower is better. Bold and underlining mark the best and second-best point estimates; blue shading identifies PACG variants.
Comparison
Metric
Effect
95% CI
DAPO − Base
Unsupported persistence
+0.0651
[0.022,0.108]
VPPO − Base
Unsupported persistence
+0.0831
[0.036,0.130]
DAPO+PACG − DAPO
Unsupported persistence
−0.0151
[−0.022,−0.008]
VPPO+PACG − VPPO
Unsupported persistence
−0.0819
[−0.115,−0.049]
DAPO − Base
Direct EFS
+0.0715
[0.025,0.118]
VPPO − Base
Direct EFS
+0.2330
[0.160,0.306]
Appendix
Table 12: Paired mechanism effects with question-clustered 95% bootstrap intervals. Each effect is the first model minus the second. Accuracy effects are on the [0,1] scale; other effects use their metric’s native scale.
Metric
GRPO
DAPO
VPPO
Positive routed-token rate, Ppos (%)
71.8
66.3
78.9
Effective positive update mass, Meff
0.35
0.93
1.00
Appendix
Table 13: Online positive-credit exposure over steps 181–200. Results concern routed direct visual claims in the parent training runs. Values are point estimates.
Figure 6: Sensitivity and unsupported-claim persistence diverge during training. DAPO and DAPO+PACG are evaluated at matched steps 0,50,100,150, and 200 . Direct EFS increases for both methods, whereas DAPO’s unsupported persistence becomes progressively less negative. PACG prevents the same drift and ends with greater unsupported retractability than DAPO. The temporal series uses 500 fixed questions sampled from CrossBench4-2000, with four Base rollouts each; subset averages can differ from full-set averages. Lines connect observed checkpoints rather than fitted curves.
Intervention
EFS ↑
Overall Pers. ↓
Unsup. Pers. ↓
PACG Effect ↓
Valid Rate
Random mask
0.31
−0.2558
−0.139
−0.015
96%
Coarse ink
0.28
−0.148
−0.125
−0.013
94%
Foreground/salient
0.33
−0.172
−0.151
−0.018
93%
Appendix
Table 14: Robustness to counterfactual corruption. Absolute DAPO+PACG scores and the matched DAPO+PACG minus DAPO change in unsupported persistence (PACG Effect). Entries are point estimates.
Figure 7: Three image interventions on the same example. Illustrative outputs of the evaluation operators using their default parameters on the original panel recovered from a saved MMK12 audit image. Random mask removes grid cells; coarse ink expands dark strokes; foreground masking removes their padded bounding rectangle. These are image-level controls, not verified claim-localized evidence removals. The illustration does not add observations to Table 14 .
Method
MMK12
Math Verse
Dyna Math
Math Vision
Geometry 3K
We-Math
Logic Vista
MMMU- Pro
Clever- Count
Avg.
GRPO
72.2
67.8
65.1
30.1
42.0
67.8
47.1
38.3
81.0
56.8
GRPO+PACG
73.5
68.1
65.0
30.2
42.3
68.8
47.3
39.0
81.7
57.3
Appendix
Table 15: PACG composed with plain GRPO. Accuracy (%); GRPO+PACG uses mean accuracy@8. Bold marks the higher displayed value in each column. Entries are point estimates.
Variant
Direct EFS ↑
Corr. Pers. ↓
Unsup. Pers. ↓
DAPO
0.2806
−0.2095
−0.1241
w/o Residual Corr.
0.4099
−0.2410
−0.1278
w/o Direct Routing
3.0366†
−0.1766
−0.0374
w/o Positive-Only
0.2046
−0.2224
−0.1301
PACG (full)
0.3096
−0.2558
−0.1392
Appendix
Table 16: Additional component mechanism metrics. Point estimates; lower persistence indicates stronger retraction. † denotes a degenerate EFS value associated with training collapse, not improved visual grounding.
Model
Avg. (%)
Δ vs. parent
p
DAPO
58.1±0.40
–
–
DAPO+PACG
59.9±0.52
+1.8
0.010
VPPO
59.8±0.40
–
–
VPPO+PACG
60.9±0.47
+1.1
0.038
Appendix
Table 17: Seed variance (Qwen2.5-VL-7B; nine-benchmark Avg., %; mean ± sample standard deviation over three training seeds). p : two-sided Welch’s t -test between each PACG composition and its parent ( n=3 per group); Δ is in percentage points.
m
α
gmin
MMK12 Accuracy (%) ↑
EFS ↑
Corrected Persistence ↓
Unsupported Persistence ↓
Positive-credit retention (%)
0.00
10
0.30
81.0
0.34
−0.10
−0.08
95
0.05
10
0.30
81.2
0.31
−0.23
−0.14
81
0.10
10
0.30
81.0
0.28
−0.28
−0.20
65
Appendix
Table 18: Sensitivity to the persistence margin. Gate strength α=10 and floor gmin=0.30 are fixed. Accuracy and retention are percentages. Shading marks the selected operating point; bold marks the best accuracy, EFS, and persistence values. Entries are point estimates.
Figure 8: PACG gate dynamics. (a) Mean span gate; (b) mean multiplier over positive direct-claim tokens; (c) fraction of these tokens downweighted; (d) fraction at the gate floor. Faint curves show observations and solid curves trailing 10-step means.
Figure 9: Online training-reward dynamics. Paired parent/PACG comparisons of reward/accuracy through step 200. Faint curves show logged values and solid curves trailing 10-step means.
Figure 10: Online VPPO visual-selector scores. Mean low-variance KL scores for (a) selected and (b) unselected tokens. The panels use different vertical scales.
Figure 11: Overlap between VPPO selection and PACG attenuation. (a) Fraction of VPPO-selected tokens downweighted by PACG; (b) fraction of PACG-downweighted tokens selected by VPPO. The two panels use different denominators.
Figure 12: Training dynamics across six runs. Accuracy reward, entropy, response length, clipping rate, gradient norm, and approximate policy-update KL. Dashed curves denote parents and solid curves their PACG compositions. Policy-update KL is distinct from reference-model KL and image-intervention EFS.
Figure 13: Online persistence and routing coverage. (a)–(c) Mean corrected delta, its 90th percentile, and margin-violation rate over eligible spans. (d)–(f) Routed-rollout rate, direct-token fraction, and intervention-valid rate. The 90th percentile describes the span distribution, not a confidence interval.
Stage
DAPO
PACG
Overhead vs. DAPO
Rollout
148.8
174.6
–
Original-image teacher forcing
46.7
49.2
–
Reference log-probability (KL)
44.9
–
–
Augmented scoring
–
51.2
–
Counterfactual teacher forcing
–
56.6
–
Blocking annotation wait
–
159.4
–
Appendix
Table 19: Training efficiency on Qwen2.5-VL-7B using 16 PPU-810E accelerators. Mean seconds per step over steps 1–200. Total time includes validation, checkpointing, and other work.
Figure 14: Training-time breakdown for PACG ( m=0.05 ). Mean seconds per step over steps 1–200. TF denotes teacher forcing. Augmented scoring corresponds to timing_s/aug .