Vision-Language (VL) reasoning requires a model to both extract relevant and accurate information from an image (visual reasoning, VR), and to infer the answer from it (language reasoning, LR). Reinforcement learning with verifiable rewards typically trains both through a single chain-of-thought with a final-answer reward. This gives every CoT token the same sequence-level advantage, failing to distinguish capability specific errors. We propose SPLIT-RL, a staged post-training approach that trains VR and LR in disjoint phases. Because a group's rollouts differ along one capability at a time, the group-relative advantage isolates it, and each phase is optimized using phase-specific reward. We further introduce Claim-Level Advantage (CLA-GRPO), which decomposes VR-phase rollouts into atomic visual claims and provides a fine-grained advantage at claim level based on visual-type group formation. Although trained in two phases, trained policy is evaluated like GRPO model, with a single CoT call at inference time. Under this protocol, SPLIT-RL improves average accuracy over GRPO by 1.4-6.1 points across Qwen3-VL models from 2B to 30B-A3B and InternVL3.5-8B. Evaluating each capability using an oracle based diagnostic shows that answer-only GRPO leaves perception unchanged, whereas SPLIT-RL improves both VR and LR.
Figures & tables
Figure 1: Entangled CoT vs. SPLIT-RL. Top: One misread visual claim produces a wrong answer, assigning every token the same nonpositive advantage including four correct claims. With a fixed oracle (Section 4.3.2), answer-only training improves LR (+3.8) but not VR (+0.0). Bottom: SPLIT-RL separates visual description from text-only reasoning, using answer correctness to assess description sufficiency. CLA-GRPO assigns claim-level advantages, and training VR before LR improves both capabilities (+1.7 VR, +6.2 LR).
Figure 2: SPLIT-RL. One policy is trained in two stages. Stage 1 (VR): Samples G descriptions per image–question pair and updates description tokens using recall, answer-correctness, and claim-level signals. Stage 2 (LR): Starting from Stage 1, generates one description d and samples G reasoning–answer rollouts from (Q,d) without the image; GRPO updates only reasoning and answer tokens. Right: CLA-GRPO verifies atomic claims against the image, groups them by visual type across rollouts, and assigns claim-level advantages to their tokens.
Backbone
Method
MVista
MVision
MVerse
WeMath
LogicVista
DynaMath
AVG
Base
60.3
18.8
38.8
56.5
34.2
55.9
44.1
GRPO
60.6
35.9
44.4
54.1
36.7
56.5
48.0
SPLIT-RL
65.4
42.8
46.7
62.1
44.1
63.8
54.1
Qwen3-VL 2B
Δ vs. GRPO
+4.8
+6.9
+2.3
+8.0
+7.4
+7.3
+6.1
Base
77.1
59.2
61.7
75.6
63.1
73.4
68.4
GRPO
78.0
61.8
66.4
79.8
56.4
74.7
69.5
Table 1: Main results. Accuracy (%) on six benchmarks. Δ rows show percentage-point gains over GRPO for the same backbone. Bold marks the best accuracy per column within each backbone.
Figure 3: Qualitative example (WeMath). Rotating AB by 90∘ . GRPO misreads A as (5,8) and hits the token limit without an answer; SPLIT-RL reads A=(6,8) and answers correctly.
Figure 5
Stage 1
Stage 2
Mean acc.
Gain
Qwen3-VL-8B
69.8
–
Qwen3-VL-8B (GRPO)
72.6
+2.8
LR
VR
72.1
+2.3
VR
LR
75.3
+5.5
Table 2: Order of staged training. Gains are relative to the base model.
Source
# examples
Math-PUMA (synthesis)
6,696
GeoQA170K
6,499
CLEVR-Math
2,000
ArxivQA ( 2× downsampled)
1,000
DOCCI ( 2× downsampled)
3,360
Total
19,555
Table 3: Training data mixture. Difficulty filtering ( pass_rate@16=1 removed) reduces 16,195→13,862 samples.
Method
MVista
MVision
MVerse
WeMath
LogicVista
DynaMath
AVG
Δ
Qwen3-VL-8B (GRPO)
77.9
64.5
67.8
84.6
65.6
75.3
72.6
—
LR → VR
80.0
61.2
68.8
82.3
62.4
77.7
72.1
−0.5
VR → LR (SPLIT-RL)
79.8
69.4
69.8
87.4
66.0
79.1
75.3
+2.7
Table 4: Order of staged training on Qwen3-VL-8B, per benchmark; the summary is Table 2 . Both orders use the same VR-stage rewards (CLA + recall). Δ is the change in average accuracy relative to GRPO. Best per column in bold .
Figure 6: The recall reward on a worked example. Left: the question and image (arc length of BC ). Center: the description produced by the VR phase. Right: a text-only judge, which never sees the image, checks whether the description alone contains every fact the question needs. It finds the circle, center, points and the “21 cm” label, but the description never states the 168∘ central angle —the one fact the arc-length formula requires—so the description is insufficient and rrecall=0 .
Method
MVista
MVision
MVerse
WeMath
LogicVista
DynaMath
AVG
Δ
Qwen3-VL-8B
77.2
63.2
63.8
79.9
60.4
74.4
69.8
—
VR → LR
79.3
65.5
68.2
84.8
63.1
77.1
73.0
+3.2
VR + Recall → LR
79.6
64.1
67.4
84.9
65.1
77.3
73.1
+3.3
VR + CLA → LR
80.1
61.2
70.9
87.9
65.3
76.9
73.7
+3.9
VR + CLA + Recall → LR
79.8
69.4
69.8
87.4
66.0
79.1
75.3
+5.4
Table 5: Contribution of each component on Qwen3-VL-8B. All staged variants share the same LR stage and differ only in the rewards used in the VR stage. Δ is the change in average accuracy relative to the untrained Qwen3-VL-8B. The best result per column is shown in bold .
Method
MVista
MVision
MVerse
WeMath
LogicVista
DynaMath
Mean
Δ
(a) Varying the describer (reasoner: Kimi-K3)
Kimi-K3 (oracle)
86.4
84.9
85.9
89.8
80.1
90.6
86.3
Qwen3-VL-8B
75.6
66.1
72.8
80.5
66.2
75.9
72.9
+0.0
Qwen3-VL-8B (GRPO)
76.3
66.8
74.6
80.3
63.5
76.0
72.9
+0.0
VR → LR
77.0
64.8
73.7
81.7
65.5
78.3
73.5
+0.6
VR + Recall → LR
76.0
64.7
74.0
81.8
64.4
75.6
72.7
−0.2
Table 6: Improvement of each capability on Qwen3-VL-8B. (a) VR: the describer is varied and the reasoner is fixed to Kimi-K3. (b) LR: the describer is fixed to Kimi-K3 and the reasoner is varied. The oracle row uses Kimi-K3 for both roles and is shared by both blocks. Δ is the change in mean accuracy relative to the base model. Within each block, best per column (excluding the oracle) in bold , second best underlined .
Method
MVista
MVision
MVerse
WeMath
LogicVista
DynaMath
Mean
Δ
Qwen3-VL-8B
Base
69.1
54.3
58.1
73.9
57.7
71.1
64.0
–
GRPO
73.4
57.2
63.7
80.3
62.4
72.2
68.2
+4.2
SPLIT-RL
76.6
62.8
66.4
85.9
65.3
75.9
72.1
+8.1
Qwen3-VL-4B
Base
68.4
45.7
51.9
64.1
50.8
70.6
58.6
–
Table 7: Two-phase (self-loop) evaluation. The model first describes the image and then answers from its own description, without the image, as in training. Δ is the change in mean accuracy relative to the base model of the same size. Best per column within each backbone in bold .
Figure 7: Qualitative example (MathVista). Is the food half eaten? The GRPO baseline hedges about the food being “mostly intact” and answers (B) No (✗). Our model identifies that only half the food remains and answers (A) Yes (✓), matching the ground truth.
Recent advancements in reinforcement learning with verifiable rewards (RLVR) have significantly improved the complex reasoning ability of vision-language models (VLMs). However, its outcome-level supervision is too coarse to diagnose and correct errors within the reasoning chain. To this end, we propose Perceval, a process reward model (PRM) that enables token-level error grounding, which can extract image-related claims from the response and compare them one by one with the visual evidence in the image, ultimately returning claims that contain perceptual errors. Perceval is trained with perception-intensive supervised training data. We then integrate Perceval into the RL training process to train the policy models. Specifically, compared to traditional GRPO, which applies sequence-level advantages, we apply token-level advantages by targeting penalties on hallucinated spans identified by Perceval, thus enabling fine-grained supervision signals. In addition to augmenting the training process, Perceval can also assist VLMs during the inference stage. Using Perceval, we can truncate the erroneous portions of the model's response, and then either have the model regenerate the response directly or induce the model to reflect on its previous output. This process can be repeated multiple times to achieve test-time scaling. Experiments show significant improvements on benchmarks from various domains across multiple reasoning VLMs trained with RL, highlighting the promise of perception-centric supervision as a general-purpose strategy. For test-time scaling, it also demonstrates consistent performance gains over other strategies, such as major voting. Our code and data will be publicly released at https://github.com/RUCAIBox/Perceval.
Yingqian Min, Kun Zhou, Yifan Li +6
Gaoling School of Artificial Intelligence, Renmin University of China. · Bytedance. · University of California, San Diego. +1
Vision-Language Models (VLMs) frequently suffer from visual perception errors and hallucinations that compromise answer accuracy in complex reasoning tasks. Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising solution by optimizing policies using answer correctness signals. Despite their effectiveness, prevailing RLVR methods face two critical limitations. First, much of the sampling budget is wasted on trajectories doomed to fail due to early visual description errors. Second, sparse rewards cannot distinguish whether failures stem from visual perception or reasoning stages. We introduce MIRL, a decoupled framework that addresses both limitations by leveraging mutual information (MI) between generated descriptions and visual inputs as a cheap pre-screening signal. This enables intelligent budget allocation toward high-potential trajectories via forking, while decoupled training provides independent MI-based rewards for visual perception optimization, resolving reward blindness. Experiments on six vision-language reasoning benchmarks demonstrate that MIRL achieves 70.22% average accuracy and successfully surpasses the performance of sampling 16 complete trajectories using only 10 pre-samples with top-6 selection (25% fewer complete trajectories). Our code is available at: https://anonymous.4open.science/r/mirl-main/.
Yin Zhang, Jiaxuan Zhao, Zonghan Wu +5
School of Mathematics, Tianjin University, Tianjin, China · Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China · Shanghai Advanced Institute of Finance (SAIFS), East China Normal University, Shanghai, China +4
Recent advances in vision-language models (VLMs) emphasize long chain-of-thought reasoning; yet, we find that their performance on visual tasks is primarily limited by a lack of visual perception as opposed to reasoning itself. In this work, we systematically study the interplay between perception and reasoning in VLM post-training by decomposing their capabilities into three separate training stages: visual perception, visual reasoning, and textual reasoning, incorporating specialized training data. We demonstrate that visual perception (a) requires targeted optimization with specialized data; (b) serves as a fundamental scaffold that should be solidified through staged training before refining visual reasoning; and (c) is more effectively learned via RL than caption-based SFT. Our experiments across multiple VLMs demonstrate that staged training consistently improves both visual perception and reasoning performance over merged training. Notably, models trained with our approach achieve 1.5% higher reasoning accuracy with 20.8% shorter reasoning traces, suggesting that superior perception reduces the need for excessive reasoning. Furthermore, we show that this capability-based staging represents a new curriculum dimension orthogonal to traditional difficulty-based curricula, and combining both yields further additive gains. Our staged-training models achieve superior performance among open-weight VLMs, establishing advanced results on several visual math and perception (e.g., +5.2% on WeMath and +3.7% on RealWorldQA) tasks compared with the base counterpart.
Juncheng Wu, Hardy Chen, Haoqin Tu +6
Amazon · UC Santa Cruz · University of Waterloo +1