cs.CVOct 7, 2026

SPLIT-RL: Staged Perception-Language Reasoning Training with Claim-Level Advantages

Authors: Raja Kumar, Rajat Koner, Ritwick Chaudhry, Zhuowei Li, Nishant Sankaran, Yifan Xing

Organizations: University of Southern California · Amazon AGI

Abstract

Vision-Language (VL) reasoning requires a model to both extract relevant and accurate information from an image (visual reasoning, VR), and to infer the answer from it (language reasoning, LR). Reinforcement learning with verifiable rewards typically trains both through a single chain-of-thought with a final-answer reward. This gives every CoT token the same sequence-level advantage, failing to distinguish capability specific errors. We propose SPLIT-RL, a staged post-training approach that trains VR and LR in disjoint phases. Because a group's rollouts differ along one capability at a time, the group-relative advantage isolates it, and each phase is optimized using phase-specific reward. We further introduce Claim-Level Advantage (CLA-GRPO), which decomposes VR-phase rollouts into atomic visual claims and provides a fine-grained advantage at claim level based on visual-type group formation. Although trained in two phases, trained policy is evaluated like GRPO model, with a single CoT call at inference time. Under this protocol, SPLIT-RL improves average accuracy over GRPO by 1.4-6.1 points across Qwen3-VL models from 2B to 30B-A3B and InternVL3.5-8B. Evaluating each capability using an oracle based diagnostic shows that answer-only GRPO leaves perception unchanged, whereas SPLIT-RL improves both VR and LR.

Figures & tables

Explore similar work

CardsList
  1. Improving Vision-language Models with Perception-centric Process Reward Models

    Apr 27, 2026Yingqian Min, Kun Zhou, Yifan Li +6Process Reward ModelsObject Hallucination in VLMs

  2. MIRL: Mutual Information-Guided Reinforcement Learning for Vision-Language Models

    May 2, 2026Yin Zhang, Jiaxuan Zhao, Zonghan Wu +5Vision-Language ModelsReinforcement Learning

  3. From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models

    May 19, 2026Juncheng Wu, Hardy Chen, Haoqin Tu +6Vision-Language ModelsVLM Reasoning