cs.AISep 29, 2026

Dual-Channel Robust Group-Relative Policy Optimization via Advantage and Sequence-Weight Estimation

Authors: Zhongyi Li, Wan Tian, Xiang Xu, Yutian Xiao, Yikun Ban, Yijie Peng, Fuzhen Zhuang

Organizations: Beihang University · Peking University · Nanjing University

Abstract

Group-relative policy optimization relies on reward-derived advantages and sequence-level likelihood weights, both of which can be sensitive to localized outliers. Extreme rewards can collapse the contrast among clean responses after group normalization, while token-level log-ratio perturbations can alter sequence weights and clipping decisions. We introduce RoVR-GSPO, a dual-channel robust optimizer that addresses these failure modes separately. Its reward channel combines robust reference estimation with bounded residual credit, while its ratio channel uses differentiable SoftRoVR aggregation to construct robust sequence weights. We provide stability and efficiency analyses for both channels. Experiments on mathematical reasoning, long-context summarization, and tool-call annotation show consistent improvements over GSPO, while controlled perturbation studies demonstrate stronger robustness to reward contamination and token-ratio anomalies.

Figures & tables

Appendix figures & tables18 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective

    May 13, 2026Feng Zhang, Xinhong Ma, Ziqiang Dong +5Reinforcement Learning With Verifiable RewardVerifiable Rewards

  2. RRPO: Reference-Relative Policy Optimization with Stratified Conditional Rollouts

    Jul 20, 2026Yuxin Xiong, Xunyi Jiang, Rohan Surana +8Reinforcement Learning With Verifiable RewardRollout