cs.LGSep 28, 2026

ReSPO: Reshaped Sequence Policy Optimization for Gradient Starvation in Off-Policy Learning

Authors: Yihang Chen, Yuanhao Ban, Cho-Jui Hsieh

Organizations: Department of Computer Science, University of California, Los Angeles

Abstract

Reinforcement learning from verifiable rewards (RLVR) frequently reuses rollouts across multiple policy updates, increasing the mismatch between the current policy and the data-generating policy. We identify a sign-dependent gradient starvation problem in clipped policy optimization: clipping suppresses under-generated positive responses at the low-importance-weight tail while permitting severely over-generated negative responses to dominate the high-weight tail. To address this, we propose ReSPO (Reshaped Sequence Policy Optimization), which replaces clipping with a smooth, two-branch sequence-level kernel derived from an αα-divergence variational objective and an exponential variance-control tilt. The positive branch preserves a nonzero gradient weight for under-generated positive responses, while the negative branch suppresses heavily over-generated negative responses. We demonstrate that ReSPO effectively learns from long positive reasoning trajectories during early training, even when accumulated policy drift relegates them to the low-importance-weight tail. On dense and MoE Qwen3 models, ReSPO accelerates early optimization, improves final training scores, and achieves higher held-out benchmark performance under a rollout reuse, validating our approach on importance-weight tail control in off-policy learning.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective

    May 13, 2026Feng Zhang, Xinhong Ma, Ziqiang Dong +5Reinforcement Learning With Verifiable RewardVerifiable Rewards

  2. Learning from the Near Future: Temporal Self-Distillation for RLVR

    Apr 22, 2026Chuanyu Qin, Chenxu Yang, Qingyi Si +6Reinforcement Learning With Verifiable RewardVerifiable Rewards