cs.LGOct 1, 2026

Range-GRPO: Policy Optimization via Pairwise Relations among Reward Intervals

Authors: Ryunyi Lee, Kangjun Noh, Somin Kim, Heedong Kim, Kyungwoo Song

Organizations: Yonsei University

Abstract

As the use of large language models (LLMs) expands, post-training has become increasingly important for adapting them to downstream tasks. However, obtaining reliable supervision remains costly, especially in domains without reference answers or executable verifiers. LLM-as-a-Judge provides scalable pseudo-rewards for unlabeled responses, but a single point score does not explicitly represent reward uncertainty. This motivates representing pseudo-rewards as conformally calibrated reward ranges. We propose Range-GRPO, a semi-supervised post-training framework that combines limited labeled data with unlabeled prompts. In Group Relative Policy Optimization (GRPO), learning signals depend on relative reward comparisons within each rollout group. The proposed objective compares reward ranges pairwise rather than reducing them to point rewards, allowing interval uncertainty to affect both the magnitude and direction of these signals. Our theoretical analysis characterizes this distinction and shows that the proposed objective recovers the Dr.GRPO advantage when all reward ranges collapse to points. Empirically, Range-GRPO achieves the highest in-distribution and out-of-distribution average performance among the evaluated semi-supervised methods while requiring fewer training resources.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. P^2O: Joint Policy and Prompt Optimization

    Mar 23, 2026Xinyu Lu, Kaiqi Zhang, Jinglin Yang +6Reinforcement Learning With Verifiable RewardLarge Language Model Alignment

  2. Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR

    Aug 4, 2026Yongshi Ye, Liang Zhang, Yidong Chen +2Reinforcement Learning With Verifiable RewardRepair-Based Group-Relative Policy Optimization