As the use of large language models (LLMs) expands, post-training has become increasingly important for adapting them to downstream tasks. However, obtaining reliable supervision remains costly, especially in domains without reference answers or executable verifiers. LLM-as-a-Judge provides scalable pseudo-rewards for unlabeled responses, but a single point score does not explicitly represent reward uncertainty. This motivates representing pseudo-rewards as conformally calibrated reward ranges. We propose Range-GRPO, a semi-supervised post-training framework that combines limited labeled data with unlabeled prompts. In Group Relative Policy Optimization (GRPO), learning signals depend on relative reward comparisons within each rollout group. The proposed objective compares reward ranges pairwise rather than reducing them to point rewards, allowing interval uncertainty to affect both the magnitude and direction of these signals. Our theoretical analysis characterizes this distinction and shows that the proposed objective recovers the Dr.GRPO advantage when all reward ranges collapse to points. Empirically, Range-GRPO achieves the highest in-distribution and out-of-distribution average performance among the evaluated semi-supervised methods while requiring fewer training resources.
Figures & tables
Figure 1: Overview of Range-GRPO and its learning signals. (a) At initialization, Range-GRPO fits a state scorer on labeled source rollouts and computes nonconformity scores on calibration rollouts. Each iteration generates policy and target DRE rollouts with G and G′ responses, respectively, and updates density ratios using the target states and reused source states. The density ratios, fixed scorer, and calibration scores are used to construct reward ranges, from which pairwise interval comparisons construct the advantages for policy updates. (b) An illustrative comparison uses identical reward intervals for both methods. Although the interval for o2 lies above that for o3 with the same width, midpoint reduction assigns it a negative advantage under Dr.GRPO. Range-GRPO instead downweights pairwise contributions involving wider, more uncertain intervals, turning this advantage positive and aligning the direction of the learning signal with the oracle preference.
Table 2: Range-GRPO achieves the highest ID and OOD averages among semi-supervised methods. We report published baseline results from TraPO ( Yang et al., 2026 ) , while methods marked with ∗ are evaluated under our experimental setting. For the reproduced TraPO configurations, values in parentheses indicate the entropy regularization coefficients. All training methods use 1,024 (1K) ID and 1,024 (1K) OOD prompts. Semi-supervised methods use labeled ID prompts and unlabeled OOD prompts, whereas the fully supervised model uses labels for all 2K prompts. GPQA † denotes GPQA-diamond. Bold and underlined scores indicate the best and second-best results among semi-supervised methods, respectively, with ties receiving the same formatting. The fully supervised result is provided for reference only and excluded from this ranking.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Notation
Meaning
GRPO and policy optimization
q,PQ
Prompt and prompt distribution.
G,G′,oi,oi,k
Policy and target DRE responses per prompt, response i , and token k of response i .
πθ,πθold,πref,πt
Parameterized policy, old rollout policy, fixed reference policy, and current policy πt:=πθt .
ri,rˉ,σr
Response reward, group-mean reward, and group reward standard deviation.
Ai,AiDr.GRPO
GRPO and Dr.GRPO response-level advantages.
Appendix
Table 4: Notation used throughout the paper. Symbols are grouped by their roles and ordered approximately by their first appearance in the paper.
Figure 2: Validation accuracy at the selected hyperparameter settings. Curves show the mean ± sample standard deviation across three seeds. We show scalar baselines with GRPO and Dr.GRPO advantages separately and include the same UARM and Range-GRPO curves in both comparisons.
In-Distribution
Out-of-Distribution
Model / Method
AIME 24/25
AMC
MATH-500
Minerva
Olympiad
Avg.
ARC-c
GPQA †
MMLU-Pro
Avg.
TraPO (entropy=0.01)
17.3 / 10.3
52.6
77.6
33.8
36.3
38.0
78.8
36.9
39.5
51.7
TraPO (entropy=0.001)
17.7 / 12.2
55.9
79.2
37.5
41.3
40.6
80.4
34.8
43.5
52.9
Appendix
Table 9: Performance of reproduced TraPO configurations with different entropy coefficients. All training methods use 1,024 (1K) ID and 1,024 (1K) OOD prompts. GPQA † denotes GPQA-diamond. Bold and underlined scores indicate the best and second-best results, respectively.
Category
TraPO
UARM
Range-GRPO
Labeled policy
122,880
92,160
92,160
Unlabeled rollouts
122,880
40,960
37,120
RM training
0
28,560
0
RM calibration
0
7,280
0
RM processing
0
76,800
0
Scorer fitting
0
0
1,280
Appendix
Table 10: Response budget for the main experiments. The table compares the response budget for TraPO, UARM, and Range-GRPO. For TraPO, unlabeled rollouts include those generated during warmup. The UARM budget includes RM processing, whereas the Range-GRPO budget includes judge generations.
Setting
Value
Model
Qwen2.5-Math-7B
Learning rate
10−6
Batch size
64 prompts
Responses per prompt
8
Epochs
15
Warmup epochs
10
Appendix
Table 11: UARM training configuration for the cross-domain setting. The table reports the policy training settings in (a) and the reward model training and calibration settings in (b).
Role
Prompts
Labeled policy
768
RM training
184
RM validation
20
Calibration
52
Unlabeled policy
1,024
Total
2,048
Appendix
Table 12: Data allocation for UARM and Range-GRPO in the main experiments. (a) summarizes the UARM allocation across policy learning and RM training, validation, and calibration. (b) details the Range-GRPO allocation across policy learning, scorer fitting, calibration, source DRE, and target DRE.
State
Criterion
Score 0
Score 0.5
Score 1
Response-level assessment
z1
Answer verification
The judge identifies a contradiction when independently checking the submitted answer against the problem.
The judge cannot verify or refute the submitted answer with sufficient confidence.
The judge independently verifies the submitted answer against the problem.
z2
Reasoning consistency
The visible reasoning contains a decisive error or conflicts with the submitted answer.
The visible reasoning is incomplete or insufficient to establish the submitted answer.
The visible reasoning consistently supports the submitted answer.
z3
Requirement compliance
The response omits or violates an explicit task requirement.
It remains unclear whether all explicit requirements are satisfied.
The response satisfies all explicit task requirements.
Group-level agreement
z4
Absolute agreement
No other response gives the same answer.
One to three other responses give the same answer.
Four to seven other responses give the same answer.
Appendix
Table 13: Five components of the judgment state Z . The first three states assess the response itself, while the last two capture agreement with other responses in the same rollout group. Each state takes a score of 0, 0.5, or 1.
Setting
Value
Scorer
XGBoost regressor
Trees / depth
200 / 2
DRE
Random forest
Trees / depth
200 / 4
DRE target per update
768 responses
Refresh
Each unlabeled update
Appendix
Table 14: Range construction settings for the main experiments. The scorer and calibration nonconformity scores remain fixed after warmup, whereas DRE is refreshed at each unlabeled update.
Setting
Value
Batch size
64 prompts
Responses per prompt
8
Warmup epochs
10
Active epochs
5
Optimizer
AdamW
PPO clipping
0.2
Appendix
Table 15: Range-GRPO policy training settings for the main experiments. (a) summarizes the policy training configuration. (b) lists the main optimization and Range-GRPO settings. LR denotes the learning rate.
Method
Active GPUs
Warmup (h)
Active (h)
Training (h)
TraPO (entropy=0.01)
4
12.8
10.6
23.4
TraPO (entropy=0.001)
4
12.6
7.0
19.7
UARM
5
6.5
9.9
16.4
Range-GRPO (ours)
4
6.3
10.0
16.3
Appendix
Table 16: Range-GRPO maintains comparable wall-clock training time despite the additional judge phase. Each run consists of 10 warmup epochs followed by 5 active epochs. Phase durations and total training times are independently rounded to one decimal place.
Group Relative Policy Optimization (GRPO) has been a key driver of recent progress in reinforcement learning with verifiable rewards (RLVR) for large language models, but it is typically trained in a low-staleness, near-on-policy regime that incurs substantial system overhead. We ask a simple question: How off-policy can GRPO be? We show that GRPO-style algorithms can tolerate substantially larger rollout staleness than previously assumed, and propose Mu-GRPO, an RL training framework that organizes training into a small number (e.g., four) of large sequential generation-optimization stages. This design induces high rollout staleness while greatly reducing rollout-optimization switching overhead. To stabilize learning under stale data, Mu-GRPO combines relaxed clipping, which preserves useful stale-rollout gradients, with negative-advantage veto, which removes destabilizing post-trigger suffix updates in negative-advantage responses. Across five language models and multiple math reasoning benchmarks, Mu-GRPO matches or exceeds the performance of standard GRPO while achieving around 2x speedup in wall-clock training time, establishing a substantially improved performance-efficiency trade-off for LLM reinforcement learning.
Reinforcement Learning with Verifiable Rewards (RLVR) enhances Large Language Model (LLM) reasoning but is suffer from advantage collapse: when all rollouts of a query receive identical rewards, the group variance vanishes, most damagingly on hard samples, where scaling rollout budgets yields little. We introduce Joint Policy and Prompt Optimization (P O) to mitigate this collapse by alternating continuous policy updates with discrete prompt evolution. P O mines hard samples with a success-rate threshold, evolves reasoning prompts for them with GEPA, and internalizes the elicited trajectories via context distillation, which optimizes each trajectory under the original query and thus removes inference-time prompting, with a Context Ratio Mask (CRM) filtering out extreme likelihood ratios. P O restores critical advantage signals and surpasses the GRPO baseline by up to 8.2 points in average accuracy on six held-out benchmarks across all training datasets and backbones, while also outperforming DAPO and other baselines. The gains are especially pronounced on hard benchmarks, reaching up to 16.3 points above GRPO on average across AIME24 and AIME25. Our findings expose the limits of standard exploration in sparse-reward environments, illuminating the potential of unifying evolutionary algorithms with reinforcement learning. This integration of discrete semantic search and continuous parameter updates provides a self-reinforcing framework that facilitates more effective LLM alignment.
Xinyu Lu, Kaiqi Zhang, Jinglin Yang +6
Chinese Information Processing Laboratory, Institute of Software, Chinese Academy of Sciences · University of Chinese Academy of Sciences · Institute of Information Engineering, Chinese Academy of Sciences +2
Reinforcement Learning with Verifiable Rewards (RLVR) improves LLM reasoning but typically relies on ground-truth (GT) answers, limiting scalability. Voting-based label-free RLVR replace gold supervision with answer-level consensus from model samples. However, collapse arises when the same answer-level signal is used both to estimate rewards and to drive token-level policy optimization, encouraging the model to directly reinforce answer tokens rather than improve reasoning. We propose OM-GRPO, a label-free RLVR framework that decouples reward estimation from policy optimization. OM-GRPO masks gradients on the answer span while retaining answer-level rewards through a soft consensus signal, shifting optimization pressure away from answer tokens. We further introduce Contrast-Augmented Reward, which refines reward estimation via low-cost pairwise comparisons over existing trajectories without additional rollouts. Across diverse reasoning benchmarks and three LLM backbones, OM-GRPO consistently outperforms existing label-free RLVR methods and matches supervised GT-reward training with stable optimization. This stability is particularly beneficial in the Test-Time Training setting, where OM-GRPO surpasses majority voting by 4.24 points.
Yongshi Ye, Liang Zhang, Yidong Chen +2
Xiamen University · Key Laboratory of Digital Protection and Intelligent Processing of Intangible Cultural Heritage of Fujian and Taiwan (Xiamen University), Ministry of Culture and Tourism