As the use of large language models (LLMs) expands, post-training has become increasingly important for adapting them to downstream tasks. However, obtaining reliable supervision remains costly, especially in domains without reference answers or executable verifiers. LLM-as-a-Judge provides scalable pseudo-rewards for unlabeled responses, but a single point score does not explicitly represent reward uncertainty. This motivates representing pseudo-rewards as conformally calibrated reward ranges. We propose Range-GRPO, a semi-supervised post-training framework that combines limited labeled data with unlabeled prompts. In Group Relative Policy Optimization (GRPO), learning signals depend on relative reward comparisons within each rollout group. The proposed objective compares reward ranges pairwise rather than reducing them to point rewards, allowing interval uncertainty to affect both the magnitude and direction of these signals. Our theoretical analysis characterizes this distinction and shows that the proposed objective recovers the Dr.GRPO advantage when all reward ranges collapse to points. Empirically, Range-GRPO achieves the highest in-distribution and out-of-distribution average performance among the evaluated semi-supervised methods while requiring fewer training resources.
Figures & tables
Figure 1: Overview of Range-GRPO and its learning signals. (a) At initialization, Range-GRPO fits a state scorer on labeled source rollouts and computes nonconformity scores on calibration rollouts. Each iteration generates policy and target DRE rollouts with G and G′ responses, respectively, and updates density ratios using the target states and reused source states. The density ratios, fixed scorer, and calibration scores are used to construct reward ranges, from which pairwise interval comparisons construct the advantages for policy updates. (b) An illustrative comparison uses identical reward intervals for both methods. Although the interval for o2 lies above that for o3 with the same width, midpoint reduction assigns it a negative advantage under Dr.GRPO. Range-GRPO instead downweights pairwise contributions involving wider, more uncertain intervals, turning this advantage positive and aligning the direction of the learning signal with the oracle preference.
Table 2: Range-GRPO achieves the highest ID and OOD averages among semi-supervised methods. We report published baseline results from TraPO ( Yang et al., 2026 ) , while methods marked with ∗ are evaluated under our experimental setting. For the reproduced TraPO configurations, values in parentheses indicate the entropy regularization coefficients. All training methods use 1,024 (1K) ID and 1,024 (1K) OOD prompts. Semi-supervised methods use labeled ID prompts and unlabeled OOD prompts, whereas the fully supervised model uses labels for all 2K prompts. GPQA † denotes GPQA-diamond. Bold and underlined scores indicate the best and second-best results among semi-supervised methods, respectively, with ties receiving the same formatting. The fully supervised result is provided for reference only and excluded from this ranking.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Notation
Meaning
GRPO and policy optimization
q,PQ
Prompt and prompt distribution.
G,G′,oi,oi,k
Policy and target DRE responses per prompt, response i , and token k of response i .
πθ,πθold,πref,πt
Parameterized policy, old rollout policy, fixed reference policy, and current policy πt:=πθt .
ri,rˉ,σr
Response reward, group-mean reward, and group reward standard deviation.
Ai,AiDr.GRPO
GRPO and Dr.GRPO response-level advantages.
Appendix
Table 4: Notation used throughout the paper. Symbols are grouped by their roles and ordered approximately by their first appearance in the paper.
Figure 2: Validation accuracy at the selected hyperparameter settings. Curves show the mean ± sample standard deviation across three seeds. We show scalar baselines with GRPO and Dr.GRPO advantages separately and include the same UARM and Range-GRPO curves in both comparisons.
In-Distribution
Out-of-Distribution
Model / Method
AIME 24/25
AMC
MATH-500
Minerva
Olympiad
Avg.
ARC-c
GPQA †
MMLU-Pro
Avg.
TraPO (entropy=0.01)
17.3 / 10.3
52.6
77.6
33.8
36.3
38.0
78.8
36.9
39.5
51.7
TraPO (entropy=0.001)
17.7 / 12.2
55.9
79.2
37.5
41.3
40.6
80.4
34.8
43.5
52.9
Appendix
Table 9: Performance of reproduced TraPO configurations with different entropy coefficients. All training methods use 1,024 (1K) ID and 1,024 (1K) OOD prompts. GPQA † denotes GPQA-diamond. Bold and underlined scores indicate the best and second-best results, respectively.
Category
TraPO
UARM
Range-GRPO
Labeled policy
122,880
92,160
92,160
Unlabeled rollouts
122,880
40,960
37,120
RM training
0
28,560
0
RM calibration
0
7,280
0
RM processing
0
76,800
0
Scorer fitting
0
0
1,280
Appendix
Table 10: Response budget for the main experiments. The table compares the response budget for TraPO, UARM, and Range-GRPO. For TraPO, unlabeled rollouts include those generated during warmup. The UARM budget includes RM processing, whereas the Range-GRPO budget includes judge generations.
Setting
Value
Model
Qwen2.5-Math-7B
Learning rate
10−6
Batch size
64 prompts
Responses per prompt
8
Epochs
15
Warmup epochs
10
Appendix
Table 11: UARM training configuration for the cross-domain setting. The table reports the policy training settings in (a) and the reward model training and calibration settings in (b).
Role
Prompts
Labeled policy
768
RM training
184
RM validation
20
Calibration
52
Unlabeled policy
1,024
Total
2,048
Appendix
Table 12: Data allocation for UARM and Range-GRPO in the main experiments. (a) summarizes the UARM allocation across policy learning and RM training, validation, and calibration. (b) details the Range-GRPO allocation across policy learning, scorer fitting, calibration, source DRE, and target DRE.
State
Criterion
Score 0
Score 0.5
Score 1
Response-level assessment
z1
Answer verification
The judge identifies a contradiction when independently checking the submitted answer against the problem.
The judge cannot verify or refute the submitted answer with sufficient confidence.
The judge independently verifies the submitted answer against the problem.
z2
Reasoning consistency
The visible reasoning contains a decisive error or conflicts with the submitted answer.
The visible reasoning is incomplete or insufficient to establish the submitted answer.
The visible reasoning consistently supports the submitted answer.
z3
Requirement compliance
The response omits or violates an explicit task requirement.
It remains unclear whether all explicit requirements are satisfied.
The response satisfies all explicit task requirements.
Group-level agreement
z4
Absolute agreement
No other response gives the same answer.
One to three other responses give the same answer.
Four to seven other responses give the same answer.
Appendix
Table 13: Five components of the judgment state Z . The first three states assess the response itself, while the last two capture agreement with other responses in the same rollout group. Each state takes a score of 0, 0.5, or 1.
Setting
Value
Scorer
XGBoost regressor
Trees / depth
200 / 2
DRE
Random forest
Trees / depth
200 / 4
DRE target per update
768 responses
Refresh
Each unlabeled update
Appendix
Table 14: Range construction settings for the main experiments. The scorer and calibration nonconformity scores remain fixed after warmup, whereas DRE is refreshed at each unlabeled update.
Setting
Value
Batch size
64 prompts
Responses per prompt
8
Warmup epochs
10
Active epochs
5
Optimizer
AdamW
PPO clipping
0.2
Appendix
Table 15: Range-GRPO policy training settings for the main experiments. (a) summarizes the policy training configuration. (b) lists the main optimization and Range-GRPO settings. LR denotes the learning rate.
Method
Active GPUs
Warmup (h)
Active (h)
Training (h)
TraPO (entropy=0.01)
4
12.8
10.6
23.4
TraPO (entropy=0.001)
4
12.6
7.0
19.7
UARM
5
6.5
9.9
16.4
Range-GRPO (ours)
4
6.3
10.0
16.3
Appendix
Table 16: Range-GRPO maintains comparable wall-clock training time despite the additional judge phase. Each run consists of 10 warmup epochs followed by 5 active epochs. Phase durations and total training times are independently rounded to one decimal place.
Chinese Information Processing Laboratory, Institute of Software, Chinese Academy of Sciences · University of Chinese Academy of Sciences · Institute of Information Engineering, Chinese Academy of Sciences +2
Xiamen University · Key Laboratory of Digital Protection and Intelligent Processing of Intangible Cultural Heritage of Fujian and Taiwan (Xiamen University), Ministry of Culture and Tourism