Reward hacking occurs when policy optimization exploits a brittle reward interface or an overly permissive proxy objective, improving the training score without improving the underlying response quality. This phenomenon is amplified in group-relative policy optimization: an unsupported reward can shift the group baseline and alter the updates of other rollouts, while post-hoc or purely relative weighting cannot represent group-wide uncertainty. We propose \emph{Self-Tuned Anchored Reliability Group-Relative Policy Optimization} (STAR-GRPO), a reliability-first advantage estimator based on paired assessments of the same rollout. STAR separates the quality signal from its learning influence: score disagreement determines rollout reliability, relative reliability enters a self-tuned robust location--scale fit before group normalization, and absolute group reliability attenuates the resulting bounded advantage. The analysis establishes coordinate and second-moment bounds, characterizes exact centering through the weighted location equation, and gives reliability-dependent attenuation guarantees for outlying rewards. We evaluate STAR-GRPO in two complementary reward-hacking regimes. In token-interface exploitation, STAR prevents runaway optimization of the deployed-interface score while improving the canonical quality signal. In rubric-proxy overoptimization for medical reasoning, STAR improves independent semantic evaluation, narrows the proxy--judge discrepancy, and reduces overclaim while optimizing the same task proxy. Together, these results show that reliability-first normalization offers a principled way to limit unsupported reward influence on both group baselines and policy updates, while retaining the task reward as the optimization target.
Figures & tables
Figure 1: Overview of token-interface reward hacking and STAR-GRPO. A deployed representation and a canonical rendering of the same rollout are scored in parallel; their discrepancy determines reliability, which enters group statistics before normalization. The final STAR advantage combines a bounded robust score with rollout- and group-level reliability, preventing unsupported rewards from dominating either the baseline or the policy update.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Input
Output
Role
Anchored quality
Rcanon,Robs,κ
r=h(Rκ)
Prespecified optimization signal
Canonical anchor
Policy output
Rcanon
Reference representation
Discrepancy calibration
Robs−Rcanon
w∈(0,1]
Rollout reliability
Weighted self-tuning
(w,r)
(μ,v,φ)
Reliability-first relative score
Absolute group factor
(wˉ,w,φ)
ASTAR
Preserve collective distrust
Appendix
Table 1: Distinct signals and their roles in STAR-GRPO.
Method
R0canon
R855canon
R855obs
Length 855
TOMPA-GRPO
–
–
9.641
2048.0
STAR-GRPO
3.121
4.902
−0.570
1872.1
Appendix
Table 2: RQ1 validation summary at step 855. STAR improves the canonical score while suppressing the runaway deployed-interface reward exploited by TOMPA-GRPO. Scores are raw reward-model outputs; length is in policy tokens.
Component
Setting
Policy and reward model
Llama-3.2-1B-Instruct and Skywork-Reward-V2-Qwen3-8B
Data
10,000 WildChat training prompts; 100 curated NoveltyBench validation prompts
STAR calibration
1,000 disjoint WildChat prompts, used only for discrepancy calibration
Length limits
512 prompt tokens; 2,048 response tokens
Batching and groups
Training batch 64; validation batch 100; eight training and eight validation rollouts per prompt
Sampling
Sampling enabled; temperature 1.0; top- p 1.0; top- k−1
Appendix
Table 3: Common base settings for the three RQ1 runs. All reward-model evaluations use a 4,096-token input limit; loss aggregation is listed per method in Table 4 . The calibration row applies only to STAR.
Parameter
TOMPA-GRPO
STAR-GRPO
Canonical-GRPO
Quality input
Robs
10tanh(Rcanon/10)
10tanh(Rcanon/10)
Advantages
Group mean/std
Reliability-weighted robust fit
Group mean/std
Loss aggregation
token-mean
token-mean
seq-mean- token-mean
RM token limit
4,096
4,096; right truncation
4,096; right truncation
Online RM calls per rollout
One observed view
Observed and canonical
Observed and canonical; observed logged only
Discrepancy calibration
None
1,000 prompts; six contexts; q=1.79285
None
Appendix
Table 4: Recorded RQ1 method settings. STAR and the canonical-only diagnostic control share the same bounded canonical quality path and 4,096-token reward-model input limit; STAR additionally uses discrepancy calibration and reliability-first robust advantages.
Parameter
Value
zA , λ , α
2.0, 1.0, 0.05
fit_fraction , min_context_size , δ
0.7, 50, 0.05
smin , smax
0.001, 100.0
εD , εw
10−6 , 10−6
vmin , vmax
0.001, 100.0
Wmin , Gmin
1.0, 2.0
Appendix
Table 5: STAR-specific numerical parameters in the token-space run.
Metric
STAR-GRPO
Canonical-GRPO
Canonical validation gain, steps 0–855
1.781
3.236
Late-window canonical reward
4.808
6.224
Late-window observed reward
−0.561
−1.865
Late-window response length (tokens)
1881.3
371.6
Training response length, step 855 (tokens)
1896.5
976.3
Training length-clip ratio, step 855
0.922
0.117
Appendix
Table 6: Additional RQ1 diagnostic control statistics. The late validation window averages the 11 checkpoints from step 805 to 855; training rows use the common step-855 endpoint.
Metric
GRPO
STAR-GRPO
Δ
Proxy score
0.5640
0.5492
−0.0148
Independent-judge score
0.2706
0.3174
+0.0469
Proxy–judge gap
0.2935
0.2318
−0.0617
Criterion pass rate
0.4752
0.5049
+0.0297
Overclaim fraction
0.2464
0.2204
−0.0260
Appendix
Table 7: RQ2 independent evaluation at the latest common eligible checkpoint. The Δ column reports STAR-GRPO minus GRPO, computed from the unrounded means. STAR improves all external-quality and reward-hacking diagnostics while optimizing the same scalar proxy reward.
Component
Setting
Policy model
Qwen3-4B
Training data
RubricHub-Medical
Evaluation data
HealthBench-Hard
Hardware
One node with four GPUs; rollout tensor parallelism 2
Rollouts per training prompt
16
Training batch size
64
Appendix
Table 8: Shared realized settings for the rubric-reward comparison.
Parameter
GRPO
STAR-GRPO
Advantage estimator
grpo ; standard deviation normalization
star_grpo ; standard normalization disabled
Training reward
Fixed rubric-conditioned proxy score
Same fixed rubric-conditioned proxy score
Proxy judge
openai/gpt-4o-mini
gpt-4o-mini
Anchor judge
Not used
gemini-2.5-flash-lite ; training only and rubric-free
Evaluation judge
anthropic/claude-sonnet-4-6
claude-sonnet-4-6 ; evaluation only
Evaluation usage
Not used in policy updates
Not used in policy updates
Appendix
Table 9: Method-specific reward and advantage configuration. Both methods optimize the identical rubric-conditioned scalar proxy reward; STAR additionally uses a rubric-free semantic anchor exclusively to compute rollout reliability.
Parameter
Value
star_gap_center
0.0
star_gap_scale
0.1
star_conformal_threshold
1.645
star_reliability_decay
2.0
star_z_a
2.0
star_scale_min , star_scale_max
0.01, 1.0
Appendix
Table 10: STAR-GRPO numerical parameters used for RQ2.
Design condition
Resulting guarantee
Role in STAR
Paired reward views
Observable discrepancy D
Measures support for the optimized score
Fixed h,κ
Auditable quality path r=h(Rκ)
Separates quality from reliability
Trusted disjoint fitting data
External discrepancy reference
Prevents current policy groups from defining their own baseline
Clean calibration exchangeability
Marginal clean downweighting ≤α
Calibrates when attenuation begins
Self-tuned robust fit
Context location and scale control
Adapts to heterogeneous discrepancy scales
Attack-score separation
Exponential expected-weight attenuation
Converts discrepancy separation into influence reduction
Appendix
Table 11: Core STAR-GRPO design conditions and the guarantees they enable.
Group Relative Policy Optimization (GRPO) eliminates the learned critic in PPO by using the mean reward of grouped rollouts as a baseline. We provide a rigorous derivation of GRPO from first principles of the policy gradient theorem, revealing a fundamental credit assignment failure: under output-only reward, every token in a rollout receives identical advantage, collapsing token-level credit to a single scalar. We prove this induces gradient sparsity that intensifies over training, and demonstrate empirically via SVD analysis of GRPO gradients on Nemotron-4B/GSM8K that the gradient matrix has effective rank ≈ 2 regardless of group size R∈{2,4,8}. We formalize this as an intrinsic rank-2 structure arising from the zero-sum constraint on advantages and derive conditions under which GRPO's baseline is optimal. Our results characterize when GRPO's simplicity is theoretically justified and identify the credit assignment bottleneck as the key limitation for multi-step reasoning.
Group Relative Policy Optimization (GRPO) is a standard algorithm for reinforcement learning from verifiable rewards, but its group-mean-centered advantage can fail under binary rewards. The failure mode is gradient starvation: when every response in a group is correct or every response is wrong, the centered advantage is exactly zero and the policy receives no learning signal. We prove that the true degeneracy rate always exceeds the i.i.d. Bernoulli prediction by Jensen's inequality, and observe a 0.69 degeneracy rate at group size four in logged Qwen3.5-9B GSM8K training. We then show that the fixed-reference Sign advantage, A=2r−1, performs pass@G failure descent by increasing the probability that at least one sample in the group succeeds. On the full GSM8K test set across seven seeds, Sign reaches 73.8% accuracy versus 28.4% for standard normalized group-mean DrGRPO at group size four, a 45.4 point gain with p<0.0001. The effect is directionally consistent on Llama-3.1-8B and positive but underpowered on a MATH-500 transfer check. Pass@k analysis indicates that the main benefit is search compression rather than large capacity expansion, aligning the empirical gains with recent RLVR ceiling observations.
Group-relative policy optimization relies on reward-derived advantages and sequence-level likelihood weights, both of which can be sensitive to localized outliers. Extreme rewards can collapse the contrast among clean responses after group normalization, while token-level log-ratio perturbations can alter sequence weights and clipping decisions. We introduce RoVR-GSPO, a dual-channel robust optimizer that addresses these failure modes separately. Its reward channel combines robust reference estimation with bounded residual credit, while its ratio channel uses differentiable SoftRoVR aggregation to construct robust sequence weights. We provide stability and efficiency analyses for both channels. Experiments on mathematical reasoning, long-context summarization, and tool-call annotation show consistent improvements over GSPO, while controlled perturbation studies demonstrate stronger robustness to reward contamination and token-ratio anomalies.
Zhongyi Li, Wan Tian, Xiang Xu +4
Beihang University · Peking University · Nanjing University