Reward hacking occurs when policy optimization exploits a brittle reward interface or an overly permissive proxy objective, improving the training score without improving the underlying response quality. This phenomenon is amplified in group-relative policy optimization: an unsupported reward can shift the group baseline and alter the updates of other rollouts, while post-hoc or purely relative weighting cannot represent group-wide uncertainty. We propose \emph{Self-Tuned Anchored Reliability Group-Relative Policy Optimization} (STAR-GRPO), a reliability-first advantage estimator based on paired assessments of the same rollout. STAR separates the quality signal from its learning influence: score disagreement determines rollout reliability, relative reliability enters a self-tuned robust location--scale fit before group normalization, and absolute group reliability attenuates the resulting bounded advantage. The analysis establishes coordinate and second-moment bounds, characterizes exact centering through the weighted location equation, and gives reliability-dependent attenuation guarantees for outlying rewards. We evaluate STAR-GRPO in two complementary reward-hacking regimes. In token-interface exploitation, STAR prevents runaway optimization of the deployed-interface score while improving the canonical quality signal. In rubric-proxy overoptimization for medical reasoning, STAR improves independent semantic evaluation, narrows the proxy--judge discrepancy, and reduces overclaim while optimizing the same task proxy. Together, these results show that reliability-first normalization offers a principled way to limit unsupported reward influence on both group baselines and policy updates, while retaining the task reward as the optimization target.
Figures & tables
Figure 1: Overview of token-interface reward hacking and STAR-GRPO. A deployed representation and a canonical rendering of the same rollout are scored in parallel; their discrepancy determines reliability, which enters group statistics before normalization. The final STAR advantage combines a bounded robust score with rollout- and group-level reliability, preventing unsupported rewards from dominating either the baseline or the policy update.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Input
Output
Role
Anchored quality
Rcanon,Robs,κ
r=h(Rκ)
Prespecified optimization signal
Canonical anchor
Policy output
Rcanon
Reference representation
Discrepancy calibration
Robs−Rcanon
w∈(0,1]
Rollout reliability
Weighted self-tuning
(w,r)
(μ,v,φ)
Reliability-first relative score
Absolute group factor
(wˉ,w,φ)
ASTAR
Preserve collective distrust
Appendix
Table 1: Distinct signals and their roles in STAR-GRPO.
Method
R0canon
R855canon
R855obs
Length 855
TOMPA-GRPO
–
–
9.641
2048.0
STAR-GRPO
3.121
4.902
−0.570
1872.1
Appendix
Table 2: RQ1 validation summary at step 855. STAR improves the canonical score while suppressing the runaway deployed-interface reward exploited by TOMPA-GRPO. Scores are raw reward-model outputs; length is in policy tokens.
Component
Setting
Policy and reward model
Llama-3.2-1B-Instruct and Skywork-Reward-V2-Qwen3-8B
Data
10,000 WildChat training prompts; 100 curated NoveltyBench validation prompts
STAR calibration
1,000 disjoint WildChat prompts, used only for discrepancy calibration
Length limits
512 prompt tokens; 2,048 response tokens
Batching and groups
Training batch 64; validation batch 100; eight training and eight validation rollouts per prompt
Sampling
Sampling enabled; temperature 1.0; top- p 1.0; top- k−1
Appendix
Table 3: Common base settings for the three RQ1 runs. All reward-model evaluations use a 4,096-token input limit; loss aggregation is listed per method in Table 4 . The calibration row applies only to STAR.
Parameter
TOMPA-GRPO
STAR-GRPO
Canonical-GRPO
Quality input
Robs
10tanh(Rcanon/10)
10tanh(Rcanon/10)
Advantages
Group mean/std
Reliability-weighted robust fit
Group mean/std
Loss aggregation
token-mean
token-mean
seq-mean- token-mean
RM token limit
4,096
4,096; right truncation
4,096; right truncation
Online RM calls per rollout
One observed view
Observed and canonical
Observed and canonical; observed logged only
Discrepancy calibration
None
1,000 prompts; six contexts; q=1.79285
None
Appendix
Table 4: Recorded RQ1 method settings. STAR and the canonical-only diagnostic control share the same bounded canonical quality path and 4,096-token reward-model input limit; STAR additionally uses discrepancy calibration and reliability-first robust advantages.
Parameter
Value
zA , λ , α
2.0, 1.0, 0.05
fit_fraction , min_context_size , δ
0.7, 50, 0.05
smin , smax
0.001, 100.0
εD , εw
10−6 , 10−6
vmin , vmax
0.001, 100.0
Wmin , Gmin
1.0, 2.0
Appendix
Table 5: STAR-specific numerical parameters in the token-space run.
Metric
STAR-GRPO
Canonical-GRPO
Canonical validation gain, steps 0–855
1.781
3.236
Late-window canonical reward
4.808
6.224
Late-window observed reward
−0.561
−1.865
Late-window response length (tokens)
1881.3
371.6
Training response length, step 855 (tokens)
1896.5
976.3
Training length-clip ratio, step 855
0.922
0.117
Appendix
Table 6: Additional RQ1 diagnostic control statistics. The late validation window averages the 11 checkpoints from step 805 to 855; training rows use the common step-855 endpoint.
Metric
GRPO
STAR-GRPO
Δ
Proxy score
0.5640
0.5492
−0.0148
Independent-judge score
0.2706
0.3174
+0.0469
Proxy–judge gap
0.2935
0.2318
−0.0617
Criterion pass rate
0.4752
0.5049
+0.0297
Overclaim fraction
0.2464
0.2204
−0.0260
Appendix
Table 7: RQ2 independent evaluation at the latest common eligible checkpoint. The Δ column reports STAR-GRPO minus GRPO, computed from the unrounded means. STAR improves all external-quality and reward-hacking diagnostics while optimizing the same scalar proxy reward.
Component
Setting
Policy model
Qwen3-4B
Training data
RubricHub-Medical
Evaluation data
HealthBench-Hard
Hardware
One node with four GPUs; rollout tensor parallelism 2
Rollouts per training prompt
16
Training batch size
64
Appendix
Table 8: Shared realized settings for the rubric-reward comparison.
Parameter
GRPO
STAR-GRPO
Advantage estimator
grpo ; standard deviation normalization
star_grpo ; standard normalization disabled
Training reward
Fixed rubric-conditioned proxy score
Same fixed rubric-conditioned proxy score
Proxy judge
openai/gpt-4o-mini
gpt-4o-mini
Anchor judge
Not used
gemini-2.5-flash-lite ; training only and rubric-free
Evaluation judge
anthropic/claude-sonnet-4-6
claude-sonnet-4-6 ; evaluation only
Evaluation usage
Not used in policy updates
Not used in policy updates
Appendix
Table 9: Method-specific reward and advantage configuration. Both methods optimize the identical rubric-conditioned scalar proxy reward; STAR additionally uses a rubric-free semantic anchor exclusively to compute rollout reliability.
Parameter
Value
star_gap_center
0.0
star_gap_scale
0.1
star_conformal_threshold
1.645
star_reliability_decay
2.0
star_z_a
2.0
star_scale_min , star_scale_max
0.01, 1.0
Appendix
Table 10: STAR-GRPO numerical parameters used for RQ2.
Design condition
Resulting guarantee
Role in STAR
Paired reward views
Observable discrepancy D
Measures support for the optimized score
Fixed h,κ
Auditable quality path r=h(Rκ)
Separates quality from reliability
Trusted disjoint fitting data
External discrepancy reference
Prevents current policy groups from defining their own baseline
Clean calibration exchangeability
Marginal clean downweighting ≤α
Calibrates when attenuation begins
Self-tuned robust fit
Context location and scale control
Adapts to heterogeneous discrepancy scales
Attack-score separation
Exponential expected-weight attenuation
Converts discrepancy separation into influence reduction
Appendix
Table 11: Core STAR-GRPO design conditions and the guarantees they enable.