Organizations: Department of Aeronautics and Astronautics, MIT · Department of Statistics, University of Wisconsin–Madison · Department of Informatics, King’s College London · Department of Computer Sciences, University of Wisconsin–Madison
Reinforcement learning (RL) has become a key driver of language model reasoning. Among RL algorithms, Group Relative Policy Optimization (GRPO) is the de facto standard, avoiding the need for a critic by using per-prompt baselines and variance normalization. Yet why and when this normalization helps remains unclear. In this work, we answer both questions through the lens of local curvature of the sequence-level policy gradient: standard deviation normalization implements an adaptive gradient. Theoretically, we prove that under mild conditions GRPO attains an improved convergence rate over unnormalized REINFORCE, with gains characterized by the average within-prompt reward standard deviation across prompts and iterations. We further introduce IS-GRPO, an importance-sampling variant of GRPO whose expected update remains aligned with the full gradient, and prove a convergence guarantee of the same form as REINFORCE's, which is tighter in practice during real training runs by a factor we measure. Empirically, on GSM8K and MATH we validate the curvature--variance link and measure the quantities that govern our bounds along real training runs. At 1.5B scale, per-prompt normalization outperforms both unnormalized and globally normalized baselines, with gains that emerge in a middle phase of training where per-prompt variances become heterogeneous. At 7B scale, the normalization schemes become statistically indistinguishable while IS-GRPO retains a small lead, in line with the theory's prediction that the gains require heterogeneity.
Figures & tables
Figure 2: Training accuracy on GSM8K (1.5B) with phase annotations; left: Easy, right: Hard. standard : per-prompt normalization; no_std : unnormalized (global std omitted for readability; see Table 3 ).
Figure 3: 7B results (Qwen2.5-Math-7B). (a) Consistent with normalization helping when many prompts remain unsolved (AMC held-out). (b) When most prompts are solved, the normalization schemes coincide and IS-GRPO leads.
Figure 4: Trajectory coefficients along the 7B IS-GRPO run. (a) Running average C(T) ; the dashed line is the speed-up condition ( 6 ). (b) Ratio D(T)/DIS(T) .
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Notations
Definition
Q
Set of questions / prompts
O
Set of possible output sequences
o∗(q)
Unique correct output for prompt q ( Assumption 1 )
ai
Index of the correct output for prompt qi
ri
Reward vector for prompt qi
πθ
Sequence-level policy of the LLM parameterized by θ
Table 3: 1.5B: final accuracy (%, mean ± std over 3 seeds) at iteration 500.
Figure 6: ct and C(T) on GSM8K Easy and Hard (1.5B). C(T) stays below 1 throughout training, ending at 0.28 (Easy) and 0.45 (Hard).
Figure 7: 7B held-out accuracy (pass@1) on AIME 2024, MATH Level 3–5 training run, one seed.
Method
Late window
Step 500
IS-GRPO
87.1±0.6
88.7
REINFORCE
86.1±0.6
86.3
Per-prompt std
85.9±0.6
87.1
Global std
85.4±0.6
85.9
Appendix
Table 4: 7B, MATH Level 1–5: greedy accuracy (%) on the training prompts, 3 seeds. Late window: mean over iterations 400/450/500 ( ± across-seed std). Step 500: single checkpoint (across-seed std ≈ 1.8 pp).
1.5B
7B
Base model
Qwen2.5-Math-1.5B
Qwen2.5-Math-7B
Training data
GSM8K Easy/Hard
MATH Lv. 1–5 / Lv. 3–5
Evaluation
GSM8K val, MATH Lv. 1–2
training prompts / AMC
LoRA r / α
64 / 128
32 / 32
LoRA target
all linear
all linear
Optimizer
AdamW ( β1=0.9,β2=0.999 )
Appendix
Table 5: Hyperparameters. For 7B, entries with “/” give the Level 1–5 run / the Level 3–5 run.
Group Relative Policy Optimization (GRPO) is widely used to train reasoning language models, where it computes advantages by centering and normalizing rewards across rollouts of the same prompt. For multiple rewards, GRPO sums the reward components and normalizes the total reward by its within-group standard deviation. The corresponding variance equals the sum of all pairwise reward covariances. For a fixed centered reward, larger aggregate covariance produces smaller advantages, and vice versa, allowing update magnitudes to adapt to reward dependence. However, correlated rewards with large scales can dominate this normalization and suppress signals from smaller-scale rewards. We propose Correlation-Normalized GRPO (CorrGRPO), which normalizes pairwise covariances into Pearson correlation coefficients. CorrGRPO keeps the centered total reward unchanged while balancing the influence of differently scaled rewards on the correlation-based normalization. This allows advantage magnitudes to adapt to reward correlations without the normalization being dominated by large-scale reward components. We compare CorrGRPO with GRPO and other variants on code generation, tool calling, and agent security, using models ranging from 0.5B to 8B parameters. These tasks all involve multiple rewards that can improve together or present tradeoffs. Results show improvements across three domains, including code generation, tool calling, and agent security. Our code is available at https://github.com/HKUST-KnowComp/CorrGRPO.
Group Relative Policy Optimization (GRPO) is the dominant reinforcement learning algorithm for training reasoning capabilities in large language models, notably adopted by DeepSeek-R1. The recent improvement Dr. GRPO (COLM 2025) identifies the response-level length bias caused by per-trajectory length normalization in GRPO and proposes removing this normalization, claiming the resulting optimizer is "unbiased." We show that this claim is incomplete. Specifically, we establish an impossibility theorem: under the standard outcome reward + GRPO setting, no length-based weighting scheme can simultaneously achieve the following two properties. (P1) Gradient unbiasedness: the gradient estimator is an unbiased estimate of the true policy gradient. (P2) Length invariance: each trajectory's effective contribution to the gradient is independent of its token length. GRPO approximately satisfies P2 but violates P1; Dr. GRPO satisfies P1 but violates P2. We characterize the complete tradeoff spectrum via the parametric family f_alpha(L) = L^{alpha - 1}, where alpha = 0 recovers GRPO, alpha = 1 recovers Dr. GRPO, and provide quantitative analysis showing that Dr. GRPO's length bias can cause longer trajectories to dominate gradient updates by a factor proportional to the length ratio. Our results reveal that neither algorithm is universally "done right"; they occupy opposite ends of a fundamental and unavoidable tradeoff.
Group Relative Policy Optimization (GRPO) has become a standard reinforcement learning method for post-training language models. Recent work shows that GRPO can reduce the base model's reasoning capacity and underperform it in Pass@k when k is large, indicating reduced coverage of reasoning paths. We find that this reduction is associated with GRPO concentrating on responses that the base model already generates with high probability. We trace this concentration to two mechanisms in the GRPO update. At the response level, high-probability responses dominate the group gradient through repeated occurrence. At the token level, GRPO's importance ratio scales gradients, further reinforcing tokens that become more likely under the current policy. We propose ReCo, a reweighting method that addresses both effects. Response contributions are normalized by their expected occurrence within the rollout group, and the token-level importance ratio is replaced with a variance-based ratio that gives larger update scale to non-saturated decision points where alternative token choices remain plausible. Across Qwen2.5-Math-1.5B/7B and Llama-3.1-8B-Instruct on five mathematical reasoning benchmarks, ReCo improves Pass@k for large values of k and is comparable to GRPO for small values of k.
Junoh Park, Junseo Hwang, Wonguk Cho +1
Graduate School of Data Science, Seoul National University