Organizations: Department of Aeronautics and Astronautics, MIT · Department of Statistics, University of Wisconsin–Madison · Department of Informatics, King’s College London · Department of Computer Sciences, University of Wisconsin–Madison
Reinforcement learning (RL) has become a key driver of language model reasoning. Among RL algorithms, Group Relative Policy Optimization (GRPO) is the de facto standard, avoiding the need for a critic by using per-prompt baselines and variance normalization. Yet why and when this normalization helps remains unclear. In this work, we answer both questions through the lens of local curvature of the sequence-level policy gradient: standard deviation normalization implements an adaptive gradient. Theoretically, we prove that under mild conditions GRPO attains an improved convergence rate over unnormalized REINFORCE, with gains characterized by the average within-prompt reward standard deviation across prompts and iterations. We further introduce IS-GRPO, an importance-sampling variant of GRPO whose expected update remains aligned with the full gradient, and prove a convergence guarantee of the same form as REINFORCE's, which is tighter in practice during real training runs by a factor we measure. Empirically, on GSM8K and MATH we validate the curvature--variance link and measure the quantities that govern our bounds along real training runs. At 1.5B scale, per-prompt normalization outperforms both unnormalized and globally normalized baselines, with gains that emerge in a middle phase of training where per-prompt variances become heterogeneous. At 7B scale, the normalization schemes become statistically indistinguishable while IS-GRPO retains a small lead, in line with the theory's prediction that the gains require heterogeneity.
Figures & tables
Figure 2: Training accuracy on GSM8K (1.5B) with phase annotations; left: Easy, right: Hard. standard : per-prompt normalization; no_std : unnormalized (global std omitted for readability; see Table 3 ).
Figure 3: 7B results (Qwen2.5-Math-7B). (a) Consistent with normalization helping when many prompts remain unsolved (AMC held-out). (b) When most prompts are solved, the normalization schemes coincide and IS-GRPO leads.
Figure 4: Trajectory coefficients along the 7B IS-GRPO run. (a) Running average C(T) ; the dashed line is the speed-up condition ( 6 ). (b) Ratio D(T)/DIS(T) .
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Notations
Definition
Q
Set of questions / prompts
O
Set of possible output sequences
o∗(q)
Unique correct output for prompt q ( Assumption 1 )
ai
Index of the correct output for prompt qi
ri
Reward vector for prompt qi
πθ
Sequence-level policy of the LLM parameterized by θ
Table 3: 1.5B: final accuracy (%, mean ± std over 3 seeds) at iteration 500.
Figure 6: ct and C(T) on GSM8K Easy and Hard (1.5B). C(T) stays below 1 throughout training, ending at 0.28 (Easy) and 0.45 (Hard).
Figure 7: 7B held-out accuracy (pass@1) on AIME 2024, MATH Level 3–5 training run, one seed.
Method
Late window
Step 500
IS-GRPO
87.1±0.6
88.7
REINFORCE
86.1±0.6
86.3
Per-prompt std
85.9±0.6
87.1
Global std
85.4±0.6
85.9
Appendix
Table 4: 7B, MATH Level 1–5: greedy accuracy (%) on the training prompts, 3 seeds. Late window: mean over iterations 400/450/500 ( ± across-seed std). Step 500: single checkpoint (across-seed std ≈ 1.8 pp).
1.5B
7B
Base model
Qwen2.5-Math-1.5B
Qwen2.5-Math-7B
Training data
GSM8K Easy/Hard
MATH Lv. 1–5 / Lv. 3–5
Evaluation
GSM8K val, MATH Lv. 1–2
training prompts / AMC
LoRA r / α
64 / 128
32 / 32
LoRA target
all linear
all linear
Optimizer
AdamW ( β1=0.9,β2=0.999 )
Appendix
Table 5: Hyperparameters. For 7B, entries with “/” give the Level 1–5 run / the Level 3–5 run.