Group Relative Policy Optimization (GRPO) is widely used to train reasoning language models, where it computes advantages by centering and normalizing rewards across rollouts of the same prompt. For multiple rewards, GRPO sums the reward components and normalizes the total reward by its within-group standard deviation. The corresponding variance equals the sum of all pairwise reward covariances. For a fixed centered reward, larger aggregate covariance produces smaller advantages, and vice versa, allowing update magnitudes to adapt to reward dependence. However, correlated rewards with large scales can dominate this normalization and suppress signals from smaller-scale rewards. We propose Correlation-Normalized GRPO (CorrGRPO), which normalizes pairwise covariances into Pearson correlation coefficients. CorrGRPO keeps the centered total reward unchanged while balancing the influence of differently scaled rewards on the correlation-based normalization. This allows advantage magnitudes to adapt to reward correlations without the normalization being dominated by large-scale reward components. We compare CorrGRPO with GRPO and other variants on code generation, tool calling, and agent security, using models ranging from 0.5B to 8B parameters. These tasks all involve multiple rewards that can improve together or present tradeoffs. Results show improvements across three domains, including code generation, tool calling, and agent security. Our code is available at https://github.com/HKUST-KnowComp/CorrGRPO.
Figures & tables
Figure 1: An overview of CorrGRPO. (1) Left: CorrGRPO replaces total-reward standard deviation normalization with a correlation-based denominator while retaining the centered total reward. (2) Right: Validation accuracy (top) and efficiency (bottom), measured by Mean@1, for Qwen2.5-Coder-7B-Instruct on LeetCodeDataset.
Figure 2: An example for comparing GRPO and CorrGRPO. (a) A group of four trajectories with three rewards. Rewards r1 and r2 are highly correlated, while r3 has a larger scale and weak correlations with both. (b) Covariance-matrix elements. In GRPO, the large scale of r3 dominates the denominator. (c) Pearson correlation coefficients matrix elements. In CorrGRPO, the strong correlation between r1 and r2 has a greater influence on the normalization denominator than r3 does.
Model
LeetCodeDataset (In-dataset Eval)
HumanEval
MBPP
LCB v6
Avg.
Efficiency
Executable
Pass@1
Pass@1
Pass@1
Pass@1
Pass@1
Qwen2.5-Coder-0.5B-Instruct
50.00
61.40
1.75
54.88
33.60
5.14
23.84
+ GRPO
60.71
87.72
3.07
57.32
40.00
3.43
25.96
+ GDPO
60.94
88.60
3.51
57.32
38.40
4.57
25.95
+ CorrGRPO
67.86
89.04
3.95
59.76
42.20
6.29
28.05
Qwen2.5-Coder-1.5B-Instruct
58.93
75.44
3.07
61.59
54.40
10.86
32.48
Table 1: Coding RL results on LeetCodeDataset, HumanEval, MBPP, and LCB v6. Efficiency is the mean percentage of eligible programs that run faster than the reference code. Executable denotes running time success rate. Avg. is the arithmetic mean of the Pass@1 evaluation results across 4 datasets.
Figure 3: Pareto frontier of correctness–efficiency tradeoff on LeetCodeDataset.
Model
RLLA-4K (In-dataset Eval)
API-Bank
Avg.
Function
Param. Name
Param. Value
All Exact
V1
V2
V3
Avg.
Qwen2.5-7B-Instruct
73.17
69.77
54.78
38.03
17.29
19.26
24.49
20.35
29.19
+ GRPO
97.65
95.60
76.91
63.38
79.45
44.44
35.92
53.27
58.33
+ CorrGRPO
95.07
94.95
82.14
67.61
81.45
46.67
35.10
54.41
61.01
Qwen3-4B-Thinking-2507
33.80
31.46
25.35
21.13
53.38
39.26
31.43
41.36
31.25
+ GRPO
94.84
92.08
72.07
57.75
79.20
52.59
52.24
61.34
59.55
Table 2: Tool-call results on RLLA-4K and API-Bank. API-Bank scores use the same all-exact metric for the correctness of final tool call. The rightmost Avg. is the arithmetic mean of RLLA-4K all-exact score and API-Bank Avg. score.
Table 3: Agent utility and security results. Utility measures the agent’s tool-calling success rate, while ASR measures the success rate of prompt injection attacks. Joint accuracy combines utility and security as (1−attack_success)×2clean_utility+utility_under_attack . In InjecAgent, all reported scores represent ASR.
Figure 4: Validation reward dynamics of GRPO, GDPO, and CorrGRPO. Curves report validation mean@1 scores with exponential moving average smoothing (decay = 0.6).
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Training dynamics on LeetCodeDataset using equally weighted runtime-success and passing rewards.
Figure 6: LeetCodeDataset: validation rewards. Columns correspond to model backbones. Rows report the task-specific validation mean@1 reward components and the total reward. The horizontal axis denotes training steps. Curves use exponential moving average smoothing (decay 0.6) and end at the last available observation.
Figure 7: LeetCodeDataset: training rewards. Columns correspond to model backbones. Rows report the mean training reward components and the mean total reward. The horizontal axis denotes training steps. Curves use exponential moving average smoothing (decay 0.6) and end at the last available observation.
Figure 8: RLLA-4K: validation rewards. Columns correspond to model backbones. Rows report the task-specific validation mean@1 reward components and the total reward. The horizontal axis denotes training steps. Curves use exponential moving average smoothing (decay 0.6) and end at the last available observation.
Figure 9: RLLA-4K: training rewards. Columns correspond to model backbones. Rows report the mean training reward components and the mean total reward. The horizontal axis denotes training steps. Curves use exponential moving average smoothing (decay 0.6) and end at the last available observation.
Figure 10: AgentDojo: validation rewards. Columns correspond to model backbones. Rows report the task-specific validation mean@1 reward components and the total reward. Joint Reward denotes the logged trajectory-level metric. The horizontal axis denotes training steps. AgentDojo panels are restricted to the overlapping recorded step ranges of GRPO and CorrGRPO. Curves use exponential moving average smoothing (decay 0.6) and end at the last available observation.
Figure 11: AgentDojo: training rewards. Columns correspond to model backbones. Rows report the mean training reward components and the mean total reward. The horizontal axis denotes training steps. AgentDojo panels are restricted to the overlapping recorded step ranges of GRPO and CorrGRPO. Curves use exponential moving average smoothing (decay 0.6) and end at the last available observation.
Figure 12: Response length and entropy on LeetCodeDataset (top), RLLA-4K (middle), and AgentDojo (bottom). Within each dataset, columns correspond to model backbones; the two rows show mean response length in tokens and policy entropy, respectively. Horizontal axes denote training steps. AgentDojo panels are restricted to the overlapping recorded step ranges of GRPO and CorrGRPO. Curves use exponential moving average smoothing (decay 0.6) and end at the last available observation within the displayed range.
Training data / method
Training
HumanEval
MBPP
LiveCode Bench
LeetCode Dataset
examples
↑
↑
↑
↑
Original SFT results: Qwen2.5-Coder-7B
24-08–25-02
24-07–25-03
Magicoder Evol-Instruct-110K
111.1K
77.4
74.1
15.1
13.7
Magicoder OSS-Instruct-75K
75.1K
73.8
76.5
15.1
12.9
Open-R1 CodeForces-CoT
9.5K
79.9
74.1
15.8
13.3
Appendix
Table 4: Code-generation Pass@1. Evaluation splits are indicated in the table.
Model
Clean utility ↑
Utility under attack ↑
Targeted ASR ↓
Claude 3 Opus
66.61±3.69
52.46±3.90
11.29±2.47
Claude 3 Sonnet
53.10±3.90
33.23±3.68
26.71±3.46
Claude 3.5 Sonnet
78.22±3.23
51.19±3.91
33.86±3.70
Command-R+
25.44±3.40
25.12±3.39
0.95±0.76
Gemini 1.5 Flash
36.09±3.75
34.18±3.71
12.24±2.56
Gemini 1.5 Pro
45.63±3.89
28.93±3.54
25.60±3.41
Appendix
Table 5: AgentDojo with original 95% confidence intervals; protocols differ.
Model
Clean utility ↑
ASR (OPI) ↓
Claude-3.5 Sonnet
100.00
59.70
LLaMA3-70B
66.50
43.70
GPT-4o
79.00
62.45
Gemma2-27B
31.50
14.20
LLaMA3.1-70B
21.25
12.10
Qwen2-7B
9.75
9.00
Appendix
Table 6: Clean Utility and OPI ASR on Agent Security Bench.
Base setting: ASR-valid ↓
Model
Direct harm
Data stealing
Overall
S1
S2
Total
total
Original prompted agents (ReAct)
Qwen-1.8B
36.1
35.1
82.6
17.6
29.7
Qwen-72B
8.7
37.9
98.4
37.1
23.2
Mistral-7B
13.4
25.0
87.8
20.1
16.7
Appendix
Table 7: InjecAgent base-setting ASR-valid. S1/S2 denote data extraction/transmission; S2 is conditional on reaching that stage.
Model
LeetCodeDataset (In-dataset Eval)
HumanEval
MBPP
LCB v6
Avg.
Efficiency
Executable
Pass@1
Pass@1
Pass@1
Pass@1
Pass@1
CISPO
34.51
83.77
21.93
76.83
79.00
23.43
50.30
CISPO + CorrGRPO
38.44
82.46
25.00
81.71
77.80
21.71
51.56
GDPO
41.28
84.21
16.23
76.83
79.60
22.86
48.88
GDPO + CorrGRPO
44.64
86.40
17.11
76.83
77.20
21.71
48.21
DAPO
40.48
81.14
20.18
81.71
77.40
23.43
50.68
Appendix
Table 8: Results of compatibility RL experiments.
Figure 13: Effect of reward scale and correlation on advantage normalization at fixed centered total reward. (a) Advantage as the third reward’s standard deviation varies. (b) Relative GRPO advantages under separate correlation perturbations. (c) Corresponding CorrGRPO responses.
Reinforcement learning (RL) has become a key driver of language model reasoning. Among RL algorithms, Group Relative Policy Optimization (GRPO) is the de facto standard, avoiding the need for a critic by using per-prompt baselines and variance normalization. Yet why and when this normalization helps remains unclear. In this work, we answer both questions through the lens of local curvature of the sequence-level policy gradient: standard deviation normalization implements an adaptive gradient. Theoretically, we prove that under mild conditions GRPO attains an improved convergence rate over unnormalized REINFORCE, with gains characterized by the average within-prompt reward standard deviation across prompts and iterations. We further introduce IS-GRPO, an importance-sampling variant of GRPO whose expected update remains aligned with the full gradient, and prove a convergence guarantee of the same form as REINFORCE's, which is tighter in practice during real training runs by a factor we measure. Empirically, on GSM8K and MATH we validate the curvature--variance link and measure the quantities that govern our bounds along real training runs. At 1.5B scale, per-prompt normalization outperforms both unnormalized and globally normalized baselines, with gains that emerge in a middle phase of training where per-prompt variances become heterogeneous. At 7B scale, the normalization schemes become statistically indistinguishable while IS-GRPO retains a small lead, in line with the theory's prediction that the gains require heterogeneity.
Cheng Ge, Caitlyn Heqi Yin, Hao Liang +1
Department of Aeronautics and Astronautics, MIT · Department of Statistics, University of Wisconsin–Madison · Department of Informatics, King’s College London +1
Multi-reward reinforcement learning trains large language models to satisfy multiple behavioral objectives simultaneously. Reward-wise normalization, as used in GDPO, preserves reward-specific relative information within rollout groups, but different objectives can still exhibit uneven learning progress. We study this behavior through advantage energy, the sum of a reward's squared advantages over a batch. Under idealized GDPO normalization, we show that this energy is proportional to active-group density: the fraction of rollout groups in which the reward provides nonzero relative advantages. This reveals a residual batch-level signal imbalance and provides a basis for calibrating reward contributions. Based on this relation, we propose Density-Aware Reward Aggregation (DARA). We derive an inverse-square-root density correction that gives greater weight to signals from less frequently active rewards. DARA computes its weights from each rollout batch, adapting to changes in reward activity throughout training without modifying the underlying policy optimization objective. Experiments on tool calling and mathematical reasoning show that DARA learns the targeted behaviors faster than GDPO, reaching high format compliance in up to 26% fewer training steps on tool calling and near-saturated length compliance in up to 65% fewer steps on mathematical reasoning, while remaining competitive in final performance. Our code is available at https://github.com/zhaihaotian/DARA.
Tong Zheng, Skylar Zhai, Zhan Cheng +7
University of Chinese Academy of Sciences · University of Minnesota Twin Cities · University of Wisconsin–Madison +5
Group Relative Policy Optimization (GRPO) remains the dominant critic-free approach for fine-tuning LLMs and VLMs, but its compatibility with constrained policy optimization (e.g. for safety-critical domains) has not been carefully examined. In this work, we introduce Constrained GRPO, a Lagrangian-based extension of GRPO for constrained policy optimization. We show that the standard practice of scalarizing rewards before normalization introduces a critical Lagrangian-specific failure mode: GRPO's within-group normalization makes constrained optimization highly sensitive to how multi-component learning signals are aggregated. We show that scalarizing rewards before normalization introduces shared-denominator coupling, so that changing one multiplier alters not only the emphasis on its corresponding constraint, but also the relative weighting of the reward and other constraints. We address this with a simple but crucial modification: scalarizing standardized advantages rather than rewards. This yields a better-conditioned update by addressing the coupling induced by reward scalarization, resulting in better-behaved multiplier dynamics and more stable constraint enforcement in practice. Empirically, across a controlled gridworld, a real-world autonomous driving benchmark, and a mathematical reasoning task, Constrained GRPO consistently achieves better adherence to specified constraints while maintaining or improving task performance.