CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning
Organizations: Hong Kong University of Science and Technology
Abstract
Group Relative Policy Optimization (GRPO) is widely used to train reasoning language models, where it computes advantages by centering and normalizing rewards across rollouts of the same prompt. For multiple rewards, GRPO sums the reward components and normalizes the total reward by its within-group standard deviation. The corresponding variance equals the sum of all pairwise reward covariances. For a fixed centered reward, larger aggregate covariance produces smaller advantages, and vice versa, allowing update magnitudes to adapt to reward dependence. However, correlated rewards with large scales can dominate this normalization and suppress signals from smaller-scale rewards. We propose Correlation-Normalized GRPO (CorrGRPO), which normalizes pairwise covariances into Pearson correlation coefficients. CorrGRPO keeps the centered total reward unchanged while balancing the influence of differently scaled rewards on the correlation-based normalization. This allows advantage magnitudes to adapt to reward correlations without the normalization being dominated by large-scale reward components. We compare CorrGRPO with GRPO and other variants on code generation, tool calling, and agent security, using models ranging from 0.5B to 8B parameters. These tasks all involve multiple rewards that can improve together or present tradeoffs. Results show improvements across three domains, including code generation, tool calling, and agent security. Our code is available at https://github.com/HKUST-KnowComp/CorrGRPO.
Figures & tables
| Model | LeetCodeDataset (In-dataset Eval) | HumanEval | MBPP | LCB v6 | Avg. | ||
|---|---|---|---|---|---|---|---|
| Efficiency | Executable | Pass@1 | Pass@1 | Pass@1 | Pass@1 | Pass@1 | |
| Qwen2.5-Coder-0.5B-Instruct | 50.00 | 61.40 | 1.75 | 54.88 | 33.60 | 5.14 | 23.84 |
| + GRPO | 60.71 | 87.72 | 3.07 | 57.32 | 40.00 | 3.43 | 25.96 |
| + GDPO | 60.94 | 88.60 | 3.51 | 57.32 | 38.40 | 4.57 | 25.95 |
| + CorrGRPO | 67.86 | 89.04 | 3.95 | 59.76 | 42.20 | 6.29 | 28.05 |
| Qwen2.5-Coder-1.5B-Instruct | 58.93 | 75.44 | 3.07 | 61.59 | 54.40 | 10.86 | 32.48 |
| Model | RLLA-4K (In-dataset Eval) | API-Bank | Avg. | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Function | Param. Name | Param. Value | All Exact | V1 | V2 | V3 | Avg. | ||
| Qwen2.5-7B-Instruct | 73.17 | 69.77 | 54.78 | 38.03 | 17.29 | 19.26 | 24.49 | 20.35 | 29.19 |
| + GRPO | 97.65 | 95.60 | 76.91 | 63.38 | 79.45 | 44.44 | 35.92 | 53.27 | 58.33 |
| + CorrGRPO | 95.07 | 94.95 | 82.14 | 67.61 | 81.45 | 46.67 | 35.10 | 54.41 | 61.01 |
| Qwen3-4B-Thinking-2507 | 33.80 | 31.46 | 25.35 | 21.13 | 53.38 | 39.26 | 31.43 | 41.36 | 31.25 |
| + GRPO | 94.84 | 92.08 | 72.07 | 57.75 | 79.20 | 52.59 | 52.24 | 61.34 | 59.55 |
| AgentDojo (In-dataset Eval) / Agent Security Bench | ||||
|---|---|---|---|---|
| Model | Clean Utility | Utility Under Attack | ASR | Joint Accuracy |
| Qwen2.5-3B-Instruct | 28.57 / 10.00 | 14.62 / 6.25 | 3.85 / 8.00 | 18.97 / 6.13 |
| + GRPO | 52.38 / 21.75 | 66.92 / 9.75 | 0.00 / 9.00 | 60.13 / 14.88 |
| + CorrGRPO | 95.24 / 26.00 | 84.36 / 13.25 | 1.54 / 9.25 | 89.62 / 18.13 |
| Qwen2.5-7B-Instruct | 38.10 / 64.75 | 30.51 / 47.00 | 11.28 / 33.25 | 28.72 / 42.00 |
| + GRPO | 90.48 / 30.00 | 80.00 / 41.75 | 0.26 / 26.50 | 84.49 / 31.38 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Training data / method | Training | HumanEval | MBPP | LiveCode Bench | LeetCode Dataset |
|---|---|---|---|---|---|
| examples | |||||
| Original SFT results: Qwen2.5-Coder-7B | |||||
| 24-08–25-02 | 24-07–25-03 | ||||
| Magicoder Evol-Instruct-110K | 111.1K | 77.4 | 74.1 | 15.1 | 13.7 |
| Magicoder OSS-Instruct-75K | 75.1K | 73.8 | 76.5 | 15.1 | 12.9 |
| Open-R1 CodeForces-CoT | 9.5K | 79.9 | 74.1 | 15.8 | 13.3 |
| Model | Clean utility | Utility under attack | Targeted ASR |
|---|---|---|---|
| Claude 3 Opus | |||
| Claude 3 Sonnet | |||
| Claude 3.5 Sonnet | |||
| Command-R+ | |||
| Gemini 1.5 Flash | |||
| Gemini 1.5 Pro |
| Model | Clean utility | ASR (OPI) |
|---|---|---|
| Claude-3.5 Sonnet | 100.00 | 59.70 |
| LLaMA3-70B | 66.50 | 43.70 |
| GPT-4o | 79.00 | 62.45 |
| Gemma2-27B | 31.50 | 14.20 |
| LLaMA3.1-70B | 21.25 | 12.10 |
| Qwen2-7B | 9.75 | 9.00 |
| Base setting: ASR-valid | |||||
| Model | Direct harm | Data stealing | Overall | ||
| S1 | S2 | Total | total | ||
| Original prompted agents (ReAct) | |||||
| Qwen-1.8B | 36.1 | 35.1 | 82.6 | 17.6 | 29.7 |
| Qwen-72B | 8.7 | 37.9 | 98.4 | 37.1 | 23.2 |
| Mistral-7B | 13.4 | 25.0 | 87.8 | 20.1 | 16.7 |
| Model | LeetCodeDataset (In-dataset Eval) | HumanEval | MBPP | LCB v6 | Avg. | ||
|---|---|---|---|---|---|---|---|
| Efficiency | Executable | Pass@1 | Pass@1 | Pass@1 | Pass@1 | Pass@1 | |
| CISPO | 34.51 | 83.77 | 21.93 | 76.83 | 79.00 | 23.43 | 50.30 |
| CISPO + CorrGRPO | 38.44 | 82.46 | 25.00 | 81.71 | 77.80 | 21.71 | 51.56 |
| GDPO | 41.28 | 84.21 | 16.23 | 76.83 | 79.60 | 22.86 | 48.88 |
| GDPO + CorrGRPO | 44.64 | 86.40 | 17.11 | 76.83 | 77.20 | 21.71 | 48.21 |
| DAPO | 40.48 | 81.14 | 20.18 | 81.71 | 77.40 | 23.43 | 50.68 |