Unveiling Implicit Advantage Symmetry: Why GRPO Struggles with Exploration and Difficulty Adaptation
Organizations: University of Hong Kong · Tsinghua University · Sun Yat-sen University
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR), particularly GRPO, has become the standard for eliciting LLM reasoning. However, its efficiency in exploration and difficulty adaptation remains an open challenge. In this work, we identify an implicit advantage symmetry inherent in Group Relative Advantage Estimation (GRAE) as a structural property that provides a new perspective for understanding these bottlenecks. This symmetry induces two critical limitations: (i) at the group level, strict symmetry in weights between correct and incorrect trajectories leaves unsampled action logits unchanged, thereby hindering the exploration of novel correct solution. (ii) at the sample level, the algorithm implicitly prioritizes medium-difficulty samples, remaining agnostic to the non-stationary demands of difficulty focus. Through controlled experiments, we reveal that this symmetric property is sub-optimal, yielding two pivotal insights: (i) asymmetrically down-weighting the advantages of correct trajectories encourages essential exploration but risks instability; (ii) learning efficiency can be boosted by a curriculum-like transition-prioritizing simpler samples initially before gradually shifting to complex ones. Motivated by these findings, we propose Asymmetric GRAE (A-GRAE), which dynamically modulates exploration incentives and sample-difficulty focus. Experiments across seven benchmarks demonstrate that A-GRAE consistently improves GRPO and its variants across both LLMs and MLLMs. Code is available at https://github.com/HKU-HealthAI/A-GRAE
Figures & tables
| Method | Pass@ | |||||||||
| 1 | 2 | 4 | 8 | 16 | 32 | 64 | 128 | 256 | Avg. | |
| MATH | ||||||||||
| Base Model | 63.4 | 74.8 | 83.2 | 88.6 | 91.2 | 93.4 | 94.1 | 95.0 | 96.3 | 86.7 |
| SEC | 77.6 | 82.5 | 86.0 | 88.9 | 90.2 | 92.2 | 92.6 | 93.3 | 94.7 | 88.7 |
| GRPO-LEAD | 77.8 | 83.0 | 86.5 | 89.2 | 90.5 | 92.3 | 92.8 | 93.6 | 95.0 | 89.0 |
| P@kT | 76.2 | 82.2 | 86.9 | 89.6 | 91.4 | 93.3 | 94.0 | 96.2 | 96.2 | 89.3 |
| ID Domain | OOD Domain | ||
| Method | Geo3K | MathVision | Mathverse |
| Task A: General Mathematical Reasoning | |||
| Base Model | 27.8 | 20.8 | 31.6 |
| GRPO | 43.5 | 23.4 | 35.2 |
| GRPO + A-GRAE | 45.7 | 24.0 | 36.8 |
| DAPO | 44.7 | 23.8 | 35.9 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Pass@ | |||||||||
| 1 | 2 | 4 | 8 | 16 | 32 | 64 | 128 | 256 | Avg. | |
| MATH | ||||||||||
| Base Model | 84.3 | 89.6 | 92.8 | 94.8 | 96.0 | 96.8 | 97.4 | 97.8 | 98.0 | 94.2 |
| GRPO | 92.4 | 95.3 | 96.5 | 97.0 | 97.3 | 97.7 | 98.0 | 98.3 | 98.6 | 96.8 |
| GRPO + A-GRAE | 92.9 | 95.7 | 96.9 | 97.5 | 97.9 | 98.3 | 98.6 | 99.0 | 99.3 | 97.3 |
| DAPO | 92.6 | 95.3 | 96.4 | 97.2 | 97.5 | 97.8 | 98.0 | 98.2 | 98.6 | 96.8 |
| Method | Pass@ | ||||||||
| 1 | 2 | 4 | 8 | 16 | 32 | 64 | 128 | 256 | |
| MATH | |||||||||
| Base Model | 63.4 | 74.8 | 83.2 | 88.6 | 91.2 | 93.4 | 94.1 | 95.0 | 96.3 |
| GRPO | 76.5 | 82.3 | 86.1 | 88.8 | 90.3 | 92.6 | 93.5 | 93.9 | 95.0 |
| GRPO + A-GRAE (sample level) | 77.8 | 84.2 | 88.4 | 90.5 | 92.1 | 94.0 | 94.3 | 94.8 | 96.0 |
| GRPO + A-GRAE (group level) | 77.6 | 84.0 | 88.1 | 91.2 | 93.0 | 94.3 | 95.3 | 95.6 | 97.0 |
| Method | ||||||
| GRPO (Baseline) | 76.5 | 82.3 | 86.1 | 88.8 | 90.3 | 92.4 |
| Negative-Dominant | 77.2 | 83.8 | 87.6 | 90.4 | 92.8 | 94.5 |
| GRPO + Entropy-Promoting | 75.8 | 80.3 | 83.6 | 86.5 | 88.4 | 89.0 |
| Comparison Pair | |||||||||
| MATH Benchmark | |||||||||
| GRPO + A-GRAE vs. GRPO | .0032 | .0018 | .0025 | .0041 | .0084 | .0115 | .0128 | .0155 | .0121 |
| DAPO + A-GRAE vs. DAPO | .0015 | .0022 | .0048 | .0055 | .0072 | .0098 | .0105 | .0112 | .0108 |
| Dr.GRPO + A-GRAE vs. Dr.GRPO | .0024 | .0009 | .0011 | .0022 | .0015 | .0028 | .0035 | .0042 | .0038 |
| AIME 2025 Benchmark | |||||||||
| GRPO + A-GRAE vs. GRPO | .0125 | .0098 | .0112 | .0105 | .0142 | .0128 | .0115 | .0092 | .0085 |
| Method / Setting | MATH | AMC23 | AIME2025 |
| GRPO (Baseline) | 76.5 | 59.2 | 10.3 |
| 77.8 | 61.8 | 10.9 | |
| 77.8 | 61.8 | 11.0 | |
| 78.3 | 62.6 | 11.3 | |
| 78.0 | 62.8 | 11.3 | |
| 76.8 | 62.3 | 10.8 |