cs.LGFeb 5, 2026

Unveiling Implicit Advantage Symmetry: Why GRPO Struggles with Exploration and Difficulty Adaptation

Authors: Zhiqi Yu, Zhangquan Chen, Mengting Liu, Heye Zhang, Liangqiong Qu

Organizations: University of Hong Kong · Tsinghua University · Sun Yat-sen University

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR), particularly GRPO, has become the standard for eliciting LLM reasoning. However, its efficiency in exploration and difficulty adaptation remains an open challenge. In this work, we identify an implicit advantage symmetry inherent in Group Relative Advantage Estimation (GRAE) as a structural property that provides a new perspective for understanding these bottlenecks. This symmetry induces two critical limitations: (i) at the group level, strict symmetry in weights between correct and incorrect trajectories leaves unsampled action logits unchanged, thereby hindering the exploration of novel correct solution. (ii) at the sample level, the algorithm implicitly prioritizes medium-difficulty samples, remaining agnostic to the non-stationary demands of difficulty focus. Through controlled experiments, we reveal that this symmetric property is sub-optimal, yielding two pivotal insights: (i) asymmetrically down-weighting the advantages of correct trajectories encourages essential exploration but risks instability; (ii) learning efficiency can be boosted by a curriculum-like transition-prioritizing simpler samples initially before gradually shifting to complex ones. Motivated by these findings, we propose Asymmetric GRAE (A-GRAE), which dynamically modulates exploration incentives and sample-difficulty focus. Experiments across seven benchmarks demonstrate that A-GRAE consistently improves GRPO and its variants across both LLMs and MLLMs. Code is available at https://github.com/HKU-HealthAI/A-GRAE

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. EP-GRPO: Entropy-Progress Aligned Group Relative Policy Optimization with Implicit Process Guidance

    May 6, 2026Song Yu, Li Li, Wenwen Zhao +1Reinforcement Learning With Verifiable RewardFlow-Grpo

  2. F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare

    Feb 6, 2026Daniil Plyusov, Alexey Gorbatovski, Boris Shaposhnikov +4Reinforcement Learning With Verifiable RewardFlow-Grpo