cs.LGSep 29, 2026

Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving

Authors: Yiming Wang, Yikang Liu, Qingyuan Tian, Xingyu Chen, Zhuosheng Zhang, Zhaopeng Tu, Rui Wang

Organizations: School of Computer Science, Shanghai Jiao Tong University · Tencent

Abstract

Self-rewarding reinforcement learning (RL) enables large language models (LLMs) to self-evolve without human labels. Existing ensemble-based methods construct reward references from rollout groups and assign rewards accordingly. However, a response's reward representation also depends on its randomly sampled group context, i.e., the other responses in its group. Using only one group-context realization may miss desired reward signals and provide unreliable guidance for policy optimization. To address this issue, we propose Group-Marginalized Advantage Estimation (GMAE), which aggregates reward realizations across possible contexts into a response-level distribution and estimates expected advantages. Experiments across eight benchmarks and four base models demonstrate strong performance and cross-domain generalization. GMAE also exhibits stable learning, low extra cost, and good applicability across training datasets and RL backbones.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output

    Jun 9, 2026Guozheng Li, Xiyan Fu, Yiwen GuoReinforcement Learning From Human FeedbackGroup-Based Reinforcement Learning

  2. Label-Free Reinforcement Learning via Cross-Model Entropy

    May 27, 2026Matt Gorbett, Hossein ShiraziReinforcement Learning Post-TrainingCross-Entropy Losses

  3. Self-Review Reinforcement Learning (SRRL) with Cross-Episode Memory and Policy Distillation

    Jul 6, 2026Muhammad Zain Amin, Kibele Sebnem YildirimReinforcement Learning With Verifiable RewardOffline Reinforcement Learning