Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving
Organizations: School of Computer Science, Shanghai Jiao Tong University · Tencent
Abstract
Self-rewarding reinforcement learning (RL) enables large language models (LLMs) to self-evolve without human labels. Existing ensemble-based methods construct reward references from rollout groups and assign rewards accordingly. However, a response's reward representation also depends on its randomly sampled group context, i.e., the other responses in its group. Using only one group-context realization may miss desired reward signals and provide unreliable guidance for policy optimization. To address this issue, we propose Group-Marginalized Advantage Estimation (GMAE), which aggregates reward realizations across possible contexts into a response-level distribution and estimates expected advantages. Experiments across eight benchmarks and four base models demonstrate strong performance and cross-domain generalization. GMAE also exhibits stable learning, low extra cost, and good applicability across training datasets and RL backbones.
Figures & tables
| Method | Mathematics | Science | Knowledge | Coding | Average | Mathematics | Science | Knowledge | Coding | Average | ||||||||
| MATH500 | AMC | AIME24 | AIME25 | AIME26 | GPQA | MMLU-Pro | LiveCode | MATH500 | AMC | AIME24 | AIME25 | AIME26 | GPQA | MMLU-Pro | LiveCode | |||
| Qwen3-1.7B-Base | Qwen3-4B-Base | |||||||||||||||||
| Raw Model | 45.7 0.22 | 17.0 0.09 | 2.3 0.07 | 0.8 0.18 | 1.3 0.19 | 7.9 0.23 | 13.3 0.29 | 6.7 0.05 | 11.9 0.07 | 50.5 0.06 | 24.3 0.18 | 8.0 0.20 | 3.8 0.11 | 5.4 0.05 | 15.8 0.27 | 14.2 0.17 | 15.3 0.06 | 17.2 0.09 |
| w/ Verifiable Reward | 64.6 0.08 | 27.1 0.13 | 7.1 0.20 | 5.4 0.18 | 4.2 0.23 | 22.6 0.24 | 31.0 0.23 | 14.2 0.07 | 22.0 0.17 | 81.3 0.16 | 48.0 0.25 | 19.3 0.03 | 17.7 0.11 | 14.8 0.18 | 33.0 0.17 | 39.4 0.32 | 19.1 0.08 | 34.1 0.13 |
| Intuitor ∗ | 52.1 0.14 | 19.8 0.10 | 1.1 0.24 | 1.2 0.13 | 0.8 0.18 | 13.3 0.24 | 15.7 0.30 | 10.7 0.22 | 14.3 0.21 | 67.1 0.18 | 32.9 0.04 | 9.2 0.21 | 2.7 0.20 | 0.7 0.20 | 30.0 0.22 | 41.1 0.23 | 15.8 0.08 | 24.9 0.21 |
| EM-RL ∗ | 57.5 0.08 | 19.2 0.23 | 2.6 0.18 | 1.7 0.13 | 0.9 0.22 | 13.2 0.32 | 14.8 0.19 | 11.6 0.08 | 15.2 0.06 | 70.9 0.19 | 30.5 0.18 | 4.4 0.09 | 6.3 0.04 | 2.6 0.08 | 28.3 0.19 | 40.8 0.25 | 16.3 0.15 | 25.0 0.21 |
| MATH500 | AMC | AIME24 | AIME25 | AIME26 | GPQA | MMLU-Pro | LiveCode | Average | |
| w/o Calibration | 78.4 | 45.2 | 16.0 | 10.3 | 9.7 | 39.2 | 63.6 | 22.4 | 35.6 |
| w/ Prompt-level Calibration | 78.0 | 44.7 | 15.5 | 10.8 | 8.6 | 39.4 | 63.8 | 21.3 | 35.3 |
| w/ Batch-level Calibration (Ours) | 79.8 | 46.2 | 16.2 | 12.5 | 9.2 | 41.7 | 67.4 | 23.5 | 37.1 |
| Method | RL Backbones | Training Datasets | ||
| GSPO | REINFORCE++ | Open-RS | MATH-8K | |
| TTRL | 35.8 | 32.8 | 34.0 | 33.2 |
| Self-Harmony | 36.1 | 33.6 | 34.9 | 33.0 |
| Co-Reward | 37.3 | 33.9 | 35.3 | 34.3 |
| SR-TTRL | 36.8 | 33.6 | 35.2 | 34.3 |
| GMAE 0 (Ours) | 38.1 | 36.4 | 36.2 | 35.5 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Mathematics | Science | Knowledge | Coding | Average | ||||
| MATH500 | AMC | AIME24 | AIME25 | AIME26 | GPQA | MMLU-Pro | LiveCode | ||
| TTRL | 78.8 | 45.3 | 14.4 | 11.3 | 7.2 | 40.3 | 66.6 | 22.3 | 35.8 |
| Self-Harmony | 78.4 | 43.2 | 15.2 | 11.0 | 6.7 | 40.9 | 70.8 | 22.8 | 36.1 |
| Co-Reward | 80.3 | 47.7 | 15.5 | 12.7 | 7.9 | 42.0 | 68.7 | 23.2 | 37.3 |
| SR-TTRL | 79.7 | 47.1 | 14.7 | 12.6 | 8.0 | 41.4 | 67.6 | 22.9 | 36.8 |
| GMAE 0 (Ours) | 81.9 | 48.7 | 17.3 | 14.0 | 9.0 | 42.1 | 67.9 | 24.2 | 38.1 |
| Method | Mathematics | Science | Knowledge | Coding | Average | ||||
| MATH500 | AMC | AIME24 | AIME25 | AIME26 | GPQA | MMLU-Pro | LiveCode | ||
| TTRL | 75.9 | 40.1 | 10.7 | 9.7 | 5.7 | 36.8 | 64.3 | 19.4 | 32.8 |
| Self-Harmony | 75.5 | 39.6 | 11.1 | 10.4 | 6.3 | 38.1 | 68.3 | 19.6 | 33.6 |
| Co-Reward | 76.7 | 41.5 | 11.9 | 11.5 | 6.6 | 37.7 | 64.5 | 20.4 | 33.9 |
| SR-TTRL | 76.9 | 40.8 | 11.7 | 10.8 | 6.1 | 37.4 | 65.0 | 20.1 | 33.6 |
| GMAE 0 (Ours) | 79.8 | 45.5 | 13.8 | 12.9 | 9.3 | 40.2 | 66.7 | 22.8 | 36.4 |
| Method | Mathematics | Science | Knowledge | Coding | Average | ||||
| MATH500 | AMC | AIME24 | AIME25 | AIME26 | GPQA | MMLU-Pro | LiveCode | ||
| TTRL | 76.4 | 42.7 | 13.5 | 10.4 | 7.6 | 38.8 | 62.1 | 20.3 | 34.0 |
| Self-Harmony | 75.6 | 41.6 | 13.7 | 10.8 | 7.2 | 39.3 | 69.3 | 21.4 | 34.9 |
| Co-Reward | 77.3 | 43.2 | 14.1 | 11.6 | 7.8 | 40.4 | 64.9 | 22.7 | 35.3 |
| SR-TTRL | 77.0 | 43.5 | 13.8 | 11.4 | 7.8 | 40.5 | 65.7 | 22.2 | 35.2 |
| GMAE 0 (Ours) | 77.8 | 44.0 | 14.9 | 12.3 | 8.3 | 41.2 | 67.9 | 22.9 | 36.2 |
| Method | Mathematics | Science | Knowledge | Coding | Average | ||||
| MATH500 | AMC | AIME24 | AIME25 | AIME26 | GPQA | MMLU-Pro | LiveCode | ||
| TTRL | 75.5 | 41.2 | 12.3 | 9.5 | 6.7 | 37.5 | 63.2 | 19.5 | 33.2 |
| Self-Harmony | 76.4 | 41.5 | 12.0 | 9.9 | 6.5 | 37.3 | 61.3 | 19.4 | 33.0 |
| Co-Reward | 77.0 | 42.3 | 13.6 | 8.9 | 7.1 | 38.2 | 66.0 | 21.2 | 34.3 |
| SR-TTRL | 77.1 | 42.3 | 13.6 | 10.4 | 7.8 | 38.6 | 63.9 | 21.0 | 34.3 |
| GMAE 0 (Ours) | 79.1 | 43.7 | 15.0 | 11.1 | 8.3 | 40.1 | 65.1 | 21.2 | 35.5 |