Self-rewarding reinforcement learning (RL) enables large language models (LLMs) to self-evolve without human labels. Existing ensemble-based methods construct reward references from rollout groups and assign rewards accordingly. However, a response's reward representation also depends on its randomly sampled group context, i.e., the other responses in its group. Using only one group-context realization may miss desired reward signals and provide unreliable guidance for policy optimization. To address this issue, we propose Group-Marginalized Advantage Estimation (GMAE), which aggregates reward realizations across possible contexts into a response-level distribution and estimates expected advantages. Experiments across eight benchmarks and four base models demonstrate strong performance and cross-domain generalization. GMAE also exhibits stable learning, low extra cost, and good applicability across training datasets and RL backbones.
Figures & tables
Figure 1: (a) GMAE Motivation: RLVR evaluates each response against external ground truth, whereas in self-rewarding RL its reward representation depends on the remaining responses in its rollout group, termed its group context . (b) GMAE Overview: Conventional methods use a single group-context realization to obtain each response’s reward representation, while GMAE aggregates rewards across possible contexts into a reward distribution and estimates the expected advantage.
Figure 2: Self-Reward Uncertainty Demonstration. (a) Reward distributions for a fixed correct response (top) and a fixed incorrect response (bottom) across group contexts under TTRL ( Zuo et al., 2026 ) , Self-Harmony ( Wang et al., 2026a ) , SCOPE ( Wang et al., 2026b ) , and SR-TTRL ( Wu et al., 2026 ) . Dots denote individual realizations and gray bars denote means. (b) Joint distributions of self-rewarding ( x -axis) and RLVR ( y -axis) advantages for the same responses. Results here are from Qwen3-8B-Base on a DAPO-14K prompt; more cases are provided in Appendix A .
Method
Mathematics
Science
Knowledge
Coding
Average
Mathematics
Science
Knowledge
Coding
Average
MATH500
AMC
AIME24
AIME25
AIME26
GPQA
MMLU-Pro
LiveCode
MATH500
AMC
AIME24
AIME25
AIME26
GPQA
MMLU-Pro
LiveCode
Qwen3-1.7B-Base
Qwen3-4B-Base
Raw Model
45.7 0.22
17.0 0.09
2.3 0.07
0.8 0.18
1.3 0.19
7.9 0.23
13.3 0.29
6.7 0.05
11.9 0.07
50.5 0.06
24.3 0.18
8.0 0.20
3.8 0.11
5.4 0.05
15.8 0.27
14.2 0.17
15.3 0.06
17.2 0.09
w/ Verifiable Reward
64.6 0.08
27.1 0.13
7.1 0.20
5.4 0.18
4.2 0.23
22.6 0.24
31.0 0.23
14.2 0.07
22.0 0.17
81.3 0.16
48.0 0.25
19.3 0.03
17.7 0.11
14.8 0.18
33.0 0.17
39.4 0.32
19.1 0.08
34.1 0.13
Intuitor ∗
52.1 0.14
19.8 0.10
1.1 0.24
1.2 0.13
0.8 0.18
13.3 0.24
15.7 0.30
10.7 0.22
14.3 0.21
67.1 0.18
32.9 0.04
9.2 0.21
2.7 0.20
0.7 0.20
30.0 0.22
41.1 0.23
15.8 0.08
24.9 0.21
EM-RL ∗
57.5 0.08
19.2 0.23
2.6 0.18
1.7 0.13
0.9 0.22
13.2 0.32
14.8 0.19
11.6 0.08
15.2 0.06
70.9 0.19
30.5 0.18
4.4 0.09
6.3 0.04
2.6 0.08
28.3 0.19
40.8 0.25
16.3 0.15
25.0 0.21
Table 1: Main Results Across Models and Domains: Mean@16 scores on eight benchmarks across four base models. Average is computed over all eight benchmarks. Bold and underlined values denote the best and second-best results among all self-rewarding methods, respectively.
Figure 3: Training Dynamics on the DAPO-500 validation set over training steps.
Figure 4: Multi-Seed Training Variance. Results show the average Mean@16 standard deviation across eight benchmarks for Qwen3-8B-Base over five training seeds.
Figure 5: Training Duration vs. Performance. Each point represents one method. Duration is normalized to RLVR, and performance is the average Mean@16 across eight benchmarks on Qwen3-8B-Base .
Figure 6: Distribution of Correct Responses per Prompt. Illustrative proportions of training prompts by the number of correct responses among G=16 sampled responses under GMAE 0 and TTRL at steps 50 and 200.
Figure 7: Auxiliary Response Size Ablation. Average Mean@16 across all eight benchmarks with different auxiliary response sizes.
MATH500
AMC
AIME24
AIME25
AIME26
GPQA
MMLU-Pro
LiveCode
Average
w/o Calibration
78.4
45.2
16.0
10.3
9.7
39.2
63.6
22.4
35.6
w/ Prompt-level Calibration
78.0
44.7
15.5
10.8
8.6
39.4
63.8
21.3
35.3
w/ Batch-level Calibration (Ours)
79.8
46.2
16.2
12.5
9.2
41.7
67.4
23.5
37.1
Table 2: Calibration Ablation on GMAE 0 . Mean@16 results for Qwen3-8B-Base .
Figure 8: Calibration Factor λ Dynamics.
Method
RL Backbones
Training Datasets
GSPO
REINFORCE++
Open-RS
MATH-8K
TTRL
35.8
32.8
34.0
33.2
Self-Harmony
36.1
33.6
34.9
33.0
Co-Reward
37.3
33.9
35.3
34.3
SR-TTRL
36.8
33.6
35.2
34.3
GMAE 0 (Ours)
38.1
36.4
36.2
35.5
Table 3: Scalability Across Training Settings. Average Mean@16 across all eight benchmarks for Qwen3-8B-Base . Full benchmark-wise results are in Appendix E .
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9: Self-Reward Uncertainty (Additional Case 2).
Figure 10: Self-Reward Uncertainty (Additional Case 3).
Figure 11: Self-Reward Uncertainty (Additional Case 4).
Figure 12: Self-Reward Uncertainty (Additional Case 5).
Method
Mathematics
Science
Knowledge
Coding
Average
MATH500
AMC
AIME24
AIME25
AIME26
GPQA
MMLU-Pro
LiveCode
TTRL
78.8
45.3
14.4
11.3
7.2
40.3
66.6
22.3
35.8
Self-Harmony
78.4
43.2
15.2
11.0
6.7
40.9
70.8
22.8
36.1
Co-Reward
80.3
47.7
15.5
12.7
7.9
42.0
68.7
23.2
37.3
SR-TTRL
79.7
47.1
14.7
12.6
8.0
41.4
67.6
22.9
36.8
GMAE 0 (Ours)
81.9
48.7
17.3
14.0
9.0
42.1
67.9
24.2
38.1
Appendix
Table 4: Scalability on RL backbone. Benchmark-wise Mean@16 results on Qwen3-8B-Base using GSPO as the RL backbone.
Method
Mathematics
Science
Knowledge
Coding
Average
MATH500
AMC
AIME24
AIME25
AIME26
GPQA
MMLU-Pro
LiveCode
TTRL
75.9
40.1
10.7
9.7
5.7
36.8
64.3
19.4
32.8
Self-Harmony
75.5
39.6
11.1
10.4
6.3
38.1
68.3
19.6
33.6
Co-Reward
76.7
41.5
11.9
11.5
6.6
37.7
64.5
20.4
33.9
SR-TTRL
76.9
40.8
11.7
10.8
6.1
37.4
65.0
20.1
33.6
GMAE 0 (Ours)
79.8
45.5
13.8
12.9
9.3
40.2
66.7
22.8
36.4
Appendix
Table 5: Scalability on RL backbone. Benchmark-wise Mean@16 results on Qwen3-8B-Base using REINFORCE++ as the RL backbone.
Method
Mathematics
Science
Knowledge
Coding
Average
MATH500
AMC
AIME24
AIME25
AIME26
GPQA
MMLU-Pro
LiveCode
TTRL
76.4
42.7
13.5
10.4
7.6
38.8
62.1
20.3
34.0
Self-Harmony
75.6
41.6
13.7
10.8
7.2
39.3
69.3
21.4
34.9
Co-Reward
77.3
43.2
14.1
11.6
7.8
40.4
64.9
22.7
35.3
SR-TTRL
77.0
43.5
13.8
11.4
7.8
40.5
65.7
22.2
35.2
GMAE 0 (Ours)
77.8
44.0
14.9
12.3
8.3
41.2
67.9
22.9
36.2
Appendix
Table 6: Scalability on training dataset. Mean@16 results on Qwen3-8B-Base trained with the Open-RS dataset.
Method
Mathematics
Science
Knowledge
Coding
Average
MATH500
AMC
AIME24
AIME25
AIME26
GPQA
MMLU-Pro
LiveCode
TTRL
75.5
41.2
12.3
9.5
6.7
37.5
63.2
19.5
33.2
Self-Harmony
76.4
41.5
12.0
9.9
6.5
37.3
61.3
19.4
33.0
Co-Reward
77.0
42.3
13.6
8.9
7.1
38.2
66.0
21.2
34.3
SR-TTRL
77.1
42.3
13.6
10.4
7.8
38.6
63.9
21.0
34.3
GMAE 0 (Ours)
79.1
43.7
15.0
11.1
8.3
40.1
65.1
21.2
35.5
Appendix
Table 7: Scalability on training dataset. Mean@16 results on Qwen3-8B-Base trained with the MATH-8K dataset.
Department of Software Engineering and IT, ´Ecole de Technologie Sup´erieure, Montreal, Canada · Centre int´egr´e de traumatologie, Hˆopital du Sacr´e-Cœur de Montr´eal, Universit´e de Montr´eal, Montreal, Canada · Desautels Faculty of Management, McGill University, Montreal, Canada.