Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization
Authors: Yixuan Wang, Yifei Chen, Haichao Zhang, Haozheng Luo, Xander Wu, Jie Ni, Yun Fu, Nuno Vasconcelos, +1 more
Organizations: University of Florida · UC San Diego · Northeastern University · Northwestern University · Stanford University, Zillion Network · Universität Innsbruck
Reinforcement learning (RL) with group-relative advantages has become the de facto standard for post-training language model reasoners. However, when optimizing multiple reward objectives, existing methods typically scalarize the reward vector with a fixed weighted sum before group-wise standardization. We show that this design leads to two fundamental problems: rollouts with distinct reward profiles can receive identical advantages, and all objectives are optimized with fixed relative weights regardless of their current level of saturation. Consequently, an objective whose rewards are already near their upper bound can retain substantial influence when its rewards still vary within rollout groups. We introduce Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization (SA-MRPO), which standardizes each reward objective independently and adaptively discounts its contribution according to a batch-level estimate of objective saturation. This changes the relative contribution of each objective according to its observed reward headroom. We further derive an exact condition under which saturation aware reweighting reverses the sign of a rollout's aggregate advantage. Across mathematical reasoning with two- and three-objective reward combinations, SA-MRPO improves the harder correctness objective over GDPO in 12 of 15 benchmark comparisons, with gains of up to 5% on AIME24. On adaptive reasoning it improves accuracy on all five benchmarks, by 3.8% on average and up to 9.2% on AMC23, and on coding benchmarks it improves pass rate by up to 2.3%, while in all settings maintaining the easier objectives near their already satisfied levels. Additional experiments characterize the accompanying reward tradeoffs and sensitivity to corrupted rewards.
Figures & tables
Figure 1 : Saturation-aware reweighting changes rollout preference. GRPO gives rollouts 2 and 3 identical advantages, while GDPO ranks rollout 3 higher despite its lower correctness reward. SA-MRPO discounts the nearly saturated format objective, reversing both the ordering and the advantage signs relative to GDPO. The illustrative example uses one query, four rollouts, continuous rewards in [0,1] displayed out of 100, equal objective weights, and γ=1 for SA-MRPO.
Figure 2 : Reward tradeoffs and generation cost in controlled adaptive reasoning. Left: mean evaluation rewards over three seeds, with sample standard deviations. SA-MRPO settings are connected in γ order. Right: mean benchmark accuracy versus generated tokens. Trained methods use a 4096 token cap and two or three seeds. The base model is evaluated with caps of 1024, 2048, and 4096 tokens. Vertical error bars show sample standard deviations across training seeds.
Benchmark
Metric
Qwen2.5-7B-Instruct
Qwen2.5-3B-Instruct
Three objectives
Base
Rcorrect+Rlength
Rcorrect+Rlength+Rformat
Base
GDPO
SA-MRPO
GDPO 2obj
SA-MRPO 2obj
GDPO 3obj
SA-MRPO 3obj
AIME24
Acc ↑
11.7%
11.5%
16.5%
0.6%
5.0%
8.5%
6.7%
8.1%
Exceed ↓
3.5%
1.2%
1.5%
6.2%
0.0%
0.6%
0.6%
1.7%
Minerva
Acc ↑
16.1%
24.2%
24.8%
6.7%
16.2%
16.6%
16.9%
18.1%
Exceed ↓
0.4%
0.1%
0.0%
0.4%
0.1%
0.1%
0.1%
0.1%
Table 1 : Comparison of SA-MRPO, GDPO, and the base model on mathematical reasoning with two and three reward objectives. Higher accuracy is better, while Exceed measures the fraction of responses exceeding the 4,000-token budget and is lower is better. The same Qwen2.5-3B-Instruct base model is used for both reward settings.
Table 4
Method
AIME24
AMC23
MATH500
Minerva
Average
Len. ↓
Base
50.0
62.6
60.0
18.4
47.8
2212
GRPO
8.3
45.8
59.6
18.9
33.2±0.7
834
GDPO
22.2
51.4
70.9
25.7
42.6±1.0
487
Focal adaptation
25.6
57.0
71.9
26.7
45.3±2.7
574
SAW
25.6
61.9
75.1
27.2
47.4±1.7
855
DVAO
11.1
51.0
71.5
26.6
40.0±2.2
508
Table 4: Adaptive reasoning with greedy decoding and a maximum generation length of 4096 tokens. Accuracy is averaged equally over AIME24, AMC23, MATH500, and Minerva. Uncertainty denotes the sample standard deviation across training seeds; Appendix B lists the number of seeds per configuration. The Focal baseline is adapted to pointwise rewards. Bold values identify the highest means among trained configurations.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Configuration
Lmin
Lmax
Sampled mathematical reasoning
0
4000
Additional adaptive comparison
1024
2048
Adaptive reward sweep
1024
2048
Corrupted reward experiment
1000
4000
Appendix
Table 5: Length reward thresholds for the mathematical comparisons.
Figure 7
Method
Correctness reward
Length reward
Base
0.327±0.004
0.419
GRPO
0.296±0.004
0.973
GDPO
0.448±0.005
0.997
GD 2 PO
0.430±0.007
0.996
DVAO
0.452±0.009
0.996
Focal adaptation
0.474±0.027
0.989
Appendix
Table 6: Adaptive DeepScaleR evaluation rewards. Trained configurations use three training seeds, with sample standard deviations of correctness. Length entries are means. Base denotes the initial policy measurements.
Figure 5 : Adaptive training with SA-MRPO at γ=0.75 over three seeds. Panels show reward means, adjusted objective weights, and mean response length.
Figure 6 : Training reward trajectories for Qwen2.5 3B and 7B under different saturation exponents. Each configuration uses one training run. Scores are displayed on a scale from zero to one hundred.
Benchmark
Base
γ=0
γ=0.25
γ=0.5
γ=0.75
γ=1
AIME24
0.6 / 6.2
5.0 / 0.0
8.5 / 0.6
9.0 / 0.7
8.7 / 0.8
7.4 / 1.1
Minerva
6.7 / 0.4
16.2 / 0.1
16.6 / 0.1
16.9 / 0.2
17.0 / 0.1
16.8 / 0.3
AMC23
10.7 / 1.7
33.2 / 0.0
34.9 / 0.1
35.6 / 0.2
35.3 / 0.2
34.8 / 0.4
MATH500
26.3 / 0.6
57.1 / 0.0
58.2 / 0.1
58.6 / 0.1
58.8 / 0.2
58.5 / 0.3
Olympiad
4.5 / 3.3
20.6 / 0.1
19.3 / 0.3
20.1 / 0.1
19.4 / 0.7
20.7 / 0.3
Appendix
Table 7: Sampled 3B exponent study. Cells show accuracy / responses exceeding 4000 tokens, both in percent. Evaluation uses 16 responses per problem and one training run per configuration.