Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization
Authors: Yixuan Wang, Yifei Chen, Haichao Zhang, Haozheng Luo, Xander Wu, Jie Ni, Yun Fu, Nuno Vasconcelos, +1 more
Organizations: University of Florida · UC San Diego · Northeastern University · Northwestern University · Stanford University, Zillion Network · Universität Innsbruck
Reinforcement learning (RL) with group-relative advantages has become the de facto standard for post-training language model reasoners. However, when optimizing multiple reward objectives, existing methods typically scalarize the reward vector with a fixed weighted sum before group-wise standardization. We show that this design leads to two fundamental problems: rollouts with distinct reward profiles can receive identical advantages, and all objectives are optimized with fixed relative weights regardless of their current level of saturation. Consequently, an objective whose rewards are already near their upper bound can retain substantial influence when its rewards still vary within rollout groups. We introduce Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization (SA-MRPO), which standardizes each reward objective independently and adaptively discounts its contribution according to a batch-level estimate of objective saturation. This changes the relative contribution of each objective according to its observed reward headroom. We further derive an exact condition under which saturation aware reweighting reverses the sign of a rollout's aggregate advantage. Across mathematical reasoning with two- and three-objective reward combinations, SA-MRPO improves the harder correctness objective over GDPO in 12 of 15 benchmark comparisons, with gains of up to 5% on AIME24. On adaptive reasoning it improves accuracy on all five benchmarks, by 3.8% on average and up to 9.2% on AMC23, and on coding benchmarks it improves pass rate by up to 2.3%, while in all settings maintaining the easier objectives near their already satisfied levels. Additional experiments characterize the accompanying reward tradeoffs and sensitivity to corrupted rewards.
Figures & tables
Figure 1 : Saturation-aware reweighting changes rollout preference. GRPO gives rollouts 2 and 3 identical advantages, while GDPO ranks rollout 3 higher despite its lower correctness reward. SA-MRPO discounts the nearly saturated format objective, reversing both the ordering and the advantage signs relative to GDPO. The illustrative example uses one query, four rollouts, continuous rewards in [0,1] displayed out of 100, equal objective weights, and γ=1 for SA-MRPO.
Figure 2 : Reward tradeoffs and generation cost in controlled adaptive reasoning. Left: mean evaluation rewards over three seeds, with sample standard deviations. SA-MRPO settings are connected in γ order. Right: mean benchmark accuracy versus generated tokens. Trained methods use a 4096 token cap and two or three seeds. The base model is evaluated with caps of 1024, 2048, and 4096 tokens. Vertical error bars show sample standard deviations across training seeds.
Benchmark
Metric
Qwen2.5-7B-Instruct
Qwen2.5-3B-Instruct
Three objectives
Base
Rcorrect+Rlength
Rcorrect+Rlength+Rformat
Base
GDPO
SA-MRPO
GDPO 2obj
SA-MRPO 2obj
GDPO 3obj
SA-MRPO 3obj
AIME24
Acc ↑
11.7%
11.5%
16.5%
0.6%
5.0%
8.5%
6.7%
8.1%
Exceed ↓
3.5%
1.2%
1.5%
6.2%
0.0%
0.6%
0.6%
1.7%
Minerva
Acc ↑
16.1%
24.2%
24.8%
6.7%
16.2%
16.6%
16.9%
18.1%
Exceed ↓
0.4%
0.1%
0.0%
0.4%
0.1%
0.1%
0.1%
0.1%
Table 1 : Comparison of SA-MRPO, GDPO, and the base model on mathematical reasoning with two and three reward objectives. Higher accuracy is better, while Exceed measures the fraction of responses exceeding the 4,000-token budget and is lower is better. The same Qwen2.5-3B-Instruct base model is used for both reward settings.
Table 4
Method
AIME24
AMC23
MATH500
Minerva
Average
Len. ↓
Base
50.0
62.6
60.0
18.4
47.8
2212
GRPO
8.3
45.8
59.6
18.9
33.2±0.7
834
GDPO
22.2
51.4
70.9
25.7
42.6±1.0
487
Focal adaptation
25.6
57.0
71.9
26.7
45.3±2.7
574
SAW
25.6
61.9
75.1
27.2
47.4±1.7
855
DVAO
11.1
51.0
71.5
26.6
40.0±2.2
508
Table 4: Adaptive reasoning with greedy decoding and a maximum generation length of 4096 tokens. Accuracy is averaged equally over AIME24, AMC23, MATH500, and Minerva. Uncertainty denotes the sample standard deviation across training seeds; Appendix B lists the number of seeds per configuration. The Focal baseline is adapted to pointwise rewards. Bold values identify the highest means among trained configurations.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Configuration
Lmin
Lmax
Sampled mathematical reasoning
0
4000
Additional adaptive comparison
1024
2048
Adaptive reward sweep
1024
2048
Corrupted reward experiment
1000
4000
Appendix
Table 5: Length reward thresholds for the mathematical comparisons.
Figure 7
Method
Correctness reward
Length reward
Base
0.327±0.004
0.419
GRPO
0.296±0.004
0.973
GDPO
0.448±0.005
0.997
GD 2 PO
0.430±0.007
0.996
DVAO
0.452±0.009
0.996
Focal adaptation
0.474±0.027
0.989
Appendix
Table 6: Adaptive DeepScaleR evaluation rewards. Trained configurations use three training seeds, with sample standard deviations of correctness. Length entries are means. Base denotes the initial policy measurements.
Figure 5 : Adaptive training with SA-MRPO at γ=0.75 over three seeds. Panels show reward means, adjusted objective weights, and mean response length.
Figure 6 : Training reward trajectories for Qwen2.5 3B and 7B under different saturation exponents. Each configuration uses one training run. Scores are displayed on a scale from zero to one hundred.
Benchmark
Base
γ=0
γ=0.25
γ=0.5
γ=0.75
γ=1
AIME24
0.6 / 6.2
5.0 / 0.0
8.5 / 0.6
9.0 / 0.7
8.7 / 0.8
7.4 / 1.1
Minerva
6.7 / 0.4
16.2 / 0.1
16.6 / 0.1
16.9 / 0.2
17.0 / 0.1
16.8 / 0.3
AMC23
10.7 / 1.7
33.2 / 0.0
34.9 / 0.1
35.6 / 0.2
35.3 / 0.2
34.8 / 0.4
MATH500
26.3 / 0.6
57.1 / 0.0
58.2 / 0.1
58.6 / 0.1
58.8 / 0.2
58.5 / 0.3
Olympiad
4.5 / 3.3
20.6 / 0.1
19.3 / 0.3
20.1 / 0.1
19.4 / 0.7
20.7 / 0.3
Appendix
Table 7: Sampled 3B exponent study. Cells show accuracy / responses exceeding 4000 tokens, both in percent. Evaluation uses 16 responses per problem and one training run per configuration.
Multi-reward reinforcement learning trains large language models to satisfy multiple behavioral objectives simultaneously. Reward-wise normalization, as used in GDPO, preserves reward-specific relative information within rollout groups, but different objectives can still exhibit uneven learning progress. We study this behavior through advantage energy, the sum of a reward's squared advantages over a batch. Under idealized GDPO normalization, we show that this energy is proportional to active-group density: the fraction of rollout groups in which the reward provides nonzero relative advantages. This reveals a residual batch-level signal imbalance and provides a basis for calibrating reward contributions. Based on this relation, we propose Density-Aware Reward Aggregation (DARA). We derive an inverse-square-root density correction that gives greater weight to signals from less frequently active rewards. DARA computes its weights from each rollout batch, adapting to changes in reward activity throughout training without modifying the underlying policy optimization objective. Experiments on tool calling and mathematical reasoning show that DARA learns the targeted behaviors faster than GDPO, reaching high format compliance in up to 26% fewer training steps on tool calling and near-saturated length compliance in up to 65% fewer steps on mathematical reasoning, while remaining competitive in final performance. Our code is available at https://github.com/zhaihaotian/DARA.
Tong Zheng, Skylar Zhai, Zhan Cheng +7
University of Chinese Academy of Sciences · University of Minnesota Twin Cities · University of Wisconsin–Madison +5
Group-based reinforcement learning objectives such as GRPO can allocate learning signal poorly across prompt difficulty: under binary rewards, group normalization induces a divergent weighting on easy prompts. We introduce Softmax Advantage Group Estimation (SoftmaxGRPO), a drop-in alternative that replaces z-score-normalized group advantages with temperature-scaled softmax advantages, keeping weights bounded regardless of prompt difficulty. For binary rewards, we derive the exact finite-group population objective and identify MaxRL as its low-temperature limit. For bounded scalar rewards, we show that the large-group update exactly optimizes a log-moment-generating-function objective, while a universal finite-group scalar objective cannot exist without additional assumptions on the reward distribution. Empirically, SoftmaxGRPO reallocates measured gradient budget away from near-solved prompts and consistently improves over GRPO under identical rewards. It reaches 51.8% on DeepMath with verifiable rewards and improves a 1.5B instruction-tuned model from 35.0% to 68.0% on Poetry using only lightweight text-similarity rewards.
Reinforcement learning with verifiable rewards (RLVR) has become an effective paradigm for improving reasoning language models on tasks such as mathematics, coding, and scientific question answering. However, widely used group-relative objectives, such as GRPO, summarize each sampled group with scalar statistics and therefore discard fine-grained relational information among candidate responses. This weakens credit assignment under sparse outcome rewards, especially when multiple generated solutions differ only subtly in reasoning quality. We propose \textbf{LamPO}, a \textbf{Lambda-Style Policy Optimization} method that replaces scalar group advantages with a \emph{Pairwise Decomposed Advantage}. LamPO aggregates pairwise reward gaps within each response group and modulates each comparison by a confidence-aware weight computed from sequence log-probability differences, while retaining the critic-free and clipped-update structure of PPO-style optimization. When reference solutions are available, we further add a lightweight ROUGE-L-based dense auxiliary reward to reduce reward sparsity. Experiments on AIME24, AIME25, MATH-500, and GPQA-Diamond with Qwen3-1.7B, Qwen3-4B, and Phi-4-mini show that LamPO consistently improves over GRPO and recent RLVR variants, with more stable training dynamics and better sample efficiency.