Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL
Organizations: University of Chinese Academy of Sciences · University of Minnesota Twin Cities · University of Wisconsin–Madison · Stony Brook University · Dalian University of Technology · Beijing Foreign Studies University · Southeast University · Kuaishou Technology
Abstract
Multi-reward reinforcement learning trains large language models to satisfy multiple behavioral objectives simultaneously. Reward-wise normalization, as used in GDPO, preserves reward-specific relative information within rollout groups, but different objectives can still exhibit uneven learning progress. We study this behavior through advantage energy, the sum of a reward's squared advantages over a batch. Under idealized GDPO normalization, we show that this energy is proportional to active-group density: the fraction of rollout groups in which the reward provides nonzero relative advantages. This reveals a residual batch-level signal imbalance and provides a basis for calibrating reward contributions. Based on this relation, we propose Density-Aware Reward Aggregation (DARA). We derive an inverse-square-root density correction that gives greater weight to signals from less frequently active rewards. DARA computes its weights from each rollout batch, adapting to changes in reward activity throughout training without modifying the underlying policy optimization objective. Experiments on tool calling and mathematical reasoning show that DARA learns the targeted behaviors faster than GDPO, reaching high format compliance in up to 26% fewer training steps on tool calling and near-saturated length compliance in up to 65% fewer steps on mathematical reasoning, while remaining competitive in final performance. Our code is available at https://github.com/zhaihaotian/DARA.
Figures & tables
| Live | Non-Live | Multi-Turn | Average | ||||||
| Size | Method | Acc. | Format | Acc. | Format | Acc. | Format | Acc. | Format |
| 1.5B | Base | 9.33 | 0.37 | 12.06 | 0.12 | 0.38 | 0.21 | 7.25 | 0.23 |
| GRPO | 64.80 | 99.88 | 78.04 | 99.91 | 2.55 | 43.24 | 48.46 | 81.01 | |
| GDPO | 67.05 | 99.94 | 79.30 | 99.97 | 4.75 | 91.40 | 50.37 | 97.11 | |
| DVAO | 64.99 | 99.75 | 76.84 | 99.62 | 4.08 | 81.19 | 48.64 | 93.52 | |
| GD 2 PO-Hard | 67.06 | 99.97 | 78.96 | 100.00 | 4.38 | 86.59 | 50.13 | 95.52 | |
| Live | Non-Live | Multi-Turn | Average | ||||||||||
| Size | Method | Acc. | Format | Len. | Acc. | Format | Len. | Acc. | Format | Len. | Acc. | Format | Len. |
| 1.5B | Base | 9.33 | 0.37 | 1.33 | 12.06 | 0.12 | 0.73 | 0.38 | 0.21 | 3.58 | 7.25 | 0.23 | 1.88 |
| GRPO | 67.80 | 100.00 | 86.34 | 78.26 | 99.96 | 78.62 | 4.06 | 64.14 | 72.59 | 50.04 | 88.03 | 79.19 | |
| GDPO | 66.62 | 100.00 | 100.00 | 80.94 | 100.00 | 99.93 | 3.81 | 83.17 | 93.97 | 50.46 | 94.39 | 97.96 | |
| DARA-Asym | 68.02 | 100.00 | 99.41 | 80.00 | 100.00 | 100.00 | 3.88 | 90.80 | 94.74 | 50.63 | 96.93 | 98.05 | |
| DARA-Sym | 67.95 | 100.00 | 100.00 | 80.21 | 100.00 | 99.99 | 3.31 | 92.30 | 96.16 | 50.49 | 97.43 | 98.72 | |
| DeepSeek-R1-1.5B | Qwen3-4B-Instruct | DeepSeek-R1-7B | |||||||
| Method | Acc | Exceed | Joint | Acc | Exceed | Joint | Acc | Exceed | Joint |
| Base | 48.93 | 65.23 | 27.75 | 68.15 | 33.60 | 51.69 | 64.66 | 56.16 | 36.32 |
| GRPO | 45.52 | 11.03 | 44.69 | 62.14 | 8.07 | 60.48 | 58.26 | 9.73 | 56.74 |
| GDPO | 44.12 | 8.66 | 43.47 | 60.82 | 7.64 | 59.19 | 57.45 | 6.82 | 56.48 |
| DVAO | 45.58 | 11.35 | 44.57 | 61.37 | 4.77 | 60.34 | 58.15 | 5.75 | 57.53 |
| GD 2 PO-Hard | 44.24 | 11.41 | 43.45 | 60.58 | 7.28 | 59.16 | 57.78 | 7.96 | 56.35 |
| Method | Acc | Exceed | Joint |
|---|---|---|---|
| DARA-Asym | 48.03 | 4.21 | 47.76 |
| Static-Asym ( ) | 46.84 | 3.14 | 46.72 |
| DARA-Sym | 46.71 | 0.89 | 46.67 |
| Static-Sym ( ) | 45.97 | 0.87 | 45.96 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Reward weighting | Advantage |
|---|---|---|
| GRPO | Rewards summed before normalization | |
| GDPO | Equal weights on normalized rewards | |
| DVAO | Within-group reward standard deviation | |
| SAW | Batch-level coefficient of variation of the offset reward | |
| SA-MRPO | Batch-level saturation | |
| GD 2 PO-Hard | Sign consistency of | over retained responses with ; the loss of query is scaled by |
| Setting | Value |
|---|---|
| Backbones | Qwen2.5-1.5B-Instruct / Qwen2.5-3B-Instruct |
| Training / rollout | FSDP / vLLM, tensor parallelism 1, bfloat16 |
| Training / validation / checkpoint | 100 steps / every 10 steps / step 100 |
| Prompts / responses per prompt / total | 512 / 4 / 2,048 |
| PPO minibatch / microbatch / epochs | 512 responses / 256 responses / 1 |
| Learning rate / gradient clip / PPO clip | / 1.0 / 0.2 |
| Setting | Value |
|---|---|
| Evaluator | BFCL-v4, gorilla commit 6ea57973 |
| Subsets | Live, Non-Live, and Multi-Turn |
| Trajectories | One per evaluation case |
| Temperature / top- | 0.6 / 0.95 |
| Maximum new tokens / context length | 8,192 / 32,768 |
| Inference seed / top- | 0 / |
| Prompts | PPO mini | PPO micro | Responses | |
|---|---|---|---|---|
| 4 | 512 | 128 | 64 | 2,048 |
| 8 | 256 | 64 | 32 | 2,048 |
| 16 | 128 | 32 | 16 | 2,048 |
| 32 | 64 | 16 | 8 | 2,048 |
| Setting | Value |
|---|---|
| Backbones | DeepSeek-R1-Distill-Qwen-1.5B / 7B |
| Qwen3-4B-Instruct-2507 / Qwen3-4B-Thinking-2507 | |
| Training data | DeepScaleR-Preview, full training collection |
| Training / rollout | FSDP / vLLM, rollout tensor parallelism 1 |
| Prompt / response length | 1,024 / 8,000 tokens |
| Accepted prompt groups / candidate groups | 512 / 768 per generation batch |
| Setting | Value |
|---|---|
| MATH | MATH-500, test split, 500 questions |
| AIME | AIME 2024, 30 questions |
| AMC | AMC 2022/2023, 83 questions |
| Minerva | Minerva Math, test split, 272 questions |
| OlympiadBench | OE_TO_maths_en_COMP , 674 questions |
| Samples per question | 16 |