Dr. MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems
Organizations: Nanyang Technological University, Singapore
Abstract
Multi-agent LLM systems enable advanced reasoning and tool use via role specialization, yet reliable reinforcement learning (RL) post-training for such systems remains difficult. In this work, we theoretically pinpoint a key reason for training instability when extending group-based RL to multi-agent LLM systems. We show that under GRPO-style optimization, a global normalization baseline may deviate from diverse agents' reward distributions, which ultimately leads to gradient-norm instability. Based on this finding, we propose Dr. MAS, a simple and stable RL training recipe for multi-agent LLM systems. Dr. MAS uses an agent-wise remedy: normalizing advantages per agent using each agent's own reward statistics, which calibrates gradient scales and dramatically stabilizes training, both theoretically and empirically. Beyond the algorithm, Dr. MAS provides an end-to-end RL training framework for multi-agent LLM systems, supporting scalable orchestration, flexible per-agent LLM serving and optimization configs, and shared resource scheduling of LLM actor backends. We evaluate Dr. MAS on multi-agent math reasoning and multi-turn search benchmarks using Qwen2.5 and Qwen3 series models. Dr. MAS achieves clear gains over vanilla GRPO (e.g., +5.6% avg@16 and +4.6% pass@16 on math, and +15.2% avg@16 and +13.1% pass@16 on search) while largely eliminating gradient spikes. Moreover, it remains highly effective under heterogeneous agent-model assignments while improving efficiency.
Figures & tables
| Benchmark | Single-Agent | Multi-Agent & LLM Sharing | Multi-Agent & LLM Non-Sharing | |||||||
| GRPO | GRPO | Dr. MAS | GRPO | Dr. MAS | ||||||
| avg@16 | pass@16 | avg@16 | pass@16 | avg@16 | pass@16 | avg@16 | pass@16 | avg@16 | pass@16 | |
| Qwen3-4B | ||||||||||
| AIME’24 | 38.8 | 63.3 | 39.3 | 64.0 | \text{39.3}_{\smash{\scriptsize{\color[rgb]{0.5898,0.5898,0.5898}0.0}\phantom{00}}} | \text{63.3}_{\smash{\scriptsize{\color[rgb]{0.8516,0.5508,0.2734}\downarrow\!0.7}\phantom{0}}} | 42.7 | 73.3 | \text{46.9}_{\smash{\scriptsize{\color[rgb]{0.457,0.2734,0.4922}\uparrow\!4.2}}} | \text{80.0}_{\smash{\scriptsize{\color[rgb]{0.457,0.2734,0.4922}\uparrow\!6.7}}} |
| AIME’25 | 33.1 | 56.7 | 31.4 | 53.3 | \text{38.1}_{\smash{\scriptsize{\color[rgb]{0.457,0.2734,0.4922}\uparrow\!6.7}\phantom{0}}} | \text{63.3}_{\smash{\scriptsize{\color[rgb]{0.457,0.2734,0.4922}\uparrow\!10.0}}} | 35.6 | 63.3 | \text{38.1}_{\smash{\scriptsize{\color[rgb]{0.457,0.2734,0.4922}\uparrow\!2.5}}} | \text{66.7}_{\smash{\scriptsize{\color[rgb]{0.457,0.2734,0.4922}\uparrow\!3.4}}} |
| AMC’23 | 83.5 | 95.0 | 85.6 | 95.0 | \text{87.3}_{\smash{\scriptsize{\color[rgb]{0.457,0.2734,0.4922}\uparrow\!1.7}\phantom{0}}} | \text{95.0}_{\smash{\scriptsize{\color[rgb]{0.5898,0.5898,0.5898}0.0}\phantom{00}}} | 83.5 | 95.0 | \text{89.5}_{\smash{\scriptsize{\color[rgb]{0.457,0.2734,0.4922}\uparrow\!6.0}}} | \text{97.5}_{\smash{\scriptsize{\color[rgb]{0.457,0.2734,0.4922}\uparrow\!2.5}}} |
| Benchmark | Single-Agent | Multi-Agent & LLM Sharing | Multi-Agent & LLM Non-Sharing | |||||||
| GRPO | GRPO | Dr. MAS | GRPO | Dr. MAS | ||||||
| avg@16 | pass@16 | avg@16 | pass@16 | avg@16 | pass@16 | avg@16 | pass@16 | avg@16 | pass@16 | |
| Qwen2.5-3B | ||||||||||
| NQ | 40.6 | 54.7 | 41.0 | 59.0 | \text{43.8}_{\smash{\scriptsize{\color[rgb]{0.457,0.2734,0.4922}\uparrow\!2.8}}} | \text{58.5}_{\smash{\scriptsize{\color[rgb]{0.8516,0.5508,0.2734}\downarrow\!0.5}}} | 43.8 | 54.5 | \text{44.6}_{\smash{\scriptsize{\color[rgb]{0.457,0.2734,0.4922}\uparrow\!0.8}\phantom{0}}} | \text{58.1}_{\smash{\scriptsize{\color[rgb]{0.457,0.2734,0.4922}\uparrow\!3.6}\phantom{0}}} |
| TriviaQA | 58.1 | 68.8 | 57.9 | 68.4 | \text{61.7}_{\smash{\scriptsize{\color[rgb]{0.457,0.2734,0.4922}\uparrow\!3.8}}} | \text{70.1}_{\smash{\scriptsize{\color[rgb]{0.457,0.2734,0.4922}\uparrow\!1.7}}} | 60.6 | 70.8 | \text{61.1}_{\smash{\scriptsize{\color[rgb]{0.457,0.2734,0.4922}\uparrow\!0.5}\phantom{0}}} | \text{71.7}_{\smash{\scriptsize{\color[rgb]{0.457,0.2734,0.4922}\uparrow\!0.9}\phantom{0}}} |
| PopQA | 44.2 | 49.6 | 43.2 | 58.0 | \text{45.0}_{\smash{\scriptsize{\color[rgb]{0.457,0.2734,0.4922}\uparrow\!1.8}}} | \text{57.6}_{\smash{\scriptsize{\color[rgb]{0.8516,0.5508,0.2734}\downarrow\!0.4}}} | 45.6 | 54.5 | \text{46.5}_{\smash{\scriptsize{\color[rgb]{0.457,0.2734,0.4922}\uparrow\!0.9}\phantom{0}}} | \text{57.4}_{\smash{\scriptsize{\color[rgb]{0.457,0.2734,0.4922}\uparrow\!2.9}\phantom{0}}} |
| Metric | Normalization Configuration | |||
| avg@16 | 28.0 | \text{39.1}_{\smash{\scriptsize{\color[rgb]{0.457,0.2734,0.4922}\uparrow\!11.1}}} | \text{42.9}_{\smash{\scriptsize{\color[rgb]{0.457,0.2734,0.4922}\uparrow\!14.9}}} | \text{43.8}_{\smash{\scriptsize{\color[rgb]{0.457,0.2734,0.4922}\uparrow\!15.8}}} |
| pass@16 | 40.5 | \text{53.5}_{\smash{\scriptsize{\color[rgb]{0.457,0.2734,0.4922}\uparrow\!13.0}}} | \text{57.6}_{\smash{\scriptsize{\color[rgb]{0.457,0.2734,0.4922}\uparrow\!17.1}}} | \text{58.3}_{\smash{\scriptsize{\color[rgb]{0.457,0.2734,0.4922}\uparrow\!17.8}}} |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Setting | AIME’24 | AIME’25 | OlympiadBench |
| Stronger-MAS | LLM sharing | 50.0 | 35.2 | 56.8 |
| Dr. MAS | LLM sharing | 54.8 | 39.4 | 59.3 |
| Stronger-MAS | LLM non-sharing | 57.0 | 40.0 | 56.6 |
| Dr. MAS | LLM non-sharing | 44.6 | 41.5 | 59.0 |
| Method | NQ | TriviaQA | PopQA | HotpotQA | 2Wiki | MuSiQue | Bamboogle | Average |
| GSPO | 42.6 | 58.6 | 46.2 | 31.2 | 28.9 | 8.3 | 22.4 | 34.0 |
| GSPO Dr. MAS | 44.2 | 60.8 | 47.0 | 35.8 | 35.1 | 10.2 | 26.4 | 37.1 |
| Strategy | Model Configuration | GPU Consumption | Est. Time |
| LLM sharing | Qwen2.5-3B | H100 (95GB) | 12h |
| LLM non-sharing | Qwen2.5-3B | H100 (95GB) | 13h |
| LLM sharing | Qwen2.5-7B | H100 (95GB) | 25h |
| LLM non-sharing | Qwen2.5-7B | H100 (95GB) | 26h |
| Heterogeneous | Llama-3.2-3B Qwen2.5-7B | H100 (95GB) | 16h |
| Strategy | Model Configuration | GPU Consumption | Est. Time |
| LLM sharing | Qwen3-4B | H100 (95GB) | 35h |
| LLM non-sharing | Qwen3-4B | H100 (95GB) | 38h |
| LLM sharing | Qwen3-8B | H100 (95GB) | 37h |
| LLM non-sharing | Qwen3-8B | H100 (95GB) | 42h |