Recent advances in multi-agent systems highlight the potential of specialized small agents that collaborate via division of labor. Existing tool-integrated reasoning systems, however, often follow a single-agent paradigm in which one large model interleaves long-horizon reasoning with precise tool operations, leading to cognitive-load interference and unstable coordination. We present MSARL, a Multi-Small-Agent Reinforcement Learning framework that explicitly decouples reasoning from tool use. In MSARL, a Reasoning Agent decomposes problems and plans tool invocations, while multiple Tool Agents specialize in specific external tools, each trained via a combination of imitation learning and reinforcement learning with role-specific rewards. On mathematical problem solving with code execution, MSARL significantly improves reasoning stability and final-answer accuracy over single-agent baselines. Moreover, the architecture generalizes to diverse tool-use tasks, demonstrating that cognitive-role decoupling with small agents is a scalable blueprint for multi-agent AI design.
Figures & tables
Figure 1: Model performance on Math-500 under two prompting regimes.
Figure 2: Overview of our method
Figure 3: Training prompt templates for MSARL
Model
AIME24
AIME25
MATH500
Olympiad
AMC23
Avg
Models based on Qwen2.5-Math-1.5B-Base
Qwen2.5-Math-1.5B-Instruct
10.0
10.0
66.0
31.0
62.5
35.9
Qwen2.5-Math-1.5B-Instruct-TIR
23.3
20.0
75.6
48.5
62.5
50.0
Models based on Qwen2.5-Math-7B-Base
Qwen2.5-Math-7B-Instruct
10.0
16.7
74.8
32.4
65.0
39.8
Qwen2.5-Math-7B-Instruct-TIR
20.0
6.7
70.4
45.0
50.0
34.2
Table 1: Comparison of different models testing accuracy on mathematical benchmarks with Pass@1. The best two performance are bold and underlined .
Model
AIME24
AIME25
MATH500
Olympiad
AMC23
pass@8
maj@8
pass@8
maj@8
pass@8
maj@8
pass@8
maj@8
pass@8
maj@8
Qwen2.5-Math-1.5B-Instruct-TIR
40.0
20.0
40.0
30.0
93.4
81.6
66.5
54.3
87.5
60
Qwen2.5-Math-7B-Instruct-TIR
46.7
26.7
26.7
13.3
88.8
78.4
63.9
48.1
80
67.5
SimpleRL-Zero
50.0
30.0
26.7
20.0
90.2
82.0
-
-
85.0
67.5
Eurus-2-7B-PRIME
46.7
20.0
36.7
16.7
90.2
73.4
60.4
40.2
85.0
57.5
MSARL -1.5B (Ours)
40.0
23.3
36.7
20.0
92
82.6
77.3
54.7
87.5
72.5
Table 2: Comparison of different models on mathematical benchmarks with Pass@8 and Maj@8. The best two performance are bold and underlined .
Figure 4: The model’s performance (Average Pass@1) at different training checkpoints. Performance saturates after 2k steps.
Figure 5: Average reward score during training. The consistent upward trend demonstrates successful and stable learning, with the policy converging in the later stages.
Agentic Reinforcement Learning (ARL) trains large language models to interleave reasoning with external tool execution to solve complex tasks. Most existing ARL methods train a single set of parameters to support both reasoning and tool-use behaviors, implicitly assuming that joint training leads to improved overall agent performance. Despite its widespread adoption, this assumption has rarely been examined empirically. In this paper, we systematically examine this assumption by introducing Capability Effect Attribution (CEA), which provides quantitative evidence of interference between reasoning and tool-use behaviors. Through an in-depth analysis, we show that these two capabilities often induce misaligned gradient directions, leading to training interference that undermines the effectiveness of joint optimization and challenges the prevailing ARL paradigm. To address this issue, we propose Disentangled Action--Reasoning Tuning (DART), a simple and efficient framework that explicitly decouples parameter updates for reasoning and tool use via separate low-rank adaptation modules. With this simple change alone, DART outperforms all joint-optimization baselines and approaches the 2-Agent upper bound across thirteen benchmarks on retrieval-augmented QA and NL2SQL, further supporting our finding of capability interference under shared optimization.
Yu Li, Mingyang Yi, Xiuyu Li +6
School of Information, Renmin University of China, Beijing, China · Bytedance Inc., Beijing, China · Bytedance Inc., San Jose, USA
Large language model (LLM)-based multi-agent systems (MAS) have become a promising paradigm for complex information-seeking and reasoning tasks by enabling collaborative problem solving among specialized agents. However, existing MAS frameworks tightly couple task reasoning with coordination operations, including task selection, role assignment, message routing, and context management. As interactions grow, using powerful LLMs for these bounded control decisions introduces substantial token overhead and latency, limiting the scalability of agentic Web services. In this paper, we investigate whether coordination can be decoupled from expensive reasoning without compromising collaborative performance. We propose S1-MAS, a token-efficient multi-agent framework based on System One-guided computational division of labor. S1-MAS assigns bounded coordination decisions to lightweight System One models while reserving open-ended reasoning for capable LLM workers. Specifically, a lightweight controller selects inspection conditions, chooses subsequent tasks, and determines termination, while a compact reader retrieves condition-relevant evidence from authorized sources to support these decisions. Through a decision-evidence loop, selected tasks dynamically determine worker roles and source access, enabling adaptive collaboration without task-specific training. Extensive experiments on seven diverse benchmarks demonstrate that S1-MAS achieves superior accuracy while substantially reducing the inference cost. Across individual comparisons with AgentVerse, DyLAN, and SelfOrg on seven benchmarks, S1-MAS reduces GPT-4o token consumption by 44.9%-97.2% and measured end-to-end latency by 37.8%-93.0%. These results highlight its potential for scalable and cost-effective agentic Web applications.
Zihan Zhou, Xinzhe Hu, Hanxu Yang +2
University of Electronic Science and Technology of China · Southwestern University of Finance and Economics
Chain-of-thought prompting has popularized step-by-step reasoning in large language models, yet model performance still degrades as problem complexity and context length grow. By decomposing difficult tasks with long contexts into shorter, manageable ones, recent multi-agent paradigms offer a promising near-term solution to this problem. However, the fundamental capacities of such systems are poorly understood. In this work, we propose a theoretical framework to analyze the expressivity of multi-agent systems. We apply our framework to three algorithmic families: state tracking, recall, and k-hop reasoning. We derive bounds on (i) the number of agents required to solve the task exactly, (ii) the quantity and structure of inter-agent communication, and (iii) the achievable speedups as problem size and context scale. Our results identify regimes where communication is provably beneficial, delineate tradeoffs between agent count and bandwidth, and expose intrinsic limitations when either resource is constrained. We complement our theoretical analysis with a set of experiments on pretrained LLMs using controlled synthetic benchmarks. Empirical outcomes confirm the tradeoffs between key quantities predicted by our theory. Collectively, our analysis offers principled guidance for designing scalable multi-agent reasoning systems.
Michael Rizvi-Martel, Satwik Bhattamishra, Neil Rathi +2
Mila & Universit´e de Montr´eal · University of Oxford · Stanford University +1