Large language model (LLM)-based multi-agent systems (MAS) have become a promising paradigm for complex information-seeking and reasoning tasks by enabling collaborative problem solving among specialized agents. However, existing MAS frameworks tightly couple task reasoning with coordination operations, including task selection, role assignment, message routing, and context management. As interactions grow, using powerful LLMs for these bounded control decisions introduces substantial token overhead and latency, limiting the scalability of agentic Web services. In this paper, we investigate whether coordination can be decoupled from expensive reasoning without compromising collaborative performance. We propose S1-MAS, a token-efficient multi-agent framework based on System One-guided computational division of labor. S1-MAS assigns bounded coordination decisions to lightweight System One models while reserving open-ended reasoning for capable LLM workers. Specifically, a lightweight controller selects inspection conditions, chooses subsequent tasks, and determines termination, while a compact reader retrieves condition-relevant evidence from authorized sources to support these decisions. Through a decision-evidence loop, selected tasks dynamically determine worker roles and source access, enabling adaptive collaboration without task-specific training. Extensive experiments on seven diverse benchmarks demonstrate that S1-MAS achieves superior accuracy while substantially reducing the inference cost. Across individual comparisons with AgentVerse, DyLAN, and SelfOrg on seven benchmarks, S1-MAS reduces GPT-4o token consumption by 44.9%-97.2% and measured end-to-end latency by 37.8%-93.0%. These results highlight its potential for scalable and cost-effective agentic Web applications.
Figures & tables
Figure 1: S1-MAS decouples coordination from reasoning: a System One controller makes bounded coordination decisions, while a lightweight reader localizes relevant evidence, reducing token overhead while preserving LLM reasoning capability.
Figure 2: Overview of S1-MAS. (a) The state St records the history of t completed worker executions; in this AQuA example, the latest response mt generated by task jt serves as the current candidate answer. (b) A lightweight System One controller makes bounded coordination decisions by selecting inspection conditions, choosing the next task, or terminating based on evidence localized by a lightweight reader. (c) When collaboration continues, an LLM worker executes jt+1 to generate mt+1 and update the state to St+1 ; otherwise, the controller returns an archived answer upon termination.
Method
MMLU
GSM8K
AQuA
MultiArith
SVAMP
HumanEval
Average
Direct Answer
Vanilla
85.19 ± 1.00
94.93 ± 1.33
75.86 ± 4.34
98.81 ± 1.36
93.72 ± 0.26
85.67 ± 3.44
89.03 ± 1.26
Single-Agent Prompting
CoT
85.40 ± 1.64
97.13 ± 0.32
76.17 ± 1.87
99.23 ± 0.81
93.75 ± 0.81
87.60 ± 4.60
89.88 ± 0.97
SC-CoT
89.11 ± 1.00
97.30 ± 0.10
82.40 ± 3.32
99.40 ± 1.03
94.41 ± 1.73
88.98 ± 1.26
91.93 ± 0.57
Fixed Multi-Agent Topologies
Table 1: Performance comparison using GPT-4o. Accuracy values are reported as percentages. Results are means with standard deviations over three seeds. Average summarizes performance across six datasets. Best and second-best means are bold and underlined , respectively.
Method
MMLU
GSM8K
AQuA
MultiArith
SVAMP
HumanEval
Average
Ptok.
Ctok.
Ptok.
Ctok.
Ptok.
Ctok.
Ptok.
Ctok.
Ptok.
Ctok.
Ptok.
Ctok.
Ptok.
Ctok.
Vanilla
117
25
80
98
105
176
62
48
62
39
161
155
98
90
MAD-M 2 (S)
4,066
1,407
2,907
1,235
4,764
2,045
2,164
729
2,052
697
6,418
2,734
3,728
1,475
SelfOrg
9,795
5,429
7,032
3,254
10,910
5,979
4,682
1,922
4,835
2,097
17,513
10,909
9,128
4,932
AgentVerse
7,018
3,562
6,780
3,205
9,645
5,206
4,156
1,688
4,332
1,826
14,070
7,728
7,667
3,869
MACNET
2,356
1,165
2,329
1,270
3,472
2,097
1,511
722
1,539
745
3,984
2,272
2,532
1,379
Table 2: Mean GPT-4o tokens per question, averaged over three seeds under the same configurations as Table 1 (lower is better). Ptok. and Ctok. denote input and generation tokens across all calls. Values are rounded to integers.
Figure 3: Mean end-to-end latency across six benchmarks, averaged over three seeds under the same configurations as Tables 1 and 2 . Lower is better; the star highlights S1-MAS.
Method
MMLU
AQuA
HumanEval
Average
GPT tokens
S1-MAS
80.17
81.15
91.18
84.17
535,918
w/ Random
78.43
78.50
90.08
82.34
542,491
w/ GPT
78.87
79.91
90.63
83.14
2,190,250
w/o Reader
79.08
78.35
89.81
82.41
482,314
Table 3: Ablation results with GPT-4o-mini, averaged over three seeds. Accuracy (%) is reported on MMLU, AQuA, and HumanEval. GPT tokens is the per-seed total over these three test sets, averaged across seeds.
Method
EM (%)
GPT tokens
Latency (s)
AgentVerse
89.78
32,386
28.78
DyLAN
89.89
9,927
15.22
SelfOrg
89.11
28,229
61.03
S1-MAS
90.11
4,847
9.47
Table 4: WebSRC performance and efficiency with GPT-4o. Exact match (EM) scores are averaged across three random seeds. GPT tokens and latency are per-question means over the same runs. Best results are bold.
Multi-agent systems (MAS) powered by large language models suffer from severe token inefficiency arising from two compounding sources: (i) unstructured parallel execution, where all agents activate simultaneously irrespective of input readiness; and (ii) unrestricted context sharing, where every agent receives the full accumulated context regardless of relevance. Existing mitigation strategies - static pruning, hierarchical decomposition, and learned routing - treat coordination as a structural allocation problem and fundamentally ignore its temporal dimension. We propose Phase-Scheduled Multi-Agent Systems (PSMAS), a framework that reconceptualizes agent activation as continuous control over a shared attention space modeled on a circular manifold. Each agent i is assigned a fixed angular phase theta_i in the range [0, 2*pi], derived from the task dependency topology; a global sweep signal phi(t) rotates at velocity omega, activating only agents within an angular window epsilon. Idle agents receive compressed context summaries, reducing per-step token consumption. We implement PSMAS on LangGraph, evaluate on four structured benchmarks (HotPotQA-MAS, HumanEval-MAS, ALFWorld-Multi, WebArena-Coord) and two unstructured conversational settings, and prove stability, convergence, and optimality results for the sweep dynamics. PSMAS achieves a mean token reduction of 27.3 percent (range 21.4-34.8 percent) while maintaining task performance within 2.1 percentage points of a fully activated baseline (p < 0.01, n = 500 per configuration), and outperforms the strongest learned routing baseline by 5.6 percentage points in token reduction with 2.0 percentage points less performance drop. Crucially, we show that scheduling and compression are independent sources of gain: scheduling alone accounts for 18-20 percentage points of reduction, robust to compression degradation up to alpha = 0.40.
Although large language model (LLM) based multi-agent systems (MAS) show their capability to solve complex tasks and achieve higher performance over single agent systems, they lead to huge computational overheads because of heavy communication between agents. Previous research has made efforts to train a sparse multi-agent graph or fine-tune a planner to orchestrate the workflow better. However, such extra training processes introduce computational costs and limit MAS to specific domains, therefore compromising their generalizability. In this paper, we propose CONCAT, a training-free multi-agent collaboration framework based on CONsensus and Confidence-driven Ad hoc Teaming to efficiently organize agent interactions. Specifically, agents are clustered based on their initial answers, and leaders of each cluster are selected based on the agents' confidence. Then, a heuristic function based on the Theory of Mind is designed to predict the collaboration benefits between every two leaders according to their answers and confidence. Finally, an ad hoc multi-agent network is organized after evicting a percentage of communications based on the predicted benefits. Experiments across three LLMs and three benchmarks show that CONCAT achieves up to 2.02x higher efficiency (accuracy/latency ratio) than LLM-Debate and outperforms training-aware methods such as AgentDropout, while reducing average latency by 50.1% on Qwen2.5-14B-Instruct, without any task-specific training.
Multi-agent systems (MAS) can scale large language model agents on long-horizon tasks by running them in parallel, yet existing designs waste much of this parallelism in bubbles: agent time spent waiting on others or redoing a peer's work. These bubbles stem from how agents communicate. Independent agents share nothing and rediscover what their peers have already found; peer-communicating agents wait at synchronous rounds; and under centralized orchestration, the main agent blocks on its sub-agents while progress is relayed. We propose Decentralized Language Models (DeLM), a MAS framework on top of existing agent harnesses that squeezes out these bubbles by replacing the main agent with a shared context and a task queue. Agents asynchronously claim tasks, publish findings as soon as they are available, and build on or correct one another's progress, with every peer's status visible to all. On long-horizon tasks from Terminal-Bench 4.0 and DeepSWE v1.1, and on SWE-bench Verified, DeLM is both more accurate and faster than Codex, Claude Code, their native subagents, and AOrchestra in every setting, improving accuracy by up to 17.5 points over the strongest baseline and running up to 2.49x faster than the harness it builds on. On ProgramBench, where agents rebuild programs from scratch, DeLM makes faster progress than Claude Code and finishes a 120-minute budget up to 19.9 points higher in test pass rate. The code is available on our project website at https://yuzhenmao.github.io/DeLM/.