Large language model (LLM)-based multi-agent systems (MAS) have become a promising paradigm for complex information-seeking and reasoning tasks by enabling collaborative problem solving among specialized agents. However, existing MAS frameworks tightly couple task reasoning with coordination operations, including task selection, role assignment, message routing, and context management. As interactions grow, using powerful LLMs for these bounded control decisions introduces substantial token overhead and latency, limiting the scalability of agentic Web services. In this paper, we investigate whether coordination can be decoupled from expensive reasoning without compromising collaborative performance. We propose S1-MAS, a token-efficient multi-agent framework based on System One-guided computational division of labor. S1-MAS assigns bounded coordination decisions to lightweight System One models while reserving open-ended reasoning for capable LLM workers. Specifically, a lightweight controller selects inspection conditions, chooses subsequent tasks, and determines termination, while a compact reader retrieves condition-relevant evidence from authorized sources to support these decisions. Through a decision-evidence loop, selected tasks dynamically determine worker roles and source access, enabling adaptive collaboration without task-specific training. Extensive experiments on seven diverse benchmarks demonstrate that S1-MAS achieves superior accuracy while substantially reducing the inference cost. Across individual comparisons with AgentVerse, DyLAN, and SelfOrg on seven benchmarks, S1-MAS reduces GPT-4o token consumption by 44.9%-97.2% and measured end-to-end latency by 37.8%-93.0%. These results highlight its potential for scalable and cost-effective agentic Web applications.
Figures & tables
Figure 1: S1-MAS decouples coordination from reasoning: a System One controller makes bounded coordination decisions, while a lightweight reader localizes relevant evidence, reducing token overhead while preserving LLM reasoning capability.
Figure 2: Overview of S1-MAS. (a) The state St records the history of t completed worker executions; in this AQuA example, the latest response mt generated by task jt serves as the current candidate answer. (b) A lightweight System One controller makes bounded coordination decisions by selecting inspection conditions, choosing the next task, or terminating based on evidence localized by a lightweight reader. (c) When collaboration continues, an LLM worker executes jt+1 to generate mt+1 and update the state to St+1 ; otherwise, the controller returns an archived answer upon termination.
Method
MMLU
GSM8K
AQuA
MultiArith
SVAMP
HumanEval
Average
Direct Answer
Vanilla
85.19 ± 1.00
94.93 ± 1.33
75.86 ± 4.34
98.81 ± 1.36
93.72 ± 0.26
85.67 ± 3.44
89.03 ± 1.26
Single-Agent Prompting
CoT
85.40 ± 1.64
97.13 ± 0.32
76.17 ± 1.87
99.23 ± 0.81
93.75 ± 0.81
87.60 ± 4.60
89.88 ± 0.97
SC-CoT
89.11 ± 1.00
97.30 ± 0.10
82.40 ± 3.32
99.40 ± 1.03
94.41 ± 1.73
88.98 ± 1.26
91.93 ± 0.57
Fixed Multi-Agent Topologies
Table 1: Performance comparison using GPT-4o. Accuracy values are reported as percentages. Results are means with standard deviations over three seeds. Average summarizes performance across six datasets. Best and second-best means are bold and underlined , respectively.
Method
MMLU
GSM8K
AQuA
MultiArith
SVAMP
HumanEval
Average
Ptok.
Ctok.
Ptok.
Ctok.
Ptok.
Ctok.
Ptok.
Ctok.
Ptok.
Ctok.
Ptok.
Ctok.
Ptok.
Ctok.
Vanilla
117
25
80
98
105
176
62
48
62
39
161
155
98
90
MAD-M 2 (S)
4,066
1,407
2,907
1,235
4,764
2,045
2,164
729
2,052
697
6,418
2,734
3,728
1,475
SelfOrg
9,795
5,429
7,032
3,254
10,910
5,979
4,682
1,922
4,835
2,097
17,513
10,909
9,128
4,932
AgentVerse
7,018
3,562
6,780
3,205
9,645
5,206
4,156
1,688
4,332
1,826
14,070
7,728
7,667
3,869
MACNET
2,356
1,165
2,329
1,270
3,472
2,097
1,511
722
1,539
745
3,984
2,272
2,532
1,379
Table 2: Mean GPT-4o tokens per question, averaged over three seeds under the same configurations as Table 1 (lower is better). Ptok. and Ctok. denote input and generation tokens across all calls. Values are rounded to integers.
Figure 3: Mean end-to-end latency across six benchmarks, averaged over three seeds under the same configurations as Tables 1 and 2 . Lower is better; the star highlights S1-MAS.
Method
MMLU
AQuA
HumanEval
Average
GPT tokens
S1-MAS
80.17
81.15
91.18
84.17
535,918
w/ Random
78.43
78.50
90.08
82.34
542,491
w/ GPT
78.87
79.91
90.63
83.14
2,190,250
w/o Reader
79.08
78.35
89.81
82.41
482,314
Table 3: Ablation results with GPT-4o-mini, averaged over three seeds. Accuracy (%) is reported on MMLU, AQuA, and HumanEval. GPT tokens is the per-seed total over these three test sets, averaged across seeds.
Method
EM (%)
GPT tokens
Latency (s)
AgentVerse
89.78
32,386
28.78
DyLAN
89.89
9,927
15.22
SelfOrg
89.11
28,229
61.03
S1-MAS
90.11
4,847
9.47
Table 4: WebSRC performance and efficiency with GPT-4o. Exact match (EM) scores are averaged across three random seeds. GPT tokens and latency are per-question means over the same runs. Best results are bold.