Multi-agent systems (MAS) have recently emerged as an effective approach for coordinating large language model (LLM)-based agents to solve complex tasks through structured interactions. In practice, MASs often handle a stream of heterogeneous and complex tasks, requiring agents to decompose each task and then self-organize and self-evolve to adapt to incoming tasks while sharing execution resources. However, most early approaches to MASs rely on centralized controllers or fixed coordination patterns, which can limit scalability or adaptability. In contrast, existing decentralized and dynamic MASs often require training dedicated routers or invoking LLMs for agent selection, resulting in substantial computational costs and coordination overhead. To address these challenges and enable efficient task-level self-organization and self-evolution for task- and workload-level collaboration, we propose Decentralized Expertise-Aware Load Serving (DEALS), a decentralized and low-complexity framework that enables agents to self-organize and dynamically route concurrent tasks for processing. Specifically, each agent maintains local queues of incoming tasks, and its router decides whether to process a task locally or forward it to a neighbor based on differences in backlog and success rate. Meanwhile, executors process independent tasks concurrently within and across agents, and partially solved tasks can be resumed by other agents. Experiments show that DEALS not only improves performance along multiple dimensions (e.g., answer accuracy and task throughput) in both homogeneous and heterogeneous agent pools, but also balances agent expertise and workload in a self-organized manner, enabling effective decentralized coordination.
Figures & tables
Backbone
System
MATH
BBH
MMLU-Pro
Average
Qwen2.5-7B
Single Agent
66.43
62.00
52.00
60.14
AgentNet
68.57
71.00
50.00
63.19
GPTSwarm
65.00
75.00
21.00
53.67
AgentPrune
65.00
68.00
34.00
55.67
DEALS (Ours)
71.43
74.00
53.00
66.14
Mistral-7B
Single Agent
6.43
43.00
34.00
27.81
Table 1: Accuracy (%) across backbones and datasets. Bold indicates the best result, and underlining indicates the second-best result. The quality weight is set to V=20 for homogeneous configurations and V=80 for the mixed-family configuration.
Setting
Group
V=0
V=20
V=80
Homogeneous
H5
68.57
68.57
68.57
H7
70.71
70.71
72.14
Heterogeneous
M5
47.86
64.29
72.86
M7
40.00
60.71
67.86
Table 2: MATH test accuracy (%) with five and seven agents. Bold marks the highest accuracy within each group, including ties.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Setting
Number of agents
3 , 5 , 7
Graph topology
Complete graph
Dirichlet concentration α
0.7
Execution capacity per agent Ci
9
Hop limit H
3
Split limit K
3
Appendix
Table 3: Implementation parameter settings for all experiments.
Agent pool
Dataset
V=0
V=5
V=10
V=20
V=40
V=80
Spread
Homogeneous Qwen2.5-7B ×3
MATH
72.14
70.00
68.57
71.43
68.57
72.86
4.29
BBH
73.00
72.00
69.00
74.00
67.00
73.00
7.00
MMLU-Pro
54.00
52.00
56.00
53.00
55.00
52.00
4.00
Heterogeneous Qwen/Mistral/Llama
MATH
44.29
44.29
49.29
60.71
57.86
73.57
29.28
BBH
56.00
66.00
63.00
66.00
69.00
73.00
17.00
MMLU-Pro
43.00
44.00
45.00
47.00
49.00
53.00
10.00
Appendix
Table 4: Accuracy (%) sensitivity to the quality weight V . Spread is the difference between the highest and lowest accuracy across the reported values of V . Bold indicates the best result within each row.
Dataset
V
Qwen
Mistral
Llama
MATH
0
46
48
46
80
138
0
2
BBH
0
37
30
33
80
88
0
12
MMLU-Pro
0
31
34
35
80
64
5
31
Appendix
Table 5: Test tasks completed by each model-backed agent in the heterogeneous configuration. The rows correspond to the selected V=0 and V=80 runs in Table 4 . Each count identifies the agent responsible for the final task completion.
Large language model (LLM)-based multi-agent systems (MAS) have become a promising paradigm for complex information-seeking and reasoning tasks by enabling collaborative problem solving among specialized agents. However, existing MAS frameworks tightly couple task reasoning with coordination operations, including task selection, role assignment, message routing, and context management. As interactions grow, using powerful LLMs for these bounded control decisions introduces substantial token overhead and latency, limiting the scalability of agentic Web services. In this paper, we investigate whether coordination can be decoupled from expensive reasoning without compromising collaborative performance. We propose S1-MAS, a token-efficient multi-agent framework based on System One-guided computational division of labor. S1-MAS assigns bounded coordination decisions to lightweight System One models while reserving open-ended reasoning for capable LLM workers. Specifically, a lightweight controller selects inspection conditions, chooses subsequent tasks, and determines termination, while a compact reader retrieves condition-relevant evidence from authorized sources to support these decisions. Through a decision-evidence loop, selected tasks dynamically determine worker roles and source access, enabling adaptive collaboration without task-specific training. Extensive experiments on seven diverse benchmarks demonstrate that S1-MAS achieves superior accuracy while substantially reducing the inference cost. Across individual comparisons with AgentVerse, DyLAN, and SelfOrg on seven benchmarks, S1-MAS reduces GPT-4o token consumption by 44.9%-97.2% and measured end-to-end latency by 37.8%-93.0%. These results highlight its potential for scalable and cost-effective agentic Web applications.
Zihan Zhou, Xinzhe Hu, Hanxu Yang +2
University of Electronic Science and Technology of China · Southwestern University of Finance and Economics
Multi-agent large language model (LLM) systems have shown promise for solving complex tasks through agent collaboration. However, existing frameworks assign tasks based on predefined roles without considering whether an agent can accurately assess its own competence boundaries, leading to overconfident execution on tasks beyond its expertise. Inspired by metacognition theory from cognitive science, we propose MetaCogAgent, a multi-agent LLM framework where each agent is equipped with a Metacognitive Self-Assessment Unit that evaluates task-capability alignment before execution. The framework introduces three contributions: (1) a self-assessment mechanism that estimates per-task confidence by combining verbalized uncertainty with historical capability profiles; (2) an adaptive delegation protocol that routes low-confidence tasks to better-suited agents through cross-agent evaluation; and (3) a capability boundary learning module that iteratively refines each agent's competence model via cybernetic feedback. Experiments on our constructed MetaCog-Eval benchmark (700 tasks across 5 cognitive dimensions) demonstrate that MetaCogAgent achieves 82.4% task accuracy -- 8.7% above the best routing baseline -- while using 5% fewer API calls than AutoGen and 34% fewer than ensemble voting. Ablation studies confirm that each metacognitive component contributes to overall system performance.
Chenyu Wang, Yang Shu
School of Computer and Artificial Intelligence, Zhengzhou University, Zhengzhou, China · Zhejiang University, Hangzhou, China
Multi-agent systems (MAS) can scale large language model agents on long-horizon tasks by running them in parallel, yet existing designs waste much of this parallelism in bubbles: agent time spent waiting on others or redoing a peer's work. These bubbles stem from how agents communicate. Independent agents share nothing and rediscover what their peers have already found; peer-communicating agents wait at synchronous rounds; and under centralized orchestration, the main agent blocks on its sub-agents while progress is relayed. We propose Decentralized Language Models (DeLM), a MAS framework on top of existing agent harnesses that squeezes out these bubbles by replacing the main agent with a shared context and a task queue. Agents asynchronously claim tasks, publish findings as soon as they are available, and build on or correct one another's progress, with every peer's status visible to all. On long-horizon tasks from Terminal-Bench 4.0 and DeepSWE v1.1, and on SWE-bench Verified, DeLM is both more accurate and faster than Codex, Claude Code, their native subagents, and AOrchestra in every setting, improving accuracy by up to 17.5 points over the strongest baseline and running up to 2.49x faster than the harness it builds on. On ProgramBench, where agents rebuild programs from scratch, DeLM makes faster progress than Claude Code and finishes a 120-minute budget up to 19.9 points higher in test pass rate. The code is available on our project website at https://yuzhenmao.github.io/DeLM/.