LLM-Based Multi-Agent Systems over Wireless Networks: A Joint Agent--Network Design Perspective
Authors: Chao Hu, Yuan Guo, Guanlin Wu, Yueling Che, Han Hu, Jie Xu
Organizations: College of Computer Science and Software Engineering (CSSE), Shenzhen University, Shenzhen 518060, China · School of Science and Engineering (SSE), the Shenzhen Future Network of Intelligence Institute (FNii), and the Guangdong Provincial Key Laboratory of Future Networks of Intelligence, The Chinese University of Hong Kong (Shenzhen), Shenzhen 518172, China · School of Information and Electronics, Beijing Institute of Technology, Beijing 100081, China
As large language models (LLMs) evolve from standalone models into collaborative agents embedded in physical systems, their reasoning and execution are increasingly distributed across wireless edge nodes. In this setting, wireless networks are experiencing a paradigm shift from only providing data connectivity to supporting the multi-agent reasoning workflow itself. The task performance of such network-constrained LLM-based multi-agent systems (MASs) is jointly affected by the multi-agent reasoning dependencies as well as the underlying network connectivity and edge resources. This coupling gives rise to various technical challenges, including the metric misalignment and message redundancy, state inconsistency and topology mismatch, as well as resource limitation and trust discontinuity. To address these challenges, this article develops a novel joint agent--network design perspective that coordinates decisions on both sides of the system. Specifically, we present the joint design of agent--interaction scheduling and resource allocation, the message selection-transmission co-design, as well as the joint agent--network topology design and workload--resource allocation. Furthermore, we consider the network-verified provenance that is linked with agent-side information-flow control to constrain how received information affects subsequent operations. An illustrative vehicle-to-everything (V2X) case study shows that jointly adapting agent-side interaction decisions and network operations improves task completion under communication and edge-resource constraints, outperforming the conventional agent-only and wireless-only separate designs.
Figures & tables
Fig. 1: From single-agent operation to multi-agent collaboration over wireless networks. Task planning forms inter-agent collaboration, while its runtime execution and adaptation depend on wireless connectivity and edge resources.
Challenge
Agent Side
Network/Edge Side
Coupling
Joint Decision
Metric misalignment
Interaction priority [ 11 ]
Resource allocation
Interaction credit ↔ service feasibility
Interaction–resource scheduling
Message redundancy and state inconsistency
Message selection/compression [ 9 ]
Delivery fidelity/timeliness [ 6 , 8 ]
Receiver-context-conditioned decision distortion
Message selection and adaptive delivery
Topology mismatch
Logical edge/endpoint selection [ 10 , 9 , 15 ]
Association/path selection [ 5 , 8 ]
Interaction credit ↔ support cost
Joint edge–network mapping
Resource limitation
Reasoning budget/workload
Communication/edge-resource allocation [ 7 , 4 ]
Workload footprint ↔ resource feedback
Workload–resource allocation
Trust discontinuity
Information-flow control [ 13 ]
Provenance verification/admission
Provenance ↔ trust requirement
Provenance-aware information flow
TABLE I: Joint design space for network-constrained LLM-based MASs.
Fig. 2: Key challenges and representative applications of LLM-based MASs over wireless networks.
Fig. 3: Illustration of case scenario and collaboration workflow.
Fig. 4: Task completion rate under two bottleneck regimes. Under the information-limited regime (left), κ measures the relative decision relevance of the key agent; under the bandwidth-limited regime (right), Bmax denotes the available uplink bandwidth.
Fig. 5: Compute and memory bottlenecks under different agent–network designs (OOM is out of memory).
The convergence of large language models (LLMs) with 6G networks is fostering a paradigm of autonomous multi-agent cooperation, which in turn is expected to substantially increase east-west traffic. Although latent-space interaction mechanisms can enable more efficient collaboration than symbolic natural-language (NL) exchanges, prior work often abstracts away the associated communication overhead under practical wireless constraints. In embodied multi-agent settings, heterogeneous interaction media incur disparate inference and transmission costs, thereby inducing an inherent end-to-end (E2E) latency trade-off. To address this, we propose a joint design that integrates communication-media selection with wireless resource allocation. Through analytical characterization and simulation-based evaluation, we show that neither token-based transmission nor key-value (KV) cache-based transmission is uniformly optimal across operating regimes, as performance depends critically on system parameters such as available computational resources and channel conditions. Accordingly, we formulate a joint optimization problem aimed at minimizing the E2E latency of multi-agent collaboration and develop a low-complexity joint media selection and resource allocation (JMSRA) algorithm. Numerical results further confirm that, by adaptively coordinating the interaction media and bandwidth allocation over heterogeneous links, the proposed scheme achieves markedly reduced E2E latency relative to conventional NL-only and KV-cache-only baselines, enabling efficient and robust multi-agent collaboration in future wireless networks.
Lipeng Dai, Luping Xiang, Kun Yang
State Key Laboratory of Novel Software Technology, Nanjing University, Nanjing, 210008, China · Institute of Intelligent Networks and Communications (NINE), Nanjing University (Suzhou Campus), Suzhou, 215163, China
LLM-based multi-agent systems have the potential to enable collective intelligence and scale toward solving highly complex tasks through coordinated ensembles of specialized agents. However, despite their theoretical potential, the architectural design space remains largely non-systematized and lacks broadly established design principles. Furthermore, the scalability characteristics of such systems are only partially understood so far. This paper makes two contributions. We first distill four design principles for scalable MAS architectures from a structured analysis of prior work: simplicity, elastic feedback, sequential workflows with optional loops, and summary-based communication. We operationalize these principles in a reference architecture whose topology is formalized as a constrained directed workflow graph, and we evaluate four configurations of increasing complexity on a standardized benchmark of terminal-based system engineering tasks using two LLMs of differing capability. Our findings show that scaling yields measurable accuracy improvements with approximately linear cost growth, but only when the underlying LLM exceeds a minimum capability threshold. Performance peaks at intermediate complexity, then degrades due to timeouts and evaluation limitations. In addition, persistent consistency issues emerge as a central challenge across all scaling levels. These results provide concrete design guidance for practitioners and highlight consistency and evaluation standardization as key targets for future research.
Linus Sander, Fengjunjie Pan, Vahid Zolfaghari +3
Robotics, Artificial Intelligence and Real-Time Systems, School of Computation, Information and Technology, Technical University of Munich, Munich, Germany
Although large language model (LLM) based multi-agent systems (MAS) show their capability to solve complex tasks and achieve higher performance over single agent systems, they lead to huge computational overheads because of heavy communication between agents. Previous research has made efforts to train a sparse multi-agent graph or fine-tune a planner to orchestrate the workflow better. However, such extra training processes introduce computational costs and limit MAS to specific domains, therefore compromising their generalizability. In this paper, we propose CONCAT, a training-free multi-agent collaboration framework based on CONsensus and Confidence-driven Ad hoc Teaming to efficiently organize agent interactions. Specifically, agents are clustered based on their initial answers, and leaders of each cluster are selected based on the agents' confidence. Then, a heuristic function based on the Theory of Mind is designed to predict the collaboration benefits between every two leaders according to their answers and confidence. Finally, an ad hoc multi-agent network is organized after evicting a percentage of communications based on the predicted benefits. Experiments across three LLMs and three benchmarks show that CONCAT achieves up to 2.02x higher efficiency (accuracy/latency ratio) than LLM-Debate and outperforms training-aware methods such as AgentDropout, while reducing average latency by 50.1% on Qwen2.5-14B-Instruct, without any task-specific training.