cs.CLAug 31, 2026

SwarmBench: Can Large Language Models Act as Agent Swarm Orchestrators?

Authors: Jinshan GaoZhuoran JinTianyi MenKang LiuJun Zhao

Organizations: The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China · School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China

Abstract

Large language model-based multi-agent systems are evolving from fixed interaction topologies toward dynamically orchestrated Agent Swarms. However, existing benchmarks are still largely based on single-agent or general-purpose agent tasks, making it difficult to systematically evaluate key orchestration capabilities. We propose SwarmBench, a benchmark that evaluates model performance from multiple perspectives, including accuracy, efficiency, cost, and process quality. Experimental results show that current models exhibit substantial differences in orchestration capability. These differences are reflected not only in final accuracy, efficiency, and cost, but also in the overall quality of the orchestration process itself. Based on these findings, we further propose SwarmExp, a simple yet effective method based on experience extraction and experience replay, which consistently improves the orchestration performance of large language models.

Explore similar work

CardsList
  1. Reward Modeling for Multi-Agent Orchestration

    Jun 11, 2026King Yeung Tsang, Zihao Zhao, Vishal Venkataramani +5Process Reward ModelingReward Modeling