cs.MASep 28, 2026

CEO Arena: Evaluating Long-Horizon Multi-Agent Decision-Making in Competitive Markets

Authors: An Yan, Yu Huo, Zhiwei Shang, Yiran Peng, Chenglin Wu

Organizations: DeepWisdom · Fudan University · The Chinese University of Hong Kong · The Chinese University of Hong Kong, Shenzhen · The Hong Kong University of Science and Technology (Guangzhou)

Abstract

Long-horizon competition tests agents' ability to coordinate business decisions under uncertainty and adapt to changing rival strategies. We introduce CEO Arena, a benchmark that uses matched replacement evaluation to assess operating returns alongside an agent's effects on rivals and the market. Each CEO agent is compared with a reference policy in the same company under the same economic seed, holding other agents' identities and assignments fixed while all agents adapt. In a shared eight-company market spanning 500 simulated days, CEOs make sequential decisions on pricing, procurement, marketing, research and development, and service using private company information and noisy market signals, under resource constraints and delayed feedback. We evaluate eight LLM-based CEO agents in 27 main runs and 26 robustness runs. In the main evaluation, most agents have negative mean returns, and private gains can accompany market losses. Robustness analyses suggest that aggregate patterns extend beyond the original rule-based baseline; four of the 56 directed pairs show relatively stable effects. Memory, action, and accounting traces suggest demand capture and rivals' pricing and spending responses as possible explanations. CEO Arena provides a controlled testbed for studying long-horizon agent competition, strategic interaction, and market externalities.

Figures & tables

Appendix figures & tables34 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. CEO-Bench: Can Agents Play the Long Game?

    Jun 16, 2026Haozhe Chen, Karthik Narasimhan, Zhuang LiuLong-Horizon AgentsLanguage-Model Agents

  2. Business Arena: Benchmarking LLM Agents in a Realistic Marketplace

    Aug 9, 2026Yijun Pan, Yukun Lian, Kunyu Shi +5Large Language Model AgentsWebarena Domains

  3. ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies

    Sep 7, 2026Xinran Zhang, Pengrui Lu, Lyumanshan Ye +1Personalized Large Language Model AgentsLarge Language Model Agents