CEO Arena: Evaluating Long-Horizon Multi-Agent Decision-Making in Competitive Markets
Organizations: DeepWisdom · Fudan University · The Chinese University of Hong Kong · The Chinese University of Hong Kong, Shenzhen · The Hong Kong University of Science and Technology (Guangzhou)
Abstract
Long-horizon competition tests agents' ability to coordinate business decisions under uncertainty and adapt to changing rival strategies. We introduce CEO Arena, a benchmark that uses matched replacement evaluation to assess operating returns alongside an agent's effects on rivals and the market. Each CEO agent is compared with a reference policy in the same company under the same economic seed, holding other agents' identities and assignments fixed while all agents adapt. In a shared eight-company market spanning 500 simulated days, CEOs make sequential decisions on pricing, procurement, marketing, research and development, and service using private company information and noisy market signals, under resource constraints and delayed feedback. We evaluate eight LLM-based CEO agents in 27 main runs and 26 robustness runs. In the main evaluation, most agents have negative mean returns, and private gains can accompany market losses. Robustness analyses suggest that aggregate patterns extend beyond the original rule-based baseline; four of the 56 directed pairs show relatively stable effects. Memory, action, and accounting traces suggest demand capture and rivals' pricing and spending responses as possible explanations. CEO Arena provides a controlled testbed for studying long-horizon agent competition, strategic interaction, and market externalities.
Figures & tables
| Capability | Evidence and action | Direct / competitive consequence | Delay |
|---|---|---|---|
| Analysis | Query private SQL and public quotes | Diagnose policy outcomes; no direct state change | Immediate |
| Pricing | Set product price and promotion | Margin versus jointly allocated demand | Next window |
| Procurement | Inspect inventory; set reorder target | Cash tied in stock; shortages redirect demand | 5–10 days |
| Marketing | Set segment–channel budgets | Attention share and persistent awareness | Daily carryover |
| Innovation | Fund daily development or a project | Product quality changes relative attractiveness | 7 days / 45–150 day medians |
| Operations | Fund reliability and customer support | Failures, queues, satisfaction, future demand | Daily carryover |
| Rank | CEO | Mean Score | Bankruptcy rate |
|---|---|---|---|
| 1 | GPT-6 Astra | 116.94 | 0.0% |
| 2 | GPT-5.6 Sol | 98.44 | 0.0% |
| 3 | Qwen 3.8 Max | 8.04 | 0.0% |
| 4 | Rule CEO | -47.30 | 0.0% |
| 5 | DeepSeek V4 Flash | -114.50 | 4.2% |
| 6 | Gemini 3.8 Flash | -119.79 | 0.0% |
| Entrant Agent | Private Score difference | Mean externality per rival | Total externality | Market Score difference | Positive market- difference seeds |
|---|---|---|---|---|---|
| GPT-6 Astra | 116.36 | 50.05 | 350.33 | 466.68 | 3/3 |
| GPT-5.6 Sol | 131.55 | -31.60 | -221.22 | -89.66 | 1/3 |
| Qwen 3.8 Max | 23.28 | 19.09 | 133.64 | 156.92 | 2/3 |
| DeepSeek V4 Flash | 13.62 | -13.97 | -97.82 | -84.20 | 2/3 |
| Gemini 3.8 Flash | -53.09 | 15.55 | 108.87 | 55.78 | 2/3 |
| DeepSeek V4 Pro | -54.96 | -74.04 | -518.29 | -573.25 | 0/3 |
Appendix figures & tables34 assets
Supplementary material from the paper’s appendix.
Appendix
| CEO agent | F11 | R1 | R2 |
|---|---|---|---|
| GPT-6 Astra | |||
| GPT-5.6 Sol | |||
| Qwen 3.8 Max | |||
| Rule CEO | |||
| DeepSeek V4 Flash | |||
| Gemini 3.8 Flash |
| Pair | F11 | F29 | F47 | R1 | R2 | Mean | SD | Sign |
|---|---|---|---|---|---|---|---|---|
| 1 2 | 3/5 | |||||||
| 1 3 | 4/5 | |||||||
| 1 4 | 5/5 | |||||||
| 1 5 | 2/5 | |||||||
| 1 6 | 5/5 | |||||||
| 1 7 | 2/5 |
| Private | Others | Market | ||||
|---|---|---|---|---|---|---|
| Entrant | R | N | R | N | R | N |
| Astra | 129.13 | 229.91 | 170.34 | -152.47 | 299.47 | 77.44 |
| Sol | 145.38 | 245.64 | -333.09 | -353.99 | -187.71 | -108.35 |
| Qwen | 33.30 | 132.62 | 68.24 | -500.26 | 101.54 | -367.64 |
| DeepSeek Flash | -78.93 | 47.83 | -235.81 | -124.26 | -314.74 | -76.42 |
| Gemini | -55.01 | 48.74 | -55.16 | -258.86 | -110.17 | -210.12 |
| Pair | Rule CEO | No Action CEO | Same sign |
|---|---|---|---|
| 1 2 | Yes | ||
| 1 3 | Yes | ||
| 1 4 | No | ||
| 1 5 | Yes | ||
| 1 6 | Yes | ||
| 1 7 | Yes |
| Tool | Arguments and semantics |
|---|---|
| query_company | Read-only SQL, parameters, row limit; returns authorized company rows. Disabled throughout finalization. |
| set_product | Product, list price, promotion, reorder target, daily development budget; omitted fields remain unchanged. |
| set_marketing | Segment, channel, recurring daily budget. |
| set_operations | Recurring reliability and support budgets. |
| start_research_project | Innovation tier 1–3; one-time company-wide quality investment with stochastic completion and gain. |
| expand_capacity | Higher capacity tier; cumulative capital difference paid now, new throughput and overhead after construction. |
| Component | Contract |
|---|---|
| Business replies | Finalization starts after 28 accounting-valid replies; at most 32 per ordinary logical decision attempt. |
| Generated output | Finalization at 98,304 tokens; hard allowance 131,072, including reasoning without double counting. |
| Active time | Finalization at 2,700 seconds; hard allowance 3,600, including automatic retry backoff but excluding queue time and manual pauses. |
| Closing phase | The first threshold starts at most four further replies within the same hard limits; SQL queries are rejected locally without changing signed tool definitions. |
| Memory and context | Fresh weekly conversation plus the first 40,000 characters of MEMORY.md ; native weekly history retained, with no automatic summary or rolling compaction. |
| Unfinished window | No unsealed draft execution; standing policies continue. Budget exhaustion and infrastructure timeout are recorded separately from valid seals. |
| Seat | API model identifier | Configured protocol | Tokens |
|---|---|---|---|
| 1 | gpt-6-astra | Responses; max effort | 128,000 |
| 2 | deepseek-v4-pro | Chat; thinking enabled, max effort | 393,216 |
| 3 | claude-sonnet-5 | Messages; adaptive thinking, max effort | 128,000 |
| 4 | qwen3.8-max | Chat; thinking enabled, xhigh effort | 65,536 |
| 5 | gpt-5.6-sol | Responses; max effort | 128,000 |
| 6 | gemini-3.8-flash | Chat; high effort | 65,536 |
| Dimension | Values |
|---|---|
| Price multiplier | 0.9 or 1.0 times the catalog initial prices (75, $135). |
| Focus | Value/A, mainstream/B, premium/C (segment/core product). |
| Spending package | Light: 50/50; standard: 100/100 for development/total marketing/reliability/support. |
| Cash floor | 150,000; gates development and marketing only. |
| Not searched | A/B/C stock targets 150/100/50; no promotion or capacity adjustment. |
| ID | m | Focus | Pack | F | Seed 7 | Seed 19 | Seed 42 | Mean Score |
|---|---|---|---|---|---|---|---|---|
| 01 | 0.9 | Value | L | 50 | -72,196.49 | -58,480.39 | -72,550.61 | -67,742.50 |
| 02 | 0.9 | Value | L | 150 | -72,196.49 | -58,480.39 | -72,550.61 | -67,742.50 |
| 03 | 0.9 | Value | S | 50 | -142,152.16 | -130,244.56 | -140,070.98 | -137,489.23 |
| 04 | 0.9 | Value | S | 150 | -142,152.16 | -130,244.56 | -140,070.98 | -137,489.23 |
| 05 | 0.9 | Mainstream | L | 50 | 20,850.49 | 48,618.48 | 29,736.59 | 33,068.52 |
| 06 | 0.9 | Mainstream | L | 150 | 20,850.49 | 48,618.48 | 29,736.59 | 33,068.52 |
| Product parameter | A | B | C |
| Base quality | 0.55 | 0.68 | 0.82 |
| Initial price / unit | $40 | $75 | $135 |
| Purchase cost / unit | $26 | $48 | $85 |
| Liquidation value / unit | $8 | $15 | $26 |
| Holding cost / unit-day | $0.02 | $0.03 | $0.06 |
| Initial inventory / units | 600 | 400 | 200 |
| Parameter | Value | Mainstream | Premium |
| Baseline daily opportunities | 240 | 300 | 150 |
| Reference price / dollars | 45 | 80 | 140 |
| Price weight | 5.8 | 4.4 | 2.2 |
| Quality weight | 1.7 | 2.6 | 2.7 |
| Service weight | 0.7 | 1.0 | 1.6 |
| Awareness weight | 0.3 | 0.4 | 0.3 |
| Phase | Demand | Mix | Weights | Quality shift |
|---|---|---|---|---|
| Balanced | 1.00 | 0 | ||
| Price pressure | .90 | 0 | ||
| Quality growth | 1.10 | .035 | ||
| Delivery peak | 1.35 | 0 |
| Block | Reference values | Purpose |
|---|---|---|
| Marketing | , , b_{0}=\100\delta_{G}=.025\gamma_{G}=.08$ ; initial awareness .20. | Concave spending response, finite attention, and carryover; goodwill half-life about 27 days. |
| Knowledge | /day, , K_{0}=\50,000$ ; daily development delay 7 days; capital multiplier 2.0. | Slow depreciation and diminishing quality gains; spending cannot yield instant products. |
| Project tiers | Costs 20k/70k/200k; log standard deviations .25 (duration), .20 (gain). | Risky, delayed investment competing with current liquidity; gains are not cash payments. |
| Reliability | , , r_{0}=\6012 remediation/failure. | Nonzero residual risk and a cost for operating near or beyond capacity. |
| Support | \eta_{0}=.015\eta_{f}=.20$ . | A separate queue that requires recurring expenditure. |
| Experience | , ; , ; for every segment; loyalty update .10. | Persistent service effects and greater weight on disappointments; numerical scales are chosen, not fitted. |
| Rank | CEO agent | Official Score | Salvage | Cash-only Score |
|---|---|---|---|---|
| 1 | GPT-6 Astra | |||
| 2 | GPT-5.6 Sol | |||
| 3 | Qwen 3.8 Max | |||
| 4 | Rule CEO | |||
| 5 | DeepSeek V4 Flash | |||
| 6 | Gemini 3.8 Flash |
| Market difference | Market seed signs | Directed sign matches | |||
|---|---|---|---|---|---|
| Entrant | Official | Cash-only | Mean | Seed-level | |
| GPT-6 Astra | 7/7 | 20/21 | |||
| GPT-5.6 Sol | 7/7 | 20/21 | |||
| Qwen 3.8 Max | 7/7 | 21/21 | |||
| DeepSeek V4 Flash | 7/7 | 20/21 | |||
| Gemini 3.8 Flash | 7/7 | 20/21 | |||
| Agent/sample | Score | Revenue | Net outlays | Fulfilled units | Marketing | |
|---|---|---|---|---|---|---|
| Astra | 24 | 116.94 | 1,460.95 | 1,329.10 | 23,252 | 35.91 |
| Sol | 24 | 98.44 | 1,839.24 | 1,726.16 | 29,766 | 81.67 |
| Qwen | 24 | 8.04 | 1,279.92 | 1,258.52 | 20,954 | 85.22 |
| Qwen: | 11 | 79.25 | 1,501.79 | 1,409.47 | 24,246 | 84.46 |
| Qwen: | 13 | 1,092.19 | 1,130.79 | 18,170 | 85.86 |
| Agent; seed/day | Note excerpt | Sealed action and settlement |
|---|---|---|
| Astra; 11/7 | Reductions pause future procurement | A/B/C reorder targets become 350/160/120 on day 8. |
| Astra; 29/147 | Tier1 is unnecessary for observed consumer volume | Downsize from tier 1 to 0, completed on day 154; C target falls from 400 to 80. |
| Astra; 47/7 | Base capacity has 25 units/day spare | Cancel pending tier-1 expansion; the receipt records a $6,600 refund. |
| Sol; 11/126 | a lead-time hedge for tier-3 innovation | Expand from tier 1 to 2, due day 161; later downsize on days 182 and 224. |
| Sol; 29/231 | Don’t keep unused capacity just for terminal value | Downsize from tier 1 to 0, completed on day 238; A target falls from 583 to 300. |
| Sol; 47/7 | Keep all nine segment/channel marketing budgets at 5k/day | Retain total marketing of $450/day while cancelling pending tier-2 capacity. |
| Own consumer | Other seven firms | |||||
|---|---|---|---|---|---|---|
| Entrant | Seed | order difference | Revenue | Net outlays | Salvage | Score |
| Astra | 11 | |||||
| Astra | 29 | |||||
| Astra | 47 | |||||
| Sol | 11 | |||||
| Sol | 29 | |||||
| Entrant rival | Seed | Revenue | Net outlays | Salvage | Score |
|---|---|---|---|---|---|
| Pro Doubao | 11 | ||||
| 29 | |||||
| 47 | |||||
| Sol Sonnet | 11 | ||||
| 29 | |||||
| 47 |
| Status | Economic treatment | Reporting consequence |
|---|---|---|
| Sealed | Execute the validated draft at the barrier; an unchanged-policy seal is valid. | Count a successful decision and its reported usage. |
| Budget exhausted | Discard unsealed changes; standing policies continue and the barrier may settle. | Record budget_exhausted , separately from a deliberate unchanged-policy seal. |
| Timeout | Exhausted infrastructure recovery applies no unsealed changes; standing policies continue. | Record timeout , failure metadata and the infrastructure-timeout count. |
| Protocol error | Correct recoverable replies in-session; a final invalid Harness result applies no draft. | Retain economic consequences and correction diagnostics separately. |
| Provider error | Retry eligible failures; terminal failures stop the affected run at its durable frontier. | Apply the bounded supervision rule where eligible; retain all attempt costs. |
| Pause / integrity error | Preserve receipts; do not settle an incomplete or inconsistent barrier. | Resolve before accepting a completed result. |
| CEO agent | Active company- windows | Sealed | Budget exhausted | Timeout | Protocol error |
| (a) Main evaluation: 27 runs | |||||
| gpt-6-astra | 1,728 | 1,715 | 9 | 4 | 0 |
| deepseek-v4-pro | 1,556 | 1,556 | 0 | 0 | 0 |
| claude-sonnet-5 | 1,680 | 1,273 | 397 | 10 | 0 |
| qwen3.8-max | 1,728 | 1,707 | 21 | 0 | 0 |
| gpt-5.6-sol | 1,728 | 1,691 | 27 | 10 | 0 |
| Artifact | Evidence |
|---|---|
| Main-evaluation manifest | Plan hash prefix 3ee34a22e881 ; scenario hash prefix 893815172625 ; model controls, seats and economic configuration. |
| Repeated/baseline records | Separate campaign records, final checkpoints, terminal analysis and verification receipts for the 18 repeated and eight No Action runs; see Appendix F.6 . |
| Main-evaluation replay | 27 verified run receipts, 216 terminal company scores, 13,500 daily transitions and 1,944 world decision windows. |
| Repeated/baseline replay | 26 verified run receipts, 208 terminal company scores, 13,000 daily transitions and 1,872 world decision windows. |
| Evaluation total | 53 verified markets, 424 terminal company scores, 26,500 daily transitions and 3,816 world decision windows. |
| Decision health | Per-agent status counts for all three experiment groups in Table 25 ; 29,748 active company-windows in total. |
| Experiment | Markets | Firms per market | Company trajectories | Simulation days | CEO decision windows |
|---|---|---|---|---|---|
| Main evaluation | 27 | 8 | 216 | 13,500 | 15,078 |
| Repeated runs | 18 | 8 | 144 | 9,000 | 10,133 |
| No Action baseline | 8 | 8 | 64 | 4,000 | 4,537 |
| Evaluation total | 53 | 8 | 424 | 26,500 | 29,748 |
| Rule parameter search | 75 | 9 | 675 | 37,500 | 48,600 |
| CEO agent | Trajectories | Counted windows | Model replies | Tool calls |
|---|---|---|---|---|
| GPT-6 Astra | 47 | 2,376 | 23,973 | 96,763 |
| GPT-5.6 Sol | 47 | 2,265 | 27,133 | 135,060 |
| Qwen 3.8 Max | 47 | 3,315 | 36,607 | 94,260 |
| Claude Sonnet 5 | 47 | 2,244 | 30,959 | 88,768 |
| DeepSeek V4 Flash | 47 | 3,273 | 35,343 | 132,440 |
| DeepSeek V4 Pro | 47 | 3,096 | 36,527 | 97,639 |