Organizations: DeepWisdom · Fudan University · The Chinese University of Hong Kong · The Chinese University of Hong Kong, Shenzhen · The Hong Kong University of Science and Technology (Guangzhou)
Long-horizon competition tests agents' ability to coordinate business decisions under uncertainty and adapt to changing rival strategies. We introduce CEO Arena, a benchmark that uses matched replacement evaluation to assess operating returns alongside an agent's effects on rivals and the market. Each CEO agent is compared with a reference policy in the same company under the same economic seed, holding other agents' identities and assignments fixed while all agents adapt. In a shared eight-company market spanning 500 simulated days, CEOs make sequential decisions on pricing, procurement, marketing, research and development, and service using private company information and noisy market signals, under resource constraints and delayed feedback. We evaluate eight LLM-based CEO agents in 27 main runs and 26 robustness runs. In the main evaluation, most agents have negative mean returns, and private gains can accompany market losses. Robustness analyses suggest that aggregate patterns extend beyond the original rule-based baseline; four of the 56 directed pairs show relatively stable effects. Memory, action, and accounting traces suggest demand capture and rivals' pricing and spending responses as possible explanations. CEO Arena provides a controlled testbed for studying long-horizon agent competition, strategic interaction, and market externalities.
Figures & tables
Figure 1: CEO Arena system overview. CEO agents use private company data and tools to make operating decisions in a shared market. Each decision window collects their committed policies before advancing the market and returning feedback. The right-hand timeline shows the seed-11 reference run.
Figure 2: Shared-market mechanics. In the seed-11 reference run, the panels show realized sales by segment, one same-version decision window, Astra’s delayed projects, and a terminal-state comparison of private and public information. The 690 daily opportunities are a baseline before market variation; Appendix D specifies delay rules and parameter ranges.
Capability
Evidence and action
Direct / competitive consequence
Delay
Analysis
Query private SQL and public quotes
Diagnose policy outcomes; no direct state change
Immediate
Pricing
Set product price and promotion
Margin versus jointly allocated demand
Next window
Procurement
Inspect inventory; set reorder target
Cash tied in stock; shortages redirect demand
5–10 days
Marketing
Set segment–channel budgets
Attention share and persistent awareness
Daily carryover
Innovation
Fund daily development or a project
Product quality changes relative attractiveness
7 days / 45–150 day medians
Operations
Fund reliability and customer support
Failures, queues, satisfaction, future demand
Daily carryover
Table 1: Company capabilities and their consequences. “Next window” means the next committed policy period; daily execution and longer project delays occur inside the engine. Appendix D gives formulas and sources; Appendix B specifies tools.
Figure 3: An illustrative shared-market run. Daily cash balances of the eight agent-controlled companies in the main-evaluation reference lineup at seed 11. The dashed line marks the initial $500,000 cash balance; markers indicate bankruptcy. Legend parentheses give terminal cash on hand or, for bankrupt companies, the bankruptcy day.
Figure 4: Cash trajectories across the main evaluation. Each LLM-based agent panel shows 24 company trajectories: three colored curves from the reference markets at seeds 11, 29, and 47, and 21 gray curves from seven lineups in which Rule replaces a rival. The Rule panel contains its 24 replacement-market trajectories. Dashed lines mark initial cash; markers indicate bankruptcy.
Rank
CEO
Mean Score
Bankruptcy rate
1
GPT-6 Astra
116.94
0.0%
2
GPT-5.6 Sol
98.44
0.0%
3
Qwen 3.8 Max
8.04
0.0%
4
Rule CEO
-47.30
0.0%
5
DeepSeek V4 Flash
-114.50
4.2%
6
Gemini 3.8 Flash
-119.79
0.0%
Table 2: Operating performance. Amounts are USD thousands; each CEO agent has 24 outcomes. Rank uses Mean Score. Bankruptcy rate counts cash-negative exits.
Figure 5: Operating performance in the main evaluation. Open circles show each agent’s eight three-seed lineup means; pale bars span their minimum and maximum. Diamonds and right-hand values give Mean Score across 24 company outcomes. Rows follow Mean Score rank.
Figure 6: Externalities across agents. A, B, and D use matched Rule comparisons. (A) Logos mark three-seed mean private and market Score differences. Pale crosses span the minimum–maximum across seeds on each axis; shading above/below y=x denotes aggregate gains/losses for rivals. (B) Entrant agents are columns and affected agents rows; blue/red denotes positive/negative directed three-seed mean differences. Diagonal cells are undefined; boxes link to D. (C) Seed-11 market differences against Rule (diamonds) and No Action (open circles); seven of eight signs agree. (D) Four of 56 directed pairs satisfy ∣mean∣>SD across three main seeds and two seed-11 repeats. Bars show sample SD; counts show sign agreement.
Entrant Agent
Private Score difference
Mean externality per rival
Total externality
Market Score difference
Positive market- difference seeds
GPT-6 Astra
116.36
50.05
350.33
466.68
3/3
GPT-5.6 Sol
131.55
-31.60
-221.22
-89.66
1/3
Qwen 3.8 Max
23.28
19.09
133.64
156.92
2/3
DeepSeek V4 Flash
13.62
-13.97
-97.82
-84.20
2/3
Gemini 3.8 Flash
-53.09
15.55
108.87
55.78
2/3
DeepSeek V4 Pro
-54.96
-74.04
-518.29
-573.25
0/3
Table 3: Private returns and externalities. Main-evaluation three-seed means (USD thousands), subtracting matched Rule outcomes from agent outcomes. Mean externality equals the total divided by seven. Private plus total externality equals the market difference before rounding. The final column counts positive market differences, not profitable markets.
Appendix figures & tables34 assets
Supplementary material from the paper’s appendix.
Appendix
CEO agent
F11
R1
R2
GPT-6 Astra
91.14(1)
54.97(1)
105.44(1)
GPT-5.6 Sol
87.21(2)
17.00(3)
66.98(2)
Qwen 3.8 Max
−20.41(3)
26.05(2)
−51.57(4)
Rule CEO
−61.69(4)
−54.80(4)
−50.10(3)
DeepSeek V4 Flash
−106.45(5)
−79.47(5)
−148.50(7)
Gemini 3.8 Flash
−152.67(7)
−97.07(6)
−136.08(5)
Appendix
Table 4: Operating performance across three seed-11 runs (USD thousands). Each value averages eight lineups; parentheses show ranks within each run.
Pair
F11
F29
F47
R1
R2
Mean
SD
Sign
1 → 2
58.85
−82.49
−7.86
2.99
180.06
30.31
97.68
3/5
1 → 3
−50.24
−90.48
−41.60
47.59
−67.60
−40.47
52.65
4/5
1 → 4
75.01
293.35
290.29
6.43
182.75
169.57
128.09
5/5
1 → 5
131.04
19.32
−39.45
−13.86
−8.97
17.62
66.75
2/5
1 → 6
233.33
149.53
29.61
59.90
45.48
103.57
86.17
5/5
1 → 7
−81.83
224.64
43.64
−28.45
−47.95
22.01
122.22
2/5
Appendix
Table 5: All 56 directed externalities across five observations (USD thousands).
Private
Others
Market
Entrant
R
N
R
N
R
N
Astra
129.13
229.91
170.34
-152.47
299.47
77.44
Sol
145.38
245.64
-333.09
-353.99
-187.71
-108.35
Qwen
33.30
132.62
68.24
-500.26
101.54
-367.64
DeepSeek Flash
-78.93
47.83
-235.81
-124.26
-314.74
-76.42
Gemini
-55.01
48.74
-55.16
-258.86
-110.17
-210.12
Appendix
Table 6: Baseline comparisons at seed 11 (USD thousands). R denotes Rule CEO and N denotes No Action CEO.
Pair
Rule CEO
No Action CEO
Same sign
1 → 2
80.63
26.20
Yes
1 → 3
−23.42
−51.68
Yes
1 → 4
88.06
−145.97
No
1 → 5
36.07
23.48
Yes
1 → 6
112.91
170.75
Yes
1 → 7
−52.74
−109.79
Yes
Appendix
Table 7: Directed externalities under the two baseline policies at seed 11 (USD thousands); agent indices match Table 5 .
Tool
Arguments and semantics
query_company
Read-only SQL, parameters, row limit; returns authorized company rows. Disabled throughout finalization.
set_product
Product, list price, promotion, reorder target, daily development budget; omitted fields remain unchanged.
set_marketing
Segment, channel, recurring daily budget.
set_operations
Recurring reliability and support budgets.
start_research_project
Innovation tier 1–3; one-time company-wide quality investment with stochastic completion and gain.
expand_capacity
Higher capacity tier; cumulative capital difference paid now, new throughput and overhead after construction.
Appendix
Table 8: The complete official tool surface. Monetary inputs are integer cents.
Figure 7: Agent interface examples. A CEO queries its private records, stages segment-by-channel budgets within one decision, and competes for an enterprise tender. The tender panel depicts the seed-11 reference market; eligibility and deliverable capacity are checked before the award.
Component
Contract
Business replies
Finalization starts after 28 accounting-valid replies; at most 32 per ordinary logical decision attempt.
Generated output
Finalization at 98,304 tokens; hard allowance 131,072, including reasoning without double counting.
Active time
Finalization at 2,700 seconds; hard allowance 3,600, including automatic retry backoff but excluding queue time and manual pauses.
Closing phase
The first threshold starts at most four further replies within the same hard limits; SQL queries are rejected locally without changing signed tool definitions.
Memory and context
Fresh weekly conversation plus the first 40,000 characters of MEMORY.md ; native weekly history retained, with no automatic summary or rolling compaction.
Unfinished window
No unsealed draft execution; standing policies continue. Budget exhaustion and infrastructure timeout are recorded separately from valid seals.
Appendix
Table 9: Default decision-v9 allowance shared across model routes. These limits do not constitute a dollar or input-context budget.
Seat
API model identifier
Configured protocol
Tokens
1
gpt-6-astra
Responses; max effort
128,000
2
deepseek-v4-pro
Chat; thinking enabled, max effort
393,216
3
claude-sonnet-5
Messages; adaptive thinking, max effort
128,000
4
qwen3.8-max
Chat; thinking enabled, xhigh effort
65,536
5
gpt-5.6-sol
Responses; max effort
128,000
6
gemini-3.8-flash
Chat; high effort
65,536
Appendix
Table 10: Evaluated model IDs, fixed seats, configured transport, and native response ceilings. Actual output is further bounded by the remaining phase allowance.
Dimension
Values
Price multiplier
0.9 or 1.0 times the catalog initial prices (40,75, $135).
Light: 50/50/25/50; standard: 100/100/50/100 for development/total marketing/reliability/support.
Cash floor
50,000or150,000; gates development and marketing only.
Not searched
A/B/C stock targets 150/100/50; no promotion or capacity adjustment.
Appendix
Table 11: Predeclared fixed-Rule grid. Dollar amounts are daily unless marked as a cash floor.
ID
m
Focus
Pack
F
Seed 7
Seed 19
Seed 42
Mean Score
01
0.9
Value
L
50
-72,196.49
-58,480.39
-72,550.61
-67,742.50
02
0.9
Value
L
150
-72,196.49
-58,480.39
-72,550.61
-67,742.50
03
0.9
Value
S
50
-142,152.16
-130,244.56
-140,070.98
-137,489.23
04
0.9
Value
S
150
-142,152.16
-130,244.56
-140,070.98
-137,489.23
05
0.9
Mainstream
L
50
20,850.49
48,618.48
29,736.59
33,068.52
06
0.9
Mainstream
L
150
20,850.49
48,618.48
29,736.59
33,068.52
Appendix
Table 12: All development results. Values are focal-company Scores in dollars; L/S denotes light/standard spending and F is the cash floor in thousands of dollars. Bold identifies the winner under the declared mean/tie rule.
Figure 8: Price and model-implied demand share. The focal firm varies one product’s price from 50% to 160% of its catalogue value; seven identical rivals keep catalogue prices. Awareness is equal across firms, and satisfaction and loyalty are set to zero. Curves A–C are conditional choice shares, not observed sales.
Product parameter
A
B
C
Base quality
0.55
0.68
0.82
Initial price / unit
$40
$75
$135
Purchase cost / unit
$26
$48
$85
Liquidation value / unit
$8
$15
$26
Holding cost / unit-day
$0.02
$0.03
$0.06
Initial inventory / units
600
400
200
Appendix
Table 14: Symmetric product and capacity configuration. Product columns refer to product tiers, not agent identifiers.
Parameter
Value
Mainstream
Premium
Baseline daily opportunities dh
240
300
150
Reference price / dollars
45
80
140
Price weight whp
5.8
4.4
2.2
Quality weight whq
1.7
2.6
2.7
Service weight whs
0.7
1.0
1.6
Awareness weight wha
0.3
0.4
0.3
Appendix
Table 15: Customer-segment parameters. All three segments draw from the same active-company offer set.
Phase
Demand
Mix
Weights
Quality shift
Balanced
1.00
(1,1,1)
(1,1,1)
0
Price pressure
.90
(1.5,1,.7)
(1.08,.90,1)
0
Quality growth
1.10
(.8,1,1.4)
(.95,1.25,1)
.035
Delivery peak
1.35
(1,1,1)
(.97,1,1.35)
0
Appendix
Table 16: Shared market phases. Mix entries are value/mainstream/premium; weight entries are price/quality/service multipliers.
A separate queue that requires recurring expenditure.
Experience
ρ=.04 , ω−=1.8 ; ωw=.15 , ωf=ωu=1 ; ωh=.12 for every segment; loyalty update .10.
Persistent service effects and greater weight on disappointments; numerical scales are chosen, not fitted.
Appendix
Table 17: Remaining dynamics and their design roles. Numerical values are frozen scenario settings.
Rank
CEO agent
Official Score
Salvage
Cash-only Score
1
GPT-6 Astra
116.94
1.09
115.85
2
GPT-5.6 Sol
98.44
1.36
97.08
3
Qwen 3.8 Max
8.04
2.63
5.41
4
Rule CEO
−47.30
4.00
−51.30
5
DeepSeek V4 Flash
−114.50
9.70
−124.21
6
Gemini 3.8 Flash
−119.79
5.03
−124.81
Appendix
Table 18: Terminal-valuation sensitivity. Means over 24 main-evaluation outcomes per CEO (USD thousands). Salvage includes on-hand and paid inbound inventory; ranks are identical under both measures.
Market difference
Market seed signs
Directed sign matches
Entrant
Official
Cash-only
Mean
Seed-level
GPT-6 Astra
+466.68
+475.40
(+,+,+)
7/7
20/21
GPT-5.6 Sol
−89.66
−48.99
(−,+,−)
7/7
20/21
Qwen 3.8 Max
+156.92
+183.76
(+,+,−)
7/7
21/21
DeepSeek V4 Flash
−84.20
−79.99
(+,+,−)
7/7
20/21
Gemini 3.8 Flash
+55.78
+62.37
(+,+,−)
7/7
20/21
Appendix
Table 19: Matched differences without terminal salvage. Three-seed market means (USD thousands), seed signs in order 11/29/47, and retained directed signs under cash-only rescoring. Each entrant has seven directed means and 21 seed-level contrasts.
Figure 9: Cash on hand in all nine main-evaluation markets at seed 11. The upper-left panel is the eight-agent reference market; the other panels replace the named agent with Rule CEO. Colors identify agents, the dashed brown line identifies Rule, and skull markers indicate firm exit days.
Figure 10: Cash on hand in all nine main-evaluation markets at seed 29. Panel order and visual encoding match Figure 9 : the reference market precedes eight same-seat Rule replacements, and skull markers indicate firm exit days.
Figure 11: Cash on hand in all nine main-evaluation markets at seed 47. Panel order and visual encoding match Figure 9 : the reference market precedes eight same-seat Rule replacements, and skull markers indicate firm exit days.
Figure 12: Cash outlays in the main evaluation. Bars pool each CEO agent’s 24 outcomes and show the share spent by category; right-hand values are total spending divided by 24 (USD thousands). These are gross cash outflows, not Score contributions.
Figure 13: Recorded behavior and operating returns. Vertical positions are main-evaluation Mean Scores over 24 outcomes. In (a), horizontal values are the percentage of active decision windows left unsealed across those outcomes. In (b), queries per window come from the seed-11 reference market; (c) counts if–then statements per note for agents with notes in that market. These comparisons do not identify causal effects.
Figure 14: Selected memory and operating events. Astra, Sonnet, and Doubao in the seed-11 reference lineup; badges report terminal Score in USD thousands, while cash values marked c in excerpts are cents. This Doubao firm exits on day 304; case T2 above describes a different seed-11 replacement lineup whose Doubao firm exits on day 110.
Figure 15: Conditional plans in private notes. Selected seed-11 reference-market excerpts and if–then statements per written note for agents with notes. These counts describe text, not plan execution or decision quality.
Agent/sample
n
Score
Revenue
Net outlays
Fulfilled units
Marketing
Astra
24
116.94
1,460.95
1,329.10
23,252
35.91
Sol
24
98.44
1,839.24
1,726.16
29,766
81.67
Qwen
24
8.04
1,279.92
1,258.52
20,954
85.22
Qwen: S>0
11
79.25
1,501.79
1,409.47
24,246
84.46
Qwen: S<0
13
−52.21
1,092.19
1,130.79
18,170
85.86
Appendix
Table 20: Accounting profiles of the three agents with positive Mean Score. Money is in USD thousands; all entries except n are means. Net outlays subtract recorded refunds. The last two rows condition on Qwen’s realized Score and are descriptive, not separate experimental groups.
Agent; seed/day
Note excerpt
Sealed action and settlement
Astra; 11/7
Reductions pause future procurement
A/B/C reorder targets become 350/160/120 on day 8.
Astra; 29/147
Tier1 is unnecessary for observed consumer volume
Downsize from tier 1 to 0, completed on day 154; C target falls from 400 to 80.
Astra; 47/7
Base capacity has 25 units/day spare
Cancel pending tier-1 expansion; the receipt records a $6,600 refund.
Sol; 11/126
a lead-time hedge for tier-3 innovation
Expand from tier 1 to 2, due day 161; later downsize on days 182 and 224.
Sol; 29/231
Don’t keep unused capacity just for terminal value
Downsize from tier 1 to 0, completed on day 238; A target falls from 583 to 300.
Sol; 47/7
Keep all nine segment/channel marketing budgets at 5k/day
Retain total marketing of $450/day while cancelling pending tier-2 capacity.
Appendix
Table 21: Memory statements linked to executed decisions. Entries identify agent, seed, and decision day. All use the reference market except Qwen: Rule replaces Pro at seeds 11/47 and Astra at seed 29. Excerpts omit Markdown emphasis and use cents; action amounts are converted to dollars. Receipts confirm execution, not the action’s causal payoff.
Own consumer
Other seven firms
Entrant
Seed
order difference
Δ Revenue
Δ Net outlays
Δ Salvage
Δ Score
Astra
11
+6,172
−1,982.84
−2,324.11
+24.60
+365.87
Astra
29
+6,221
−199.92
−749.27
−30.79
+518.56
Astra
47
+7,179
+872.10
+694.86
−10.68
+166.55
Sol
11
+16,399
−3,009.41
−2,707.13
−30.28
−332.57
Sol
29
+15,539
−483.51
−759.61
−79.35
+196.75
Appendix
Table 22: Accounting decomposition of aggregate externalities. Each row compares the reference market with the same-seed Rule replacement of the entrant. Rival columns sum the other seven firms; monetary differences are in USD thousands. Net outlays subtract refunds.
Entrant → rival
Seed
Δ Revenue
Δ Net outlays
Δ Salvage
Δ Score
Pro → Doubao
11
+549.63
+550.01
−150.81
−151.19
29
−984.90
−929.82
+0.02
−55.05
47
−999.76
−940.06
−4.92
−64.62
Sol → Sonnet
11
−169.19
+198.26
+27.30
−340.16
29
−324.02
−245.91
+5.09
−73.02
47
−238.33
−66.83
−2.31
−173.80
Appendix
Table 23: Main-evaluation accounting for the four directed pairs. Differences are reference minus matched Rule-replacement outcomes for the affected agent, in USD thousands. These three-seed observations provide the mechanism evidence; the five-observation stability summary remains in Figure 6 D.
Status
Economic treatment
Reporting consequence
Sealed
Execute the validated draft at the barrier; an unchanged-policy seal is valid.
Count a successful decision and its reported usage.
Budget exhausted
Discard unsealed changes; standing policies continue and the barrier may settle.
Record budget_exhausted , separately from a deliberate unchanged-policy seal.
Timeout
Exhausted infrastructure recovery applies no unsealed changes; standing policies continue.
Record timeout , failure metadata and the infrastructure-timeout count.
Protocol error
Correct recoverable replies in-session; a final invalid Harness result applies no draft.
Retain economic consequences and correction diagnostics separately.
Provider error
Retry eligible failures; terminal failures stop the affected run at its durable frontier.
Apply the bounded supervision rule where eligible; retain all attempt costs.
Pause / integrity error
Preserve receipts; do not settle an incomplete or inconsistent barrier.
Resolve before accepting a completed result.
Appendix
Table 24: Decision outcomes and their economic treatment.
CEO agent
Active company- windows
Sealed
Budget exhausted
Timeout
Protocol error
(a) Main evaluation: 27 runs
gpt-6-astra
1,728
1,715
9
4
0
deepseek-v4-pro
1,556
1,556
0
0
0
claude-sonnet-5
1,680
1,273
397
10
0
qwen3.8-max
1,728
1,707
21
0
0
gpt-5.6-sol
1,728
1,691
27
10
0
Appendix
Table 25: Final decision status by CEO agent and experiment group. Active company-windows count decisions while the firm remains active; the status columns sum to that count. All settled decision records are included.
Figure 16: Recorded actions and window outcomes. Left: mean settings written per active decision window across each CEO’s 24 main-evaluation outcomes, grouped by policy type. Right: sealed versus unsealed windows; unsealed includes budget exhaustion, timeout, and protocol error. Unsealed drafts do not change standing policies.
Artifact
Evidence
Main-evaluation manifest
Plan hash prefix 3ee34a22e881 ; scenario hash prefix 893815172625 ; model controls, seats and economic configuration.
Repeated/baseline records
Separate campaign records, final checkpoints, terminal analysis and verification receipts for the 18 repeated and eight No Action runs; see Appendix F.6 .
Main-evaluation replay
27 verified run receipts, 216 terminal company scores, 13,500 daily transitions and 1,944 world decision windows.
Repeated/baseline replay
26 verified run receipts, 208 terminal company scores, 13,000 daily transitions and 1,872 world decision windows.
Evaluation total
53 verified markets, 424 terminal company scores, 26,500 daily transitions and 3,816 world decision windows.
Decision health
Per-agent status counts for all three experiment groups in Table 25 ; 29,748 active company-windows in total.
Appendix
Table 26: Retained evidence for the 53 evaluation runs. Main-evaluation and repeated/baseline records are archived separately; full hashes accompany the research artifacts.
Experiment
Markets
Firms per market
Company trajectories
Simulation days
CEO decision windows
Main evaluation
27
8
216
13,500
15,078
Repeated runs
18
8
144
9,000
10,133
No Action baseline
8
8
64
4,000
4,537
Evaluation total
53
8
424
26,500
29,748
Rule parameter search
75
9
675
37,500
48,600
Appendix
Table 27: Experimental scale. Simulation days count market-level transitions. Shared controls are counted once; replay does not add new experimental runs.
CEO agent
Trajectories
Counted windows
Model replies
Tool calls
GPT-6 Astra
47
2,376
23,973
96,763
GPT-5.6 Sol
47
2,265
27,133
135,060
Qwen 3.8 Max
47
3,315
36,607
94,260
Claude Sonnet 5
47
2,244
30,959
88,768
DeepSeek V4 Flash
47
3,273
35,343
132,440
DeepSeek V4 Pro
47
3,096
36,527
97,639
Appendix
Table 28: Recorded workload across the 53 evaluation runs. Counted windows are successfully sealed decisions with complete reported usage. Model replies per trajectory are the mean ± sample SD of the counted replies across each agent’s company trajectories. Token quantities are millions (M).
Language model agents are becoming proficient executors at isolated, short-horizon tasks such as software engineering and customer service. Yet real-world challenges require a combination of sophisticated skills that remain largely untested in agents: (1) navigating long horizons amid uncertainty; (2) acquiring information in noisy environments; (3) adapting to a changing world; (4) orchestrating multiple moving parts toward a coherent goal. We introduce CEO-Bench, which evaluates these capabilities together by simulating a representative real-world task: operating a startup for 500 days. An agent manages pricing, marketing, budgeting, and many other aspects of a fictional company through a programmable Python interface, operating in the same environment and facing the same challenges as a human CEO. Success demands analyzing noisy, interconnected business databases, translating signals into sound strategy, and coordinating many decisions with programming. The strongest agents write sophisticated code that simulates customer cohorts to forecast future cash and mines negotiation history to uncover hidden customer preferences. Even so, most state-of-the-art models struggle in this environment. Only Claude Opus 4.8 and GPT-5.5 finish above the $1M starting balance, and neither consistently turns a profit. CEO-Bench takes a first step toward measuring the intelligence required to drive sustained, adaptive progress over time.
Running a business is a challenging form of intelligent work. Operators must infer opportunities from partial signals, commit capital under uncertainty, adapt to delayed outcomes in a changing market, and satisfy regulatory obligations before trading legally. Frontier LLM agents can increasingly complete complex workflows, yet business-related capabilities are rarely evaluated in existing agent benchmarks. We introduce \textbf{Business Arena}, a controlled environment where an AI agent runs a cross-border shop, buying from suppliers and selling to buyers over a long horizon. We ground the arena in real Alibaba.com sourcing data and market conditions calibrated from authoritative sources. Delayed and coupled consequences make individual business decisions difficult to judge, but their combined outcome is measurable through profit. Because profit alone cannot explain why an agent succeeds or fails, we compare agents with human-designed strategies to estimate available opportunity, use skill-level metrics to reveal underlying strengths and weaknesses, and trace realized gains and losses to the actions that produced them. We use mechanism ablations to establish that strong results reflect genuine business intelligence rather than neglect or simulator-specific shortcuts. We evaluate 15 frontier models and find a ninefold difference in mean final net worth. Even the best model falls behind human-designed strategies, indicating that business operation remains challenging for LLM agents. Skill-level analysis reveals operating styles, from margin-focused premium sellers to high-turnover wholesalers and customer-service specialists, while action-level attribution identifies the sourcing, pricing, and recovery decisions that create or destroy value. Together, Business Arena takes a first step toward a realistic and trustworthy testbed for evaluating end-to-end business agents.
Large language model (LLM) agents are increasingly proposed for enterprise workflows, yet existing evaluations rarely test whether business-decision conclusions transfer across competitive market ecologies. We introduce ERPBench, an execution-instrumented benchmark for enterprise decision agents in a six-round Enterprise Resource Planning (ERP) simulation with coupled pricing, production, procurement, inventory, finance, and shared-market competition. ERPBench evaluates the same 100 fixed problems in two matched competitive market ecologies: Solo, where each evaluated LLM agent competes against fixed rule-based opponents, and Arena, where six evaluated LLM agents compete in a shared market. Across six model families, this yields 1,200 model-level trajectories spanning 7,200 decision rounds. Under the observed service configuration, the leading model differs between ecologies: DeepSeek leads in Solo (252.29M mean valuation; mean rank 1.67), whereas Gemini leads in Arena (263.95M; 1.76). The two ecologies identify the same task-level winner on only 21 of 100 problems, and Gemini's bottom-rank rate falls from 22 % to 0 % in Arena. ERPBench supports paired evaluation of whether enterprise-agent rankings transfer across competitive market ecologies, supplemented by aggregate execution-intervention analysis. Code and benchmark resources are available in our https://github.com/GAIR-NLP/erp-bench.
Xinran Zhang, Pengrui Lu, Lyumanshan Ye +1
1Shanghai Innovation Institute · 2Beijing Institute of Technology · 4GAIR Lab +1