The last mile toward enterprise AGI is a company that runs itself. Training and adapting such agents require longitudinal enterprise data, which remain scarce, costly to acquire, and often restricted by privacy constraints. Historical archives are also frequently incomplete and record only what actually happened. They cannot show the outcomes of alternative decisions. We introduce MiniCorp, an office simulator for studying how agents can collectively run a company while generating enterprise data at scale. Using an e-commerce company as a demonstration, MiniCorp connects two interacting worlds. The external world models customers, dynamic competitors, and market mechanisms. The internal world consists of agents that observe events, discuss their options, and make strategic decisions. These decisions have lasting effects on the market, and the resulting feedback informs the firm's later decisions. As the firm and market interact, MiniCorp continuously records the agents' communications and decisions. These records preserve the information available at the time and the business results that followed. Checkpointing allows the same situation to be replayed under different decisions, providing comparisons unavailable in static archives. We evaluate end-to-end fidelity against patterns reported in empirical studies of real markets. These evaluations provide agents with realistic market feedback and reduce the risk that they learn to exploit flaws in the simulator. Our experiments show agents coordinating across roles and adapting their decisions to market feedback. With explicit long-term strategic guidance, they also sustain advertising exploration despite weak early returns. MiniCorp thus provides an environment for studying AI-run companies and a scalable source of longitudinal and counterfactual enterprise data for agent training and evaluation.
Figures & tables
Figure 2 : Overview of MiniCorp’s interaction between the external world model and the internal AI agent team for an autonomous e-commerce company.
Test
Level
Key result
Why this supports fidelity
Immediate price response
Outcome
A 10% price reduction and increase yielded absolute arc elasticities of 2.9 and 3.2, respectively.
Demand increased when price fell and decreased when price rose. The observed elasticities were also close to the empirical benchmark [ Bijmolt et al., 2005 ] .
Temporary discount and price restoration
Pathway
Prices were reduced by 20% for four weeks and then restored to their original levels. Units increased by 71% during the discount, remained 31% above the control for the first four weeks after restoration, and remained 11% above it for the following six weeks. Nearly half of the incremental units were sold after the discount ended.
The observed response followed the pathway: temporary price reduction → higher sales → higher recent sales velocity → more appearances in organic search results → more impressions → persistently higher sales after price restoration. This pathway is consistent with evidence that past demand can influence subsequent marketplace traffic [ Cooprider and Nassiri, 2023 ; Quaker, 2025 ] .
Temporary stockout
Pathway
The sales loss persisted after a two-week stockout: 59% of the cumulative unit shortfall occurred after restocking, and sales remained 36% below the control in week 26.
The stockout produced the opposite pathway: temporary stockout → no sales while inventory was unavailable → lower recent sales velocity → fewer appearances in organic search results after restocking → fewer impressions → persistently lower sales. This pathway is consistent with evidence that stockouts can reduce future demand and that past demand can influence subsequent marketplace traffic [ Anderson et al., 2006 ; Cooprider and Nassiri, 2023 ] .
Advertising
Pathway
Ad-attributed units increased by 23–32%, while the organic gain grew from 0.5% to 5.8%. Each additional advertising dollar returned $0.53 in contribution within the observation window, and 84% of incremental units came from competitor displacement.
The observed response followed the pathway: higher advertising expenditure → more paid sales immediately → higher recent sales velocity → more appearances in organic search results → more organic impressions → a delayed increase in organic sales. This pathway is consistent with Amazon’s description of how sales and sponsored advertising affect search visibility [ Quaker, 2025 ] . The response also involved an immediate advertising cost but a delayed organic benefit, while most incremental sales came at competitors’ expense rather than from market expansion.
Table 1 : Summary of the main external-world fidelity results.
Provenance
Description
Share
Carried-over
Unfinished in the prior cycle; re-enters the queue
58%
Message-triggered
Derived from a message sent by a colleague
21%
Interrupt-driven
Injected by an urgent message that preempts current work
10%
Self-planned
Scheduled by the agent at the start of the cycle
11%
Table 2 : Provenance of work items.
Work category
Items
Msgs./item
Solo
Dec./item
Rejected
Quality incidents
5,891
0.78
80%
0.24
13%
Customer service
1,936
0.65
82%
0.65
20%
Reporting and bookkeeping
2,060
0.51
85%
0.23
18%
Replenishment and logistics
11,610
0.51
81%
0.34
29%
Advertising and pricing
4,445
0.48
80%
0.61
25%
Supplier management
534
0.50
86%
0.20
19%
Table 3: Process signatures by work category. Msgs./item and Dec./item are category means; Solo is the share of work items whose owner sends no message; and Rejected is the share of proposed decisions declined by the CEO.
A support backlog mixing genuine defects with false alarms
16 aged claims on orders that arrived damaged and 16 on orders that arrived intact
Answer the genuine claims; refuse payment on the rest
53.1%
Advertising that continues after the product has sold out
Inventory zeroed on the best-selling half of the catalogue, campaigns untouched
Cut the spend, restock, or both
12.0%
A supplier shipping a high rate of defective units
Worst vendor’s defect propensity raised to 25% and bound to a third of the catalogue
Hold, audit, or replace the vendor
7.6%
A seasonal dip in demand
Category demand pool cut by 40%, leaving advertising efficiency untouched
None; the recovery is exogenous
5.5%
Rivals bidding up the cost of advertising
Competing offers improved, so impressions cost more at unchanged click-through
Defend share; do not cut the budget
3.3%
A product whose customer ratings have collapsed
One-star reviews seeded, lowering click-through at unchanged cost per impression
Repair the product’s standing, not the budget
2.8%
Appendix
Table 5: The ten situations, ordered by the share of archived actions submitted to the world model accounted for by the action types each one exercises. Each situation initializes selected state variables or latent parameters while preserving the structure of the world model.
Situation
Found
What the company did
Collaboration
Why
Support backlog
2/2
Refunded the genuine claims, paid nothing on the false alarms
Irreplaceable, in part
Refund or replacement depends on a colleague’s pending hold
Listings paused stale
2/2
Resumed all 6 listings in week 1
Irreplaceable
The agent was not looking that way until a colleague wrote
Seasonal dip
2/2
Left advertising budgets and prices untouched
Unhelpful
Deadlock; right outcome anyway
Shared part shortage
1/2
Aimed purchase orders at the affected SKUs
Irreplaceable
The component call and the order sit in two roles
Advertising after sell-out
1/2
Paused all 15 sold-out listings once, none the second time
Helpful, then unhelpful
The second request was ignored
Cheap specification
1/2
Proposed and approved the component switch once
Helpful
—
Appendix
Table 6: Response to each situation, and what collaboration contributed. Found counts the repetitions that reached the open route of Table 5 . Collaboration is audited chain by chain: irreplaceable means the agent could not have decided alone, helpful means it happened and was cited but the agent could have decided alone, unhelpful means it happened normally and did not produce the outcome, and absent means it did not happen.
Enterprise AI agents act across many apps whose data changes continuously, so an answer is correct only relative to what data existed and who could see it at the moment it was asked. Offline evaluation today grades against a single static snapshot, effectively the end of the episode. So, it can only evaluate one situation, the final one, even though every earlier moment of the episode is a different situation that invites its own realistic questions with its own correct answers. Recreating each of those moments as a separate snapshot would mean re-provisioning a whole tenant per instant, which is prohibitively costly; and even a single snapshot leaks future state hidden inside records and cannot represent the multi-app, time-ordered way real work happens. Our system closes two gaps at once: it generates a realistic, persona-driven, temporally-evolving enterprise world from real research, and replays that world at any chosen moment to evaluate any pluggable agent. A schema-inferred temporal description drives a deterministic-plus-LLM rebuild of each record's past state; because the queryable moments are finite, all rebuilds are precomputed into a compact difference cache, making evaluation a fast, reproducible lookup with no model in the path. We describe the design, an architecture spanning both flows, and early experience evaluating enterprise agents.
Language model agents are becoming proficient executors at isolated, short-horizon tasks such as software engineering and customer service. Yet real-world challenges require a combination of sophisticated skills that remain largely untested in agents: (1) navigating long horizons amid uncertainty; (2) acquiring information in noisy environments; (3) adapting to a changing world; (4) orchestrating multiple moving parts toward a coherent goal. We introduce CEO-Bench, which evaluates these capabilities together by simulating a representative real-world task: operating a startup for 500 days. An agent manages pricing, marketing, budgeting, and many other aspects of a fictional company through a programmable Python interface, operating in the same environment and facing the same challenges as a human CEO. Success demands analyzing noisy, interconnected business databases, translating signals into sound strategy, and coordinating many decisions with programming. The strongest agents write sophisticated code that simulates customer cohorts to forecast future cash and mines negotiation history to uncover hidden customer preferences. Even so, most state-of-the-art models struggle in this environment. Only Claude Opus 4.8 and GPT-5.5 finish above the $1M starting balance, and neither consistently turns a profit. CEO-Bench takes a first step toward measuring the intelligence required to drive sustained, adaptive progress over time.
Enterprise agents increasingly operate inside workspaces: they read heterogeneous files, invoke tools, and deliver business artifacts. We introduce EnterpriseClawBench, an enterprise agent benchmark constructed from proprietary, real-world agent sessions. Starting from a large archive of workplace sessions, the EnterpriseClawBench produces 852 reproducible tasks, each paired with recovered fixtures, rewritten prompts, role classes, skill subclasses, hard rules, and semantic rubrics. Because the sessions contain internal enterprise content, we do not release the benchmark data; instead, our reusable contribution is the construction and evaluation protocol. On EnterpriseClawBench, the best configuration reaches only 0.663 (Codex with GPT-5.5). These results show that enterprise agent evaluation must report harness--model combinations, artifact delivery, visual quality, cost, runtime, and skill-transfer behavior, rather than collapsing performance into a single score. Code: https://github.com/FrontisAI/EnterpriseClawBench