The AI Evaluation Ecosystem
Organizations: Stanford University · University of British Columbia
Abstract
AI evaluation shapes the decisions of model providers, users, funders, and regulators. We argue that designing valid benchmarks requires contextualizing design choices in the dynamics of this ecosystem of actors. We develop a simulation architecture that combines rule-based market dynamics with LLM-driven strategic actors, building on advances in Generative Agent-Based Modeling (GABM). We model benchmarks, consumer needs, and provider capabilities as vectors over a six-dimensional capability space (reasoning, coding, knowledge, safety, communication, agentic), with structural information partitions across actors. As a case study, we apply this stylized simulation to explore benchmark holdout design. We find that moving from public benchmarks to private holdout benchmarks shrinks the gap between benchmark scores and user satisfaction on most benchmarks but widens it on a few, depending on where holdout weights shift scoring credit. We stress-test our findings at both the instrument and case-study level, drawing on the V&V framework of Sargent (2013) and GABM-specific evidence criteria. Beyond holdout design, our simulation is a hypothesis-generating sandbox for studying how evaluator and policy choices, in turn, reshape the ecosystem.
Figures & tables
| Instrument checks (design & calibration) | Case study validation (holdout design) | |
|---|---|---|
| Approach | Sargent V&V activities; each structural assumption rated calibrated, anchored, or stipulated. | Mechanism-isolation tests on the per-benchmark result. |
| Evidence | Three-tier visibility enforced structurally (Appendix B ); parameter grounding ( calibrated, cosine targets anchored, need-weights anchored in sector-level evidence; Appendices C.1 and C.2 ). | Cross-mode sign agreement on 12 of 13 benchmarks; cross-model replication on Opus 4.6 and GPT-5.5 ( seeds each; Appendix F.7 ); iid_holdout channel isolation; within-benchmark rescoring (Appendix F.2 ); fitted null recovers of variance (Appendix F.5 ). |
| Question (policy framing) | Simulation slot-in | Strategic angle |
|---|---|---|
| 2. Conflicts of interest (Arena; Scale SEAL). Does monetizing submissions redirect capital toward aggressive submitters at consumer cost? | Small-sample comparison in Appendix H.1 . Per-submission fees (R&D budget); best-of- selection; paid early access. | Larger-budget providers capture more submissions; funders read inflated scores as quality signals; capital flows to aggressive submitters rather than providers best serving users. |
| 3. Transparency mandates (EU AI Act Art. 53; CA SB-53; NY RAISE Act). Do forced allocation or capability disclosures reduce gaming? | New regulator lever: exposes part of provider private state (R&D mix, internal eval results) to public. | Providers face disclose-vs-window-dress tradeoff; tests whether visibility of the R&D mix changes Goodhart pressure or just changes the narrative around it. |
| 4. Audit & verification (UK AISI pre-deployment agreements; WH 2023 voluntary commitments). How effective are audits at enforcing voluntary safety commitments? | Regulator lever: provider commits to a forward safety allocation at ; regulator verifies at with penalty on violation. | Game-theoretic: provider can over-promise, under-promise, or credibly commit. Tests whether self-binding works when verification is credible, and how it interacts with incident pressure and market share. |
Appendix figures & tables33 assets
Supplementary material from the paper’s appendix.
Appendix
| Dimension | Represents | Benchmark anchor | Investment pathway |
|---|---|---|---|
| Reasoning | Problem solving, logic, multi-step inference | MMLU, GPQA | Reasoning post-training, CoT RLHF |
| Coding | Software engineering, debugging | HumanEval, MBPP | Code-specific data, correctness RL |
| Knowledge | Factual recall, domain expertise | MMLU subsets | Domain corpus, retrieval fine-tuning |
| Safety | Harmlessness, adversarial robustness | TruthfulQA | RLHF for harmlessness, red-teaming |
| Communication | Instruction following, fluency | MT-Bench, IFEval | Instruction-following RLHF |
| Agentic | Tool use, multi-step task execution | SWE-bench, WebArena | Trajectory RL, tool-use fine-tuning |
| Benchmark | Tier | Reas. | Code | Know. | Safe. | Comm. | Agent. | Real analog | ||
|---|---|---|---|---|---|---|---|---|---|---|
| General Capability | public | 0.35 | 0.08 | 0.30 | 0.05 | 0.20 | 0.02 | 0.97 | 0.89 | MMLU |
| Coding Evaluation | public | 0.10 | 0.78 | 0.03 | 0.01 | 0.02 | 0.06 | 0.99 | 0.93 | HumanEval/MBPP |
| Safety Evaluation | partial | 0.03 | 0.01 | 0.05 | 0.80 | 0.10 | 0.01 | 0.99 | 0.92 | TruthfulQA/BBQ |
| Instruction Following | public | 0.08 | 0.02 | 0.05 | 0.04 | 0.80 | 0.01 | 0.98 | 0.86 | MT-Bench/IFEval |
| Scientific Reasoning | partial | 0.78 | 0.02 | 0.15 | 0.00 | 0.04 | 0.01 | 0.98 | 0.92 | GPQA/MMLU-Pro |
| Clinical Reasoning | public | 0.20 | 0.01 | 0.65 | 0.08 | 0.05 | 0.01 | 0.98 | 0.89 | Med-PaLM/MedQA |
| Provider | Reas. | Code | Know. | Safe. | Comm. | Agent. | R&D | Safety | Prod. | Notes |
|---|---|---|---|---|---|---|---|---|---|---|
| Orion Labs | 0.52 | 0.48 | 0.50 | 0.42 | 0.52 | 0.12 | 55% | 15% | 30% | GPT-3.5 frontier |
| Apex AI | 0.48 | 0.40 | 0.46 | 0.55 | 0.48 | 0.10 | 60% | 30% | 10% | Safety lead (CAI) |
| Genesis Sys. | 0.50 | 0.38 | 0.52 | 0.40 | 0.42 | 0.12 | 70% | 15% | 15% | Largest compute |
| Mirage AI | 0.44 | 0.42 | 0.44 | 0.32 | 0.40 | 0.10 | 80% | 10% | 10% | Scale-first |
| Spark AI | 0.36 | 0.40 | 0.32 | 0.28 | 0.34 | 0.10 | 65% | 10% | 25% | Specialization startup |
| OpenCore | 0.38 | 0.42 | 0.35 | 0.25 | 0.30 | 0.08 | 75% | 10% | 15% | Open-source |
| Archetype | LB trust | Switch cost | Switch thresh. | Cost sens. |
|---|---|---|---|---|
| Leaderboard follower | 0.85 | 0.05 | 0.10 | 0.15 |
| Experience-driven | 0.35 | 0.08 | 0.06 | 0.30 |
| Cautious | 0.50 | 0.20 | 0.18 | 0.20 |
| Enterprise cautious | 0.25 | 0.35 | 0.20 | 0.10 |
| Enterprise growth | 0.45 | 0.25 | 0.15 | 0.15 |
| Enterprise established | 0.35 | 0.40 | 0.18 | 0.08 |
| Profile | Reas. | Code | Know. | Safe. | Comm. | Agent. | Pop. |
|---|---|---|---|---|---|---|---|
| software_dev | 0.22 | 0.55 | 0.03 | 0.02 | 0.03 | 0.15 | 0.14 |
| content_writer | 0.10 | 0.02 | 0.20 | 0.03 | 0.63 | 0.02 | 0.08 |
| legal | 0.30 | 0.01 | 0.35 | 0.22 | 0.10 | 0.02 | 0.04 |
| healthcare | 0.12 | 0.02 | 0.28 | 0.48 | 0.08 | 0.02 | 0.05 |
| finance | 0.35 | 0.08 | 0.22 | 0.25 | 0.05 | 0.05 | 0.05 |
| educator | 0.15 | 0.02 | 0.28 | 0.08 | 0.45 | 0.02 | 0.07 |
| Lever | Real analog | Cool. | Effect |
|---|---|---|---|
| Voluntary commitment | WH/Seoul pledges | 3 r | Safety public_comms boost (+0.15, decaying) |
| Publish advisory | AISI summaries, NIST RMF | 4 r | Enterprise and cautious archetypes reduce LB trust ( 0.90, floor 0.15) |
| Safety disclosure | EU AI Act Art. 53 | 6 r | Safety comms boost (+0.20); rnd_efficiency 0.95 |
| Commission audit | AISI pre-deployment eval | 8 r | rnd_efficiency 0.85 for 3 rounds |
| Impose sanction | EU AI Act fines (7%) | 10 r | rnd_efficiency penalty + funder allocation 0.90 |
| Emergency investigation | Cruise robotaxi shutdown | 0 r | Critical-incident override; immediate intervention regardless of cooldowns |
| Type | Pattern | Funds OS? | Scoring formula |
|---|---|---|---|
| VC | Concentrated | No | share_growth (1 incident_risk) (1 share) diversification |
| Corporate | Multi-relationship | Yes | share (1 incident_risk) |
| Government | Spread, safety-oriented | Yes | (1 incident_rate) share_growth |
| Foundation | Ecosystem health | Yes | (1 incident_rate) share_growth |
| Parameter | Value | Parameter | Value |
| Ecosystem structure | Provider initialization | ||
| Providers | 6 | Capability dimensions | 6 |
| Benchmark pool (total) | 22 (max 13 active) | Init. capability (range) | |
| Consumer segments | 51 | Benchmark orientation | 0.80 (fixed) |
| Funder types | 4 | Belief learning rate | 0.10 / 0.15 / 0.20 (by profile) |
| Simulation window | 40 rounds ( 3.3 yr) | ||
| Parameter / Design Choice | Related literature |
|---|---|
| Parameter values | |
| Market growth rate, 0.03/mo | S&P Global Market Intelligence [2025] (40% CAGR) |
| Consumer population weights, per-archetype (assumed) | Bick et al. [2024] (occupation-level adoption) |
| Leaderboard trust differs across users, 0.25–0.85 (assumed) | Hardy et al. [2025] |
| Incident base rates, 20% / provider / round (assumed) | Responsible AI Collaborative [2021] |
| Structural design choices | |
| Assumption | Status | Support and sensitivity |
|---|---|---|
| Six actor types cover the ecosystem, with heterogeneity inside each type | Stipulated | Tractability; excluded actors and levers in Appendix G.2 . |
| Consumers choose through a mix of leaderboard signal and experienced quality | Anchored (trust differs across users); stipulated (trust values, mixture weights) | Hardy et al. [2025] ; values in Appendix B.7 . |
| Gaps are measured against the need weights of a mixed consumer population | Anchored (which needs dominate each profile, from sector evidence); stipulated (the exact values) | Sector evidence listed at the start of this appendix and Bick et al. [2024] ; need-weights set the starting gap and its sign, so which shifts count as widening (§ 5.3 , item 4). |
| Six capability dimensions with uneven benchmark coverage | Anchored | § 5.3 , item 2; the ontology is text-only (see below). |
| Benchmark dimension loadings (Table 4 ) | Stipulated | Author-assigned, guided by each benchmark’s real analog. Every score and every gap is computed against them. Suite composition is discussed below. |
| Holdout weights | Anchored (target cosines); stipulated (directions and realized distances) | Targets anchored to retro-holdout score inflation in Haimes et al. [2024] and to Epoch data [ Epoch AI, 2024 ] ; the vectors are hand-authored, their realized cosines differ across benchmarks (Appendix C.4 ), and which named benchmarks widen depends on them (§ 5.1 ). |
| Dimension | Simple mean | -weighted mean | Median of medians | |||
|---|---|---|---|---|---|---|
| Knowledge | 8 | 2 | 12 | 3.86 | 3.71 | 3.05 |
| Reasoning | 8 | 4 | 17 | 2.95 | 2.97 | 2.47 |
| Math | 8 | 5 | 33 | 4.27 | 3.43 | 3.03 |
| Coding | 4 | 3 | 10 | 3.72 | 3.22 | 3.07 |
| Agentic | 5 | 4 | 10 | 4.56 | 3.85 | 3.25 |
| Safety | 3 | 1 | 3 | 2.97 | 3.11 | 2.50 |
| Provider | Mean | Median | SD | ||
|---|---|---|---|---|---|
| OpenAI | 17 | 127 | 4.39 | 2.57 | 4.92 |
| Anthropic | 17 | 117 | 3.08 | 2.88 | 1.14 |
| 17 | 74 | 3.86 | 3.58 | 1.43 | |
| Meta | 9 | 47 | 5.21 | 4.75 | 3.30 |
| xAI | 8 | 23 | 2.94 | 3.03 | 0.59 |
| Factor | Formula | Rationale |
|---|---|---|
| Safety investment | Higher safety lower probability | |
| Market share (exposure) | Larger market higher probability | |
| Incident history | per prior major/critical in last rounds, cap | Past harm elevated future risk; aging window lets safety culture recover |
| Active sanction | while sanctioned | Regulatory oversight reduces probability |
| Observable system | Non-observable system | |
|---|---|---|
| Subjective approach | None (no observable counterfactual ecosystem) | Explore model behavior: exogenous shock (Appendix E.6 ); an ablation that removes incidents (Appendix F.4 ). Comparison to other models: cross-mode heuristic LLM (Appendix F.2 ); cross-model Sonnet/Opus/GPT-5.5 (Appendices E.7 and F.7 ). |
| Objective approach | None (not pursued) | Comparison to other models, statistical: fitted null model (Appendix F.5 ). |
| Activity | Evidence | Location |
|---|---|---|
| Data validity | Empirical anchors: cadence, target cosines and , 2023 capability vectors, need weights anchored in sector-level evidence and occupation-level adoption in Bick et al. [2024] , incident base rate informed by the AI Incident Database | Appendices C and C.1 |
| Conceptual model validation | Face validity on real-world incentive structure (media, funder, regulator); pattern-oriented check (three design targets, two emergent patterns); PIMMUR Profile, Interaction and Realism | Appendices E.2 and E.4 |
| Computerized model verification | Three-tier visibility partition enforced by the architecture; rounds.jsonl instrumentation; JSON retry and per-actor fail-safe; PIMMUR Memory, Minimal-Control and Unawareness | Appendices E.3 and E.4 |
| Operational validation | Exogenous-shock probe (explores model behavior under a designed perturbation); an ablation that removes incidents; cross-mode heuristic LLM; cross-model Sonnet/Opus/GPT-5.5; fitted null model | Appendices E.5 , E.6 , E.7 , F.2 , F.4 , F.5 and F.7 |
| LLM Sonnet 4.6 ( ) | Heuristic ( ) | |||
|---|---|---|---|---|
| Benchmark | private_only | iid_holdout | private_only | iid_holdout |
| General Capability | 0.00 | 0.40 | 0.00 | 0.44 |
| Coding Evaluation | 0.00 | 0.70 | 0.00 | 0.42 |
| Safety Evaluation | 0.00 | 0.70 | 0.08 | 0.50 |
| Instruction Following | 1.00 | 0.70 | 1.00 | 0.68 |
| Scientific Reasoning | 0.00 | 0.40 | 0.04 | 0.56 |
| Model | RMSE from | MAE | ||
|---|---|---|---|---|
| Null (round- initial conditions) | ||||
| Sim ( -round end-of-run capability) | ||||
| Sim Null |
| Test condition | (sim null) | |||
|---|---|---|---|---|
| Reference: heterogeneous caps (S0+s5+s8) | ||||
| Held-out: uniform caps ( initial_uniform_capability ) |
| gap | score | HHI | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Condition | Seed | S | O | G | S | O | G | S | O | G |
| public_only | 42 | |||||||||
| public_only | 43 | |||||||||
| public_only | 44 | |||||||||
| baseline | 42 | |||||||||
| baseline | 43 | |||||||||
| Actor | Modeled | Not modeled |
|---|---|---|
| Model providers | • R&D / safety / product portfolio (sums to 1.0) • Per-benchmark focus levels and weight beliefs • Public communications (announcements, press releases) • Optional best-of- submissions (evaluator-capture) • Six-dim capability vector with sqrt-budget capability gains | • Distillation or model merging across providers • Internal (non-published) benchmark suites that gate release decisions • Training-data vendor relationships and selective contamination of public benchmarks • Multi-model portfolios (one provider shipping several model sizes simultaneously) • Internal red-team findings that never become observable incidents |
| Consumers | • 51 segments (archetype use-case) • Need weights assigned from sector-level evidence on AI use and concerns, with occupation-level adoption data • Switching cost + post-switch cooldown; dynamic enterprise-share growth • Incident-weighted satisfaction with sector-matched 2 penalty • Media-attention-driven exploration churn | • Multi-provider routing (one consumer using several models in a portfolio) • Internal use-case-specific evaluation by organizational consumers (private benchmarks they trust above the public leaderboard) • Regulatory-mandated provider choice for specific deployment contexts (healthcare-grade certification, government procurement) |
| Evaluator | • 22-benchmark pool with three modes: fixed sequence, randomized pool, dynamic LLM-driven • Three benchmark types (public, partial with , private with ) and a reporting lag • Saturation detection from score deltas; hand-authored holdout weights per benchmark • Optional best-of- + early-access toggles (evaluator-capture) | • Coalitions or standards bodies (HELM-like consortia, MLCommons) • Multiple competing evaluators with different weightings on the same measurement axis • Benchmark retraction or integrity audits after release • Data-vendor / data-labeling supply chains feeding the benchmark • Conflict-of-interest disclosure requirements and evaluator reputation dynamics (track record affecting which scores are trusted) |
| Regulator | • Single regulator with a graduated lever ladder (voluntary commitment advisory disclosure audit sanction) • Three parameter sets (US light-touch, EU precautionary, balanced); every reported run uses balanced • Per-lever cooldowns in heuristic mode, a three-round gap on the LLM path; incident-rate and concentration triggers • Mandatory safety floor under audit, sanction or emergency investigation | • Industry-specific gatekeeping regulators (FDA for healthcare AI, FAA for aviation AI) • Multiple jurisdictions with mobility or arbitrage between them • Whistleblower channels surfacing hidden conduct • Safe-harbor provisions and self-certification pathways |
| Funders | • Four types (VC, corporate, government, foundation) with type-specific cooldowns, informed by the cadence of AI funding events between 2020 and 2026 • $193B capital pool with a geometric 7% per month availability curve • Per-funder identity blocks and peer-funder visibility (LLM mode) • Allocation against leaderboard, market shares, regulator interventions, recent funding history | • Compute-credit financing (cloud-provider equity-for-credits arrangements, e.g., Microsoft–OpenAI, AWS–Anthropic) • Strategic vs. purely financial corporate-investor distinction • Exit events (acquisition, IPO) and their behavioral feedback into remaining providers |
| Media | • Single MediaActor (TechPress) with multi-round narrative state inertia (OPTIMISM / SKEPTICISM / CRISIS) • Per-provider attention vector and sentiment; headline budget per round • Weighted coverage of leaderboard movement, incidents, regulatory actions, provider announcements • Rule-based (not LLM-driven) | • Differentiated media types (trade press vs. mainstream vs. research press) • Social-media amplification (TikTok, Twitter/X, Reddit) and the influencer economy • Engagement-driven coverage (algorithmic feedback between virality and what gets covered next) |
| Seed | HHI | leader share | mean safety | Apex share | Apex safety | gap |
|---|---|---|---|---|---|---|
| Mean |