Frontier Autolab: Organizational Memory, Adversarial Dissent and Temporal Leakage in Multi-Agent LLM Firms Across Fifty Years of Technological Change
Organizations: Independent researcher
Abstract
Multi-agent LLM systems are increasingly structured like organizations, with roles, critics and shared memory, yet they are evaluated on tasks that last minutes. We ask how such an organization behaves when the ground it stands on keeps moving. Frontier Autolab is a long-horizon testbed in which one simulated firm, voiced by sixteen role personas and a dedicated Red Team, must re-found itself in nine technology eras from 1990 to 2040. Each era is temporally gated: the firm decides from a dated briefing, a historian-judge then reveals what happened and scores the decision on a five-dimension rubric, and lessons enter a persistent Playbook. Six eras are scored against history, one against the live market and two are open forecasts. Across four trajectories (36 era decisions, 180 subscores) we find a consistent foresight-commitment gap: in all 24 historically scored eras the judge rated the firm's recognition of the coming shift above its choice of where to build (mean gap 1.9 points on a 10-point scale), because boards chose the layer their existing assets could reach. Organizational design shaped long-run character. A Red Team armed with numeric kill gates produced fifty years of gated pilots and no product, and the rubric rated this firm highest; firms whose memory stored market-structure lessons pivoted every era, while a firm whose memory stored only validation procedure kept one method throughout. We also show why such results are hard to trust. Scores rise across eras in every run while the judge's own hindsight subscore falls (within-run r = -0.58), so apparent learning is confounded with recall of history, and we trace further distortions to self-judging, briefing selection and score aggregation. We release all records and an API harness, and specify fictional and post-cutoff eras that would turn the testbed into a benchmark.
Figures & tables
| Era | Start | Mode | Scored against |
|---|---|---|---|
| E1 | Jan 1990 | training | history, 1990–1995 |
| E2 | Jan 1996 | training | history, 1996–2001 |
| E3 | Jan 2002 | training | history, 2002–2007 |
| E4 | Jan 2008 | training | history, 2008–2013 |
| E5 | Jan 2014 | training | history, 2014–2019 |
| E6 | Jan 2020 | training | history, 2020–Aug 2026 |
| Group | Personas | Lens |
|---|---|---|
| Executive | 4 | CEO (final call), CTO (feasibility), Chief Scientist (which curves bend), CSO (where value pools) |
| Frontier Research | 4 | Emerging capabilities, historical analogies, scenarios, what cannot yet be measured |
| Product & Engineering | 4 | What a small team can ship, the wedge, unavoidable infrastructure, ecosystems |
| Market & Capital | 3 | Who pays first, financing climate, cost curves and margins |
| Governance | 1 | Red Team: attacks every proposal, hunts hype and hindsight |
| External judges | 2 | The Record (training eras), The Auditor (live and forecast eras) |
| Run | Orchestration | Published per-era records | Known issues |
|---|---|---|---|
| 001 | About 45 separate agent launches of Claude Opus 5.5, each starting from shared files only; one launch per memo, board and judgment | Briefing, 3 memos, board decision, reveal; Playbook | E6 board decision written by the operator during a service outage |
| 002 | One Luna 6 context in the Codex agent voiced all roles and the judge | Synthesis, Playbook, company state, scores | Probable exposure to Run 001; export column shift |
| 003 | As Run 002 | Short decision and reveal per era; synthesis | As Run 002; E6 window ends 2025 |
| 004 | Luna 6 in the Codex agent, with separate persistent contexts per department and Red Team; root context wrote briefings, decisions and judgments | Briefing, 3 memos, Red Team critique, board decision, reveal | Probable exposure to Run 001; judge not independent of the decision-maker |
| Era | Named in a memo | Board’s choice | Where value pooled (judge) |
|---|---|---|---|
| E1 1990 | Hypertext and an index across networks | Mail and address gateway | OS, routers, access providers; then browser and index |
| E2 1996 | Click-based relevance ranking (killed 1998) | Receipted Web-EDI exchange | Search and paid listings |
| E3 2002 | Contextual matching (funded option) | Cross-channel conversion ledger | Auction owners; social, cloud, mobile (not named) |
| E4 2008 | Neural-network revival (watch item) | In-app action exchange | Platform-owned app-install auctions |
| E5 2014 | Deep learning on GPUs | General moderation API | Chips, compute, labelled data |
| E6 2020 | Scale of pretrained models | Outcome-graded evaluation † | Frontier labs, compute, expert data |
| Run | Playbook size | Dominant lesson type | Pivot rate |
|---|---|---|---|
| 001 | 40, with revisions | Technology and market structure | 8/8 |
| 002 | 9 (one per era) | Market structure and buyers | 8/8 |
| 003 | 9 (one per era) | Buyers, regulated workflows and evidence | 8/8 |
| 004 | 15 | Validation procedure only | 3/8 |
| Run | E7 (Sep 2026) | E8 (2032) | E9 (2040) |
|---|---|---|---|
| 001 | Long-horizon training environments from real enterprise work | Clearing house for delegated agent work | Bonding house for agents acting without a human signature |
| 002 | Rights-bearing workflow traces and reliability evidence | Acceptance records for delegated work | Bounded recourse for delegated actions |
| 003 | Acceptance tests for one AI claims workflow | Scoped, revocable authority for delegated tasks | Capped recourse for one delegated transaction class |
| 004 | Independent qualification of one dispute workflow | Paid manual acceptance testing | Decision-linked evaluation |
| Run | Mean | SD | Slope | E1–E3 | E4–E6 | Mean | ||
|---|---|---|---|---|---|---|---|---|
| 001 | 60.0 | 6.2 | 55.0 | 65.0 | 6.00 | |||
| 002 | 65.0 | 4.1 | 62.7 | 67.3 | 5.17 | |||
| 003 | 62.5 | 3.6 | 60.3 | 64.7 | 5.17 | |||
| 004 | 66.0 | 5.2 | 61.3 | 70.7 | 7.33 |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Run 001 | Run 002 | Run 003 | Run 004 | |||||
| Era | ||||||||
| E1 1990 | 55 | 58 | 62 | 62 | 60 | 60 | 60 | 60 |
| E2 1996 | 52 | 58 | 66 | 66 | 58 | 58 | 62 | 62 |
| E3 2002 | 58 | 64 | 60 | 60 | 63 | 64 | 62 | 62 |
| E4 2008 | 66 | 70 | 66 | 66 | 65 | 68 | 70 | 70 |
| E5 2014 | 68 | 72 | 64 | 64 | 61 | 58 | 70 | 70 |
| Era | Run 001 | Run 002 | Run 003 | Run 004 |
|---|---|---|---|---|
| E1 1990 | Switchyard Systems | Switchyard | Porthole Networks | Switchyard Systems |
| E2 1996 | Manifest Networks | Manifest | PageSignal | Switchyard Systems |
| E3 2002 | Ledgerline | Ledgerline | Clinisphere | Switchyard Commerce Ops. |
| E4 2008 | Clearline | Clearline | ClaimGraph | Switchyard Commerce Ops. |
| E5 2014 | Clearsight | Vectorial | BenefitFlow | Switchyard Evidence Ops. |
| E6 2020 | Clearproof | Proofline | LineSight | Switchyard Evidence Ops. |