Multi-agent LLM systems are increasingly structured like organizations, with roles, critics and shared memory, yet they are evaluated on tasks that last minutes. We ask how such an organization behaves when the ground it stands on keeps moving. Frontier Autolab is a long-horizon testbed in which one simulated firm, voiced by sixteen role personas and a dedicated Red Team, must re-found itself in nine technology eras from 1990 to 2040. Each era is temporally gated: the firm decides from a dated briefing, a historian-judge then reveals what happened and scores the decision on a five-dimension rubric, and lessons enter a persistent Playbook. Six eras are scored against history, one against the live market and two are open forecasts. Across four trajectories (36 era decisions, 180 subscores) we find a consistent foresight-commitment gap: in all 24 historically scored eras the judge rated the firm's recognition of the coming shift above its choice of where to build (mean gap 1.9 points on a 10-point scale), because boards chose the layer their existing assets could reach. Organizational design shaped long-run character. A Red Team armed with numeric kill gates produced fifty years of gated pilots and no product, and the rubric rated this firm highest; firms whose memory stored market-structure lessons pivoted every era, while a firm whose memory stored only validation procedure kept one method throughout. We also show why such results are hard to trust. Scores rise across eras in every run while the judge's own hindsight subscore falls (within-run r = -0.58), so apparent learning is confounded with recall of history, and we trace further distortions to self-judging, briefing selection and score aggregation. We release all records and an API harness, and specify fictional and post-cutoff eras that would turn the testbed into a benchmark.
Figures & tables
Era
Start tk
Mode
Scored against
E1
Jan 1990
training
history, 1990–1995
E2
Jan 1996
training
history, 1996–2001
E3
Jan 2002
training
history, 2002–2007
E4
Jan 2008
training
history, 2008–2013
E5
Jan 2014
training
history, 2014–2019
E6
Jan 2020
training
history, 2020–Aug 2026
Table 1: Eras. The E6 outcome window extends to the live era and is longer than the others.
Group
Personas
Lens
Executive
4
CEO (final call), CTO (feasibility), Chief Scientist (which curves bend), CSO (where value pools)
Frontier Research
4
Emerging capabilities, historical analogies, scenarios, what cannot yet be measured
Product & Engineering
4
What a small team can ship, the wedge, unavoidable infrastructure, ecosystems
Market & Capital
3
Who pays first, financing climate, cost curves and margins
Governance
1
Red Team: attacks every proposal, hunts hype and hindsight
External judges
2
The Record (training eras), The Auditor (live and forecast eras)
Table 2: Role charter.
Run
Orchestration
Published per-era records
Known issues
001
About 45 separate agent launches of Claude Opus 5.5, each starting from shared files only; one launch per memo, board and judgment