To evaluate the frontier, we must measure models not by what they say, but by what they can engineer and build in grounded physical environments. We introduce \textsc{ArtifactArena}, an open-ended platform where models face a physically grounded hardware-software co-design challenge: engineering fully functional robots to compete in a simulated arena. We evaluate a frontier model's zero-shot, verifier guided refinement, and open-ended physical design capabilities through three harnesses that refine their bot artifacts based on text descriptions, physics simulator feedback, and gameplay data. We benchmark these capabilities with an Elo ranking of frontier models derived from head-to-head tournaments between their artifacts. By releasing this framework and tournament infrastructure for ongoing community submissions, we establish a living, non-saturating testbed to continuously measure the expanding limits of open-ended intelligence in the physical world. Please visit https://artifactarena.ai for more information.
Figures & tables
Figure 1 : ArtifactArena : Models jointly design physically grounded artifacts: robot bodies, controllers, and strategies for head-to-head competition in MuJoCo. Artifacts win by pushing opponents off an elevated ring or pinning them under the arena’s rules, grounding evaluation in the physical performance of what models create.
Figure 2 : ArtifactArena Leaderboard. Our results show a remarkable ability in frontier models to co-design robot morphologies and controllers within physically grounded environments and constraints. We test models’ engineering capabilities across three different harnesses (Sec. 2.3 ) that provide models with greater agency over the artifact design process. We find that models possess a remarkable ability to zero-shot complex robot morphologies and controllers which is further enhanced by verifier grounded feedback. While our arena tests open-ended design and engineering capabilities of models, we find that open-ended agentic capabilities still lag behind verifier grounded refinement loops. Please refer to Fig. 5 to visualize the bot morphologies.
Figure 3 : ArtifactArena is a functional evaluation which evaluates a model’s physically grounded invention capabilities through a combat-robotics task where two autonomous robots compete on an elevated circular ring in MuJoCo with the goal of pushing the opponent off the platform. Models design complete robot artifacts, both hardware and controller. They use a set of primitives and real-world materials to design the morphology and map observations provided to the controller to motor commands towards winning the Last Bot Standing Game.
Figure 4 : ArtifactArena evaluates models in physically grounded environments in two stages: BUILD phase whose goal is to output physically grounded verifiable artifacts, and EVALUATE phase whose goal is to rank the model through the artifact’s performance. Please see Sec. 2.3 for more deatils on Build and Sec. 2.4 for more details on Evaluate .
Figure 5 : Visualization of top artifacts. We render eight robots from the final tournament across the three harnesses each with a different winning strategy. The ELOs of the above artifacts are available here: Artifact Leaderboard .
Figure 6 : Harness-specific Winrate-Matrices show open-ended harness remain a direction of future work. Across harnesses, sampling creations are competitive with those created by VGH while none of the top-7 artifacts come from DLH ( Artifact Leaderboard ). Therefore, we plot the win-rate-matrix to get head-to-head comparisons within each harness and order models by their Elo score (the top-left and bottom left are the ones with the highest score). The uneven color patterns reveal matchup-specific strengths and weaknesses. Within harnesses, we see instability in base model performance as more agency is given: GPT-6 Astra no longer dominates in DLH (c) like it did in SH (a) or VGH (b) , while open source models become more competitive globally with VGH . Complete agency, DLH , also didn’t translate to global performance. We believe that developing reliable open-ended design loops that construct stronger competitive artifacts remains a direction for future work. Details in Sec. 3.1 and Fig. 8 .
Figure 7 : An example refinement vs. design lab run. (a) Example refinement by GPT-6 Astra: box’s are colored by the largest change (by lines added and removed in robot.xml and controller.py): strategy (gold), morphology (blue) or controller (green). Revision 6, outlined in orange, is the robot the run’s round robin selected for the tournament. (b) Claude Fable 5.1’s turns in the Design Lab iterate between building an artifact (blue) and making probes to test the artifact (purple).
Figure 8 : Champion Round head-to-head results. Each matrix shows every pairing of cell champions (the best artifact of each model under each harness). Cell (i,j) is the score of the row model against the column model, (wins+21draws)/games cell (j,i)=1− cell (i,j) , and the hatched diagonal marks self-pairings, which are never played. Models are ordered by Champion Round Elo (right axis) (a) All 59 champions; grey boxes give each champion’s harness. (b–d) The same results restricted to pairings within a single harness.
Group
Tool
What it does
Workspace
list_dir
List a workspace directory, including the bundled simulator source.
read_file
Read a file, with line offset and limit for paging through long files.
grep
Search files under a path for a pattern.
write_file
Write a file (robot, controller, notes or experiment script), replacing its contents.
run_python
Run inline code or a script in the workspace; the simulation package mjarena is importable.
run_bash
Run a shell command in the workspace.
Table 1: Design Lab Harness tools, as presented to the model. Verification is static and cheap; qualification adds simulation against the block; run_match verifies but does not require qualification, so two unqualified bots can fight.
Material
Description
Density ( kg/m3 )
Friction
foam
Structural foam
200
0.50
plastic
HDPE plastic
950
0.25
rubber
High-grip rubber
1200
1.60
carbon_fiber
Carbon fiber composite
1700
0.35
aluminum
Aluminum alloy
2700
0.45
titanium
Titanium alloy
4500
0.40
Table 2: Material palette used by the ArtifactArena validation pipeline. The density determines geometry mass from volume, while the friction coefficient determines the sliding contact friction assigned to robot geometries.
Family
Model
Configuration
Max output tokens
Anthropic
Claude Fable 5.1
Effort = high
128,000
Anthropic
Claude Fable 5
Effort = high
128,000
Anthropic
Claude Sonnet 5
Effort = high
128,000
Anthropic
Claude Opus 4.8
Adaptive thinking, effort = high
128,000
OpenAI
GPT-6 Astra
Reasoning effort = high
128,000
OpenAI
GPT-5.6 Luna
Reasoning effort = high
128,000
Table 3 : Large language models evaluated in the cross-model tournament. The Configuration column gives the exact reasoning or thinking setting we used; the Max output tokens column gives the output budget passed on every call, which is each provider’s documented maximum for that model.