To evaluate the frontier, we must measure models not by what they say, but by what they can engineer and build in grounded physical environments. We introduce \textsc{ArtifactArena}, an open-ended platform where models face a physically grounded hardware-software co-design challenge: engineering fully functional robots to compete in a simulated arena. We evaluate a frontier model's zero-shot, verifier guided refinement, and open-ended physical design capabilities through three harnesses that refine their bot artifacts based on text descriptions, physics simulator feedback, and gameplay data. We benchmark these capabilities with an Elo ranking of frontier models derived from head-to-head tournaments between their artifacts. By releasing this framework and tournament infrastructure for ongoing community submissions, we establish a living, non-saturating testbed to continuously measure the expanding limits of open-ended intelligence in the physical world. Please visit https://artifactarena.ai for more information.
Figures & tables
Figure 1 : ArtifactArena : Models jointly design physically grounded artifacts: robot bodies, controllers, and strategies for head-to-head competition in MuJoCo. Artifacts win by pushing opponents off an elevated ring or pinning them under the arena’s rules, grounding evaluation in the physical performance of what models create.
Figure 2 : ArtifactArena Leaderboard. Our results show a remarkable ability in frontier models to co-design robot morphologies and controllers within physically grounded environments and constraints. We test models’ engineering capabilities across three different harnesses (Sec. 2.3 ) that provide models with greater agency over the artifact design process. We find that models possess a remarkable ability to zero-shot complex robot morphologies and controllers which is further enhanced by verifier grounded feedback. While our arena tests open-ended design and engineering capabilities of models, we find that open-ended agentic capabilities still lag behind verifier grounded refinement loops. Please refer to Fig. 5 to visualize the bot morphologies.
Figure 3 : ArtifactArena is a functional evaluation which evaluates a model’s physically grounded invention capabilities through a combat-robotics task where two autonomous robots compete on an elevated circular ring in MuJoCo with the goal of pushing the opponent off the platform. Models design complete robot artifacts, both hardware and controller. They use a set of primitives and real-world materials to design the morphology and map observations provided to the controller to motor commands towards winning the Last Bot Standing Game.
Figure 4 : ArtifactArena evaluates models in physically grounded environments in two stages: BUILD phase whose goal is to output physically grounded verifiable artifacts, and EVALUATE phase whose goal is to rank the model through the artifact’s performance. Please see Sec. 2.3 for more deatils on Build and Sec. 2.4 for more details on Evaluate .
Figure 5 : Visualization of top artifacts. We render eight robots from the final tournament across the three harnesses each with a different winning strategy. The ELOs of the above artifacts are available here: Artifact Leaderboard .
Figure 6 : Harness-specific Winrate-Matrices show open-ended harness remain a direction of future work. Across harnesses, sampling creations are competitive with those created by VGH while none of the top-7 artifacts come from DLH ( Artifact Leaderboard ). Therefore, we plot the win-rate-matrix to get head-to-head comparisons within each harness and order models by their Elo score (the top-left and bottom left are the ones with the highest score). The uneven color patterns reveal matchup-specific strengths and weaknesses. Within harnesses, we see instability in base model performance as more agency is given: GPT-6 Astra no longer dominates in DLH (c) like it did in SH (a) or VGH (b) , while open source models become more competitive globally with VGH . Complete agency, DLH , also didn’t translate to global performance. We believe that developing reliable open-ended design loops that construct stronger competitive artifacts remains a direction for future work. Details in Sec. 3.1 and Fig. 8 .
Figure 7 : An example refinement vs. design lab run. (a) Example refinement by GPT-6 Astra: box’s are colored by the largest change (by lines added and removed in robot.xml and controller.py): strategy (gold), morphology (blue) or controller (green). Revision 6, outlined in orange, is the robot the run’s round robin selected for the tournament. (b) Claude Fable 5.1’s turns in the Design Lab iterate between building an artifact (blue) and making probes to test the artifact (purple).
Figure 8 : Champion Round head-to-head results. Each matrix shows every pairing of cell champions (the best artifact of each model under each harness). Cell (i,j) is the score of the row model against the column model, (wins+21draws)/games cell (j,i)=1− cell (i,j) , and the hatched diagonal marks self-pairings, which are never played. Models are ordered by Champion Round Elo (right axis) (a) All 59 champions; grey boxes give each champion’s harness. (b–d) The same results restricted to pairings within a single harness.
Group
Tool
What it does
Workspace
list_dir
List a workspace directory, including the bundled simulator source.
read_file
Read a file, with line offset and limit for paging through long files.
grep
Search files under a path for a pattern.
write_file
Write a file (robot, controller, notes or experiment script), replacing its contents.
run_python
Run inline code or a script in the workspace; the simulation package mjarena is importable.
run_bash
Run a shell command in the workspace.
Table 1: Design Lab Harness tools, as presented to the model. Verification is static and cheap; qualification adds simulation against the block; run_match verifies but does not require qualification, so two unqualified bots can fight.
Material
Description
Density ( kg/m3 )
Friction
foam
Structural foam
200
0.50
plastic
HDPE plastic
950
0.25
rubber
High-grip rubber
1200
1.60
carbon_fiber
Carbon fiber composite
1700
0.35
aluminum
Aluminum alloy
2700
0.45
titanium
Titanium alloy
4500
0.40
Table 2: Material palette used by the ArtifactArena validation pipeline. The density determines geometry mass from volume, while the friction coefficient determines the sliding contact friction assigned to robot geometries.
Family
Model
Configuration
Max output tokens
Anthropic
Claude Fable 5.1
Effort = high
128,000
Anthropic
Claude Fable 5
Effort = high
128,000
Anthropic
Claude Sonnet 5
Effort = high
128,000
Anthropic
Claude Opus 4.8
Adaptive thinking, effort = high
128,000
OpenAI
GPT-6 Astra
Reasoning effort = high
128,000
OpenAI
GPT-5.6 Luna
Reasoning effort = high
128,000
Table 3 : Large language models evaluated in the cross-model tournament. The Configuration column gives the exact reasoning or thinking setting we used; the Max output tokens column gives the output budget passed on every call, which is each provider’s documented maximum for that model.
World models have emerged as a central paradigm for embodied intelligence, enabling agents to predict action-conditioned future and reason about environmental dynamics. However, existing embodied world model benchmarks are still largely confined to vision-only prediction, offline embodied applications, and simulator-based evaluation, making them insufficient for assessing increasingly comprehensive world models. In this work, we introduce WorldArena 2.0, an expanded benchmark that systematically broadens embodied world model evaluation along three dimensions: modality, functionality, and platform. Along the modality dimension, WorldArena 2.0 extends evaluation from vision-only to visuotactile modalities, enabling assessment of multimodal perception and prediction. Along the functionality dimension, it extends beyond policy evaluation and planning to assess world models as interactive RL environments for policy optimization. Along the platform dimension, it moves beyond simulator-only evaluation to a diverse suite of simulated and real-world robotic settings across multiple embodiments. Under a standardized protocol, WorldArena 2.0 comprehensively evaluates perceptual quality, interactive utility, and cross-platform performance, providing a comprehensive testbed for tracking progress toward embodied world models. The benchmark is available at: https://world-arena.ai.
Yu Shang, Yinzhou Tang, Yiding Ma +22
Tsinghua University · Shanghai Jiao Tong University · Zhejiang University +7
General purpose language models are increasingly able to control robotic hardware. Understanding the capabilities and safety of these models when embodied is therefore increasingly important for understanding their societal impact and risks. To this end, we introduce Inspect Robots, a modular, open-source framework for developing and running evaluations of embodied agents. Inspect Robots pairs customizable, reusable abstractions for specifying physical evaluations and analyzing their results with infrastructure that automates evaluation setup, execution and termination. We demonstrate Inspect Robots by using it to evaluate the capabilities and safety of six policies based on frontier language models. Inspect Robots has seen significant early uptake, receiving nearly 100,000 downloads in the three months since its release.
Christopher Leet, Achu Menon, Sravanthi Machcha +12
Robocurve · University of Southern California · San Jose State University +7
World models are central to building agents that can reason, plan, and generalize beyond their training data. However, research on world models is currently fragmented, with disparate codebases, data pipelines, and evaluation protocols hindering reproducibility and fair comparison. Current practice is further limited by three key bottlenecks: fragile one-off codebases, slow video data loading, and the lack of standardized generalization benchmarks. We present stable-worldmodel (swm), an open-source platform for standardized and reproducible world modeling research and evaluation. It delivers (1) a high-performance Lance-based data layer with native support and conversion tools for MP4, HDF5, and LeRobot datasets, (2) clean, well-tested implementations of modern world model baselines and planning solvers, and (3) a broad suite of environments and tasks extended with controllable visual, geometric, and physical factors of variation for systematic in-silico evaluation of dynamics understanding, control performance, representation quality, and out-of-distribution generalization. By unifying the full pipeline under a single, scalable framework, \texttt{swm} dramatically reduces research overhead and accelerates trustworthy progress toward reliable world models.
Lucas Maes, Quentin Le Lidec, Luiz Facury +9
Mila & Université de Montréal · New York University · Universidade Federal de Minas Gerais +4