cs.ROOct 5, 2026

ArtifactArena: Evaluating Models by What They Build in the Physical World

Authors: Kushagra Tiwary*, David Mayo*, Nikhil Behari, Xiangzhou Sun, Abdulrahman Alabdulkareem, Isaac Galatzer-Levy, Boris Katz, Brian Cheung

Organizations: InfoLab, MIT CSAIL · Camera Culture, MIT Media Lab · Department of Psychiatry, NYU Grossman School of Medicine · Discovery Lab, UCSF

Abstract

To evaluate the frontier, we must measure models not by what they say, but by what they can engineer and build in grounded physical environments. We introduce \textsc{ArtifactArena}, an open-ended platform where models face a physically grounded hardware-software co-design challenge: engineering fully functional robots to compete in a simulated arena. We evaluate a frontier model's zero-shot, verifier guided refinement, and open-ended physical design capabilities through three harnesses that refine their bot artifacts based on text descriptions, physics simulator feedback, and gameplay data. We benchmark these capabilities with an Elo ranking of frontier models derived from head-to-head tournaments between their artifacts. By releasing this framework and tournament infrastructure for ongoing community submissions, we establish a living, non-saturating testbed to continuously measure the expanding limits of open-ended intelligence in the physical world. Please visit https://artifactarena.ai for more information.

Figures & tables

Explore similar work

CardsList
  1. WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform

    May 18, 2026Yu Shang, Yinzhou Tang, Yiding Ma +22Embodied Artificial IntelligenceWorld Models

  2. Inspect Robots: Evaluating the Capabilities and Safety of Embodied AI

    Oct 5, 2026Christopher Leet, Achu Menon, Sravanthi Machcha +12Agentic Evaluations

  3. stable-worldmodel: A Platform for Reproducible World Modeling Research and Evaluation

    May 20, 2026Lucas Maes, Quentin Le Lidec, Luiz Facury +9World ModelsLarge-Scale Robot Demonstration Datasets