cs.SESep 29, 2026

Zero2Repo: Can Coding Agents Build Repositories from Scratch?

Authors: Pei Yang, Tianyu Shi, Yuhang Yao, Wanyi Chen, Tongyun Yang, Dun Pei, Haonan Wang, Pengbin Feng, +18 more

Organizations: Gradient Data · McGill University · Carnegie Mellon University · Independent Researcher · University of Southern California · Queen’s University · MIT · University College London, University of London · University of California, San Diego

Abstract

Coding agents are increasingly asked to build software rather than patch it, yet benchmarks for from-scratch repository construction are mostly limited to a single language and depend on manually curated tasks. We introduce Zero2Repo, a benchmark in which an agent receives a product requirements document, an interface contract, and an empty workspace, and must deliver a complete repository in the project's native ecosystem. Tasks are produced by a language-agnostic authoring pipeline that converts real, version-pinned open-source projects into behavioral specifications, reproducible environments, and hidden acceptance tests. Each task is validated by execution: a reference implementation derived from the upstream project must pass, and adversarial validation must show that the tests reject incorrect implementations. Evaluation runs production coding agents in isolated containers, withholds the acceptance tests until an explicit submission, and assigns a binary reward only when every test passes, with no LLM judge. The pipeline and harness make no language-specific assumptions and apply to mainstream programming ecosystems; the current release contains Python, TypeScript, Go, and C++ tasks. Even on 11 tasks drawn from repositories that frontier models have very likely seen during training, the strongest agent solves only 10, and every failing submission passes 90-99% of the hidden tests; for the two strongest agents, 67-100% of failed tests trace to a single omission or a low-frequency rule stated in the specification rather than to a missing subsystem, so each failure is a concrete target for improvement.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. E2E-SWE: Benchmarking LLMs on Building Working Codebases from Scratch

    Sep 29, 2026Hantian Ding, Chloe Bi, Jiacheng Zhu +5CodebasesModel Auditing

  2. Change2Task: From Repository Changes to Executable Coding Agent Tasks and Environments

    Jul 30, 2026Haomin Qi, Xingliang Wang, Xuanqi Gao +9Coding AgentsEnvironment

  3. CodeTeam: An LLM-Powered Multi-Agent Framework for Repository-Level Code Generation

    Jun 20, 2026Yifei Wang, Ruiyin Li, Peng Liang +4Repository-Level Code UnderstandingSketches