Zero2Repo: Can Coding Agents Build Repositories from Scratch?
Organizations: Gradient Data · McGill University · Carnegie Mellon University · Independent Researcher · University of Southern California · Queen’s University · MIT · University College London, University of London · University of California, San Diego
Abstract
Coding agents are increasingly asked to build software rather than patch it, yet benchmarks for from-scratch repository construction are mostly limited to a single language and depend on manually curated tasks. We introduce Zero2Repo, a benchmark in which an agent receives a product requirements document, an interface contract, and an empty workspace, and must deliver a complete repository in the project's native ecosystem. Tasks are produced by a language-agnostic authoring pipeline that converts real, version-pinned open-source projects into behavioral specifications, reproducible environments, and hidden acceptance tests. Each task is validated by execution: a reference implementation derived from the upstream project must pass, and adversarial validation must show that the tests reject incorrect implementations. Evaluation runs production coding agents in isolated containers, withholds the acceptance tests until an explicit submission, and assigns a binary reward only when every test passes, with no LLM judge. The pipeline and harness make no language-specific assumptions and apply to mainstream programming ecosystems; the current release contains Python, TypeScript, Go, and C++ tasks. Even on 11 tasks drawn from repositories that frontier models have very likely seen during training, the strongest agent solves only 10, and every failing submission passes 90-99% of the hidden tests; for the two strongest agents, 67-100% of failed tests trace to a single omission or a low-frequency rule stated in the specification rather than to a missing subsystem, so each failure is a concrete target for improvement.
Figures & tables
| Benchmark | Initial artifact | Required deliverable | Languages | Task construction |
| Long-horizon agent tasks | ||||
| SWE-Together | Repository and first user turn | Repository changes | Mixed | 109 tasks from recorded sessions |
| TerminalWorld | Container and instruction | Completed terminal task | Mixed | Derived from command recordings |
| FrontierSWE | Prepared environment | Open-ended solution | Mixed | 17 expert-curated tasks |
| From-scratch repository construction | ||||
| DevBench | Stage-dependent repository context | Per-stage artifacts | Python, Java, JS, C/C++ | 22 curated repositories |
| Model | Harness | Solved | Pass rate | Submitted | Time (min) | Cost ($) |
|---|---|---|---|---|---|---|
| GPT-6 Astra | Codex | 10/11 | 90.9% | 11/11 | 14.4 | 47.04 |
| Claude Opus 5.5 | Claude Code | 9/11 | 81.8% | 11/11 | 19.7 | 49.37 |
| Grok 4.7 High | Cursor CLI | 7/11 | 63.6% | 11/11 | 51.5 | 45.75 |
| Failure mode | Grok 4.7 High | Claude Opus 5.5 | GPT-6 Astra |
|---|---|---|---|
| Long-tail edge-case rule | 27 (31%) | 12 (67%) | 8 (38%) |
| Less common subsystem non-functional | 45 (52%) | 2 (11%) | 0 |
| Single omission breaking a test group | 5 (6%) | 0 | 13 (62%) |
| Environment-dependent behavior | 10 (11%) | 2 (11%) | 0 |
| Crash | 0 | 2 (11%) | 0 |
| Total | 87 | 18 | 21 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | Product | Source | Language | Domain | Acceptance modules |
|---|---|---|---|---|---|
| 001 | Tomlparse | tomli | Python | TOML v1.1 parser | 5 |
| 002 | python-envfile | python-dotenv | Python | .env loading and expansion | 8 |
| 003 | ymlcodec | js-yaml | TypeScript | YAML codec with schemas and tags | 7 |
| 004 | Signtoken | itsdangerous | Python | HMAC signing and serialization | 4 |
| 005 | httpwire | h11 | Python | HTTP/1.1 protocol state machine | 7 |
| 006 | Optlyn | Click | Python | CLI framework with shell completion | 14 |
| Task | 001 | 002 | 003 | 004 | 005 | 006 | 007 | 008 | 009 | 010 | 011 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Rounds | 3 | 1 | 3 | 3 | 2 | 4 | 2 | 2 | 3 | 3 | 2 |
| GPT-6 Astra | Claude Opus 5.5 | Grok 4.7 High | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Task | Result | Time | Cost | Result | Time | Cost | Result | Time | Cost |
| 001 | Pass | 391 | 1.32 | Pass | 398 | 1.68 | Pass | 2093 | 2.95 |
| 002 | Pass | 399 | 4.36 | Pass | 399 | 1.58 | Pass | 2640 | 3.20 |
| 003 | Pass | 786 | 5.42 | Pass | 1581 | 5.94 | 499/507 | 4234 | 6.10 |
| 004 | Pass | 217 | 1.65 | Pass | 290 | 1.43 | Pass | 2095 | 2.40 |
| 005 | Pass | 530 | 3.90 | Pass | 595 | 2.15 | Pass | 3285 | 4.55 |
| Model | Python | TypeScript (003) | C++ (008) | Go (009) |
|---|---|---|---|---|
| GPT-6 Astra | 8/8 | Pass | Pass | Fail |
| Claude Opus 5.5 | 7/8 | Pass | Pass | Fail |
| Grok 4.7 High | 7/8 | Fail | Fail | Fail |
| Task | Grok 4.7 High | Claude Opus 5.5 | GPT-6 Astra |
|---|---|---|---|
| 003 YAML codec | 499/507 (98.4%) | Pass | Pass |
| 006 CLI framework | 467/476 (98.1%) | 473/476 (99.4%) | Pass |
| 008 URL parser | 223/248 (89.9%) | Pass | Pass |
| 009 Large-file client | 453/498 (91.0%) | 483/498 (97.0%) | 477/498 (95.8%) |