Zero2Repo: Can Coding Agents Build Repositories from Scratch?
Authors: Pei Yang, Tianyu Shi, Yuhang Yao, Wanyi Chen, Tongyun Yang, Dun Pei, Haonan Wang, Pengbin Feng, +18 more
Organizations: Gradient Data · McGill University · Carnegie Mellon University · Independent Researcher · University of Southern California · Queen’s University · MIT · University College London, University of London · University of California, San Diego
Coding agents are increasingly asked to build software rather than patch it, yet benchmarks for from-scratch repository construction are mostly limited to a single language and depend on manually curated tasks. We introduce Zero2Repo, a benchmark in which an agent receives a product requirements document, an interface contract, and an empty workspace, and must deliver a complete repository in the project's native ecosystem. Tasks are produced by a language-agnostic authoring pipeline that converts real, version-pinned open-source projects into behavioral specifications, reproducible environments, and hidden acceptance tests. Each task is validated by execution: a reference implementation derived from the upstream project must pass, and adversarial validation must show that the tests reject incorrect implementations. Evaluation runs production coding agents in isolated containers, withholds the acceptance tests until an explicit submission, and assigns a binary reward only when every test passes, with no LLM judge. The pipeline and harness make no language-specific assumptions and apply to mainstream programming ecosystems; the current release contains Python, TypeScript, Go, and C++ tasks. Even on 11 tasks drawn from repositories that frontier models have very likely seen during training, the strongest agent solves only 10, and every failing submission passes 90-99% of the hidden tests; for the two strongest agents, 67-100% of failed tests trace to a single omission or a low-frequency rule stated in the specification rather than to a missing subsystem, so each failure is a concrete target for improvement.
Figures & tables
Figure 1: Data flow for one Zero2Repo task (Tomlparse). The Phys. Rev. Dand Interface Contract from the task assets are given to the agent as task context in the agent image, where it builds the repository in /app and submits by writing the submit file. The workspace is then copied into a fresh, offline judge container, the hidden acceptance tests are injected, and the judge writes a binary reward.
Benchmark
Initial artifact
Required deliverable
Languages
Task construction
Long-horizon agent tasks
SWE-Together
Repository and first user turn
Repository changes
Mixed
109 tasks from recorded sessions
TerminalWorld
Container and instruction
Completed terminal task
Mixed
Derived from command recordings
FrontierSWE
Prepared environment
Open-ended solution
Mixed
17 expert-curated tasks
From-scratch repository construction
DevBench
Stage-dependent repository context
Per-stage artifacts
Python, Java, JS, C/C++
22 curated repositories
Table 1: Benchmarks closest to Zero2Repo.
Figure 2: Zero2Repo evaluation architecture. Asset layer : the environment recipe, Phys. Rev. D, and Interface Contract are built into the agent image; the test code is used only by the judge. Image layer : a base image shared by all tasks, a per-task agent image built from the environment recipe without hidden tests, and the judge environment, which is a fresh container started from the agent image with the hidden tests injected rather than a separately built image. Solve layer : the agent receives the task context (specification paths, build rules, and submission rules), works in an empty project directory, and submits by writing a submit file. Judge layer : the submission is verified and checked against the package denylist, the workspace is copied into the isolated judge container, the hidden tests are run, and a binary score is recorded.
Figure 3: Zero2Repo pipeline. Authoring : a repository is pinned, its environment recipe is built, a Phys. Rev. Dand Interface Contract are reverse-engineered from its behavior, test code is written from the Phys. Rev. Dand the repository, public names are neutralized, and the task is verified, reviewed, and released. Release : the environment recipe, Phys. Rev. D, and Interface Contract are public; the test code is hidden. Evaluation : the base image is pulled and built once, the agent image is built from the task’s environment recipe, the agent writes code in an empty directory, and the submission is checked. Across the isolation boundary, the judge runs the hidden tests in a fresh container started from the agent image and records a binary score with logs.
Model
Harness
Solved
Pass rate
Submitted
Time (min)
Cost ($)
GPT-6 Astra
Codex
10/11
90.9%
11/11
14.4
47.04
Claude Opus 5.5
Claude Code
9/11
81.8%
11/11
19.7
49.37
Grok 4.7 High
Cursor CLI
7/11
63.6%
11/11
51.5
45.75
Table 2: Results on the 11 released tasks, one adopted run per task. Time is the mean time from start to submission; cost is the total token cost over the 11 tasks.
Figure 4: Time to submission per task for each configuration. Marker shade encodes the token cost of the run: each configuration keeps its own hue, and darker means more expensive. The gray color bar gives the mapping from shade to cost.
Failure mode
Grok 4.7 High
Claude Opus 5.5
GPT-6 Astra
Long-tail edge-case rule
27 (31%)
12 (67%)
8 (38%)
Less common subsystem non-functional
45 (52%)
2 (11%)
0
Single omission breaking a test group
5 (6%)
0
13 (62%)
Environment-dependent behavior
10 (11%)
2 (11%)
0
Crash
0
2 (11%)
0
Total
87
18
21
Table 3: Failed hidden tests by failure mode, summed over each configuration’s failing tasks. Each failed test is assigned to one mode.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Product
Source
Language
Domain
Acceptance modules
001
Tomlparse
tomli
Python
TOML v1.1 parser
5
002
python-envfile
python-dotenv
Python
.env loading and expansion
8
003
ymlcodec
js-yaml
TypeScript
YAML codec with schemas and tags
7
004
Signtoken
itsdangerous
Python
HMAC signing and serialization
4
005
httpwire
h11
Python
HTTP/1.1 protocol state machine
7
006
Optlyn
Click
Python
CLI framework with shell completion
14
Appendix
Table 4: Tasks in the first Zero2Repo release. The product name is the neutralized name given to the agent; the source is the upstream project the task was derived from. An acceptance module groups related test cases.
Task
001
002
003
004
005
006
007
008
009
010
011
Rounds
3
1
3
3
2
4
2
2
3
3
2
Appendix
Table 5: Validation rounds per task before release.
GPT-6 Astra
Claude Opus 5.5
Grok 4.7 High
Task
Result
Time
Cost
Result
Time
Cost
Result
Time
Cost
001
Pass
391
1.32
Pass
398
1.68
Pass
2093
2.95
002
Pass
399
4.36
Pass
399
1.58
Pass
2640
3.20
003
Pass
786
5.42
Pass
1581
5.94
499/507
4234
6.10
004
Pass
217
1.65
Pass
290
1.43
Pass
2095
2.40
005
Pass
530
3.90
Pass
595
2.15
Pass
3285
4.55
Appendix
Table 6: Per-task results. “Pass” means that every hidden test passed; otherwise the number of passed hidden tests is shown. Time is seconds from start to submission; cost is in US dollars.
Model
Python
TypeScript (003)
C++ (008)
Go (009)
GPT-6 Astra
8/8
Pass
Pass
Fail
Claude Opus 5.5
7/8
Pass
Pass
Fail
Grok 4.7 High
7/8
Fail
Fail
Fail
Appendix
Table 7: Results by language. Python has eight tasks; TypeScript, C++, and Go have one task each.
Task
Grok 4.7 High
Claude Opus 5.5
GPT-6 Astra
003 YAML codec
499/507 (98.4%)
Pass
Pass
006 CLI framework
467/476 (98.1%)
473/476 (99.4%)
Pass
008 URL parser
223/248 (89.9%)
Pass
Pass
009 Large-file client
453/498 (91.0%)
483/498 (97.0%)
477/498 (95.8%)
Appendix
Table 8: Hidden tests passed by the failing submissions.
School of Computer Science, Wuhan University, China · School of Computer Science, Nanjing University of Science and Technology, China · School of Computer Science, Central China Normal University, China +2