Organizations: Peking University · Fudan University · Shanghai Jiao Tong University · Tsinghua University · The Chinese University of Hong Kong, Shenzhen · Zhongguancun Academy
Coding agents are increasingly deployed for iterative development on real repositories, yet existing evaluation barely answers a basic question: \emph{do coding agents reuse existing code or reinvent the wheel?} The question matters: every duplicated implementation is a fix applied twice and agents produce code far faster than humans can audit, so redundancy accumulates unsupervised. Thus, we present \textbf{RepoReuse}, a multi-turn benchmark for auditing code reuse in real repositories, where requirements are revealed turn by turn and the workspace accumulates across turns. It is built by a fully automated pipeline combining AST-based dependency graphs, guided evidence collection, and execution-verified task synthesis, and scales readily to new repositories. Beyond pass rates, we measure the reuse rate together with recall and cross-turn structural redundancy. An audit over 3{,}000 turns shows that agents progressively stop exploring relevant repository code, reuse their own history less even when it is fully in the workspace, and leave duplicated logic in 50.8% of task chains by turn~5---all while pass rates barely move. Such deficiencies are invisible to pass rates, underscoring the need to evaluate code generation beyond functional correctness.
Figures & tables
Figure 1: Overview of RepoReuse. Left: the benchmark protocol—turn-wise requirements over a workspace that accumulates the agent’s own code, scored per turn by functional tests and by recall, reuse, and redundancy. Right: an illustrative example—the agent reuses one repository symbol and one of its own earlier functions but rewrites the two remaining targets; tests pass while both reuse rates are 1/2 .
Statistic
Value
Tasks
75
Avg. turns per task
5.0
Avg. test cases per turn
10.8
Avg. requirement length (chars)
∼ 2,570
New functions per turn
3
Reused symbols per turn (repo / self)
1.9 / 2.4
Table 1: Statistics of RepoReuse.
Upstream
Downstream
Reuse
Recall
Agent Behavior
Correctness
Redundancy
Harness
Model
repo
self
repo
self
Files
Lines
Resolve
Pass
Cdup
mini-SWE-agent
GPT-5.6 Terra
62.8
71.1
52.7
99.6
20.5
1150
48.8
81.2
42.7
DeepSeek-v4.1-flash
68.8
84.2
61.4
99.9
45.4
1267
66.4
90.8
21.1
Qwen3.7-plus
48.9
67.7
36.2
99.6
22.0
680
27.7
71.1
37.1
GLM-5.3
55.3
71.7
47.2
99.2
36.5
1104
44.3
81.2
39.7
Table 2: Main results on RepoReuse. Columns are organized around reuse: upstream metrics describe how the agent explores, and downstream metrics describe the outcome. Reuse, recall, correctness, and Cdup are in %; Cdup is averaged over the five cumulative turn-wise rates. Best reuse, recall, and correctness per column in bold .
recallrepo
reuserepo
reuseself
Cdup
Harness
Model
T1
T5
Δ
T1
T5
Δ
T2
T5
Δ
T1
T5
Δ
mini-SWE-agent
GPT-5.6 Terra
86.3
38.8
− 47.5
61.7
64.3
+2.6
77.1
69.0
− 8.1
14.7
60.0
+45.3
mini-SWE-agent
DeepSeek-v4.1-flash
88.1
49.8
− 38.3
74.4
72.2
− 2.2
92.2
79.3
− 12.9
10.7
33.3
+22.6
mini-SWE-agent
Qwen3.7-plus
80.1
21.7
− 58.4
52.5
50.1
− 2.4
84.2
61.3
− 22.9
14.7
57.3
+42.6
mini-SWE-agent
GLM-5.3
82.4
33.6
− 48.8
63.9
54.7
− 9.2
80.9
66.2
− 14.7
16.0
58.7
+42.7
OpenCode
GPT-5.6 Terra
76.3
30.7
− 45.6
62.4
57.2
− 5.2
77.1
65.3
− 11.8
10.7
52.0
+41.3
Table 3: Recall, reuse, and redundancy at the first and last turn (%). Self reuse is defined from turn 2; Δ is the change from the first to the last turn. Self reuse declines while the share of task chains containing re-implemented targets grows over turns.
Upstream
Downstream
Reuse
Recall
Agent Behavior
Correctness
Redundancy
Memory
repo
self
repo
self
Files
Lines
Resolve
Pass
Cdup
No memory
58.9
30.0
55.5
85.5
27.7
878
30.4
72.4
65.1
Interface memory
50.5
67.8
35.1
99.7
13.5
576
30.4
71.9
46.7
Source memory
51.3
29.2
38.5
84.4
16.8
567
27.7
70.2
70.7
Table 4: Effect of historical memory on OpenCode with Qwen3.7-plus (75 task chains, 375 turns per setting). Reuse, recall, correctness, and Cdup are in %; agent behavior and Cdup are averaged over turns. Best reuse, recall, and correctness per column in bold .
recallrepo
reuserepo
reuseself
Cdup
Memory
T1
T5
Δ
T1
T5
Δ
T2
T5
Δ
T1
T5
Δ
No memory
77.0
46.5
− 30.5
56.4
61.4
+5.0
34.2
32.1
− 2.1
22.7
89.3
+66.6
Interface memory
77.4
15.8
− 61.6
54.4
50.8
− 3.6
82.9
58.9
− 24.0
21.3
69.3
+48.0
Source memory
77.8
20.6
− 57.2
53.4
50.1
− 3.3
29.6
29.1
− 0.5
17.3
94.7
+77.4
Table 5: Effect of historical memory across turns on OpenCode with Qwen3.7-plus (%). Self reuse is defined from turn 2; Δ is the change from the first to the last turn.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
mini-SWE-agent
OpenCode
Tools
bash only
native tools
Step limit
120 commands
none
Cost limit
$5
none
Per-command timeout
120 s
harness default
Wall-time limit
2,400 s
2,400 s
Temperature
0
provider default
Appendix
Table 6: Per-turn budgets and settings of the two harnesses.
Most coding-agent benchmarks ask whether generated code behaves correctly. That remains essential, but repository-level engineering is increasingly agent-managed: one agent writes a repository, and later agents inspect, audit, or extend it as working context. In that setting, a generated repository is not only an answer to a task but also a communication artifact for future work. Even when strong agents nearly satisfy the visible behavioral objective, repositories can differ in how clearly they expose the intended behavior and design choices behind that behavior. We introduce BUILD-AND-FIND, a protocol for evaluating whether downstream agents can recover those intended choices from generated repositories, and how much inspection that recovery requires. For each task, a builder sees a hidden repository specification and creates a codebase; a finder sees only the codebase and a specification-traced multiple-choice question bank. The protocol separates behavioral correctness from artifact-side recovery and reports recovery accuracy, repeatability, implementation coverage, and inspection effort. Accuracy and stability act as gates: effort is interpreted only when recovery succeeds reliably. Among artifacts from which the same intent can be recovered, lower effort by the same finder suggests that the artifact makes that intent easier to locate. Question-only and spec-only controls quantify generic priors and specification access, while audits separate omitted claims from finder failures and check whether correct answers cite artifact evidence. In the released high-prior task pack, recovery accuracy is near saturation, so inspection effort and finder-specific effects provide the main panel-local comparison.
Coding agents are increasingly asked to build software rather than patch it, yet benchmarks for from-scratch repository construction are mostly limited to a single language and depend on manually curated tasks. We introduce Zero2Repo, a benchmark in which an agent receives a product requirements document, an interface contract, and an empty workspace, and must deliver a complete repository in the project's native ecosystem. Tasks are produced by a language-agnostic authoring pipeline that converts real, version-pinned open-source projects into behavioral specifications, reproducible environments, and hidden acceptance tests. Each task is validated by execution: a reference implementation derived from the upstream project must pass, and adversarial validation must show that the tests reject incorrect implementations. Evaluation runs production coding agents in isolated containers, withholds the acceptance tests until an explicit submission, and assigns a binary reward only when every test passes, with no LLM judge. The pipeline and harness make no language-specific assumptions and apply to mainstream programming ecosystems; the current release contains Python, TypeScript, Go, and C++ tasks. Even on 11 tasks drawn from repositories that frontier models have very likely seen during training, the strongest agent solves only 10, and every failing submission passes 90-99% of the hidden tests; for the two strongest agents, 67-100% of failed tests trace to a single omission or a low-frequency rule stated in the specification rather than to a missing subsystem, so each failure is a concrete target for improvement.
Pei Yang, Tianyu Shi, Yuhang Yao +23
Gradient Data · McGill University · Carnegie Mellon University +6
Modern coding agents are usually evaluated by whether they eventually produce a correct patch, but patch generation depends on an earlier context-acquisition stage: finding the repository files needed for the task. We introduce Agent Retrieval Bench, a file-level benchmark for this upstream retrieval problem. Samples are built from real coding-workflow signals and evaluated against frozen base-commit repositories, with relevance defined by what an agent needs next rather than direct query-file semantic similarity. The benchmark covers four positive-retrieval tasks: code2test, comment2context, trace2code, and edit2ripple; a fifth subset evaluates selective retrieval using natural evidence-backed no-gold cases and counterfactual wrong-repository controls. Agent Retrieval Bench contains 427 samples across 25 repositories: 345 positive examples, 50 natural no-gold examples, and 32 counterfactual controls. The corpus includes 308 base-commit snapshots, 392,000 files, and 7.9 million chunks. We evaluate lexical retrieval, RepoMap, open-source embeddings, selective abstention, and logged agent context selection. No single retrieval family dominates: Qwen3-Embedding-4B has the best sample-weighted MRR on positive samples, Qwen3-Embedding-8B the best Recall@20, and RepoMap the best budgeted context yield at 8K tokens, with task-level winners differing substantially. Selective thresholds calibrated with counterfactual controls do not improve selective success on natural no-gold cases, revealing a calibration gap. Logged trajectories also miss every gold file on 27-35 percent of samples. A controlled seed-intervention pilot finds that retrieval-derived initial context yields higher file F1 with less post-seed exploration than random non-gold context, while oracle gold context shows substantial remaining headroom.
Bowen Qin, Yi Xie
1National University of Singapore (NUS) · 2Peking University