Organizations: Peking University · Fudan University · Shanghai Jiao Tong University · Tsinghua University · The Chinese University of Hong Kong, Shenzhen · Zhongguancun Academy
Coding agents are increasingly deployed for iterative development on real repositories, yet existing evaluation barely answers a basic question: \emph{do coding agents reuse existing code or reinvent the wheel?} The question matters: every duplicated implementation is a fix applied twice and agents produce code far faster than humans can audit, so redundancy accumulates unsupervised. Thus, we present \textbf{RepoReuse}, a multi-turn benchmark for auditing code reuse in real repositories, where requirements are revealed turn by turn and the workspace accumulates across turns. It is built by a fully automated pipeline combining AST-based dependency graphs, guided evidence collection, and execution-verified task synthesis, and scales readily to new repositories. Beyond pass rates, we measure the reuse rate together with recall and cross-turn structural redundancy. An audit over 3{,}000 turns shows that agents progressively stop exploring relevant repository code, reuse their own history less even when it is fully in the workspace, and leave duplicated logic in 50.8% of task chains by turn~5---all while pass rates barely move. Such deficiencies are invisible to pass rates, underscoring the need to evaluate code generation beyond functional correctness.
Figures & tables
Figure 1: Overview of RepoReuse. Left: the benchmark protocol—turn-wise requirements over a workspace that accumulates the agent’s own code, scored per turn by functional tests and by recall, reuse, and redundancy. Right: an illustrative example—the agent reuses one repository symbol and one of its own earlier functions but rewrites the two remaining targets; tests pass while both reuse rates are 1/2 .
Statistic
Value
Tasks
75
Avg. turns per task
5.0
Avg. test cases per turn
10.8
Avg. requirement length (chars)
∼ 2,570
New functions per turn
3
Reused symbols per turn (repo / self)
1.9 / 2.4
Table 1: Statistics of RepoReuse.
Upstream
Downstream
Reuse
Recall
Agent Behavior
Correctness
Redundancy
Harness
Model
repo
self
repo
self
Files
Lines
Resolve
Pass
Cdup
mini-SWE-agent
GPT-5.6 Terra
62.8
71.1
52.7
99.6
20.5
1150
48.8
81.2
42.7
DeepSeek-v4.1-flash
68.8
84.2
61.4
99.9
45.4
1267
66.4
90.8
21.1
Qwen3.7-plus
48.9
67.7
36.2
99.6
22.0
680
27.7
71.1
37.1
GLM-5.3
55.3
71.7
47.2
99.2
36.5
1104
44.3
81.2
39.7
Table 2: Main results on RepoReuse. Columns are organized around reuse: upstream metrics describe how the agent explores, and downstream metrics describe the outcome. Reuse, recall, correctness, and Cdup are in %; Cdup is averaged over the five cumulative turn-wise rates. Best reuse, recall, and correctness per column in bold .
recallrepo
reuserepo
reuseself
Cdup
Harness
Model
T1
T5
Δ
T1
T5
Δ
T2
T5
Δ
T1
T5
Δ
mini-SWE-agent
GPT-5.6 Terra
86.3
38.8
− 47.5
61.7
64.3
+2.6
77.1
69.0
− 8.1
14.7
60.0
+45.3
mini-SWE-agent
DeepSeek-v4.1-flash
88.1
49.8
− 38.3
74.4
72.2
− 2.2
92.2
79.3
− 12.9
10.7
33.3
+22.6
mini-SWE-agent
Qwen3.7-plus
80.1
21.7
− 58.4
52.5
50.1
− 2.4
84.2
61.3
− 22.9
14.7
57.3
+42.6
mini-SWE-agent
GLM-5.3
82.4
33.6
− 48.8
63.9
54.7
− 9.2
80.9
66.2
− 14.7
16.0
58.7
+42.7
OpenCode
GPT-5.6 Terra
76.3
30.7
− 45.6
62.4
57.2
− 5.2
77.1
65.3
− 11.8
10.7
52.0
+41.3
Table 3: Recall, reuse, and redundancy at the first and last turn (%). Self reuse is defined from turn 2; Δ is the change from the first to the last turn. Self reuse declines while the share of task chains containing re-implemented targets grows over turns.
Upstream
Downstream
Reuse
Recall
Agent Behavior
Correctness
Redundancy
Memory
repo
self
repo
self
Files
Lines
Resolve
Pass
Cdup
No memory
58.9
30.0
55.5
85.5
27.7
878
30.4
72.4
65.1
Interface memory
50.5
67.8
35.1
99.7
13.5
576
30.4
71.9
46.7
Source memory
51.3
29.2
38.5
84.4
16.8
567
27.7
70.2
70.7
Table 4: Effect of historical memory on OpenCode with Qwen3.7-plus (75 task chains, 375 turns per setting). Reuse, recall, correctness, and Cdup are in %; agent behavior and Cdup are averaged over turns. Best reuse, recall, and correctness per column in bold .
recallrepo
reuserepo
reuseself
Cdup
Memory
T1
T5
Δ
T1
T5
Δ
T2
T5
Δ
T1
T5
Δ
No memory
77.0
46.5
− 30.5
56.4
61.4
+5.0
34.2
32.1
− 2.1
22.7
89.3
+66.6
Interface memory
77.4
15.8
− 61.6
54.4
50.8
− 3.6
82.9
58.9
− 24.0
21.3
69.3
+48.0
Source memory
77.8
20.6
− 57.2
53.4
50.1
− 3.3
29.6
29.1
− 0.5
17.3
94.7
+77.4
Table 5: Effect of historical memory across turns on OpenCode with Qwen3.7-plus (%). Self reuse is defined from turn 2; Δ is the change from the first to the last turn.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
mini-SWE-agent
OpenCode
Tools
bash only
native tools
Step limit
120 commands
none
Cost limit
$5
none
Per-command timeout
120 s
harness default
Wall-time limit
2,400 s
2,400 s
Temperature
0
provider default
Appendix
Table 6: Per-turn budgets and settings of the two harnesses.