Large-scale Repository Engineering via Agent-Native Reusable Code Primitives
Organizations: University of Illinois Urbana-Champaign, USA
Abstract
Large language models equipped with development environments have moved code generation toward repository-scale construction, yet building complete repositories remains difficult because interacting modules, interfaces, configurations, tests, and dependencies must work together. We introduce Code Primitives, agent-native reusable executable components with interface contracts, dependency closures, validation tests, and provenance. Each primitive uses a resident LLM to assess relevance and adapt its implementation, interfaces, and dependencies to the target repository, and we organize 1,424 validated primitives in CodeFace, a searchable library for repository construction. We introduce LEGO (Large-scale repository Engineering via aGent-native reusable cOde primitives), which activates task-relevant primitives, integrates their adapted implementations with task-specific code while resolving cross-component constraints, and revises the result against executed tests. To measure construction end to end, we build LEGO-REPO, a benchmark of 522 executable reconstruction tasks spanning seven software domains, 22 capability tracks, and five difficulty levels, scored against native test suites between an empty-package floor and original-source ceiling. The strongest of 13 evaluated backbones reaches a delivery score of 0.318 and scores zero on 41.0% of tasks; LEGO improves all 13 by 0.1474 on average and raises GPT-5.6-terra from 0.3180 to 0.5134 (+61.4%). In controlled comparisons, adapted primitives outperform retrieved code supplied as context or vendored unchanged. The effect persists against independent repository agents, across three external benchmarks, and with a disjointly re-mined CodeFace; GPT-OSS-20B for adaptation and diagnosis retains 95.1% of the homogeneous score at 24.0% lower cost.
Figures & tables
| Normalized delivery score by band | Outcome rate | |||||||
| Backbone | D1 | D2 | D3 | D4 | D5 | All | ||
| proprietary API models | ||||||||
| GPT-5.6-terra | .525 | .382 | .274 | .206 | .129 | .3180 | 8.8% | 41.0% |
| Claude Sonnet 5 | .521 | .379 | .260 | .194 | .091 | .3052 | 8.6% | 40.2% |
| GPT-5.4 | .485 | .301 | .216 | .123 | .061 | .2528 | 6.1% | 49.4% |
| Claude Opus 5 | .465 | .294 | .199 | .107 | .042 | .2371 | 5.9% | 53.4% |
| Score by difficulty band | Outcome rate | |||||||||
| System | Backbone | D1 | D2 | D3 | D4 | D5 | All | $/task | ||
| OpenHands | GPT-5.6-terra | .556 | .401 | .269 | .198 | .129 | .3263 | 2.41 | ||
| Claude Code | Claude Sonnet 5 | .543 | .394 | .271 | .190 | .104 | .3170 | 2.16 | ||
| SWE-agent ‡ | GPT-5.6-terra | .482 | .345 | .238 | .152 | .081 | .2750 | 1.86 | ||
| Agentless ‡ | GPT-5.6-terra | .421 | .276 | .181 | .098 | .043 | .2181 | 0.94 | ||
| Feedback (ours) | GPT-5.6-terra | .525 | .382 | .274 | .206 | .129 | .3180 | 1.55 | ||
| band | mod. | cov. | query | cand. | retain | adapt | share | used |
| D1 | 0, 721 | 75.3% | 0, 437 | 0, 874 | 0, 357 | 0, 332 | 46.0% | 61.1% |
| D2 | 2,064 | 71.2% | 0, 770 | 1,540 | 0, 613 | 0, 557 | 27.0% | 37.9% |
| D3 | 3,654 | 65.8% | 1,021 | 2,042 | 0, 818 | 0, 731 | 20.0% | 30.4% |
| D4 | 6,823 | 61.2% | 1,425 | 2,850 | 1,135 | 0, 989 | 14.5% | 23.7% |
| D5 | 9,723 | 56.7% | 1,862 | 3,724 | 1,489 | 1,264 | 13.0% | 22.9% |
| all | 22,985 | 61.4% | 5,515 | 11,030 | 4,412 | 3,873 | 16.9% | 27.5% |
| Configuration | Stages | D1 | D5 | All | $/task | |
| Single-shot | ---- | — | — | .2172 | 57.1% | 0.86 |
| Retrieval, one attempt | R--- | — | — | .2314 | 58.6% | 0.87 |
| Feedback | ---V | .525 | .129 | .3180 | 41.0% | 1.55 |
| Retrieval Feedback | R--V | .544 | .141 | .3352 | 34.7% | 1.56 |
| Feedback Diagnosis | --DV | .569 | .135 | .3469 | 30.1% | 1.79 |
| Retrieval Diagnosis | R-DV | .649 | .127 | .3980 | 31.6% | 1.80 |
| Mechanism added | All | step | % tot. |
| Feedback ( ---V ), ref. | .3180 | — | — |
| diagnosis revision ( --DV ) | .3469 | ||
| File RAG, editing ( FE DV ) | .3528 | ||
| constructed prim. ( RE DV ) | .4093 | ||
| primitive adaptation ( RADV ) | .5134 | ||
| roll-up | |||
| Configuration | D1 | D5 | All | --DV |
| Feedback Diagnosis ( --DV ) | .569 | .135 | .3469 | ref. |
| Context-only ( RC DV ) | .620 | .115 | .3544 | |
| Import Call ( RI DV ) | .627 | .060 | .3430 | |
| Structured-schema adaptation | .679 | .327 | .4974 | |
| LEGO ( RADV ) | .686 | .338 | .5134 |
| Benchmark | publ. | fb. | LEGO | |
| RepoZero, Py Py | 51.2 | 53.8 | 62.9 | |
| RepoGenesis | 38.2 | 41.0 | 48.2 | |
| RepoCraft | 0.412 | 0.436 | 0.501 |
Appendix figures & tables30 assets
Supplementary material from the paper’s appendix.
Appendix
| count | stage | transition into this stage |
| candidate registry | initial repository discovery | |
| attempted task candidates | repository screening: domain/security eligibility (including offensive-security holds), importable package, runnable native suite | |
| frozen task directories | successful materialization with clone/build/environment pinned | |
| LEGO-REPO evaluation set | stable same-pipeline ceiling and measurable span |
| category | attributable to | Feedback | LEGO |
| incomplete composition: modules never written | system under test | 84 | 38 |
| import or contract failure at collection | system under test | 61 | 27 |
| environment or dependency resolution | harness/environment | 35 | 18 |
| attempt budget exhausted, tests failing | system under test | 34 | 15 |
| zero-score tasks ( Dead@0 ) | 214 | 98 | |
| as a share of the benchmark tasks |
| domain | D1 | D2 | D3 | D4 | D5 | all |
| machine-learning | 15 | 25 | 35 | 39 | 44 | 158 |
| databases-storage | 39 | 31 | 37 | 30 | 12 | 149 |
| compilers-languages | 42 | 31 | 32 | 16 | 10 | 131 |
| graphics-vision | 5 | 11 | 6 | 4 | 6 | 32 |
| scientific-computing | 5 | 7 | 3 | 5 | 6 | 26 |
| quant-finance | 6 | 3 | 2 | 1 | 2 | 14 |
| capability track | D1 | D2 | D3 | D4 | D5 | all |
| agent-orchestration | 4 | 12 | 11 | 20 | 26 | 73 |
| learning | 2 | 10 | 12 | 11 | 11 | 46 |
| parsing-languages | 12 | 14 | 6 | 8 | 5 | 45 |
| web-protocol | 8 | 6 | 18 | 11 | 1 | 44 |
| serialization-wire | 7 | 8 | 9 | 6 | 4 | 34 |
| text-strings | 14 | 7 | 6 | 2 | 2 | 31 |
| quantity | min | p25 | median | p75 | max | mean |
| identity ceiling | 6 | 47 | 118 | 341 | 287.4 | |
| empty-package floor | 0 | 0 | 0 | 0 | 9.6 | |
| span | 5 | 46 | 114 | 335 | 277.8 | |
| floor share | 0 | 0 | 0 | 0 | 0.873 | 0.021 |
| tasks with span | tasks, mean LEGO score | |||||
| tasks with | tasks, mean LEGO score | |||||
| binning | quantity | B1 | B2 | B3 | B4 | B5 |
| (reported) | tasks | 117 | 111 | 116 | 96 | 82 |
| Feedback | .525 | .382 | .274 | .206 | .129 | |
| LEGO Feedback | ||||||
| , cuts at | tasks | 124 | 108 | 113 | 99 | 78 |
| Feedback | .515 | .383 | .280 | .199 | .121 | |
| LEGO Feedback |
| partition | tasks | modules | Feedback | LEGO | coverage | |
| D1 (mean modules) | 117 | 0, 721 | .525 | .686 | ||
| D2 (mean modules) | 111 | 2,064 | .382 | .583 | ||
| D3 (mean modules) | 116 | 3,654 | .274 | .483 | ||
| D4 (mean modules) | 96 | 6,823 | .206 | .409 | ||
| D5 (mean modules) | 82 | 9,723 | .129 | .338 | ||
| machine-learning | 158 | 9,649 | .2498 | .4368 |
| band | tasks | modules | covered | coverage | queries | retained | adapted |
| D1 | 117 | 0, 721 | 0, 543 | 0, 437 | 0, 357 | 0, 332 | |
| D2 | 111 | 2,064 | 1,470 | 0, 770 | 0, 613 | 0, 557 | |
| D3 | 116 | 3,654 | 2,404 | 1,021 | 0, 818 | 0, 731 | |
| D4 | 96 | 6,823 | 4,176 | 1,425 | 1,135 | 0, 989 | |
| D5 | 82 | 9,723 | 5,513 | 1,862 | 1,489 | 1,264 | |
| all | 522 | 22,985 | 14,106 | 5,515 | 4,412 | 3,873 |
| change from LEGO | resulting configuration | all | $/task | |
| none | RADV | .5134 | ref. | 2.08 |
| remove the whole library | --DV | .3469 | 1.79 | |
| remove primitive adaptation | R-DV | .3980 | 1.80 | |
| remove diagnosis | RA-V | .4795 | 1.83 | |
| replace primitive machinery with generic editing | RE DV | .4093 | 2.01 | |
| also replace primitives with raw files | FE DV | .3528 | 2.03 |
| Configuration | all | RE DV | LEGO |
| RAG Edit ( RE DV ) | .4093 | ref. | |
| carried tests | .4431 | ||
| contract and provenance | .4818 | ||
| dependency closure | .5039 | ||
| primitive without | .4796 | ||
| LEGO ( RADV ) | .5134 | ref. |
| Configuration | D1 | D2 | D3 | D4 | D5 | All |
| Feedback ( ---V ) | .525 | .382 | .274 | .206 | .129 | .3180 |
| LEGO, full library | .686 | .583 | .483 | .409 | .338 | .5134 |
| LEGO, same-repo donors excluded | .683 | .574 | .461 | .337 | .221 | .4743 |
| gain, full library | ||||||
| gain, same-repo excluded | ||||||
| lost to donor exclusion |
| Comparison | Mechanism | Win | Tie | Loss | score | |
| Retrieval Feedback | raw retrieved code | 148 | 269 | 105 | 2.7 | |
| Feedback Diagnosis | diagnosis, no library | 162 | 258 | 102 | 3.7 | |
| Import Call | vendored unchanged, invoked | 162 | 256 | 104 | 3.6 | |
| File RAG Edit | raw files generic editing | 174 | 247 | 101 | 4.4 | |
| Context-only | primitive as read-only reference | 161 | 260 | 101 | 3.7 | |
| Retrieval Diagnosis | retrieval diagnosis | 214 | 214 | 0 94 | 6.8 |
| primitive | score | kept | $/task | $/pt |
| GPT-5.6-terra | .5134 | 2.08 | 0.41 | |
| Claude Sonnet 5 | .5081 | 2.30 | 0.45 | |
| GPT-5.4 | .5060 | 2.08 | 0.41 | |
| GPT-5.6-luna | .5015 | 1.93 | 0.38 | |
| DeepSeek-V4 | .4992 | 1.83 | 0.37 | |
| Qwen3-Coder-480B | .4972 | 1.87 | 0.38 |
| diagnosis | score | kept | $/task | $/pt |
| GPT-5.6-terra | .5134 | 2.08 | 0.41 | |
| Claude Sonnet 5 | .5123 | 2.30 | 0.45 | |
| GPT-5.4 | .5096 | 2.08 | 0.41 | |
| GPT-5.6-luna | .5061 | 1.95 | 0.39 | |
| DeepSeek-V4 | .5058 | 1.86 | 0.37 | |
| Qwen3-Coder-480B | .5020 | 1.90 | 0.38 |
| backbone (with LEGO) | D1 | D2 | D3 | D4 | D5 | all | scratch | Dead@0 | ||
| GPT-5.6-terra | .686 | .583 | .483 | .409 | .338 | .5134 | .3180 | 18.8% | ||
| Claude Sonnet 5 | .665 | .573 | .473 | .392 | .301 | .4953 | .3052 | 30.1% | ||
| GPT-5.4 | .616 | .525 | .439 | .343 | .239 | .4478 | .2528 | 37.9% | ||
| Claude Opus 5 | .622 | .565 | .294 | .270 | .151 | .3982 | .2371 | 34.3% | ||
| GPT-5.6-luna | .553 | .471 | .428 | .174 | .060 | .3606 | .2203 | 33.0% | ||
| Grok-4.3 | .596 | .420 | .241 | .154 | .090 | .3190 | .1718 | 49.0% |
| library / retrieval setting | entries | all | vs default | |
| mined only | 871 | .5021 | ||
| mined web-sourced | 1,081 | .5108 | ||
| mined harvested | 1,214 | .5069 | ||
| full library (default) | 1,424 | .5134 | ||
| same-repo donor exclusion | 1,424 | .4743 | ||
| top- | 1,424 | .4948 |
| Feedback ( ---V ) | LEGO ( RADV ) | ||||
| budget | all | $/task | all | $/task | |
| .2172 | 0.86 | .3954 | 1.13 | ||
| .2420 | 1.05 | .4412 | 1.38 | ||
| .2785 | 1.21 | .4663 | 1.61 | ||
| .3066 | 1.36 | .5071 | 1.82 | ||
| (default) | .3180 | 1.55 | .5134 | 2.08 | |
| backbone | primitive | diagnosis | all | $/task | $/pt | % of LEGO |
| terra | terra | terra | .5134 | 2.08 | 0.41 | |
| terra | Sonnet 5 | Sonnet 5 | .5081 | 2.53 | 0.50 | |
| terra | DeepSeek-V4 | DeepSeek-V4 | .4934 | 1.62 | 0.33 | |
| terra | GPT-OSS-20B | GPT-OSS-20B | .4880 | 1.58 | 0.32 | |
| Sonnet 5 | Sonnet 5 | Sonnet 5 | .4869 | 4.26 | 0.87 | |
| Sonnet 5 | GPT-OSS-20B | GPT-OSS-20B | .4661 | 3.31 | 0.71 |
| Model | Provider | Open | Size | $/1M in | $/1M out | Used in Table |
| GPT-5.6-terra | OpenAI | No | Undisclosed | 1.25 | 10.00 | default backbone, primitive, diagnosis |
| Claude Opus 5 | Anthropic | No | Undisclosed | 5.00 | 25.00 | 1 , 24 |
| Claude Sonnet 5 | Anthropic | No | Undisclosed | 3.00 | 15.00 | 1 , 2 , 24 , 22 , 23 , 27 |
| Claude Haiku 4.5 | Anthropic | No | Undisclosed | 1.00 | 5.00 | 1 , 24 |
| GPT-5.4 | OpenAI | No | Undisclosed | 1.25 | 10.00 | 1 , 24 , 22 , 23 |
| GPT-5.6-luna | OpenAI | No | Undisclosed | 0.60 | 4.80 | 1 , 24 , 22 , 23 |
| (a) configuration means over decoding seeds | ||||
| Configuration | seed 1 | seed 2 | seed 3 | mean s.d. |
| Feedback ( ---V ) | .3180 | .3202 | .3150 | |
| Feedback Diagnosis ( --DV ) | .3469 | .3434 | .3493 | |
| RAG Edit ( RE DV ) | .4093 | .4036 | .4121 | |
| LEGO ( RADV ) | .5134 | .5125 | .5089 | |
| (b) paired contrasts on seed 1 | ||||
| quantity | D1 | D2 | D3 | D4 | D5 | all |
| tasks ( ) | 117 | 111 | 116 | 96 | 82 | 522 |
| improved tasks ( ) | 70 | 74 | 68 | 56 | 58 | 326 |
| regressed tasks ( ) | 24 | 18 | 17 | 14 | 11 | 84 |
| tied tasks ( ) | 23 | 19 | 31 | 26 | 13 | 112 |
| regressed as a share of the band | 20.5% | 16.2% | 14.7% | 14.6% | 13.4% | 16.1% |
| mean gain on improved tasks | .310 | .326 | .378 | .364 | .304 | .336 |
| successful adaptations on the task | – | – | – | Spearman | |
| tasks ( ) | 63 | 118 | 191 | 150 | |
| mean Feedback score | .3612 | .3438 | .3107 | .2888 | |
| mean LEGO score | .3735 | .4535 | .5365 | .5899 | |
| mean | |||||
| mean , difficulty-matched |
| library eligibility | entries | coverage | all | vs Feedback |
| all entries (default) | 1,424 | .5134 | ||
| same-repository donors | 1,424 | .4743 | ||
| same-org donors | 1,424 | .4597 | ||
| any fork or near-duplicate | 1,365 | .4455 | ||
| re-mined from disjoint corpus | 518 | .4528 |
| Configuration | definition | stages | score | |
| Single-shot † | no library, single pass | ---- | .2172 | |
| Retrieval, one attempt † | library, single pass | R--- | .2314 | |
| Feedback | no library | ---V | .3180 | |
| Retrieval Feedback | library, no adaptation or diagnosis | R--V | .3352 | |
| Feedback Diagnosis | no library, diagnosis enabled | --DV | .3469 | |
| Retrieval Diagnosis | library and diagnosis, no adaptation | R-DV | .3980 |
| Repository | band | mod. | cov. | donors | floor | ceil. | Feedback | LEGO |
| pandera | D5 | |||||||
| markdown | D3 | |||||||
| haystack | D5 |
| benchmark | unit the model produces | correctness signal | ceiling-rel. | accum. |
| HumanEval ( Chen et al., 2021 ) | a standalone function | hidden unit tests | no | no |
| MBPP ( Austin et al., 2021 ) | a standalone function | assertion triples | no | no |
| BigCodeBench ( Zhuo et al., 2025 ) | a function calling libraries | unit tests | no | no |
| CrossCodeEval ( Ding et al., 2023 ) | a line, cross-file context | identifier match | no | no |
| RepoBench ( Liu et al., 2024 ) | a line, retrieved context | text match, edit sim. | no | no |
| DevEval ( Li et al., 2024 ) | a function inside a repo | repository tests | no | no |
| scoring convention | mean | ? |
| reference-relative count ratio | 0.3191 | yes ( ) |
| vs. installed baseline | 0.2972 | yes |
| vs. ceiling, no floor | 0.3220 | no |
| identity ceiling (ours) | 0.3180 | no |
| audit outcome, random sample of scored reconstructions | Feedback | LEGO |
| genuine implementation of the target capability | 74 | 100 |
| partial implementation, tests pass on covered paths | 33 | 14 |
| test-targeted special-casing | 9 | 5 |
| degenerate stub that still passes graded tests | 4 | 1 |
| special-casing rate, Wilson CI | ||
| degenerate-stub rate, Wilson CI |
| license class | primitives | share | propagates into generated source? |
| permissive (MIT, BSD, Apache-2.0) | 1,149 | 80.7% | attribution only |
| weak copyleft (LGPL, MPL) | 118 | 8.3% | yes, file-level |
| strong copyleft (GPL, AGPL) | 96 | 6.7% | yes, whole-work |
| unlicensed or unresolved | 61 | 4.3% | unknown |
| total | 1,424 | 100% |