Code generation has emerged as a central capability of large language models, with coding agents now able to produce functionally correct software projects from natural language prompts. However, functional correctness alone does not capture a critical dimension of generation quality: environment specification, defined as the accurate identification of the dependencies required to execute generated code, is equally critical. We develop an agent protocol for environment specification and introduce a three-layer framework comprising declared, runtime-installed, and necessary-and-sufficient dependencies to systematically assess coding agents for environment specification. Using this protocol, we evaluate the extent to which coding agents systematically misspecify software environment dependencies and how this misspecification varies across three agents, four languages, and fifty programming tasks. Our results show that current coding agents exhibit systematic generalization failures along this dimension, producing dependency specifications that are inconsistent, redundant, or incomplete in ways that functional tests do not detect. Across agents, dependency set agreement is as low as 7% for identical tasks, and newer agents show no meaningful improvement, suggesting the failure is not resolved by scale or recency. The largest divergence occurs between the declared and runtime dependency layers, implicating environment priors learned from the models' training distributions as the primary driver. Our findings establish environment specification as a distinct, measurable axis of code generation quality that current benchmarks do not capture, and motivate training objectives and evaluation protocols that jointly optimize for functional correctness and environmental portability.
Figures & tables
Figure 1: Relationships among the three dependency layers. D(1) is extracted from the initial manifest, D(2) from the final successfully installed environment, and D(3) from independent runtime tracing via strace . No fixed set-inclusion is assumed between D(1) and D(2) . Pairwise gaps define phantom dependencies (declared but not needed), hidden dependencies (needed but not declared or installed), and bloat (installed but unused).
Figure 2: Environment Evaluation Pipeline. Left: the agent A exercises its Gen skill to produce an initial project (C(0),M(0)) , which is executed in an isolated environment I . On failure, the Repair skill is invoked with the runtime error e(i) for up to k rounds, yielding the final artifact (C(k),M(k)) . Right: three extraction primitives operate on the pipeline outputs to produce the three dependency layers: D(1) via manifest parsing of the initial declaration M(0) ; D(2) via package-manager resolution of the final manifest M(k) ; and D(3) via strace -based runtime tracing of C(k) in I , serving as ground truth.
Syscall READ s of target/dependency/*.jar via classpath execution
JavaScript
Keys of "dependencies" in package.json
npm list --all --json ; full node_modules/ tree
Syscall READ s under node_modules/
C++
FetchContent_Declare and find_package names in CMakeLists.txt
CMake configure log ( D(2)≈D(1) ; transitive expansion typically absent)
Syscall READ s of .so files; seven base system libraries excluded
Table 3: Per-ecosystem realization of ParseManifestℓ , Resolveℓ , and Traceℓ for the three dependency layers.
Figure 4: Gen skill: first-attempt success across all 600 instances (3 agents × 4 languages × 50 tasks). Each cell is one (agent, language, task) triple; green indicates first-attempt success and red indicates failure.
Figure 5: Repair skill: distribution of repair cycles needed to converge for each (agent, language) pair, bucketed into first-attempt success, 1–3, 4–7, and 8–10 cycles, and failure within the budget.
Figure 6: Protocol convergence by language. Solid bars show the first-attempt success rate SR(0) and hatched extensions show the self-correction gain contributed by Repair .
Figure 7: Phantom, hidden, and installed-but-unused bloat rates by ecosystem, averaged over successful projects with zero-denominator cases included as zero.
Figure 8: Initial-to-final environment growth by ecosystem. Each box summarizes the distribution of the inflation ratio ρ of Section 3 over all successful projects in that ecosystem (log scale).
Figure 9: Cross-agent agreement on declared dependencies. Each bar reports the mean Jaccard similarity between the declared dependency sets extracted from two agents’ initial manifests for the same task in the same ecosystem; the dotted line marks the level at which half of the declared packages would be shared.
Language
Mean Jˉintra
Median
UCR
Vˉ∪
Vˉ∩
Python
0.113
0.083
0/50 (0.0%)
5.5
0.2
Java
0.138
0.111
0/48 (0.0%)
5.6
0.1
JavaScript
0.068
0.000
0/50 (0.0%)
5.9
0.0
C++
0.093
0.067
0/35 (0.0%)
5.3
0.1
Table 5: Stochastic variability: Claude’s three independent trials produce near-zero agreement.
Agent
Fail
SysLib
SLAR
Rec.
EGAR
Claude
32
28
87.5%
28
100.0%
Codex
28
5
17.9%
5
100.0%
Gemini
10
0
0.0%
0
–
Table 6: Environment Gap: C++ system library assumption failures and recovery. “Fail” counts first-attempt failures; “SysLib” counts those whose terminal error is a missing system library; “Rec.” counts SysLib failures eventually repaired.
Figure 10: Protocol convergence by task domain, with solid and hatched bars as in Figure 6 .
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Ecosystem
ρS
p -value
Python
−0.276
0.0006
Java
−0.260
0.0013
JavaScript
−0.093
0.2564
C++
−0.410
0.0000
Appendix
Table 7: Spearman correlation between manifest size and first-attempt success.
ID
Metric
Key quantity
Results
M1
Success rate under self-correction
SR(0) , SR(≤k)
§ 6.1
M2
Environment growth
ρ
§ 6.2
M3
Phantom / hidden / bloat rates
ϕ , η , β
§ 6.2
M4
Manifest accuracy
P, R, F1
§ 6.2
M5
Cross-agent consistency
Jˉℓ , CDR
§ 6.3
M6
Stochastic variability
Jˉintra , UCR
§ 6.3
Appendix
Table 8: Metric summary: definitions, key quantities, and result section.
Large Language Model (LLM) agents demonstrate strong performance in autonomous code generation under loose specifications. However, production-grade software requires strict adherence to structural constraints, such as architectural patterns, databases, and object-relational mappings. Existing benchmarks often overlook these non-functional requirements, rewarding functionally correct but structurally arbitrary solutions. We present a systematic study evaluating how well agents handle structural constraints in multi-file backend generation. By fixing a unified API contract across 80 greenfield generation tasks and 20 feature-implementation tasks spanning eight web frameworks, we isolate the effect of structural complexity using a dual evaluation with end-to-end behavioral tests and static verifiers. Our findings reveal a phenomenon of constraint decay: as structural requirements accumulate, agent performance exhibits a substantial decline. Capable configurations lose 30 points on average in assertion pass rates from baseline to fully specified tasks, while some weaker configurations approach zero. Framework sensitivity analysis exposes significant performance disparities: agents succeed in minimal, explicit frameworks (e.g., Flask) but perform substantially worse on average in convention-heavy environments (e.g., FastAPI, Django). Finally, error analysis identifies data-layer defects (e.g., incorrect query composition and ORM runtime violations) as the leading root causes. This work highlights that jointly satisfying functional and structural requirements remains a key open challenge for coding agents.
Code generation aims to automatically generate source code from task requirements and has attracted significant attention with the rapid advancement of large language models (LLMs). Despite remarkable progress, LLMs often struggle to generate correct code for complex software engineering tasks because task descriptions are frequently incomplete, ambiguous, or lack critical contextual information. Existing approaches primarily improve the capabilities of coding agents through more sophisticated tools, skills, and workflows, while largely overlooking the quality of the task requirements themselves. To address this limitation, we draw inspiration from software requirements engineering and propose WiseSpec, a novel requirements-driven agent framework for repository-level code generation. WiseSpec automatically constructs structured and information-rich requirements, assesses their quality through execution-based evaluation, and iteratively refines them to better guide code generation. Experimental results show that WiseSpec consistently outperforms all baselines, achieving an average improvement of 13.17% in %Resolved.
Zhao Tian
School of Computer Software, Tianjin University Tianjin, China
Coding agents powered by large language models (LLMs) are evolving from making localized code changes to developing complete software repositories. However, evaluating repository-scale generation remains challenging: tasks must demand system-level reasoning while ensuring that all evaluated behaviors are precisely specified and independent of any particular implementation. We introduce E2E-SWE, a benchmark for evaluating whether coding agents can build complete, functional software repositories end to end. E2E-SWE contains 186 whole-repository generation tasks spanning 11 programming languages. Given only a natural-language specification and an empty workspace, an agent must implement a complete, installable project that satisfies a comprehensive suite of hidden tests. Each task is constructed by a software engineer in collaboration with an LLM; together, they develop the test suite and a corresponding implementation-independent specification. To ensure that tasks are well specified and practically solvable, we further subject them to an iterative verification process in which autonomous agents audit and repair task defects using static inspection and failures observed from real model rollouts. Evaluating 13 frontier models, we find substantial variation in end-to-end repository generation ability, with pass@1 ranging from 11.7% to 67.7%, providing strong model differentiation while leaving considerable headroom for future progress. Analysis of agent trajectories further reveals long, front-loaded reasoning patterns, highlighting the planning and system-level reasoning required to construct working codebases from scratch.