cs.SESep 30, 2026

Code That Works, Environments That Don't: Measuring Environment Reproducibility in AI-Generated Software

Authors: Bhanu Prakash Vangala, Tanu Malik

Organizations: Department of Electrical Engineering and Computer Science, University of Missouri–Columbia, Missouri, USA

Abstract

Code generation has emerged as a central capability of large language models, with coding agents now able to produce functionally correct software projects from natural language prompts. However, functional correctness alone does not capture a critical dimension of generation quality: environment specification, defined as the accurate identification of the dependencies required to execute generated code, is equally critical. We develop an agent protocol for environment specification and introduce a three-layer framework comprising declared, runtime-installed, and necessary-and-sufficient dependencies to systematically assess coding agents for environment specification. Using this protocol, we evaluate the extent to which coding agents systematically misspecify software environment dependencies and how this misspecification varies across three agents, four languages, and fifty programming tasks. Our results show that current coding agents exhibit systematic generalization failures along this dimension, producing dependency specifications that are inconsistent, redundant, or incomplete in ways that functional tests do not detect. Across agents, dependency set agreement is as low as 7% for identical tasks, and newer agents show no meaningful improvement, suggesting the failure is not resolved by scale or recency. The largest divergence occurs between the declared and runtime dependency layers, implicating environment priors learned from the models' training distributions as the primary driver. Our findings establish environment specification as a distinct, measurable axis of code generation quality that current benchmarks do not capture, and motivate training objectives and evaluation protocols that jointly optimize for functional correctness and environmental portability.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Constraint Decay: The Fragility of LLM Agents in Backend Code Generation

    May 7, 2026Francesco Dente, Dario Satriani, Paolo PapottiCode GenerationLarge Language Model Agents

  2. WiseSpec: Requirements-Driven Agents for Code Generation

    Sep 1, 2026Zhao TianCode GenerationSoftware Engineering

  3. E2E-SWE: Benchmarking LLMs on Building Working Codebases from Scratch

    Sep 29, 2026Hantian Ding, Chloe Bi, Jiacheng Zhu +5CodebasesModel Auditing