cs.SESep 28, 2026

From Dead Code and Static Requirements to Working Engines: Software Revival with Coding Agents

Authors: Tianyu Liu, Dingyuan Dai, Yufan Du, Zhen Yang

Organizations: Tsinghua University · UCLA

Abstract

Can coding agents restore software that no longer runs while preserving its underlying methods, and reconstruct industrial software engines from open specifications? Here we introduce ReviveBench, a benchmark with two task families evaluated by hidden verifiers calibrated against native execution environments, established engineering tools, or purpose-built reference implementations. The revival family comprises ten tasks involving dependency incompatibilities, deleted core modules, legacy builds, and a GPU-based foundation model. Every starting workspace fails verification, and the strongest evaluated model passes all ten tasks in at least one run each. In contamination-control experiments, identifier obfuscation reduces line similarity to the original implementations from 0.51--0.96 to 0.03--0.44 without reducing the observed pass rate of any evaluated model. On repositories created after the stated knowledge cutoffs, the strongest model passes eight of nine runs. The reconstruction family comprises thirteen tasks spanning numerical, geometric, hardware, and transactional systems (e.g. CAD and CRM). Two models meet the benchmark's pass criteria on all thirteen, although our audit shows that the CFD task cannot establish numerical-solver capability. Benchmark construction and auditing uncover 28 verifier defects, including 24 false negatives and two false positives. These findings show that executable verification can itself introduce substantial measurement error. We present three practical checks: test whether prescribed methods can reach the grading thresholds, investigate agreement among independently generated candidates, and recompute diagnostics from submitted artifacts. ReviveBench thus provides both an evaluation of software revival and engine reconstruction, and cases in validating the verifiers used to measure coding agents for software design.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. EvoCode-Bench: Evaluating Coding Agents in Multi-Turn Iterative Interactions

    May 22, 2026Haiyang Shen, Xuanzhong Chen, Wendong Xu +3Coding AgentsMulti-Turn Interactions

  2. SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents

    Sep 8, 2026Pujun Zheng, Zixin Shang, Shufan Jiang +5Swe-Bench VerifiedRepository-Level Code Understanding

  3. LoopsBench: From Harness Engineering to Loop Engineering in Coding Agent Evaluation

    Jul 31, 2026Han Li, Zhemin Fang, Rili Feng +8Coding AgentsAgent Loop