Can coding agents restore software that no longer runs while preserving its underlying methods, and reconstruct industrial software engines from open specifications? Here we introduce ReviveBench, a benchmark with two task families evaluated by hidden verifiers calibrated against native execution environments, established engineering tools, or purpose-built reference implementations. The revival family comprises ten tasks involving dependency incompatibilities, deleted core modules, legacy builds, and a GPU-based foundation model. Every starting workspace fails verification, and the strongest evaluated model passes all ten tasks in at least one run each. In contamination-control experiments, identifier obfuscation reduces line similarity to the original implementations from 0.51--0.96 to 0.03--0.44 without reducing the observed pass rate of any evaluated model. On repositories created after the stated knowledge cutoffs, the strongest model passes eight of nine runs. The reconstruction family comprises thirteen tasks spanning numerical, geometric, hardware, and transactional systems (e.g. CAD and CRM). Two models meet the benchmark's pass criteria on all thirteen, although our audit shows that the CFD task cannot establish numerical-solver capability. Benchmark construction and auditing uncover 28 verifier defects, including 24 false negatives and two false positives. These findings show that executable verification can itself introduce substantial measurement error. We present three practical checks: test whether prescribed methods can reach the grading thresholds, investigate agreement among independently generated candidates, and recompute diagnostics from submitted artifacts. ReviveBench thus provides both an evaluation of software revival and engine reconstruction, and cases in validating the verifiers used to measure coding agents for software design.
Figures & tables
Figure 1: ReviveBench. Agents receive software that does not run or does not exist (left), work headlessly in an isolated sandbox (middle), and are graded by hidden verifiers that use three forms of correctness checking (right). Auditing the verifiers identified 28 defects; a fixed defect triggers re-scoring if the specification is unchanged and a re-run if the specification gained information.
Category
Software
What is broken
Fable 5.1
Sonnet 5
Haiku 4.5
Dependency rot
deepimpute
TF 2.16 / Keras 3 / NumPy 2
2/3
1/1
1/1
DCA
removed Keras and TF-1 APIs
3/4
1/1
0/1
Deleted core module
deepimpute
core module deleted
3/3
1/1
0/1
deepimpute (obf.)
same, identifiers renamed
3/3
2/2
1/1
Build and reproduce
PyMOL 2.3
2019 C++ build, Python 3.12
2/6
0/1
0/1
QuPath 0.2.3
Java 14 build, JDK 21
3/4
1/1
1/1
Table 1: Software revival. Each cell reports successful runs divided by all runs recorded for that model and task. Every starting workspace fails its verifier. The models are the same Fable 5.1, Sonnet 5 and Haiku 4.5 used in the reconstruction suite, reached through Anthropic’s own endpoint or, for eight later runs, Amazon Bedrock.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Anthropic
OpenAI GPT-5.6
Zhipu
Engine
Commercial analog
Fable 5.1
Opus 5
Sonnet 5
Haiku 4.5
sol
luna
terra
GLM-5.3 flash
Statistics engine
SAS, SPSS
✓ 43/43
✓ 43/43
✗ 34/43
✗ 23/43
✓ 43/43
✓ 40/43
✓ 41/43
✓ 41/43
Circuit simulator
HSPICE
✓ 10/10
✓ 9/10
✓ 9/10
✗ 0/10
✗ 7/10
✗ 5/10
✗ 6/10
✗ 0/10
PLC runtime
TIA Portal
✓ 9/10
✓ 9/10
✓ 10/10
✗ 6/10
✓ 10/10
✗ 2/10
✓ 10/10
✓ 8/10
FE solver
Nastran
✓ 4/4
✓ 4/4
✓ 4/4
✗ 2/4
✗ 2/4
✓ 4/4
✗ 2/4
✓ 4/4
2D CAD kernel
AutoCAD
✓ 31/31
✓ 31/31
✓ 29/31
✗ 0/31
✓ 29/31
✗ 0/31
✗ 21/31
✗ 0/31
Appendix
Table 2: Clean-room reconstruction of thirteen industrial software engines (the data behind Figure 2 a). Each cell is the hidden-verifier score of the best run under the protocol condition; scores are comparable within a row only; ✓ denotes a pass. GLM-5.3 flash cells use the runs described in Section 9 . Pass marks denote task-level thresholds, not necessarily perfect scores. CFD scores measure interface compliance and do not establish numerical-solver capability (Section 8 ); MES and PLM do not distinguish the evaluated models.
Model
Engines passed
Valid runs
Mean turns
Cost (USD)
Fable 5.1
13/13
13
20.8
95
Opus 5
13/13
9
74.1
76
Sonnet 5
10/13
13
60.1
35
Haiku 4.5
2/13
15
80.3
15
sol
9/13
13
30.2
n/a †
luna
6/13
13
35.3
n/a †
Appendix
Table 3: Per-model totals on the reconstruction suite. “Engines passed” is taken from Table 2 ; runs, turns and cost are over all valid runs of that model.
Task
Effort
Score
Turns
Wall (s)
Cost (USD)
ERP core
low
5/5
9
204
1.00
medium
5/5
18
411
1.85
high
5/5
20
672
2.71
xhigh
5/5
30
1070
4.37
max
5/5
31
1464
4.79
2D CAD kernel
low
31/31
65
1426
6.95
Appendix
Table 4: Reasoning-effort ablation (Opus 5, five levels). Scores remain unchanged. Cost, turns, and wall time increase monotonically for ERP but fluctuate at intermediate settings for CAD. Both tasks are saturated at all five settings, limiting conclusions about improvements in accuracy (Appendix B ).
Coding agents are increasingly used as iterative development partners, but most benchmarks still evaluate one specification followed by one final assessment. This leaves out a basic question: can an agent keep its own codebase working as requirements change? We introduce EvoCode-Bench, a benchmark of 26 stateful coding tasks and 227 evaluated rounds. Each task preserves the agent's workspace for 5-15 rounds, states requirements through observable behavior, and uses cumulative executable tests to check new requirements and still-active prior ones. We evaluate 13 coding agents with two metrics: MT@4, a four-attempt fail-stop multi-round score, and SR, a single-round score from a reference-completed prior state. For most agents, SR exceeds MT@4 by 22-40 points. The gap also changes rankings: the highest-SR agent (78.9) ranks only third in persistent execution (44.0 MT@4). Even the strongest agents achieve only about 50% success on multi-turn metrics, and aggregate pass rate drops below half of round-1 performance by round 5. Failure analysis shows tier-dependent behavior: weaker agents fail early, while stronger agents survive long enough to expose specification-tracking and regression failures. We release the benchmark data and Harbor multi-turn infrastructure.
Haiyang Shen, Xuanzhong Chen, Wendong Xu +3
1UniPat AI · 2Peking University · 3Tsinghua University +1
SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that its evaluation is undermined by two sources of unreliability: reward hacking, enabled by leakage of gold solutions or hidden evaluation information, and task quality issues, including misleading problem statements and improperly scoped tests. These issues can inflate benchmark performance and obscure agents' true coding ability. We present SWE-Bench Pro Verified, a verified version of SWE-Bench Pro that addresses both problems. Our approach combines anti-hacking safeguards that eliminate major leakage channels without disrupting normal agent functionality, with task refinement that minimally corrects inconsistencies within flawed instances. Evaluations on SWE-Bench Pro Verified reveal that some models perform substantially worse than previously evaluated, suggesting that existing results on SWE-Bench Pro may overestimate real software engineering capability. SWE-Bench Pro Verified offers a more trustworthy benchmark for assessing software engineering agents.
Pujun Zheng, Zixin Shang, Shufan Jiang +5
East China Normal University · Shanghai Artificial Intelligence Laboratory · Fudan University
Coding agent infrastructure is shifting from harness engineering toward loop engineering as coding agents are deployed for sustained long-horizon software development. Existing benchmarks often center on localized tasks or end-state outcomes, offering limited insight into sustained execution. We introduce LOOPSBENCH, a long-horizon benchmark for loop engineering in coding agent evaluation. Each task is a dependency DAG over separately testable development units with source-evidenced prerequisite edges. LOOPSBENCH comprises 112 tasks from authentic sources spanning 8 programming languages and 9 domains. Its flow-aware runtime releases tests along the ready frontier and retains completed nodes as regression obligations. We evaluate frontier coding agents paired with widely used loop implementations. The strongest configuration, Opus-4.7 with Claude Code and outer continuation, resolves 25.00% of tasks. Recorded plans recover only part of the source-recovered prerequisite DAG, and regression events remain visible across the evaluated loop profiles. We open source the benchmark data and code, including all tasks, more than 5,300 development units, and executable tests, at microsoft/Loopsbench.
Han Li, Zhemin Fang, Rili Feng +8
Microsoft · Nanjing University · Shanghai Jiao Tong University +1