Can coding agents restore software that no longer runs while preserving its underlying methods, and reconstruct industrial software engines from open specifications? Here we introduce ReviveBench, a benchmark with two task families evaluated by hidden verifiers calibrated against native execution environments, established engineering tools, or purpose-built reference implementations. The revival family comprises ten tasks involving dependency incompatibilities, deleted core modules, legacy builds, and a GPU-based foundation model. Every starting workspace fails verification, and the strongest evaluated model passes all ten tasks in at least one run each. In contamination-control experiments, identifier obfuscation reduces line similarity to the original implementations from 0.51--0.96 to 0.03--0.44 without reducing the observed pass rate of any evaluated model. On repositories created after the stated knowledge cutoffs, the strongest model passes eight of nine runs. The reconstruction family comprises thirteen tasks spanning numerical, geometric, hardware, and transactional systems (e.g. CAD and CRM). Two models meet the benchmark's pass criteria on all thirteen, although our audit shows that the CFD task cannot establish numerical-solver capability. Benchmark construction and auditing uncover 28 verifier defects, including 24 false negatives and two false positives. These findings show that executable verification can itself introduce substantial measurement error. We present three practical checks: test whether prescribed methods can reach the grading thresholds, investigate agreement among independently generated candidates, and recompute diagnostics from submitted artifacts. ReviveBench thus provides both an evaluation of software revival and engine reconstruction, and cases in validating the verifiers used to measure coding agents for software design.
Figures & tables
Figure 1: ReviveBench. Agents receive software that does not run or does not exist (left), work headlessly in an isolated sandbox (middle), and are graded by hidden verifiers that use three forms of correctness checking (right). Auditing the verifiers identified 28 defects; a fixed defect triggers re-scoring if the specification is unchanged and a re-run if the specification gained information.
Category
Software
What is broken
Fable 5.1
Sonnet 5
Haiku 4.5
Dependency rot
deepimpute
TF 2.16 / Keras 3 / NumPy 2
2/3
1/1
1/1
DCA
removed Keras and TF-1 APIs
3/4
1/1
0/1
Deleted core module
deepimpute
core module deleted
3/3
1/1
0/1
deepimpute (obf.)
same, identifiers renamed
3/3
2/2
1/1
Build and reproduce
PyMOL 2.3
2019 C++ build, Python 3.12
2/6
0/1
0/1
QuPath 0.2.3
Java 14 build, JDK 21
3/4
1/1
1/1
Table 1: Software revival. Each cell reports successful runs divided by all runs recorded for that model and task. Every starting workspace fails its verifier. The models are the same Fable 5.1, Sonnet 5 and Haiku 4.5 used in the reconstruction suite, reached through Anthropic’s own endpoint or, for eight later runs, Amazon Bedrock.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Anthropic
OpenAI GPT-5.6
Zhipu
Engine
Commercial analog
Fable 5.1
Opus 5
Sonnet 5
Haiku 4.5
sol
luna
terra
GLM-5.3 flash
Statistics engine
SAS, SPSS
✓ 43/43
✓ 43/43
✗ 34/43
✗ 23/43
✓ 43/43
✓ 40/43
✓ 41/43
✓ 41/43
Circuit simulator
HSPICE
✓ 10/10
✓ 9/10
✓ 9/10
✗ 0/10
✗ 7/10
✗ 5/10
✗ 6/10
✗ 0/10
PLC runtime
TIA Portal
✓ 9/10
✓ 9/10
✓ 10/10
✗ 6/10
✓ 10/10
✗ 2/10
✓ 10/10
✓ 8/10
FE solver
Nastran
✓ 4/4
✓ 4/4
✓ 4/4
✗ 2/4
✗ 2/4
✓ 4/4
✗ 2/4
✓ 4/4
2D CAD kernel
AutoCAD
✓ 31/31
✓ 31/31
✓ 29/31
✗ 0/31
✓ 29/31
✗ 0/31
✗ 21/31
✗ 0/31
Appendix
Table 2: Clean-room reconstruction of thirteen industrial software engines (the data behind Figure 2 a). Each cell is the hidden-verifier score of the best run under the protocol condition; scores are comparable within a row only; ✓ denotes a pass. GLM-5.3 flash cells use the runs described in Section 9 . Pass marks denote task-level thresholds, not necessarily perfect scores. CFD scores measure interface compliance and do not establish numerical-solver capability (Section 8 ); MES and PLM do not distinguish the evaluated models.
Model
Engines passed
Valid runs
Mean turns
Cost (USD)
Fable 5.1
13/13
13
20.8
95
Opus 5
13/13
9
74.1
76
Sonnet 5
10/13
13
60.1
35
Haiku 4.5
2/13
15
80.3
15
sol
9/13
13
30.2
n/a †
luna
6/13
13
35.3
n/a †
Appendix
Table 3: Per-model totals on the reconstruction suite. “Engines passed” is taken from Table 2 ; runs, turns and cost are over all valid runs of that model.
Task
Effort
Score
Turns
Wall (s)
Cost (USD)
ERP core
low
5/5
9
204
1.00
medium
5/5
18
411
1.85
high
5/5
20
672
2.71
xhigh
5/5
30
1070
4.37
max
5/5
31
1464
4.79
2D CAD kernel
low
31/31
65
1426
6.95
Appendix
Table 4: Reasoning-effort ablation (Opus 5, five levels). Scores remain unchanged. Cost, turns, and wall time increase monotonically for ERP but fluctuate at intermediate settings for CAD. Both tasks are saturated at all five settings, limiting conclusions about improvements in accuracy (Appendix B ).