Runtime error healing lets a crashed program continue by generating code that repairs its live runtime state. Recent work shows that LLMs can generate such healing code, but it is evaluated only on small competition programs, and executing LLM-generated code inside a live process raises safety concerns that remain unaddressed. In this paper, we take LLM-based runtime healing toward practical use in real-world repositories. We first build HealBench, a benchmark of 265 runtime errors from 18 real-world repositories, each paired with a reference execution on the patched version. HealBench also provides a unified framework that lets LLM agents heal with cross-file context and live runtime state. We then design HealGuard, which requires healing code to be written in HealCore, an analyzable subset of Python, and uses static and dynamic taint analysis to check whether state changed by healing reaches operations protected by developers. We evaluate a dedicated healing method and three general coding agents with three backbone LLMs. The best setting resumes execution in 38.11% of instances and passes the target test in 28.68%, showing that existing agents can already heal a meaningful share of real repository-level crashes. However, among executions that pass, HealGuard flags 17.4% whose healing-changed state may reach a protected operation. On 684 controlled cases, HealGuard detects all unsafe cases, at the cost of a 68.42% false positive rate.
Figures & tables
Figure 1: Overview of HealBench construction and its runtime error healing framework.
Table 1: Error types and repositories in HealBench. Counts are shown in parentheses.
Method
Setting
GPT-5.6-Terra
DeepSeek-V4-Flash
GLM-5.2
PR
CR
TS
PR
CR
TS
PR
CR
TS
Healer
Vanilla
29.43
16.98
64.02
24.91
10.94
55.09
7.92
4.91
47.90
+ HealGuard
26.95
15.57
57.93
28.97
13.08
39.23
5.12
3.54
34.13
mini-SWE-agent
Vanilla
36.60
24.53
65.75
20.38
16.23
65.18
32.83
25.28
56.79
+ HealGuard
32.16
20.10
56.19
17.74
13.31
40.84
23.35
16.75
38.73
OpenHands
Vanilla
29.81
21.51
60.60
9.43
8.30
75.35
15.47
12.08
67.03
Table 2: Results of different methods on HealBench. PR, CR, and TS are in %; higher is better.
Figure 5
Figure 6: Static and dynamic decisions on simulated cases. Colors indicate safety labels, solid blocks indicate Blocked, and dashed outlines indicate Allowed. Dynamic analysis checks the cases allowed by Static.
Figure 7: An unresolved external API dependency produces an Unknown static verdict. Dynamic analysis returns Block because argument independence cannot be established.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Repository-level runtime error healing in matplotlib #23174. The agent uses cross-file context to locate the renderer in the parent Figure and recover program execution. Red marks the runtime error, and green marks the healing code.
Category
Excluded cases
Runtime Error Outside Repository Code
Test code
Runtime error occurs in test code.
Third-party dependency
Runtime error occurs in an external library.
Environment-related
FileNotFoundError , PermissionError , SSLError .
Import-related
ModuleNotFoundError , ImportError .
Version mismatch
Required library version does not match the environment.
Appendix
Table 3: Candidate errors excluded when building HealBench.
Item
Setting
Evaluated instances
265 instances from 18 repositories
Execution environment
Repository-specific Docker environment
Reference execution
Target test on the patched repository version
Healing interface
Runtime state inspection and healing submission through HTTP
Recorded budget
max_call_llm=100 , with adapter-specific counting
Current temperature
0.0 where explicitly set for Healer, mini-SWE-agent, and OpenHands
Appendix
Table 4: Recorded settings and current implementation defaults.
Method
Setting
GPT-5.6-Terra
DeepSeek-V4-Flash
GLM-5.2
Allowed
Blocked
PR
CR
TS
Allowed
Blocked
PR
CR
TS
Allowed
Blocked
PR
CR
TS
Healer
Vanilla
265
0
29.43
16.98
55.25
265
0
24.91
10.94
45.07
265
0
7.92
4.91
33.04
Static independently
208
57
28.37
16.83
57.67
108
157
29.63
13.89
39.92
258
7
5.43
3.88
35.22
Dynamic independently
219
46
28.31
15.53
55.24
200
65
29.00
12.50
40.40
258
7
6.59
3.88
32.31
+ HealGuard
167
98
26.95
15.57
57.93
107
158
28.97
13.08
39.23
254
11
5.12
3.54
34.13
mini-SWE-agent
Vanilla
265
0
36.60
24.53
55.88
265
0
20.38
16.23
43.02
265
0
32.83
25.28
40.67
Appendix
Table 5: Healing outcomes under independent and sequential checks on saved executions. Allowed and Blocked are counts. PR, CR, and TS are percentages.
Method
GPT-5.6-Terra
DeepSeek-V4-Flash
GLM-5.2
PR
CR
TS
In
Out
$
PR
CR
TS
In
Out
$
PR
CR
TS
In
Out
$
Healer
29.43
16.98
55.25
0.10
0.17
0.20
24.91
10.94
45.07
0.42
0.52
0.03
7.92
4.91
33.04
0.53
0.61
0.74
mini-SWE-agent
36.60
24.53
55.88
0.31
1.29
0.63
20.38
16.23
43.02
0.47
8.61
0.03
32.83
25.28
40.67
1.75
9.80
2.50
OpenHands
29.81
21.51
43.42
0.12
0.66
0.25
9.43
8.30
63.66
0.17
2.84
0.01
15.47
12.08
57.15
0.82
6.40
1.17
Codex
38.11
28.68
41.80
0.16
1.75
0.34
20.00
15.47
41.47
0.40
7.95
0.03
26.42
20.75
43.07
1.42
7.62
2.02
Appendix
Table 6: Healing Results without HealGuard. In, Out, and $ denote mean input tokens (M), output tokens (K), and cost (USD). PR, CR, and TS are percentages. Bold marks the highest PR, CR, and TS for each backbone LLM.
Type
Instances
GPT-5.6-Terra
DeepSeek-V4-Flash
GLM-5.2
H
M
O
C
H
M
O
C
H
M
O
C
TypeError
90
14.44
24.44
23.33
27.78
7.78
16.67
6.67
16.67
6.67
25.56
11.11
20.00
AttributeError
65
23.08
26.15
23.08
32.31
12.31
13.85
10.77
13.85
3.08
21.54
12.31
15.38
ValueError
60
13.33
25.00
15.00
21.67
15.00
15.00
10.00
15.00
3.33
28.33
15.00
26.67
IndexError
24
8.33
16.67
12.50
25.00
8.33
16.67
0.00
8.33
4.17
20.83
8.33
29.17
KeyError
9
33.33
33.33
44.44
55.56
22.22
33.33
22.22
22.22
0.00
33.33
22.22
11.11
Appendix
Table 7: Correct Rate of Healing by Error Type. Values are percentages of the selected instances in each error category, including missing runs in the denominator. Overall is weighted by instance count. H, M, O, and C denote Healer, mini-SWE-agent, OpenHands, and Codex.
Repository
Instances
GPT-5.6-Terra
DeepSeek-V4-Flash
GLM-5.2
H
M
O
C
H
M
O
C
H
M
O
C
pandas
84
3.57
10.71
9.52
17.86
3.57
2.38
0.00
5.95
4.76
7.14
2.38
14.29
numpy
80
7.50
17.50
12.50
11.25
6.25
11.25
0.00
3.75
2.50
31.25
7.50
20.00
pillow
31
25.81
41.94
25.81
54.84
29.03
19.35
12.90
29.03
0.00
0.00
0.00
0.00
matplotlib
14
50.00
28.57
57.14
57.14
7.14
0.00
7.14
21.43
14.29
42.86
14.29
35.71
scikit-learn
12
25.00
33.33
33.33
41.67
16.67
41.67
25.00
41.67
8.33
50.00
41.67
41.67
Appendix
Table 8: Correct Rate of Healing by Repository. Values are percentages of the selected instances in each repository, including missing runs in the denominator. Overall is weighted by instance count. H, M, O, and C denote Healer, mini-SWE-agent, OpenHands, and Codex.
Figure 9: Action shares during runtime error healing without HealGuard. Panels show GPT-5.6-Terra, DeepSeek-V4-Flash, and GLM-5.2. Explore denotes repository exploration, Inspect State denotes runtime state inspection, Heal denotes healing submission, and Reraise denotes recorded re-raising.
Figure 10: Recorded crash counts per execution without HealGuard. Panels show GPT-5.6-Terra, DeepSeek-V4-Flash, and GLM-5.2. Boxes show the interquartile range and median, with observations overlaid. The vertical axis ends at 75, so larger observations are not shown, including the maximum count of 130.
Setting
N
TP
FP
FN
TN
Accuracy
Precision
Recall
F1
Static independently
684
177
12
165
330
74.12
93.65
51.75
66.67
Dynamic independently
684
342
234
0
108
65.79
59.38
100.00
74.51
Dynamic after static allow
495
165
222
0
108
55.15
42.64
100.00
59.78
HealGuard
684
342
234
0
108
65.79
59.38
100.00
74.51
Appendix
Table 9: Independent and sequential decisions on simulated cases. Labels are provisional. Accuracy, precision, recall, and F1 are percentages.
Setting
N
TP
FP
FN
TN
Accuracy
Precision
Recall
F1
Static independently
30
24
2
2
2
86.67
92.31
92.31
92.31
Dynamic independently
30
26
2
0
2
93.33
92.86
100.00
96.30
Dynamic after static allow
4
2
2
0
0
50.00
50.00
100.00
66.67
HealGuard
30
26
4
0
0
86.67
86.67
100.00
92.86
Appendix
Table 10: Human validation on 30 selected cases using manual review labels. Accuracy, precision, recall, and F1 are percentages.
Setting
N
TP
FP
FN
TN
Accuracy
Precision
Recall
F1
Static independently
265
7
16
10
232
90.19
30.43
41.18
35.00
Dynamic independently
265
17
35
0
213
86.79
32.69
100.00
49.28
Dynamic after static allow
242
10
33
0
199
86.36
23.26
100.00
37.74
HealGuard
265
17
49
0
199
81.51
25.76
100.00
40.96
Appendix
Table 11: Human validation on 265 mini-SWE-agent cases with GPT-5.6-Terra using manual review labels. Accuracy, precision, recall, and F1 are percentages.
Figure 11: Examples of unsafe healing. Left shows unintended file deletion. Right shows unauthorized subprocess execution.
Figure 12: Unnecessary cookie output during runtime error healing. The healing code copies cookie objects but also prints their names and values. The reference patch iterates over cookie objects directly without adding this output.
Figure 13: Wrong and correct healing for a missing function and an invalid input.
We measure the rate at which code RL environments accept incorrect solutions as correct. On a 49-task sample of SWE-bench Verified, 28.5% of tasks have test suites weak enough that a Docker-verified incorrect patch passes them. On 20 R2E-Gym tasks across 6 repositories, the same pipeline at single-shot exploit generation yields 25.0%. A random-effects meta-analysis over 134 frontier model submissions to SWE-bench Verified finds, within the same human-rated difficulty stratum, model Pass@1 is +14.14 percentage points higher on flagged-hackable tasks than on robust ones (95% CI [+11.80, +16.48]; one-sided p < 10^-6; I^2 = 0%; 123 of 134 models positive). We then describe a procedure for hardening the broken tasks. An inline LLM judge with a Docker gold-sanity gate runs each generated test against the gold solution before the judge is consulted. On the 11 broken tasks in the audit, the gate flags 65 of 105 decisive LLM-generated tests as failing on the gold patch itself, a 61.9% per-augmentation defect rate the LLM judge alone misses. With diversity-biased retry, the loop converges 9 of 11 tasks to a gated upgrade.
An LLM agent acts on the world by emitting actions: shell commands to run, edits to apply. A wrong action does not always fail loudly; it can fail silently, producing a plausible but incorrect effect that raises no error. We argue that a cheap deterministic check, run before an action takes effect, is an effective and underused form of agent oversight, and we study it across two action modalities in one framework. The idea is to fix an action's correct effect by construction, before any executor runs, so that silent failure is measured directly and the verifier may abstain rather than guess. For shell commands, a static verifier over 9930 commands and 482 tools catches 95.8% of invalid commands at a 10.0% false-positive rate. Its syntax and binary checks are oracle-exact, giving zero false positives while catching half of all errors; the flag check is bounded only by help-text coverage and accounts for every false positive. For code edits, a benchmark of 640 edits over 224 files isolating the apply step exposes a sharp split. Content-anchored formats such as search/replace and diff fail cleanly, whereas location-anchored formats fail silently: line numbers corrupt 99.1% of files under a one-line shift, and function-name edits hit the wrong function 12.7% of the time. In both settings a refuse-when-unsure policy turns silent failures into recoverable ones at a tunable cost in applicability: selective grounding reaches 0.958 recall at 7.0% false positives, and an anchor-and-verify applier records one silent misapplication in 8320 trials (0.01%). We release both benchmarks, the verifiers, and the guards.
Large language models (LLMs) are increasingly used in coding tasks, but their ability to reason about code execution remains unclear. Existing repository-level QA benchmarks mainly evaluate static code understanding and often rely on LLM-based evaluation, while execution-reasoning benchmarks are mostly limited to snippets or functions. We introduce SWE-Flux, a repository-level benchmark for dynamic execution reasoning containing 480 execution-grounded instances across 12 real Python repositories, with gold answers automatically harvested from instrumented test executions rather than written manually or judged by LLMs. The benchmark covers singletest and multi-test questions over control flow, loops, program state, dataflow, exceptions, and program invariants. Evaluating five LLMs shows that this task remains challenging. The best model achieves only 37% accuracy. Models perform better on localized behavior such as invariants, intra-procedural control flow, exceptions, and simple loops, but struggle with dataflow, inter-procedural execution, precise state reasoning, and suite-level aggregation. Finally, we show that the oracle-harvesting pipeline can generate fresh benchmark variants using input perturbation. It successfully harvests valid variants for almost 90% of the selected instances, and the resulting variants are substantially more challenging for the evaluated models.
Hamed Taherkhani, Mohammad Abdollahi, Melika Sepidband +3
York University Toronto, Canada · The University of Texas at Dallas Texas, USA