Runtime error healing lets a crashed program continue by generating code that repairs its live runtime state. Recent work shows that LLMs can generate such healing code, but it is evaluated only on small competition programs, and executing LLM-generated code inside a live process raises safety concerns that remain unaddressed. In this paper, we take LLM-based runtime healing toward practical use in real-world repositories. We first build HealBench, a benchmark of 265 runtime errors from 18 real-world repositories, each paired with a reference execution on the patched version. HealBench also provides a unified framework that lets LLM agents heal with cross-file context and live runtime state. We then design HealGuard, which requires healing code to be written in HealCore, an analyzable subset of Python, and uses static and dynamic taint analysis to check whether state changed by healing reaches operations protected by developers. We evaluate a dedicated healing method and three general coding agents with three backbone LLMs. The best setting resumes execution in 38.11% of instances and passes the target test in 28.68%, showing that existing agents can already heal a meaningful share of real repository-level crashes. However, among executions that pass, HealGuard flags 17.4% whose healing-changed state may reach a protected operation. On 684 controlled cases, HealGuard detects all unsafe cases, at the cost of a 68.42% false positive rate.
Figures & tables
Figure 1: Overview of HealBench construction and its runtime error healing framework.
Table 1: Error types and repositories in HealBench. Counts are shown in parentheses.
Method
Setting
GPT-5.6-Terra
DeepSeek-V4-Flash
GLM-5.2
PR
CR
TS
PR
CR
TS
PR
CR
TS
Healer
Vanilla
29.43
16.98
64.02
24.91
10.94
55.09
7.92
4.91
47.90
+ HealGuard
26.95
15.57
57.93
28.97
13.08
39.23
5.12
3.54
34.13
mini-SWE-agent
Vanilla
36.60
24.53
65.75
20.38
16.23
65.18
32.83
25.28
56.79
+ HealGuard
32.16
20.10
56.19
17.74
13.31
40.84
23.35
16.75
38.73
OpenHands
Vanilla
29.81
21.51
60.60
9.43
8.30
75.35
15.47
12.08
67.03
Table 2: Results of different methods on HealBench. PR, CR, and TS are in %; higher is better.
Figure 5
Figure 6: Static and dynamic decisions on simulated cases. Colors indicate safety labels, solid blocks indicate Blocked, and dashed outlines indicate Allowed. Dynamic analysis checks the cases allowed by Static.
Figure 7: An unresolved external API dependency produces an Unknown static verdict. Dynamic analysis returns Block because argument independence cannot be established.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Repository-level runtime error healing in matplotlib #23174. The agent uses cross-file context to locate the renderer in the parent Figure and recover program execution. Red marks the runtime error, and green marks the healing code.
Category
Excluded cases
Runtime Error Outside Repository Code
Test code
Runtime error occurs in test code.
Third-party dependency
Runtime error occurs in an external library.
Environment-related
FileNotFoundError , PermissionError , SSLError .
Import-related
ModuleNotFoundError , ImportError .
Version mismatch
Required library version does not match the environment.
Appendix
Table 3: Candidate errors excluded when building HealBench.
Item
Setting
Evaluated instances
265 instances from 18 repositories
Execution environment
Repository-specific Docker environment
Reference execution
Target test on the patched repository version
Healing interface
Runtime state inspection and healing submission through HTTP
Recorded budget
max_call_llm=100 , with adapter-specific counting
Current temperature
0.0 where explicitly set for Healer, mini-SWE-agent, and OpenHands
Appendix
Table 4: Recorded settings and current implementation defaults.
Method
Setting
GPT-5.6-Terra
DeepSeek-V4-Flash
GLM-5.2
Allowed
Blocked
PR
CR
TS
Allowed
Blocked
PR
CR
TS
Allowed
Blocked
PR
CR
TS
Healer
Vanilla
265
0
29.43
16.98
55.25
265
0
24.91
10.94
45.07
265
0
7.92
4.91
33.04
Static independently
208
57
28.37
16.83
57.67
108
157
29.63
13.89
39.92
258
7
5.43
3.88
35.22
Dynamic independently
219
46
28.31
15.53
55.24
200
65
29.00
12.50
40.40
258
7
6.59
3.88
32.31
+ HealGuard
167
98
26.95
15.57
57.93
107
158
28.97
13.08
39.23
254
11
5.12
3.54
34.13
mini-SWE-agent
Vanilla
265
0
36.60
24.53
55.88
265
0
20.38
16.23
43.02
265
0
32.83
25.28
40.67
Appendix
Table 5: Healing outcomes under independent and sequential checks on saved executions. Allowed and Blocked are counts. PR, CR, and TS are percentages.
Method
GPT-5.6-Terra
DeepSeek-V4-Flash
GLM-5.2
PR
CR
TS
In
Out
$
PR
CR
TS
In
Out
$
PR
CR
TS
In
Out
$
Healer
29.43
16.98
55.25
0.10
0.17
0.20
24.91
10.94
45.07
0.42
0.52
0.03
7.92
4.91
33.04
0.53
0.61
0.74
mini-SWE-agent
36.60
24.53
55.88
0.31
1.29
0.63
20.38
16.23
43.02
0.47
8.61
0.03
32.83
25.28
40.67
1.75
9.80
2.50
OpenHands
29.81
21.51
43.42
0.12
0.66
0.25
9.43
8.30
63.66
0.17
2.84
0.01
15.47
12.08
57.15
0.82
6.40
1.17
Codex
38.11
28.68
41.80
0.16
1.75
0.34
20.00
15.47
41.47
0.40
7.95
0.03
26.42
20.75
43.07
1.42
7.62
2.02
Appendix
Table 6: Healing Results without HealGuard. In, Out, and $ denote mean input tokens (M), output tokens (K), and cost (USD). PR, CR, and TS are percentages. Bold marks the highest PR, CR, and TS for each backbone LLM.
Type
Instances
GPT-5.6-Terra
DeepSeek-V4-Flash
GLM-5.2
H
M
O
C
H
M
O
C
H
M
O
C
TypeError
90
14.44
24.44
23.33
27.78
7.78
16.67
6.67
16.67
6.67
25.56
11.11
20.00
AttributeError
65
23.08
26.15
23.08
32.31
12.31
13.85
10.77
13.85
3.08
21.54
12.31
15.38
ValueError
60
13.33
25.00
15.00
21.67
15.00
15.00
10.00
15.00
3.33
28.33
15.00
26.67
IndexError
24
8.33
16.67
12.50
25.00
8.33
16.67
0.00
8.33
4.17
20.83
8.33
29.17
KeyError
9
33.33
33.33
44.44
55.56
22.22
33.33
22.22
22.22
0.00
33.33
22.22
11.11
Appendix
Table 7: Correct Rate of Healing by Error Type. Values are percentages of the selected instances in each error category, including missing runs in the denominator. Overall is weighted by instance count. H, M, O, and C denote Healer, mini-SWE-agent, OpenHands, and Codex.
Repository
Instances
GPT-5.6-Terra
DeepSeek-V4-Flash
GLM-5.2
H
M
O
C
H
M
O
C
H
M
O
C
pandas
84
3.57
10.71
9.52
17.86
3.57
2.38
0.00
5.95
4.76
7.14
2.38
14.29
numpy
80
7.50
17.50
12.50
11.25
6.25
11.25
0.00
3.75
2.50
31.25
7.50
20.00
pillow
31
25.81
41.94
25.81
54.84
29.03
19.35
12.90
29.03
0.00
0.00
0.00
0.00
matplotlib
14
50.00
28.57
57.14
57.14
7.14
0.00
7.14
21.43
14.29
42.86
14.29
35.71
scikit-learn
12
25.00
33.33
33.33
41.67
16.67
41.67
25.00
41.67
8.33
50.00
41.67
41.67
Appendix
Table 8: Correct Rate of Healing by Repository. Values are percentages of the selected instances in each repository, including missing runs in the denominator. Overall is weighted by instance count. H, M, O, and C denote Healer, mini-SWE-agent, OpenHands, and Codex.
Figure 9: Action shares during runtime error healing without HealGuard. Panels show GPT-5.6-Terra, DeepSeek-V4-Flash, and GLM-5.2. Explore denotes repository exploration, Inspect State denotes runtime state inspection, Heal denotes healing submission, and Reraise denotes recorded re-raising.
Figure 10: Recorded crash counts per execution without HealGuard. Panels show GPT-5.6-Terra, DeepSeek-V4-Flash, and GLM-5.2. Boxes show the interquartile range and median, with observations overlaid. The vertical axis ends at 75, so larger observations are not shown, including the maximum count of 130.
Setting
N
TP
FP
FN
TN
Accuracy
Precision
Recall
F1
Static independently
684
177
12
165
330
74.12
93.65
51.75
66.67
Dynamic independently
684
342
234
0
108
65.79
59.38
100.00
74.51
Dynamic after static allow
495
165
222
0
108
55.15
42.64
100.00
59.78
HealGuard
684
342
234
0
108
65.79
59.38
100.00
74.51
Appendix
Table 9: Independent and sequential decisions on simulated cases. Labels are provisional. Accuracy, precision, recall, and F1 are percentages.
Setting
N
TP
FP
FN
TN
Accuracy
Precision
Recall
F1
Static independently
30
24
2
2
2
86.67
92.31
92.31
92.31
Dynamic independently
30
26
2
0
2
93.33
92.86
100.00
96.30
Dynamic after static allow
4
2
2
0
0
50.00
50.00
100.00
66.67
HealGuard
30
26
4
0
0
86.67
86.67
100.00
92.86
Appendix
Table 10: Human validation on 30 selected cases using manual review labels. Accuracy, precision, recall, and F1 are percentages.
Setting
N
TP
FP
FN
TN
Accuracy
Precision
Recall
F1
Static independently
265
7
16
10
232
90.19
30.43
41.18
35.00
Dynamic independently
265
17
35
0
213
86.79
32.69
100.00
49.28
Dynamic after static allow
242
10
33
0
199
86.36
23.26
100.00
37.74
HealGuard
265
17
49
0
199
81.51
25.76
100.00
40.96
Appendix
Table 11: Human validation on 265 mini-SWE-agent cases with GPT-5.6-Terra using manual review labels. Accuracy, precision, recall, and F1 are percentages.
Figure 11: Examples of unsafe healing. Left shows unintended file deletion. Right shows unauthorized subprocess execution.
Figure 12: Unnecessary cookie output during runtime error healing. The healing code copies cookie objects but also prints their names and values. The reference patch iterates over cookie objects directly without adding this output.
Figure 13: Wrong and correct healing for a missing function and an invalid input.