Large language models (LLMs) have significantly advanced automated program repair (APR), yet existing evaluations remain largely result-centric and provide limited insight into hallucination during repair. In APR, hallucination may arise not only in final patches but also in the intermediate artifacts that guide patch generation. To address this gap, we perform a multi-layered analysis of hallucination throughout the APR process. Specifically, we characterize hallucination as the production of patches or intermediate artifacts that are not faithfully grounded in the available repair evidence. We examine repair hallucination in final patches and understanding hallucination in intermediate artifacts through three tasks, namely triggering testcase identification, line coverage prediction, and additional testcase generation. We then evaluate three representative LLMs on 832 Defects4J bugs through automatic evaluation and manual analysis. Our results show that both repair and understanding hallucinations remain prevalent. Across models and settings, only 21.0%-55.9% of generated patches pass the developer-written test suite. Moreover, although more accurate intermediate artifacts are generally associated with successful repairs, this relationship does not always hold. Manual analysis of 812 sampled repairs identifies repair hallucinations in 72.7% of cases, including patches that pass all available tests; incorrect causal localization and incorrect repair strategies account for 45.9% and 18.5% of these hallucinations, respectively. Meanwhile, models frequently misidentify triggering testcases, mispredict line coverage involving branching control flow, and generate additional testcases with missing bug-triggering conditions or incorrect expected behavior.
Figures & tables
Figure 1 : Conceptual overview of our study on hallucination in LLM-based APR.
Figure 2 : Overall workflow of our study.
Figure 3 : Evaluation process for triggering testcase identification.
Figure 4 : Evaluation process for line coverage prediction.
Figure 5 : Evaluation process for additional testcase generation.
Figure 6 : Human annotation workflow of our study.
Figure 7 : Distribution of relevant test methods and triggering test methods per bug.
Task
Model
Population Size N
Sample Size n
Triggering Testcase Identification
Claude
1716
94
DeepSeek
1716
94
GPT-5
1716
94
Line Coverage Prediction
Claude
832
87
DeepSeek
832
87
GPT-5
832
87
Table 1 : Population size and sampled cases for human annotation.
Model
Baseline
Triggering Testcase Identification
Line Coverage Prediction
Additional Testcase Generation
Pass
Fail
Pass
Fail
Pass
Fail
Pass
Fail
Claude
178 (21.4%)
654
222 (26.7%)
610
308 (37.0%)
524
293 (35.2%)
539
DeepSeek
175 (21.0%)
657
314 (37.7%)
518
403 (48.4%)
429
310 (37.3%)
522
GPT-5
219 (26.3%)
613
378 (45.4%)
454
465 (55.9%)
367
387 (46.5%)
445
Table 2 : Number of plausible patches across the baseline setting and three tasks.
Figure 8 : Overlap of plausible patches across the baseline setting and the three task settings.
Model
Outcome
#Bugs
Precision
Recall
F1
Claude
Pass
222
0.478
0.564
0.482
Fail
610
0.194
0.273
0.195
DeepSeek
Pass
314
0.444
0.502
0.445
Fail
518
0.178
0.245
0.180
GPT-5
Pass
378
0.619
0.635
0.607
Fail
454
0.232
0.240
0.214
Table 3 : Precision, recall, and F1 scores for triggering testcase identification.
Figure 9 : F1-score distributions for the triggering testcase identification task. Dots indicate mean F1 scores, and error bars indicate 95% confidence intervals.
Model
Outcome
# Bugs
Buggy
Model-Patched
P
R
F1
P
R
F1
Claude
Pass
308
0.706
0.881
0.751
0.745
0.885
0.779
Not Pass
416
0.664
0.860
0.714
0.647
0.845
0.697
Uncompilable/Timeout
108
0.671
0.864
0.716
–
–
–
DeepSeek
Pass
403
0.888
0.881
0.866
0.955
0.911
0.922
Not Pass
296
0.783
0.799
0.767
0.757
0.809
0.757
Table 4 : Precision (P), recall (R), and F1 scores for line coverage prediction on the buggy and model-patched programs.
Figure 10 : F1-score distributions for line coverage prediction on the buggy and model-patched programs. Dots indicate mean F1 scores, and error bars indicate 95% confidence intervals.
Figure 11 : Distributions of generated triggering testcase validity and repair outcome across models. Rectangle areas are proportional to the corresponding percentages among all generated testcases for each model.
Category
Tpatch
Triggering Testcase
Line Coverage
Additional Testcase
Identification
Prediction
Generation
Non-existent Symbol Reference
Uncompilable
4 / 5 / 6
8 / 12 / 0
16 / 8 / 6
Missing Import or Dependency
0 / 3 / 0
2 / 1 / 0
5 / 0 / 0
Syntax or Structural Error
1 / 2 / 2
1 / 1 / 1
0 / 4 / 2
Incorrect Repair Strategy
Pass , Not Pass
15 / 13 / 12
11 / 11 / 18
7 / 12 / 10
Incorrect Causal Localization
40 / 44 / 36
28 / 31 / 23
17 / 27 / 25
Table 5 : Distribution of repair hallucination categories across the three APR artifact-generation tasks. Each cell reports counts in the order of Claude / DeepSeek / GPT-5.
Figure 12 : Example of incorrect API usage in the patch generated by Claude for Closure-7 in the additional testcase generation task.
Figure 13 : Example of partial semantic repair in the patch generated by Claude for JacksonDatabind-3 in the line coverage prediction task.
Figure 14 : Example of extraneous conditional logic in the patch generated by GPT-5 for Closure-19 in the triggering testcase identification task.
Figure 15 : Example of incorrect boundary check in the patch generated by Claude for Closure-113 in the line coverage prediction task.
Figure 16 : Example of test-specific heuristic overfitting in the patch generated by Claude for Chart-17 in the additional testcase generation task.
Category
Condition
# Bugs
Correct Identification
T^b=Tb
191 / 276 / 340
Partial Identification
∅⊂T^b⊂Tb
40 / 40 / 72
Trigger Omission
T^b=∅
54 / 71 / 14
Spurious Identification
(Tb⊆T^b)∧(T^b∖Tb=∅)
151 / 168 / 102
Mixed Misidentification
(Tb∖T^b=∅)∧(T^b∖Tb=∅)
396 / 277 / 304
Total
–
832 / 832 / 832
Table 6 : Distribution of understanding hallucinations in the Triggering Testcase Identification task. Counts in the # Bugs column are reported in the order of Claude / DeepSeek / GPT-5.
Category
Ab
Aˉb
# Bugs
Trigger-insensitive success
Pass
Pass
71 / 91 / 113
Trigger-dependent success
Pass
Fail
17 / 47 / 59
Trigger-absent-only success
Fail
Pass
50 / 40 / 36
Universally failed bugs
Fail
Fail
266 / 226 / 196
Total
–
–
404 / 404 / 404
Table 7 : Relationship between trigger testcase artifact availability and repair success in the Triggering Testcase Identification task. Restricted to bugs that have both a trigger-present ( Ab ) and a trigger-absent ( Aˉb ) variant, so that both conditions are actually tested. Counts in the # Bugs column are reported in the order of Claude / DeepSeek / GPT-5.
Category
Code Structure Example
# Predictions
Branch
\sccode if, \sccode else, \sccode switch-case, …
25 / 11 / 3
Exception Flow
\sccode try, \sccode catch, …
3 / 2 / 1
Return Statement
\sccode return
2 / 0 / 0
Assignment Statement
Variable assignments and state updates
0 / 0 / 2
Total
–
28 / 11 / 5
Table 8 : Distribution of code-structure-based understanding hallucinations among low-scoring line coverage predictions. Counts in the # Predictions column are reported in the order of Claude / DeepSeek / GPT-5. The categories are not mutually exclusive.
Category
Execution-based Condition
# Cases
Non-existent Symbol Reference
Tbuggy=\textscUncompilable
4 / 4 / 5
Missing Import or Dependency
Tbuggy=\textscUncompilable
3 / 12 / 7
Duplicated Declaration
Tbuggy=\textscUncompilable
5 / 2 / 3
Missing Test Method Wrapper
Tbuggy=\textscUncompilable
0 / 0 / 7
Malformed Escape Sequences
Tbuggy=\textscUncompilable
0 / 4 / 1
Test-framework/ Language Version Mismatch
Tbuggy=\textscUncompilable
5 / 1 / 1
Table 9 : Distribution of understanding hallucinations in the additional testcase generation task. Counts in the # Cases column are reported in the order of Claude / DeepSeek / GPT-5.
Figure 17 : Example of malformed escape sequences in the testcase generated for Closure-161.
Figure 18 : Example of incomplete bug-triggering condition in the testcase generated by Claude for Mockito-22 in the additional testcase generation task.
Contributing Factor
Hallucination Types
Interpretation
Insufficient Project-Context Grounding
Non-existent Symbol Reference Missing Import or Dependency Duplicated Declaration Test-framework/ Language Version Mismatch Incorrect API Usage
The model fails to ground its generated patch or testcase in the concrete project context, including available symbols, imports, dependencies, test framework conventions, and API semantics.
The model captures some bug-related signal, but cannot precisely distinguish causal testcases or causal code locations from merely related but non-causal context.
Unstable Execution-Path Reasoning
Branch Exception Flow
The model has difficulty predicting which branch, guard, or exception-handling path is actually exercised by the triggering testcase.
Table 10 : Potential factors contributing to hallucinations in LLM-based APR, synthesized from repair hallucinations and intermediate understanding hallucinations.