Large language models (LLMs) are increasingly used for automated code generation, but generated programs can appear syntactically plausible while still failing execution-based correctness checks. Existing validation methods, such as testing and program analysis, remain essential but are often incomplete, costly, or applied only after generation. Model-derived uncertainty is therefore a natural early reliability signal. This paper studies the dilemma of overconfidence in code LLMs where incorrect programs are often generated with token-level confidence comparable to correct programs. We study this dilemma across four open-source code models and three execution-based benchmarks. Our analysis begins by investigating whether existing uncertainty metrics provide reliable proxies for execution correctness in code generation. We then characterize overconfidence at both global and local token levels, asking whether incorrect programs remain indistinguishable from correct ones under confidence and entropy summaries, including selective generation and the limits of instruction tuning. Finally, we evaluate whether common mitigation strategies reduce this failure mode. Our study yields four findings. First, existing uncertainty signals provide only partial and model-dependent evidence of execution failure. Second, overconfidence persists at both program and token levels, and uncertainty-based selection does not consistently improve accepted-set accuracy. Third, instruction tuning can increase certainty on failing generations without consistently improving correctness discrimination. Fourth, common mitigation techniques improve specific aspects of reliability but do not reliably resolve overconfident failure. Our exploratory latent analysis suggests that hidden representations may encode correctness-related signals that output confidence does not expose.
Figures & tables
Figure 1 . Motivation for uncertainty-aware code generation. Since generalization oracle cannot be known a priori, practical systems use model-derived uncertainty as a proxy reliability signal. The example illustrates the failure mode studied in this paper: an incorrect program can receive confidence comparable to a passing program.
Model
CodeLlama
DeepSeek-Coder
OpenCoder
QwenCoder
HumanEval+
34.1
70.1
78.0
71.3
MBPP+
46.0
66.1
71.2
65.9
BigCodeBench
6.5
20.6
23.8
8.2
Table 1 . Pass@1 Accuracy for Each Benchmark and Models
Family
Metrics
Signal captured
White-box token scores
MSP, Perplexity, Entropy, PMI, CPMI, R-Div, F-R Dist ( Darrin et al., 2022 ; Van der Poel et al., 2022 ; Takayama and Arase, 2019 ; Fomicheva et al., 2020 ; Vashurin et al., 2024 )
Measure token likelihood, surprise, distributional concentration, and prompt dependence.
White-box multi-sample scores
MCEnt, MSPS, RMI, Ens, MI ( Kuhn et al., 2023 ; Vashurin et al., 2024 )
Measure probability support and distributional disagreement across sampled candidates.
Black-box output similarity
LexSim, EigV, DegGlobal ( Lin et al., 2023 ; Fomicheva et al., 2020 )
Measure diversity or disagreement among sampled programs without using model probabilities.
Judge-based scores
Judge, M_Judge ( Gu et al., 2024 ; Vashurin et al., 2024 )
Measure LLM-estimated pass likelihood from the prompt and generated solution.
Table 2 . Uncertainty signals evaluated in the preliminary study.
Metric
HumanEval+
MBPP+
CodeLlama
DeepSeek
OpenCoder
QwenCoder
CodeLlama
DeepSeek
OpenCoder
QwenCoder
Prefill uncertainty
MSP
0.690
0.590
0.530
0.530
0.460
0.490
0.480
0.610
Entropy
0.700
0.590
0.530
0.540
0.470
0.600
0.500
0.540
Margin
0.690
0.590
0.520
0.520
0.470
0.450
0.480
0.620
Energy
0.590
0.420
0.540
0.560
0.530
0.540
0.510
0.520
Table 3 . AUC ( ↑ ) results of prefill-stage and generative uncertainty metrics.
Task
Model
Sample Size
Average Confidence
Bottom-10% Confidence
Correct n
Incorrect n
Correct
Incorrect
t-value
p-value
Correct
Incorrect
t-value
p-value
Mean ± Std.
Median
Mean ± Std.
Median
Mean ± Std.
Median
Mean ± Std.
Median
HumanEval+
CodeLlama
56
108
0.962±0.012
0.965
0.956±0.014
0.957
3.194
0.002
0.703±0.076
0.721
0.651±0.083
0.654
4.054
0.000
DeepSeek
115
49
0.951±0.021
0.957
0.943±0.021
0.946
2.233
0.028
0.642±0.090
0.658
0.607±0.075
0.610
2.548
0.012
OpenCoder
128
36
0.956±0.024
0.960
0.948±0.023
0.953
1.941
0.057
0.675±0.119
0.677
0.631±0.090
0.635
2.403
0.019
QwenCoder
117
47
0.947±0.025
0.953
0.935±0.036
0.943
2.047
0.045
0.619±0.101
0.640
0.559±0.116
0.570
3.106
0.003
Table 4 . Statistics on average and bottom-10% token confidence for passing and failing generated programs.
Figure 2 . Entropy distributions for OpenCoder on MBPP+.
Figure 3 . Selective accuracy under uncertainty-based ranking. At each coverage level, we retain the most certain generations according to average confidence or average entropy and report their execution accuracy. Lower coverage corresponds to accepting a smaller subset of generations.
Figure 4 . Uncertainty before and after instruction tuning on MBPP+. Each plot compares the distributions of passing and failing generations for the corresponding base and instruction-tuned model.
MBPP+
BigCodeBench
Model
Method
Avg. Conf.
Avg. Ent.
Bottom-10% Conf.
Top-10% Ent.
Avg. Conf.
Avg. Ent.
Bottom-10% Conf.
Top-10% Ent.
CodeLlama
Before
0.459 / -0.849
0.345 / -0.387
0.252 / -0.015
0.284 / -0.142
0.856 / -12.926
0.699 / -10.379
0.548 / -7.922
0.343 / -4.578
PS
0.248 / +0.001
0.247 / +0.005
0.248 / +0.004
0.246 / +0.008
0.061 / +0.001
0.061 / -0.000
0.060 / +0.017
0.062 / -0.002
Iso
0.253 / -0.020
0.251 / -0.012
0.257 / -0.035
0.255 / -0.025
0.062 / -0.002
0.062 / -0.001
0.062 / -0.001
0.061 / +0.002
DeepSeekCoder
Before
0.300 / -0.341
0.253 / -0.130
0.228 / -0.019
0.321 / -0.435
0.659 / -3.024
0.481 / -1.940
0.252 / -0.541
0.169 / -0.035
PS
0.224 / +0.001
0.222 / +0.007
0.223 / +0.004
0.222 / +0.008
0.164 / +0.000
0.163 / +0.001
0.163 / +0.003
0.163 / +0.001
Table 5 . Calibration results measured by Brier Score (BS; ↓ ) and Brier Skill Score (BSS; ↑ ). Each entry reports BS / BSS. For each model and metric, the lowest BS and highest BSS across Before, PS, and Iso are shown in bold.
Figure 5 . Sampling outcomes under EvalPlus+. (a) Remaining-unsolved tasks after the first k generations. (b) Distribution of how many successes each task attains within K=50 generations.