Large language models (LLMs) are increasingly used for automated code generation, but generated programs can appear syntactically plausible while still failing execution-based correctness checks. Existing validation methods, such as testing and program analysis, remain essential but are often incomplete, costly, or applied only after generation. Model-derived uncertainty is therefore a natural early reliability signal. This paper studies the dilemma of overconfidence in code LLMs where incorrect programs are often generated with token-level confidence comparable to correct programs. We study this dilemma across four open-source code models and three execution-based benchmarks. Our analysis begins by investigating whether existing uncertainty metrics provide reliable proxies for execution correctness in code generation. We then characterize overconfidence at both global and local token levels, asking whether incorrect programs remain indistinguishable from correct ones under confidence and entropy summaries, including selective generation and the limits of instruction tuning. Finally, we evaluate whether common mitigation strategies reduce this failure mode. Our study yields four findings. First, existing uncertainty signals provide only partial and model-dependent evidence of execution failure. Second, overconfidence persists at both program and token levels, and uncertainty-based selection does not consistently improve accepted-set accuracy. Third, instruction tuning can increase certainty on failing generations without consistently improving correctness discrimination. Fourth, common mitigation techniques improve specific aspects of reliability but do not reliably resolve overconfident failure. Our exploratory latent analysis suggests that hidden representations may encode correctness-related signals that output confidence does not expose.
Figures & tables
Figure 1 . Motivation for uncertainty-aware code generation. Since generalization oracle cannot be known a priori, practical systems use model-derived uncertainty as a proxy reliability signal. The example illustrates the failure mode studied in this paper: an incorrect program can receive confidence comparable to a passing program.
Model
CodeLlama
DeepSeek-Coder
OpenCoder
QwenCoder
HumanEval+
34.1
70.1
78.0
71.3
MBPP+
46.0
66.1
71.2
65.9
BigCodeBench
6.5
20.6
23.8
8.2
Table 1 . Pass@1 Accuracy for Each Benchmark and Models
Family
Metrics
Signal captured
White-box token scores
MSP, Perplexity, Entropy, PMI, CPMI, R-Div, F-R Dist ( Darrin et al., 2022 ; Van der Poel et al., 2022 ; Takayama and Arase, 2019 ; Fomicheva et al., 2020 ; Vashurin et al., 2024 )
Measure token likelihood, surprise, distributional concentration, and prompt dependence.
White-box multi-sample scores
MCEnt, MSPS, RMI, Ens, MI ( Kuhn et al., 2023 ; Vashurin et al., 2024 )
Measure probability support and distributional disagreement across sampled candidates.
Black-box output similarity
LexSim, EigV, DegGlobal ( Lin et al., 2023 ; Fomicheva et al., 2020 )
Measure diversity or disagreement among sampled programs without using model probabilities.
Judge-based scores
Judge, M_Judge ( Gu et al., 2024 ; Vashurin et al., 2024 )
Measure LLM-estimated pass likelihood from the prompt and generated solution.
Table 2 . Uncertainty signals evaluated in the preliminary study.
Metric
HumanEval+
MBPP+
CodeLlama
DeepSeek
OpenCoder
QwenCoder
CodeLlama
DeepSeek
OpenCoder
QwenCoder
Prefill uncertainty
MSP
0.690
0.590
0.530
0.530
0.460
0.490
0.480
0.610
Entropy
0.700
0.590
0.530
0.540
0.470
0.600
0.500
0.540
Margin
0.690
0.590
0.520
0.520
0.470
0.450
0.480
0.620
Energy
0.590
0.420
0.540
0.560
0.530
0.540
0.510
0.520
Table 3 . AUC ( ↑ ) results of prefill-stage and generative uncertainty metrics.
Task
Model
Sample Size
Average Confidence
Bottom-10% Confidence
Correct n
Incorrect n
Correct
Incorrect
t-value
p-value
Correct
Incorrect
t-value
p-value
Mean ± Std.
Median
Mean ± Std.
Median
Mean ± Std.
Median
Mean ± Std.
Median
HumanEval+
CodeLlama
56
108
0.962±0.012
0.965
0.956±0.014
0.957
3.194
0.002
0.703±0.076
0.721
0.651±0.083
0.654
4.054
0.000
DeepSeek
115
49
0.951±0.021
0.957
0.943±0.021
0.946
2.233
0.028
0.642±0.090
0.658
0.607±0.075
0.610
2.548
0.012
OpenCoder
128
36
0.956±0.024
0.960
0.948±0.023
0.953
1.941
0.057
0.675±0.119
0.677
0.631±0.090
0.635
2.403
0.019
QwenCoder
117
47
0.947±0.025
0.953
0.935±0.036
0.943
2.047
0.045
0.619±0.101
0.640
0.559±0.116
0.570
3.106
0.003
Table 4 . Statistics on average and bottom-10% token confidence for passing and failing generated programs.
Figure 2 . Entropy distributions for OpenCoder on MBPP+.
Figure 3 . Selective accuracy under uncertainty-based ranking. At each coverage level, we retain the most certain generations according to average confidence or average entropy and report their execution accuracy. Lower coverage corresponds to accepting a smaller subset of generations.
Figure 4 . Uncertainty before and after instruction tuning on MBPP+. Each plot compares the distributions of passing and failing generations for the corresponding base and instruction-tuned model.
MBPP+
BigCodeBench
Model
Method
Avg. Conf.
Avg. Ent.
Bottom-10% Conf.
Top-10% Ent.
Avg. Conf.
Avg. Ent.
Bottom-10% Conf.
Top-10% Ent.
CodeLlama
Before
0.459 / -0.849
0.345 / -0.387
0.252 / -0.015
0.284 / -0.142
0.856 / -12.926
0.699 / -10.379
0.548 / -7.922
0.343 / -4.578
PS
0.248 / +0.001
0.247 / +0.005
0.248 / +0.004
0.246 / +0.008
0.061 / +0.001
0.061 / -0.000
0.060 / +0.017
0.062 / -0.002
Iso
0.253 / -0.020
0.251 / -0.012
0.257 / -0.035
0.255 / -0.025
0.062 / -0.002
0.062 / -0.001
0.062 / -0.001
0.061 / +0.002
DeepSeekCoder
Before
0.300 / -0.341
0.253 / -0.130
0.228 / -0.019
0.321 / -0.435
0.659 / -3.024
0.481 / -1.940
0.252 / -0.541
0.169 / -0.035
PS
0.224 / +0.001
0.222 / +0.007
0.223 / +0.004
0.222 / +0.008
0.164 / +0.000
0.163 / +0.001
0.163 / +0.003
0.163 / +0.001
Table 5 . Calibration results measured by Brier Score (BS; ↓ ) and Brier Skill Score (BSS; ↑ ). Each entry reports BS / BSS. For each model and metric, the lowest BS and highest BSS across Before, PS, and Iso are shown in bold.
Figure 5 . Sampling outcomes under EvalPlus+. (a) Remaining-unsolved tasks after the first k generations. (b) Distribution of how many successes each task attains within K=50 generations.
Large language models (LLMs) are increasingly deployed as code generators, where silently wrong programs pose real safety and reliability risks. Reliable uncertainty estimation (UE) is essential for selective prediction, human-in-the-loop review, and downstream agentic decisions. Yet most existing code UE methods are inherited from natural language (NL) generation and ignore properties that make code distinct. We argue that code differs from NL in three ways: a single wrong token can break an entire program (token fragility); algorithmic intent and concrete implementation can disagree independently (intent-code gap); and programs can be executed (executability). We instantiate these properties as three orthogonal uncertainty axes: lexical (Top-K token entropy), algorithmic (pseudo-code consistency), and functional (behavioral consistency). Across five code LLMs, our three-axis ensemble improves average AUROC from 0.696 for the strongest NL-derived baseline to 0.776 (+8.1 points). Notably, on Qwen3-14B, our single-pass Top-K token entropy matches the strongest multi-pass baseline while being over 3x cheaper; across models, it remains a competitive low-cost signal. These results suggest that code UE deserves code-specific design rather than direct NL ports.
Yuling Shi, Caiqi Zhang, Yuexian Li +4
Shanghai Jiao Tong University · University of Cambridge
Prediction sets provide a theoretically grounded framework for quantifying uncertainty in machine learning models. Adapting them to structured generation tasks, in particular, large language model (LLM) based code generation, remains a challenging problem. An existing attempt proposes PAC prediction sets but is limited by its strong monotonicity assumption on risk and single-label classification framework, which severely limits the space of candidate programs and cannot accommodate the multiple valid outputs inherent to code generation. To address these limitations, we propose an approach RisCoSet that leverages multiple hypothesis testing to construct risk-controlling predictions for LLM-based code generation. Given a trained code generation model, we produce a prediction set represented by a partial program, which is guaranteed to contain a correct solution with high confidence. Extensive experiments on three LLMs demonstrate the effectiveness of the proposed method. For instance, compared with the state-of-the-art, our method can significantly reduce the code removal by up to 24.5%, at the same level of risk.
Senrong Xu, Yuhao Tan, Yanke Zhou +6
State Key Lab of Novel Software Technology, Nanjing University · ETH Zürich · Birkbeck, University of London
Code generated by modern language models often reads naturally. Yet, it also often fails to implement what was asked. This should be no surprise, as research shows the models' own confidence signals are poorly calibrated with actual correctness. A promising way to assess correctness looks inside the model: by contrasting the hidden states of correct and incorrect programs, recent work captured an internal signal of code correctness that is able to judge candidate solutions better than the model's token-level or stated confidence, with no test execution. However, this signal was captured under one particular way, leaving open an important question: whether it reflects a robust property of the model or an artifact of that choice. We study this question systematically, varying how the signal is extracted from the model internals. Besides this, we also ask if the signal's quality is limited by the data used to extract it, by constructing program pairs that differ only in the fault that makes them incorrect. Our results show that no single configuration is best, and that isolating the fault does not help.
Francisco Ribeiro, Sohaila Abdulsattar, Renata Gonzalez +2
New York University Abu Dhabi Abu Dhabi, United Arab Emirates