cs.SEOct 8, 2026

Characterizing Overconfident Failure in LLM-Based Code Generation

Authors: Ravishka Rathnasuriya, Wei Yang

Organizations: Institute of Systems for Advanced Computing, Fudan University, Shanghai, China

Abstract

Large language models (LLMs) are increasingly used for automated code generation, but generated programs can appear syntactically plausible while still failing execution-based correctness checks. Existing validation methods, such as testing and program analysis, remain essential but are often incomplete, costly, or applied only after generation. Model-derived uncertainty is therefore a natural early reliability signal. This paper studies the dilemma of overconfidence in code LLMs where incorrect programs are often generated with token-level confidence comparable to correct programs. We study this dilemma across four open-source code models and three execution-based benchmarks. Our analysis begins by investigating whether existing uncertainty metrics provide reliable proxies for execution correctness in code generation. We then characterize overconfidence at both global and local token levels, asking whether incorrect programs remain indistinguishable from correct ones under confidence and entropy summaries, including selective generation and the limits of instruction tuning. Finally, we evaluate whether common mitigation strategies reduce this failure mode. Our study yields four findings. First, existing uncertainty signals provide only partial and model-dependent evidence of execution failure. Second, overconfidence persists at both program and token levels, and uncertainty-based selection does not consistently improve accepted-set accuracy. Third, instruction tuning can increase certainty on failing generations without consistently improving correctness discrimination. Fourth, common mitigation techniques improve specific aspects of reliability but do not reliably resolve overconfident failure. Our exploratory latent analysis suggests that hidden representations may encode correctness-related signals that output confidence does not expose.

Figures & tables

Explore similar work

CardsList
  1. Code Is More Than Text: Uncertainty Estimation for Code Generation

    Jun 8, 2026Yuling Shi, Caiqi Zhang, Yuexian Li +4Confidence Estimation in Language ModelsLLM Reliability

  2. Uncertainty Quantification for LLM-based Code Generation

    May 12, 2026Senrong Xu, Yuhao Tan, Yanke Zhou +6Uncertainty QuantificationLLM Uncertainty Estimation

  3. On the Robustness of LLMs' Internal Representation of Code Correctness

    Aug 8, 2026Francisco Ribeiro, Sohaila Abdulsattar, Renata Gonzalez +2LLM ReliabilityLanguage Model Robustness