LLM Uncertainty Estimation

LLM: Large Language Model

Latest papers 298

All topics
CardsList
  1. Beyond Confidence: Rethinking Self-Assessments for Performance Prediction in LLMs

    May 8, 2026Sree Bhattacharyya, Samarth Khanna, Leona Chen +3LLM EvaluationConfidence Estimation in Language Models

  2. Tracing Uncertainty in Language Model "Reasoning"

    May 8, 2026Nils Grünefeld, Bertram Højer, Philipp Mondorf +5LLM EvaluationCoT Reasoning

  3. POETS: Uncertainty-Aware LLM Optimization via Compute-Efficient Policy Ensembles

    May 8, 2026Nicolas Menet, Andreas Krause, Abbas RahimiThompson SamplingBlack-Box Optimization

  4. Confidence-Aware Alignment Makes Reasoning LLMs More Reliable

    May 8, 2026Kejia Chen, Jiawen Zhang, Yihong Wu +5LLM AlignmentDirect Preference Optimization

  5. Can LLMs Take Retrieved Information with a Grain of Salt?

    May 7, 2026Behzad Shayegh, Mohamed Osama Ahmed, Fred Tung +1Retrieval-Augmented GenerationLLM Evaluation

  6. LLMs are not (consistently) Bayesian: Quantifying internal (in)consistencies of LLMs' probabilistic beliefs

    May 7, 2026Chacha Chen, Matthew Jörke, Adam Goliński +4Belief RevisionLLM Auditing

  7. Rethinking Vacuity for OOD Detection in Evidential Deep Learning

    May 7, 2026Claire McNamaraLLM EvaluationEvidential Deep Learning

  8. Measuring Black-Box Confidence via Reasoning Trajectories: Geometry, Coverage, and Verbalization

    May 7, 2026Marc Boubnovski Martell, Josefa Lia Stoisser, Kaspar Märtens +4Confidence Estimation in Language ModelsCoT Reasoning

  9. Estimating the Black-box LLM Uncertainty with Distribution-Aligned Adversarial Distillation

    May 7, 2026Huizi Cui, Huan Ma, Qilin Wang +2Language Model DistillationLLM Uncertainty Estimation

  10. The First Token Knows: Single-Decode Confidence for Hallucination Detection

    May 6, 2026Mina GabrielUncertainty QuantificationSelf-Consistency Decoding

  11. Elicitation Matters: How Prompts and Query Protocols Shape LLM Surrogates under Sparse Observations

    May 6, 2026Ge Lei, Samuel J. CooperSurrogate ModelingBayesian Optimization

  12. Gradients with Respect to Semantics Preserving Embeddings Tell the Uncertainty of Large Language Models

    May 6, 2026Mingda Li, Rundong Lv, Xinyu Li +2Uncertainty QuantificationLLM Uncertainty Estimation

  13. LLMs Uncertainty Quantification via Adaptive Conformal Semantic Entropy

    May 5, 2026Hamed Karimi, Vaishali Meyappan, Reza SamaviSelective PredictionSemantic Entropy

  14. Hallucinations Undermine Trust; Metacognition is a Way Forward

    May 2, 2026Gal Yona, Mor Geva, Yossi MatiasLLM Hallucination MitigationHallucination in Language Models

  15. CLEAR: Revealing How Noise and Ambiguity Degrade Reliability in LLMs for Medicine

    May 1, 2026Kevin H. Guo, Chao Yan, Avinash Baidya +5HealthcareLLM Reliability

  16. "I Don't Know" -- Towards Appropriate Trust with Certainty-Aware Retrieval Augmented Generation

    May 1, 2026Daan Di Scala, Maaike de Boer, Pınar YolumRetrieval-Augmented GenerationLLM Reliability

  17. Confidence Estimation in Automatic Short Answer Grading with LLMs

    Apr 30, 2026Longwei Cong, Sonja Hahn, Sebastian Gombert +3Uncertainty QuantificationSelective Prediction

  18. Belief-Guided Inference Control for Large Language Model Services via Verifiable Observations

    Apr 30, 2026Wenhao Yuan, Chenchen Lin, Jian Chen +3LLM InferenceCost-Aware Inference

  19. Entropy Centroids as Intrinsic Rewards for Test-Time Scaling

    Apr 28, 2026Wenshuo Zhao, Qi Zhu, Xingshan Zeng +4LLM InferenceTest-Time Scaling

  20. Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models

    Apr 28, 2026Chun-Yi Kuan, Wei-Ping Huang, Hung-yi LeeSemantic EntropyAudio Understanding

  21. Representational Curvature Modulates Behavioral Uncertainty in Large Language Models

    Apr 27, 2026Jack King, Evelina Fedorenko, Eghbal A. HosseiniNeural Representation GeometryAutoregressive Language Modeling