LLM Uncertainty Estimation

LLM: Large Language Model

Latest papers 298

All topics
CardsList
  1. RL-ARC: Calibrating Large Reasoning Models via Reasoning-guided Uncertainty

    Oct 8, 2026Gukhyeon Lee, SangKeun LeeLanguage Model CalibrationLLM Uncertainty Estimation

  2. U-Space: Uncovering When and Why Uncertainty Arises in Language Models

    Oct 6, 2026Tobias Braun, Nils Loose, Alexander Herzog +4LLM Uncertainty EstimationConfidence Estimation in Language Models

  3. Sequential Probabilistic Uncertainty Estimation for Parallel Multi-Agent Reasoning Systems

    Oct 6, 2026Tunyu Zhang, Zihao Zhao, Yusong Zhao +5LLM Uncertainty EstimationMulti-Agent LLM Systems

  4. Single-Pass Uncertainty Heads for Claim-Level Hallucination Detection in Persian Medical Language Models

    Oct 2, 2026Mehrdad Ghassabi, Pedram Rostami, Hamidreza Baradaran Kashani +2Hallucination DetectionLLM Hallucination Detection

  5. Bayesian Fine-tuning Yields Language Models that are as Bayesian as their Beliefs Allow

    Sep 30, 2026Polina Tsvilodub, Andreas Waldis, Linlu Qiu +2Fine-TuningLLM Fine-Tuning

  6. Referential Uncertainty in Human--AI Collaboration

    Sep 30, 2026Christian Poelitz, Finale Doshi-Velez, Siân LindleyLLM Uncertainty EstimationHuman-AI Decision Making

  7. Also Small Models Can Reasonably Self-Evaluate Their Confidence

    Sep 30, 2026Idil Kapikiran, Thomas Decker, Thomas RunklerConfidence Estimation in Language ModelsSmall Language Models

  8. Probability is Not Enough: Exploring and Counting Divergent Tokens for Reasoning Uncertainty Quantification in LLMs

    Sep 29, 2026Feiyang Li, Shengjing Liu, Qi Zhan +7Uncertainty QuantificationConfidence Estimation in Language Models

  9. HARISSA: Inference-Time Self-Checks for Efficient and Safe Local Language Model Deployment

    Sep 29, 2026Kenan Alkiek, Moontae Lee, David Jurgens +1On-Device Language Model InferenceSelective Prediction

  10. BITEM at the NTCIR-19 R2C2 Task: Predicting Confidence from Agentic RAG Pipeline Signals

    Sep 29, 2026Julien Knafou, Luc Mottin, Alexandre Flament +3Retrieval-Augmented GenerationAgentic RAG

  11. XU-RS: Explaining Credal Width in Random-Set Language Models

    Sep 29, 2026David Achara, Maryam Sultana, Alexander D. Rast +1LLM InterpretabilityLLM Uncertainty Estimation

  12. Risk-Aware Semantic Grounding for Trustworthy LLM-Based Robot Planning

    Sep 29, 2026Łukasz Sobczak, Nur Keleşoğlu, Sławomir Piotr NowakLarge Language Model-Based Robot PlanningRobot Navigation

  13. TRACE: Single-Pass Decoding-Trace Risk Localization for Generation Calibration

    Sep 28, 2026Yuebin Xu, Xuemei Peng, Junlan Chen +2Confidence Estimation in Language ModelsLanguage Model Calibration

  14. Jailbreaks for Black-Box Uncertainty Quantification in Large Reasoning Models

    Sep 28, 2026Lucas Biechy, Cédric Eichler, Adrien Boiret +1Language Model CalibrationLLM Uncertainty Estimation

  15. When Confidence Rises Too Early: Detecting Shortcut Reasoning via Premature Answer Commitment

    Sep 28, 2026Zhaohan Zhang, Junjie Liu, Chengzhengxu Li +6Faithfulness of Language Model ExplanationsConfidence Estimation in Language Models

  16. Semantic Uncertainty Quantification Needs Factual Equivalence

    Sep 28, 2026Joseph Hoche, Quentin Guimard, Gianni FranchiConfidence Estimation in Language ModelsFactual Consistency Evaluation

  17. When Words Fall Short: Iterative Synergy Between Verbalized Reasoning and Hidden Features for LLM Confidence Estimation

    Sep 28, 2026Yekun Xu, Ante Wang, Jingyi Ren +3Confidence Estimation in Language ModelsLLM Reliability

  18. Toward a Graded Measure of Belief Stability in Large Language Models

    Sep 28, 2026Samantha Dies, Branden Fitelson, Tina Eliassi-RadLLM ReliabilityLLM Uncertainty Estimation

  19. LLMs learn different forms of metacognition when trained to predict their own accuracy

    Sep 27, 2026Nicolas Yax, Stefano Palminteri, Pierre-Yves OudeyerConfidence Estimation in Language ModelsLanguage Model Calibration

  20. Shared Experience, Separate Learning: Companion Confidence Calibration for LLMs

    Sep 27, 2026Shiyu Ni, Keping Bi, Jiafeng Guo +4Confidence Estimation in Language ModelsLanguage Model Calibration

  21. Calibration, Not Answer Selection: Distilling Internal Confidence in Reasoning Models

    Sep 27, 2026Yadong Xi, Rongsheng Zhang, Tangjie Lv +2Confidence Estimation in Language ModelsLanguage Model Calibration

  22. JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

    Sep 22, 2026Yubo Li, Yidi Miao, Ramayya Krishnan +1LLM-as-a-JudgeLanguage Model Generation Evaluation

  23. Pinocchio: Fast Uncertainty Estimates for Black-Box Language Models

    Sep 21, 2026Kevin David Hayes, Arka Pal, Haosong Zhang +2Language Model Error DetectionConfidence Estimation in Language Models