cs.CLJan 16, 2026

NOVA: NOise-aware Verbal Confidence CAlibration for Robust Large Language Models in RAG Systems

Authors: Jiayu Liu, Rui Wang, Qing Zong, Yumeng Wang, Cheng Qian, Qingcheng Zeng, Tianshi Zheng, Haochen Shi, +4 more

Organizations: HKUST · UIUC · Northwestern University

Abstract

Accurately assessing model confidence is essential for deploying large language models (LLMs) in mission-critical factual domains. While retrieval-augmented generation (RAG) is widely adopted to improve grounding, confidence calibration in RAG settings remains poorly understood. We conduct a systematic study across four benchmarks, revealing that LLMs exhibit poor calibration performance especially when noisy contexts are retrieved. Specifically, contradictory or irrelevant evidence tends to exacerbate the model's overconfidence issue. To address this, we propose NOVA Rules (NOise-Aware Verbal Confidence CAlibration Rules) to provide a principled foundation for resolving overconfidence under noise. We further design NOVA, a noise-aware calibration framework that synthesizes supervision from ~2K HotpotQA examples guided by these rules. By performing supervised fine-tuning (SFT) with this data, NOVA equips models with intrinsic noise awareness without relying on stronger teacher models. Empirical results show that NOVA yields substantial gains, improving ECE scores by 10.9% in-domain and 8.0% out-of-domain. By bridging the gap between retrieval noise and verbal calibration, NOVA paves the way for both accurate and epistemically reliable LLMs.

Figures & tables

Appendix figures & tables30 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ORCE: Order-Aware Alignment of Verbalized Confidence in Large Language Models

    May 12, 2026Chen Li, Xiaoling Hu, Songzhu Zheng +2Confidence EstimationFeature Alignment

  2. Retrieval-Augmented Linguistic Calibration

    May 19, 2026Yi-Fan Yeh, Linwei Tao, Minjing Dong +4LinguisticsAudience

  3. BalanceRAG: Joint Risk Calibration for Cascaded Retrieval-Augmented Generation

    May 19, 2026Zijun Jia, Yuanchang Ye, Sen Jia +6Conformal Risk ControlLarge Language Models(Llms