When Words Fall Short: Iterative Synergy Between Verbalized Reasoning and Hidden Features for LLM Confidence Estimation
Authors: Yekun Xu, Ante Wang, Jingyi Ren, Xuanyi Chen, Weizhi Ma, Yang Liu
Organizations: College of AI, Tsinghua University, Beijing, China · Institute for AI Industry Research (AIR), Tsinghua University, Beijing, China · Dept. of Comp. Sci. & Tech., Institute for AI, Tsinghua University, Beijing, China
Confidence estimation is crucial for developing trustworthy large language models (LLMs), with most methods following estimator-based or verbalization-based paradigms. While recent research increasingly focuses on improving verbalized self-reports of confidence, we challenge the prevailing view that this approach surpasses independent confidence estimators. Our empirical study shows that a dedicated confidence estimator can substantially outperform verbalized confidence, indicating that LLMs' internal representations contain richer confidence signals. Building on this finding, we propose Iterative Policy-Estimator Training (IPoET), a framework that synergizes the complementary strengths of verbalized reasoning traces and informative representations. IPoET alternates policy optimization with estimator updating, integrating estimator-derived confidence feedback into policy learning and refreshing the estimator on new policy rollouts. Experiments across diverse datasets and Qwen and Llama backbones demonstrate that, by iteratively exploiting richer hidden features and adapting to the evolving policy distribution, IPoET consistently outperforms both estimator- and verbalization-based baselines in-domain and achieves superior or comparable results across all out-of-domain metrics. For more details, refer to https://github.com/xyk829/ipoet.
Figures & tables
Figure 1: (a) In-domain and out-of-domain comparison of confidence estimator and verbalized confidence using AUROC and Brier score. Percentages denote relative improvements over verbalized confidence. (b) Estimator performance across training budgets: Brier scores on a training subset, an in-domain validation set, and an out-of-domain test set illustrate overfitting with extended training.
Figure 2: Overview of IPoET. The framework alternates between policy and estimator optimization. In the policy-update stage , the confidence estimator is frozen and provides confidence estimates to form an estimator-derived term, combined with format and accuracy rewards for GRPO. In the estimator-update stage , the policy is frozen to generate correctness-labeled rollouts for updating the estimator via supervised regression. Across iterations, the policy refines reasoning-oriented generation while the estimator adapts to the evolving policy distribution.
HotpotQA
OOD Averaged
Acc. ↑
AUROC ↑
Brier ↓
ECE ↓
Acc. ↑
AUROC ↑
Brier ↓
ECE ↓
Qwen3-8B-Base
51.6%
0.555
0.363
0.343
56.5%
0.556
0.359
0.345
└ RLVR
62.2%
0.523
0.376
0.376
62.3%
0.518
0.357
0.358
└ Answer-Prob
62.2%
0.661
0.359
0.360
62.3%
0.548
0.340
0.327
└ P(True)
62.2%
0.543
0.308
0.260
62.3%
0.588
0.304
0.314
└ RLVR + Estimator
62.2%
0.757
0.195
0.087
62.3%
0.724
0.184
0.184
Table 1: Main results for models trained on HotpotQA and Big-Math. We report accuracy and confidence-estimation metrics, including AUROC, Brier score, and ECE, for Qwen and Llama backbones. For models trained on HotpotQA, we report results on HotpotQA and a six-dataset OOD average. For models trained on Big-Math, we report a Math average over MATH-500, GSM8K, and Big-Math, together with a five-dataset OOD average. Best and second-best results for all metrics under each backbone are marked in bold and underlined , respectively.
Setting
Strategy
Accuracy ↑
AUROC ↑
Brier ↓
In-domain
Joint
60.5%
0.773
0.209
Iterative
63.3% (+2.8 pp)
0.800 (+0.027)
0.184 (-0.025)
Out-of-domain
Joint
58.4%
0.737
0.180
Iterative
62.3% (+3.9 pp)
0.738 (+0.001)
0.170 (-0.010)
Table 2: Comparison between joint and iterative training on HotpotQA. Iterative training improves accuracy, AUROC, and Brier score in both in-domain and out-of-domain settings. Values in parentheses denote changes relative to joint training.
Figure 3: Step-wise effects of IPoET, showing that both policy and estimator updates progressively improve confidence estimation.
Iteration [-0.5pt] Setting
In-domain
Out-of-domain
AUROC ↑
Brier ↓
AUROC ↑
Brier ↓
2×200
0.878
0.100
0.698
0.206
4×100
0.873
0.102
0.696
0.198
Table 3: Effect of alternating-update granularity under Big-Math training with the same update budget and data allocation. Best results are shown in bold .
Figure 4: Round-wise comparison of in-domain and out-of-domain AUROC and Brier score under Big-Math training. Blue shading marks the configuration used in our main experiments (Round 2).
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Policy training prompt template.
Model
AUROC ↑
Brier ↓
ECE ↓
GPT-5 mini
0.603
0.306
0.298
Claude Haiku 4.5
0.671
0.279
0.266
DeepSeek-V4-Flash
0.537
0.332
0.328
Gemini 3 Flash
0.509
0.357
0.357
IPoET (ours)
0.800
0.184
0.107
Appendix
Table 4: Comparison between IPoET and frontier commercial models. IPoET is trained on HotpotQA for the HotpotQA results and on Big-Math for the math average over MATH-500, GSM8K, and Big-Math. Best and second-best results are marked in bold and underlined , respectively.
Reliable confidence estimation is essential for deploying large language models (LLMs) in confidence-aware systems, where downstream decisions such as retrieval, tool use, and adaptive computation depend on accurately estimating answer reliability. Existing approaches, however, largely treat confidence as a property of completed responses, overlooking how confidence-related information evolves throughout the answering process. In this work, we investigate confidence from a temporal perspective by comparing pre-solution Feeling-of-Knowing (FOK) and post-solution Judgement-of-Learning (JOL) confidence estimates across frontier and open-source LLMs. We show that post-solution confidence is consistently better calibrated and more discriminative than pre-solution confidence, while linear probes trained on hidden representations recover substantially richer confidence-related information than models explicitly verbalise. Building on this observation, we introduce future confidence distillation, which trains predictors operating on pre-solution hidden representations using teacher confidence estimates produced by post-solution correctness probes. Despite requiring only pre-solution representations for inference, distilled predictors recover much of the calibration improvement achieved by post-solution confidence, remain highly sample efficient, and transfer across datasets within the same domain. Together, our findings demonstrate that confidence-related information evolves throughout the answering process and can be anticipated before answer generation is complete, enabling significantly more reliable yet low-cost confidence estimation.
Sahil Kale
University of California, Los Angeles Los Angeles, USA
Large language models (LLMs) often produce answers with high certainty even when they are incorrect, making reliable confidence estimation essential for deployment in real-world scenarios. Verbalized confidence, where models explicitly state their confidence in natural language, provides a flexible and user-facing uncertainty signal that can be applied even when token logits are unavailable. However, existing verbalized-confidence methods often optimize answer generation and confidence generation jointly, which can cause confidence-alignment objectives to interfere with answer accuracy. In this work, we propose a decoupled and order-aware framework for verbalized confidence calibration. Our method first generates an answer and then estimates confidence conditioned on the fixed question--answer pair, allowing confidence optimization without directly perturbing the answer-generation process. To align confidence with correctness likelihood, we construct a sampling-based surrogate from multiple model completions and optimize rank-based reinforcement learning objectives that encourage responses with higher estimated correctness likelihood to receive higher verbalized confidence. Experiments on reasoning and knowledge-intensive benchmarks show that our method improves calibration and failure prediction performance while largely preserving answer accuracy. These results demonstrate that verbalized confidence can be more reliably aligned by decoupling confidence estimation from answer generation and optimizing the relative ordering of confidence across responses.
Chen Li, Xiaoling Hu, Songzhu Zheng +2
Stony Brook University, NY, USA · Massachusetts General Hospital and Harvard Medical School, MA, USA · Morgan Stanley, NY, USA
Multimodal Large Language Models (MLLMs) have demonstrated exceptional capabilities in various perception and reasoning tasks. Despite this success, ensuring their reliability in practical deployment necessitates robust confidence estimation. Prior works have predominantly focused on text-only LLMs, often relying on computationally expensive self-consistency sampling. In this paper, we extend this to multimodal settings and conduct a comprehensive evaluation of MLLMs' response confidence estimation. Our analysis reveals a significant instinct-reflection misalignment: the model's implicit token-level support frequently diverges from its verbal self-assessment confidence. To address this misalignment, we propose a monotone confidence fusion framework to merge dual-channel signals and cross-channel consistency to estimate correctness. Subsequently, an order-preserving mean alignment step is applied to correct global bias, which improves calibration while preserving the risk-coverage trade-off for selective prediction. Experiments on diverse open-source and closed-source MLLMs show that our method consistently yields more reliable confidence estimates and improves both calibration and failure prediction. Code will be available at https://github.com/Yunkaidang/Instinct-vs.-Reflection.
Yunkai Dang, Yifan Jiang, Yizhu Jiang +3
School of Artificial Intelligence, Nanjing University.