When Words Fall Short: Iterative Synergy Between Verbalized Reasoning and Hidden Features for LLM Confidence Estimation
Authors: Yekun Xu, Ante Wang, Jingyi Ren, Xuanyi Chen, Weizhi Ma, Yang Liu
Organizations: College of AI, Tsinghua University, Beijing, China · Institute for AI Industry Research (AIR), Tsinghua University, Beijing, China · Dept. of Comp. Sci. & Tech., Institute for AI, Tsinghua University, Beijing, China
Confidence estimation is crucial for developing trustworthy large language models (LLMs), with most methods following estimator-based or verbalization-based paradigms. While recent research increasingly focuses on improving verbalized self-reports of confidence, we challenge the prevailing view that this approach surpasses independent confidence estimators. Our empirical study shows that a dedicated confidence estimator can substantially outperform verbalized confidence, indicating that LLMs' internal representations contain richer confidence signals. Building on this finding, we propose Iterative Policy-Estimator Training (IPoET), a framework that synergizes the complementary strengths of verbalized reasoning traces and informative representations. IPoET alternates policy optimization with estimator updating, integrating estimator-derived confidence feedback into policy learning and refreshing the estimator on new policy rollouts. Experiments across diverse datasets and Qwen and Llama backbones demonstrate that, by iteratively exploiting richer hidden features and adapting to the evolving policy distribution, IPoET consistently outperforms both estimator- and verbalization-based baselines in-domain and achieves superior or comparable results across all out-of-domain metrics. For more details, refer to https://github.com/xyk829/ipoet.
Figures & tables
Figure 1: (a) In-domain and out-of-domain comparison of confidence estimator and verbalized confidence using AUROC and Brier score. Percentages denote relative improvements over verbalized confidence. (b) Estimator performance across training budgets: Brier scores on a training subset, an in-domain validation set, and an out-of-domain test set illustrate overfitting with extended training.
Figure 2: Overview of IPoET. The framework alternates between policy and estimator optimization. In the policy-update stage , the confidence estimator is frozen and provides confidence estimates to form an estimator-derived term, combined with format and accuracy rewards for GRPO. In the estimator-update stage , the policy is frozen to generate correctness-labeled rollouts for updating the estimator via supervised regression. Across iterations, the policy refines reasoning-oriented generation while the estimator adapts to the evolving policy distribution.
HotpotQA
OOD Averaged
Acc. ↑
AUROC ↑
Brier ↓
ECE ↓
Acc. ↑
AUROC ↑
Brier ↓
ECE ↓
Qwen3-8B-Base
51.6%
0.555
0.363
0.343
56.5%
0.556
0.359
0.345
└ RLVR
62.2%
0.523
0.376
0.376
62.3%
0.518
0.357
0.358
└ Answer-Prob
62.2%
0.661
0.359
0.360
62.3%
0.548
0.340
0.327
└ P(True)
62.2%
0.543
0.308
0.260
62.3%
0.588
0.304
0.314
└ RLVR + Estimator
62.2%
0.757
0.195
0.087
62.3%
0.724
0.184
0.184
Table 1: Main results for models trained on HotpotQA and Big-Math. We report accuracy and confidence-estimation metrics, including AUROC, Brier score, and ECE, for Qwen and Llama backbones. For models trained on HotpotQA, we report results on HotpotQA and a six-dataset OOD average. For models trained on Big-Math, we report a Math average over MATH-500, GSM8K, and Big-Math, together with a five-dataset OOD average. Best and second-best results for all metrics under each backbone are marked in bold and underlined , respectively.
Setting
Strategy
Accuracy ↑
AUROC ↑
Brier ↓
In-domain
Joint
60.5%
0.773
0.209
Iterative
63.3% (+2.8 pp)
0.800 (+0.027)
0.184 (-0.025)
Out-of-domain
Joint
58.4%
0.737
0.180
Iterative
62.3% (+3.9 pp)
0.738 (+0.001)
0.170 (-0.010)
Table 2: Comparison between joint and iterative training on HotpotQA. Iterative training improves accuracy, AUROC, and Brier score in both in-domain and out-of-domain settings. Values in parentheses denote changes relative to joint training.
Figure 3: Step-wise effects of IPoET, showing that both policy and estimator updates progressively improve confidence estimation.
Iteration [-0.5pt] Setting
In-domain
Out-of-domain
AUROC ↑
Brier ↓
AUROC ↑
Brier ↓
2×200
0.878
0.100
0.698
0.206
4×100
0.873
0.102
0.696
0.198
Table 3: Effect of alternating-update granularity under Big-Math training with the same update budget and data allocation. Best results are shown in bold .
Figure 4: Round-wise comparison of in-domain and out-of-domain AUROC and Brier score under Big-Math training. Blue shading marks the configuration used in our main experiments (Round 2).
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Policy training prompt template.
Model
AUROC ↑
Brier ↓
ECE ↓
GPT-5 mini
0.603
0.306
0.298
Claude Haiku 4.5
0.671
0.279
0.266
DeepSeek-V4-Flash
0.537
0.332
0.328
Gemini 3 Flash
0.509
0.357
0.357
IPoET (ours)
0.800
0.184
0.107
Appendix
Table 4: Comparison between IPoET and frontier commercial models. IPoET is trained on HotpotQA for the HotpotQA results and on Big-Math for the math average over MATH-500, GSM8K, and Big-Math. Best and second-best results are marked in bold and underlined , respectively.