cs.CLSep 28, 2026

When Words Fall Short: Iterative Synergy Between Verbalized Reasoning and Hidden Features for LLM Confidence Estimation

Authors: Yekun Xu, Ante Wang, Jingyi Ren, Xuanyi Chen, Weizhi Ma, Yang Liu

Organizations: College of AI, Tsinghua University, Beijing, China · Institute for AI Industry Research (AIR), Tsinghua University, Beijing, China · Dept. of Comp. Sci. & Tech., Institute for AI, Tsinghua University, Beijing, China

Abstract

Confidence estimation is crucial for developing trustworthy large language models (LLMs), with most methods following estimator-based or verbalization-based paradigms. While recent research increasingly focuses on improving verbalized self-reports of confidence, we challenge the prevailing view that this approach surpasses independent confidence estimators. Our empirical study shows that a dedicated confidence estimator can substantially outperform verbalized confidence, indicating that LLMs' internal representations contain richer confidence signals. Building on this finding, we propose Iterative Policy-Estimator Training (IPoET), a framework that synergizes the complementary strengths of verbalized reasoning traces and informative representations. IPoET alternates policy optimization with estimator updating, integrating estimator-derived confidence feedback into policy learning and refreshing the estimator on new policy rollouts. Experiments across diverse datasets and Qwen and Llama backbones demonstrate that, by iteratively exploiting richer hidden features and adapting to the evolving policy distribution, IPoET consistently outperforms both estimator- and verbalization-based baselines in-domain and achieves superior or comparable results across all out-of-domain metrics. For more details, refer to https://github.com/xyk829/ipoet.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Future Confidence Distillation in Large Language Models

    Jul 8, 2026Sahil KaleConfidence EstimationConfidence Calibration

  2. ORCE: Order-Aware Alignment of Verbalized Confidence in Large Language Models

    May 12, 2026Chen Li, Xiaoling Hu, Songzhu Zheng +2Confidence EstimationFeature Alignment

  3. Instinct vs. Reflection: Unifying Token and Verbalized Confidence in Multimodal Large Models

    Apr 19, 2026Yunkai Dang, Yifan Jiang, Yizhu Jiang +3Confidence EstimationMultimodal Large Language Models