cs.CLSep 27, 2026

Shared Experience, Separate Learning: Companion Confidence Calibration for LLMs

Authors: Shiyu Ni, Keping Bi, Jiafeng Guo, Yilong Xu, Jingtong Wu, Zengxin Han, Xueqi Cheng

Organizations: State Key Laboratory of AI Safety · Institute of Computing Technology, Chinese Academy of Sciences · University of Chinese Academy of Sciences

Abstract

Reliable self-assessment is essential for large language models (LLMs), yet they often remain highly confident when their answers are wrong. We study \emph{concurrent confidence calibration}, where confidence is learned alongside capability improvement rather than calibrated only after training. Reinforcement learning from verifiable rewards (RLVR) provides a natural setting for this paradigm, as it continuously produces responses paired with verifiable correctness feedback. Existing concurrent methods, however, learn both capability and confidence through reinforcement learning within shared policy parameters, potentially coupling two fundamentally different learning problems. We instead propose \emph{shared experience but separate learning}: capability and confidence learn from the same trajectories, but through separate optimization mechanisms and parameters. Based on this principle, we introduce \textbf{CoCal (Companion Confidence Calibration)}, which trains a lightweight companion from rollout hidden states and verifier-derived correctness supervision while leaving task optimization unchanged. Experiments on Qwen3-8B and Qwen3-14B show that CoCal improves confidence estimation without sacrificing task performance, outperforming both RL-based concurrent methods and matched post-hoc calibration. The learned companion further generalizes across domains and policy shifts, while the benefits of CoCal persist at both scales.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Scaling with Confidence: Calibrating Confidence of LLMs for Adaptive Test Time Scaling

    Jul 2, 2026Xuqing Yang, Yi Yuan, Shanzhe Lei +1Confidence EstimationLarge Language Model Reliability

  2. On Calibration of Large Language Models: From Response To Capability

    Feb 14, 2026Sin-Han Yang, Cheng-Kuang Wu, Chieh-Yen Lin +3Confidence Estimation