cs.AIJul 2, 2026

Scaling with Confidence: Calibrating Confidence of LLMs for Adaptive Test Time Scaling

Authors: Xuqing YangYi YuanShanzhe LeiXuhong Wang

Organizations: Shanghai Jiao Tong University · Southeast University · Shanghai AI Laboratory

Abstract

Training large language models (LLMs) with reinforcement learning (RL) has significantly advanced their performance on reasoning and question-answering tasks. However, prevailing RL reward designs typically prioritize response correctness, neglecting to incentivize models to express their confidence accurately. This leads to a critical problem: performance gains are often accompanied by poor calibration between confidence and accuracy, misleading models to overconfidently hallucinate when uncertain. To address this limitation, we propose C\textbf{C}orrectness and C\textbf{C}onfidence C\textbf{C}alibration R\textbf{R}einforcement L\textbf{L}earning (C3RL\textbf{C3RL}), a novel RL algorithm integrating correctness, calibration and dataset-informed reference accuracy rewards together. Comprehensive evaluation across 8 text and multimodal datasets demonstrates that C3RL enhances calibration without sacrificing accuracy, outperforming the current state-of-the-art method in both performance and calibration metrics. Utilizing the well-calibrated verbalized confidence from C3RL, we further introduce C\textbf{C}onfidence-based A\textbf{A}daptive Test Time S\textbf{S}caling (CAS\textbf{CAS}), an adjustable inference-time strategy that allocates computational resources based on response confidence. Experiments show that CAS surpasses majority voting on both in-domain and out-of-domain datasets while reducing the inference budget by up to 12.33 times. We believe the synergy of C3RL and CAS paves the way for deploying more reliable and resource-efficient LLMs. The code, data and models will be released.

Explore similar work

CardsList