cs.LGApr 9, 2026

Less Data Approximates More: Earning Faithful Confidence in High-Stakes Domains

Authors: Haokai Ma, Lee Yan Zhen, Gang Yang, Yunxiang Chen, Yunshan Ma, Tat-Seng Chua, Ee-Chien Chang

Organizations: National University of Singapore, Singapore · Singapore Management University, Singapore

Abstract

Large language models are increasingly deployed in high-stakes domains, where confident yet incorrect inferences may cause severe real-world harm, bringing the long-overlooked issue of confidence faithfulness to the forefront. A promising solution jointly optimizes unsupervised Reinforcement Learning from Internal Feedback (RLIF) with reasoning-trace-guided Reasoning Distillation (RD), yet it faces three persistent challenges, namely the scarcity of high-quality training corpora, factually unwarranted overconfidence, and erroneous updates amplified by indiscriminate fusion. Inspired by how human confidence accumulates from uncertainty to certainty, we propose Progressive Reasoning Gain (PRG) to measure whether reasoning steps progressively strengthen confidence in the final answer. Building on PRG, we introduce HyTuning, a hybrid post-training framework that adaptively reweights RD and RLIF, using scarce supervised reasoning traces as a stable anchor while exploiting abundant unlabeled queries for scalability. Experiments on several domain-specific and general benchmarks demonstrate that HyTuning improves accuracy while achieving confidence faithfulness under limited supervision, supporting a practical ``Less Data Approximates More'' effect. Our code will be released upon acceptance.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Understanding and Mitigating Premature Confidence for Better LLM Reasoning

    May 23, 2026Jingchu Gai, Guanning Zeng, Christina Baek +4LLM Reasoning StrategiesOverconfidence

  2. Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape Pre-, Intra-, and Post-CoT Calibration

    Jul 15, 2026Shuhao Li, Guodong Du, Anhao Zhao +3Post-TrainingSupervised Finetuning

  3. Quantifying Faithful Confidence Expression in Large Reasoning Models

    Jun 2, 2026Areeb Gani, Asal Meskin, Gabrielle Kaili-May Liu +1Confidence CalibrationLarge Reasoning Models