cs.CLSep 28, 2026

When Confidence Rises Too Early: Detecting Shortcut Reasoning via Premature Answer Commitment

Authors: Zhaohan Zhang, Junjie Liu, Chengzhengxu Li, Chen Shen, Xiaoming Liu, Chao Shen, Jieping Ye, Ziquan Liu, +1 more

Organizations: Queen Mary University of London · Tongyi Lab, Alibaba Group · Xi’an Jiaotong University

Abstract

The reasoning trajectory of a Large Language Model (LLM) is often treated as a verbalized description of its internal reasoning. However, such trajectories can be unfaithful: a model may rely on shortcuts to reach an answer and then post-rationalize the decision with a seemingly coherent chain of thought. Detecting this shortcut reasoning is challenging because existing monitors and verifiers mainly inspect textual traces or final outcomes, rather than how the model's belief in its answer develops during generation. We introduce ConfLens, a framework that tracks the evolution of confidence in the final answer throughout reasoning. Across three shortcut reasoning settings, we observe a common pattern of premature confidence, where shortcut samples become highly confident in the final answer at early reasoning stages. Existing confidence estimation methods, however, show limited generalizability, reliability, or efficiency for detecting this behavior. We therefore propose the Distributional Answer Commitment Score (DACS), a distributional confidence estimator that measures the entropy of the model's probability distribution over answer commitment at each reasoning step. DACS captures how concentrated the model's answer belief is without requiring ground-truth answers or task-specific verifiers. We further convert ConfLens detection results into interpretable signals for reward models to reduce their preference for shortcut reasoning. Experiments on mathematical and code reasoning tasks show that ConfLens with DACS improves shortcut reasoning detection by over 4.3% F1 compared with strong baselines and reduces the mismatch between faithfulness and correctness in reward model preferences.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Understanding and Mitigating Premature Confidence for Better LLM Reasoning

    May 23, 2026Jingchu Gai, Guanning Zeng, Christina Baek +4LLM Reasoning StrategiesOverconfidence

  2. Selective Latent Thinking: Adaptive Compression of LLM Reasoning Chains

    May 25, 2026Hui Xie, Jie Liu, Ziyue Qiao +1Latent ThoughtsChain-of-Thought Reasoning

  3. Diagnosing Multi-step Reasoning Failures in Black-box LLMs via Stepwise Confidence Attribution

    May 19, 2026Xiaoou Liu, Tiejin Chen, Dengjia Zhang +3LLM Reasoning StrategiesReasoning Traces