The reasoning trajectory of a Large Language Model (LLM) is often treated as a verbalized description of its internal reasoning. However, such trajectories can be unfaithful: a model may rely on shortcuts to reach an answer and then post-rationalize the decision with a seemingly coherent chain of thought. Detecting this shortcut reasoning is challenging because existing monitors and verifiers mainly inspect textual traces or final outcomes, rather than how the model's belief in its answer develops during generation. We introduce ConfLens, a framework that tracks the evolution of confidence in the final answer throughout reasoning. Across three shortcut reasoning settings, we observe a common pattern of premature confidence, where shortcut samples become highly confident in the final answer at early reasoning stages. Existing confidence estimation methods, however, show limited generalizability, reliability, or efficiency for detecting this behavior. We therefore propose the Distributional Answer Commitment Score (DACS), a distributional confidence estimator that measures the entropy of the model's probability distribution over answer commitment at each reasoning step. DACS captures how concentrated the model's answer belief is without requiring ground-truth answers or task-specific verifiers. We further convert ConfLens detection results into interpretable signals for reward models to reduce their preference for shortcut reasoning. Experiments on mathematical and code reasoning tasks show that ConfLens with DACS improves shortcut reasoning detection by over 4.3% F1 compared with strong baselines and reduces the mismatch between faithfulness and correctness in reward model preferences.
Figures & tables
Figure 1: Three circumstances where shortcut reasoning arises. In the persuasive prompt setting, the model receives an authoritative hint about the final answer. In reward hacking, the model learns during RLVR to exploit defects in either the training data, corresponding to the implicit prompt setting, or the reward model, corresponding to the reward bias setting. The CoT monitor, implemented with GPT-4o in our experiments, fails to reliably identify shortcut reasoning responses.
Figure 2: Reward trajectories across training steps under implicit hint, reward bias, and non-hacking settings in GRPO training, and performance comparison on math reasoning between the shortcut-taking and non-hacking models for Qwen2.5-3B-Instruct. Left) Reward curve for math reasoning task. Middle) Reward curve for code reasoning task. Right) Performance comparison between shortcut-taking model and non-hacking model in implicit hint setting. The reward of the model that learns to take shortcuts boosts during training, but it fails when the hint is removed or misleading.
Figure 3: The confidence trajectory over CoT percentage. Current confidence estimation methods cannot perform consistently well in all circumstances where shortcut reasoning arises. SL completely fails in detecting shortcut reasoning under reward bias setting. SC brings high latency. P(pass) and Verbal is not reliable in all settings.
Explicit Prompt
Implicit Prompt
Reward Bias
Math Reasoning
SL
0.748
0.947
0.100
SC
0.702
0.923
0.941
P(pass)
0.283
0.437
0.467
Verbal
0.763
0.406
0.475
DACS
0.784
0.874
0.904
Table 1: AUROC for shortcut reasoning detection with different confidence estimation methods in ConfLens .
Methods
Math Reasoning
Code Reasoning
Explicit Hint
Implicit Hint
Reward Bias
Latency (s)
Explicit Hint
Implicit Hint
Reward Bias
Latency (s)
AUROC
F1
AUROC
F1
AUROC
F1
AUROC
F1
AUROC
F1
AUROC
F1
Qwen2.5-3B-Instruct
CoT Monitor
-
0.641
-
0.622
-
0.602
0.18
-
0.627
-
0.384
-
0.522
0.18
TRACE
0.702
0.658
0.923
0.902
0.941
0.908
3.82
0.601
0.626
0.708
0.658
0.759
0.738
3.71
ConfLens
0.784
0.732
0.874
0.904
0.928
0.906
0.35
0.637
0.633
0.774
0.748
0.773
0.796
0.21
Table 2: Shortcut reasoning detection results across different methods. For TRACE and ConfLens , we optimize the decision threshold on the validation set that yields the best F1 score, thereby converting their outputs into binary detection results. The best results are bolded .
Figure 4: The accuracy and average faithfulness of accepted samples by different reward models and signal strategies. Adopting the detection results of ConfLens help the reward model identify the shortcut reasoning responses while maintaining the accuracy.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Category
Configuration
Data
Training Set Size: 20194 Validation Set Size: 3000 Test Set Size: 5799
Sequence Lengths
Max prompt length: 512 Max response length: 1024 Overlong prompts: filtered
Batching
Total epochs: 4 Train Batch Size: 1024 Micro Batch Size Per GPU: 16
Optimization
Learning rate: 1×10−6
GRPO
Rollout: 5 KL_coef: 0.001
Appendix
Table 3: Training configuration for math reasoning task.
Category
Configuration
Data
Training Set Size: 400 Validation Set Size: 80 Test Set Size: 320
Sequence Lengths
Max prompt length: 512 Max response length: 1024 Overlong prompts: filtered
Batching
Total epochs: 10 Train Batch Size: 1024 Micro Batch Size Per GPU: 16
Optimization
Learning rate: 1×10−6
GRPO
Rollout: 5 KL_coef: 0.001
Appendix
Table 4: Training configuration for code reasoning task.
Figure 5: The F1 score for shortcut reasoning detection in math reasoning task (left) and code reasoning task (right) under implicit hint setting. ConfLens consistently surpress the comparison methods.
Long chains of thought (CoT) from current language models frequently contain logical gaps and unjustified leaps, limiting the gains from additional test-time compute. Improving reasoning quality directly would require process reward models, but the step-level annotations needed to train them are expensive and scarce. We find such a signal in how the model's confidence evolves during reasoning: premature confidence, the tendency to commit to an answer early and use the remaining tokens to rationalize it, strongly predicts flawed reasoning across tasks and model scales. We exploit this in progressive confidence shaping, a reinforcement learning objective that trains models to update their confidence as they reason rather than commit early -- rewarding gradual confidence growth and penalizing early commitment, with no external labels or reward models. The method improves accuracy and reasoning quality from 1.5B to 8B parameters across arithmetic (Countdown), math (DAPO, AIME), and science (ScienceQA): on Countdown, accuracy improves 3.2x (+42.0pp) and flawed reasoning drops 48pp; on AIME, Pass@64 improves 6.6pp. Consistent with this mechanism, the method also improves faithfulness: on a safety benchmark, our models more transparently surface misleading content in their reasoning traces rather than concealing it. Controlled experiments reveal that the problem and its remedy scale together: premature confidence grows with model size and task difficulty, and so do the gains from addressing it.
Jingchu Gai, Guanning Zeng, Christina Baek +4
1Carnegie Mellon University · 2Tsinghua University
Explicit chain-of-thought (CoT) reasoning substantially improves the reasoning ability of large language models (LLMs), but incurs high inference cost due to lengthy autoregressive traces. Existing latent reasoning methods offer a promising alternative, yet they often treat reasoning as uniformly compressible, causing precision-critical intermediate steps to be overly compressed and thereby degrading reasoning accuracy. In this work, we propose Selective Latent Thinking (SLT), a framework that selectively compresses redundant reasoning spans into latent representations while preserving precision-critical spans as explicit CoT within the same reasoning trajectory. Specifically, SLT first uses a lightweight decoder to anticipate a short upcoming reasoning span, and then applies confidence-based gating to determine the longest span that can be reliably compressed. The accepted span is encoded into a compact latent representation to improve reasoning efficiency, while uncertain or precision-critical reasoning remains in explicit CoT form to preserve accuracy. To learn this selective compression policy, SLT adopts a three-stage training strategy that combines span-level latent compression, reliability-aware future reasoning prediction, and trajectory-level reinforcement learning to optimize the trade-off between answer correctness and reasoning cost. Extensive experiments across four mathematical reasoning benchmarks demonstrate that SLT achieves 22.7% higher accuracy than latent reasoning baselines at comparable compression ratios, while reducing reasoning chain length by 58.4% with only 2.8% accuracy degradation compared to explicit CoT,Our code can be found in https://github.com/hunshi34/SLT.
Hui Xie, Jie Liu, Ziyue Qiao +1
Eindhoven University of Technology, Netherlands · School of Computing and Information Technology, Great Bay University, China
Large Language Models have achieved strong performance on reasoning tasks with objective answers by generating step-by-step solutions, but diagnosing where a multi-step reasoning trace might fail remains difficult. Confidence estimation offers a diagnostic signal, yet existing methods are restricted to final answers or require internal model access. In this paper, we introduce Stepwise Confidence Attribution (SCA), a framework for closed-source LLMs that assigns step-level confidence based only on generated reasoning traces. SCA applies the Information Bottleneck principle: steps aligning with consensus structures across correct solutions receive high confidence, while deviations are flagged as potentially erroneous. We propose two complementary methods: (1) NIBS, a non-parametric IB approach measuring consistency without graph structures, and (2) GIBS, a graph-based IB model that learns subgraphs through a differentiable mask to capture logical variability. Extensive experiments on mathematical reasoning and multi-hop question answering show that SCA reliably identifies low-confidence steps strongly correlated with reasoning errors. Moreover, using step-level confidence to guide self-correction improves the correction success rate by up to 13.5% over answer-level feedback.
Xiaoou Liu, Tiejin Chen, Dengjia Zhang +3
1Arizona State University · 2Johns Hopkins University · 3Purdue +1