The reasoning trajectory of a Large Language Model (LLM) is often treated as a verbalized description of its internal reasoning. However, such trajectories can be unfaithful: a model may rely on shortcuts to reach an answer and then post-rationalize the decision with a seemingly coherent chain of thought. Detecting this shortcut reasoning is challenging because existing monitors and verifiers mainly inspect textual traces or final outcomes, rather than how the model's belief in its answer develops during generation. We introduce ConfLens, a framework that tracks the evolution of confidence in the final answer throughout reasoning. Across three shortcut reasoning settings, we observe a common pattern of premature confidence, where shortcut samples become highly confident in the final answer at early reasoning stages. Existing confidence estimation methods, however, show limited generalizability, reliability, or efficiency for detecting this behavior. We therefore propose the Distributional Answer Commitment Score (DACS), a distributional confidence estimator that measures the entropy of the model's probability distribution over answer commitment at each reasoning step. DACS captures how concentrated the model's answer belief is without requiring ground-truth answers or task-specific verifiers. We further convert ConfLens detection results into interpretable signals for reward models to reduce their preference for shortcut reasoning. Experiments on mathematical and code reasoning tasks show that ConfLens with DACS improves shortcut reasoning detection by over 4.3% F1 compared with strong baselines and reduces the mismatch between faithfulness and correctness in reward model preferences.
Figures & tables
Figure 1: Three circumstances where shortcut reasoning arises. In the persuasive prompt setting, the model receives an authoritative hint about the final answer. In reward hacking, the model learns during RLVR to exploit defects in either the training data, corresponding to the implicit prompt setting, or the reward model, corresponding to the reward bias setting. The CoT monitor, implemented with GPT-4o in our experiments, fails to reliably identify shortcut reasoning responses.
Figure 2: Reward trajectories across training steps under implicit hint, reward bias, and non-hacking settings in GRPO training, and performance comparison on math reasoning between the shortcut-taking and non-hacking models for Qwen2.5-3B-Instruct. Left) Reward curve for math reasoning task. Middle) Reward curve for code reasoning task. Right) Performance comparison between shortcut-taking model and non-hacking model in implicit hint setting. The reward of the model that learns to take shortcuts boosts during training, but it fails when the hint is removed or misleading.
Figure 3: The confidence trajectory over CoT percentage. Current confidence estimation methods cannot perform consistently well in all circumstances where shortcut reasoning arises. SL completely fails in detecting shortcut reasoning under reward bias setting. SC brings high latency. P(pass) and Verbal is not reliable in all settings.
Explicit Prompt
Implicit Prompt
Reward Bias
Math Reasoning
SL
0.748
0.947
0.100
SC
0.702
0.923
0.941
P(pass)
0.283
0.437
0.467
Verbal
0.763
0.406
0.475
DACS
0.784
0.874
0.904
Table 1: AUROC for shortcut reasoning detection with different confidence estimation methods in ConfLens .
Methods
Math Reasoning
Code Reasoning
Explicit Hint
Implicit Hint
Reward Bias
Latency (s)
Explicit Hint
Implicit Hint
Reward Bias
Latency (s)
AUROC
F1
AUROC
F1
AUROC
F1
AUROC
F1
AUROC
F1
AUROC
F1
Qwen2.5-3B-Instruct
CoT Monitor
-
0.641
-
0.622
-
0.602
0.18
-
0.627
-
0.384
-
0.522
0.18
TRACE
0.702
0.658
0.923
0.902
0.941
0.908
3.82
0.601
0.626
0.708
0.658
0.759
0.738
3.71
ConfLens
0.784
0.732
0.874
0.904
0.928
0.906
0.35
0.637
0.633
0.774
0.748
0.773
0.796
0.21
Table 2: Shortcut reasoning detection results across different methods. For TRACE and ConfLens , we optimize the decision threshold on the validation set that yields the best F1 score, thereby converting their outputs into binary detection results. The best results are bolded .
Figure 4: The accuracy and average faithfulness of accepted samples by different reward models and signal strategies. Adopting the detection results of ConfLens help the reward model identify the shortcut reasoning responses while maintaining the accuracy.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Category
Configuration
Data
Training Set Size: 20194 Validation Set Size: 3000 Test Set Size: 5799
Sequence Lengths
Max prompt length: 512 Max response length: 1024 Overlong prompts: filtered
Batching
Total epochs: 4 Train Batch Size: 1024 Micro Batch Size Per GPU: 16
Optimization
Learning rate: 1×10−6
GRPO
Rollout: 5 KL_coef: 0.001
Appendix
Table 3: Training configuration for math reasoning task.
Category
Configuration
Data
Training Set Size: 400 Validation Set Size: 80 Test Set Size: 320
Sequence Lengths
Max prompt length: 512 Max response length: 1024 Overlong prompts: filtered
Batching
Total epochs: 10 Train Batch Size: 1024 Micro Batch Size Per GPU: 16
Optimization
Learning rate: 1×10−6
GRPO
Rollout: 5 KL_coef: 0.001
Appendix
Table 4: Training configuration for code reasoning task.
Figure 5: The F1 score for shortcut reasoning detection in math reasoning task (left) and code reasoning task (right) under implicit hint setting. ConfLens consistently surpress the comparison methods.