Superficial Reflection or Genuine Thought? A Fine-Grained Cognitive Analysis of Large Reasoning Models
Authors: Yuxiang Chen, Zuohan Wu, Ziwei Wang, Xiangning Yu, Xujia Li, Linyi Yang, Mengyue Yang, Jun Wang, +1 more
Organizations: University College London · AI Lab, the Yangtze River Delta, China · The Hong Kong University of Science and Technology (Guangzhou) · Tianjin University · The Hong Kong University of Science and Technology · Southern University of Science and Technology · University of Bristol
Motivated by the observed human-like behaviours in Large Reasoning Models (LRMs), this paper introduces a comprehensive taxonomy to characterise atomic reasoning steps and analyse the reasoning behaviours of LRMs. Grounded in human cognitive processes, we propose a taxonomy comprising five groups and seventeen categories. Through this taxonomy, we conduct an in-depth analysis of contemporary LRMs and distil four actionable takeaways for model optimisation. Most notably, we reveal that prevailing post-answer ``doublechecks'' are largely superficial and rarely yield substantive revisions. A targeted intervention further shows that explicitly eliciting richer reflection processes can substantially improve failed self-correction. To support this largescale study, we propose CAPO, an automated annotation method used to construct a dataset of 277,534 reasoning steps with strong agreement with human expert annotations. We further validate the main behavioural patterns on a newer reasoning model and a coding domain, demonstrating the broader applicability of the proposed taxonomy. All source code and data are available at https://github.com/hehepig4/psyche.
Figures & tables
Figure 1: A brief illustration of the proposed taxonomy from human cognitive perspectives.
Figure 2: Annotation consistency of CAPO and the RAG baseline. The boxes show the consistency distribution of candidate prompts on the training population, while the red line denotes the test consistency of the best validation-selected prompt.
Process
Human interval
Machine MD
Within interval
A.IO
[−0.067,+0.094]
+0.011
✓
S.HG
[−0.048,+0.018]
−0.041
✓
S.AR
[−0.061,+0.008]
−0.030
✓
R.SME
[−0.073,+0.092]
−0.018
✓
Table 1: Comparison of the principal behavioural effects estimated from human and CAPO annotations. MD denotes the mean difference between correct and incorrect CoTs. Full category-level results are provided in the appendix.
Name
Correct
Incorrect
Steps
MATH
963 (96.3%)
37 (3.7%)
90,006
AIME
810 (87.0%)
121 (13.0%)
177,687
ComS
1000
7,710
HMMT
9 (64.3%)
5 (35.7%)
4,375
AIME ⋆
6 (37.5%)
10 (62.5%)
5,466
ComS ⋆
30
254
Table 2: Number of annotated CoTs and reasoning steps. The top three rows use CAPO annotations, whereas the bottom three use human annotations. The core mathematical analysis excludes the two ComS rows.
Figure 3: Proportional differences between correct and incorrect CoTs. Red dots indicate statistically significant differences. The purple line denotes the overall proportion of each mental process.
Name
Pos. co.
Pos. inco.
P-value
S.HG
0.35
0.47 (+0.12)
4.2×10−7
S.AR
0.40
0.48 (+0.08)
6.3×10−3
Table 3: Average relative positions of Hypothesis Generation and Analogy Recall. Both occur significantly later in incorrect CoTs ( p<0.01 ).
Figure 4: Post-answer check statistics. An A → B transition denotes correctness before (A) and after (B) post-answer checking; for example, I → C denotes an incorrect answer corrected to a correct one.
CoTs
Max
Min
Average
Before Intervention
0.60
0.21
0.41
After Intervention
1.00
0.67
0.88
Table 4: PNS ( ↑ ) statistics before and after intervention. Higher PNS indicates that the retained steps have greater estimated necessity for the successful reasoning trajectory.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: The CAPO framework, where each prompt comprises separated constant, variable, and mutable areas.
Correctness
CoTs
Steps
Correct
1,803
740,399
Incorrect
118
87,965
Appendix
Table 5: Statistics of the MiniMax-M2 reasoning trajectories.
Correctness
CoTs
Steps
Correct
644
225,761
Incorrect
322
195,396
Appendix
Table 6: Statistics of the CodeForces reasoning trajectories.
Name
Pos. co.
Pos. inco.
P-value
S.HG
0.239
0.346
<10−6
S.AR
0.298
0.410
<10−6
Appendix
Table 7: Relative positions of S.HG and S.AR on CodeForces.
Category
Correct
Incorrect
Diff.
Sig.
Analysis
0.2480
0.2117
+0.0363
Yes
Inference
0.6291
0.6467
−0.0176
No
Judgment
0.1372
0.0799
+0.0573
Yes
Suggestion
0.1474
0.1507
−0.0033
No
Reflection
0.2330
0.2214
+0.0116
No
Appendix
Table 8: Correct–incorrect comparison after aggregating the seventeen fine-grained processes into five parent categories.
Large Reasoning Models (LRMs) improve performance by generating explicit intermediate reasoning traces through increased test-time compute, yet the assumption that longer reasoning is consistently beneficial remains under-examined. While recent evidence shows that additional reasoning can lead models to overthink, we ask: "Once a model has reached the correct answer, does further reasoning refine the solution, or deviate from it?" To study the dynamics after correctness, we introduce a prefix-level trajectory evaluation protocol grounded in reasoning sufficiency, defining the minimum reasoning budget required for a model to first generate the correct answer. This allows us to disentangle verbose overthinking, where additional reasoning is redundant but harmless, from harmful overthinking, where continued reasoning destabilizes an already-correct trajectory. Starting from multimodal benchmarks, we find that many instances considered reasoning-intensive require surprisingly little reasoning. Moreover, stopping at the first correct prefix improves accuracy over standard reasoning up to 21%, revealing that current models are limited not only by their ability to reason, but also by their inability to stop at the right time. Furthermore, while common efficiency strategies like early stopping substantially reduce verbose overthinking (up to 50%), they fail to mitigate harmful overthinking. Failure analysis reveals that correctness deviations are mainly driven by logical drift and visual reinterpretation. Finally, we show that our findings generalize to language-only reasoning benchmarks, highlighting harmful overthinking as a broader reliability risk. Code available at https://simonecaldarella.github.io/thinking-past-the-answer.
Simone Caldarella, Davide Talon, Rahaf Aljundi +2
University of Trento · 3Fondazione Bruno Kessler · 2Toyota Motor Europe
Large reasoning models (LRMs) have achieved remarkable success on complex tasks, yet their tendency to "overthink" leads to inefficiencies. Although "save-thinking" prompts are intended to mitigate this issue, we find that LRMs still frequently enter the "Still-thinking" mode instead of the expected "No-thinking" mode, especially on difficult queries. To analyze this behavioral divergence, we examine LRMs from three perspectives: confidence at the thinking-termination boundary, divergence in internal attention distributions, and attention allocation across prompt segments. We find that high perplexity is associated with later Still-thinking behavior, and that Still-thinking cases allocate more attention to the original question. Based on these observations, we propose an attention intervention method to regulate this behavior. While this intervention suppresses explicit thinking, it also causes a drop in accuracy, suggesting that the suppressed reasoning behavior is often useful for correctness. Our work provides confidence- and attention-level evidence for this behavior, highlighting the trade-off between instruction following, inference efficiency, and reasoning correctness.
Rongzhi Zhu, Yi Liu, Jiancheng Wang +6
State Key Laboratory for Novel Software Technology, Nanjing University, China · China Unicom Software Research Institute, China · National Institute of Healthcare Data Science, Nanjing University, China +1
Large reasoning models (LRMs) produce reasoning traces with non-linear structures, such as backtracking and self-correction, that complicate the evaluation and monitoring of the reasoning process. We introduce ReasoningFlow, a framework that captures the discourse structures of LRM reasoning traces into fine-grained directed acyclic graphs (DAGs). We develop and validate our annotation schema through careful manual annotation of 31 traces (2.1k steps), achieving high inter-annotator agreement, then scale to automatic annotation of 1,260 traces (247.7k steps) spanning three tasks (math, science, argumentation) and five models (Qwen2.5-32B-Inst, QwQ-32B, DeepSeek-V3, DeepSeek-R1, GPT-oss-120B). By analyzing ReasoningFlow graphs, we find: (1) LRMs exhibit structurally similar traces, despite being trained from different base models and potentially non-overlapping post-training data. (2) ReasoningFlow reveals diverse fine-grained reasoning behaviors (e.g., local verification, self-reflection, and assumptions) that can be used for better reasoning trace monitorability. (3) In LRMs, most of the erroneous steps are not used to derive final answers. (4) Mechanistic causal dependencies between steps do not reflect the language-level discourse structure. We release the dataset and code in: https://github.com/jinulee-v/reasoningflow.
Jinu Lee, Shivam Agarwal, Amruta Parulekar +3
The Grainger College of Engineering University of Illinois Urbana-Champaign