Large language models often improve problem-solving performance by generating multi-step reasoning paths, yet how to analyze the hidden states along these paths remains unclear. Existing approaches typically assign each intermediate state the final-answer correctness label and train probes across heterogeneous questions. We argue that this protocol obscures reasoning dynamics in two ways: (1) correctness prediction can exploit question-level variation rather than path quality, and (2) states aligned by absolute step indices may correspond to different functional phases of reasoning. In this work, we propose to view reasoning paths as phase-structured trajectories within fixed questions. We instantiate this view as PAIR, short for Phase-Aligned Intra-question Reasoning. PAIR samples multiple trajectories for each question, maps variable-length paths into shared relative phases based on normalized trajectory progress, and compares successful and unsuccessful trajectories only within the same question and phase. This yields phase-specific path-quality directions that better isolate path-quality signals from question-level variation. Empirically, we find that standard across-question correctness probes lose much of their predictive power under within-question evaluation, suggesting that these probes partly rely on question-level information. PAIR improves within-question trajectory ranking and Best-of-N trajectory selection across models and benchmarks. Phase-wise steering further shows that the learned directions can change generation outcomes, providing causal evidence that they capture trajectory-relevant information.
Figures & tables
Figure 1: Overview of our motivation. (a) Prior probing labels intermediate states by final-answer correctness across different questions. (b) Across-question prediction can exploit question identity rather than path quality. (c) Absolute step indices can misalign reasoning phases across questions. (d) We instead align trajectories by phase and compare paths within the same question.
Figure 2: Overview of PAIR. (a) For each question, PAIR samples multiple reasoning trajectories and aligns variable-length paths into a fixed set of relative phase slots. (b) Within each phase, PAIR constructs successful–unsuccessful trajectory pairs from the same question and learns a phase-specific path-quality direction using a pairwise ranking objective. (c) The learned directions are used for phase-wise activation steering, where hidden states are shifted toward successful trajectory directions during generation.
Model
GSM8K ( Cobbe et al., 2021 )
MATH ( Lightman et al., 2023 )
BBH ( Suzgun et al., 2022 )
AUC
Δ
Balanced Acc.
Δ
AUC
Δ
Balanced Acc.
Δ
AUC
Δ
Balanced Acc.
Δ
LLaMA-3.1-8B-Instruct
0.72 / 0.56
-0.16
0.66 / 0.51
-0.15
0.65 / 0.49
-0.16
0.64 / 0.50
-0.14
0.72 / 0.53
-0.19
0.66 / 0.51
-0.15
Qwen3-4B
0.72 / 0.53
-0.19
0.73 / 0.64
-0.09
0.64 / 0.51
-0.13
0.60 / 0.57
-0.03
0.77 / 0.53
-0.24
0.76 / 0.51
-0.25
Gemma3-4B
0.65 / 0.50
-0.15
0.63 / 0.52
-0.11
0.68 / 0.54
-0.14
0.61 / 0.58
-0.03
0.71 / 0.51
-0.20
0.71 / 0.41
-0.30
DeepSeek-R1-Distill-Llama-8B
0.63 / 0.62
-0.01
0.62 / 0.59
-0.03
0.65 / 0.51
-0.14
0.64 / 0.56
-0.08
0.68 / 0.51
-0.17
0.66 / 0.57
-0.09
Average
0.68 / 0.55
-0.13
0.66 / 0.57
-0.10
0.66 / 0.51
-0.14
0.62 / 0.55
-0.07
0.72 / 0.52
-0.20
0.70 / 0.50
-0.20
Table 1: Across-question and within-question evaluation of correctness probes. For each benchmark, we report across-question (AQ) and within-question (WQ) AUROC and balanced accuracy, together with Δ=WQ−AQ . AQ/WQ entries are formatted as AQ / WQ .
Model
Method
GSM8K ( Cobbe et al., 2021 )
MATH ( Lightman et al., 2023 )
BBH ( Suzgun et al., 2022 )
WQ-AUC ↑
BoN@32 ↑
WQ-AUC ↑
BoN@32 ↑
WQ-AUC ↑
BoN@32 ↑
LLaMA-3.1 -8B-Instruct Meta (2024)
Pre-Ans token
0.71
83.3
0.83
46.6
0.65
84.0
Pooled correctness
0.66
84.6
0.69
36.0
0.51
77.3
Absolute-step
0.63
81.6
0.63
36.6
0.57
81.3
Normalized-step
0.67
85.3
0.71
40.0
0.58
68.0
PAIR
0.76
87.3
0.77
46.6
0.74
92.0
Table 2: Evaluation of trajectory-quality estimation and Best-of-32 trajectory selection. We highlight the best value in green with bold text and the second best value in blue .
Intervention
GSM8K
MATH
BBH
Acc.
W→C
C→W
Tok.
Acc.
W→C
C→W
Tok.
Acc.
W→C
C→W
Tok.
Baseline
79.3
–
–
234 ± 70
42.0
–
–
383 ± 129
74.6
–
–
146 ± 28
Single P1
83.3
32.2
3.3
242 ± 71
47.3
12.6
4.7
399 ± 181
76.0
57.8
17.9
167 ± 25
Single P2
82.6
25.8
2.5
233 ± 67
44.7
6.9
3.2
386 ± 121
70.6
5.2
7.1
187 ± 34
Single P3
80.0
3.2
0
234 ± 70
42.0
2.3
3.2
387 ± 170
69.3
36.8
19.6
205 ± 66
Single P4
79.3
0.0
0.0
234 ± 70
42.7
1.1
0
386 ± 123
76.0
10.5
1.8
198 ± 54
Table 3: Phase-wise steering results across benchmarks. Baseline denotes generation without intervention. Single Pb applies the phase- b direction only at phase Pb , while PAIR applies phase-specific directions progressively along the trajectory. We report final-answer accuracy, wrong-to-correct and correct-to-wrong flip rates, and generated-token length after steering.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
α
λmax
GSM8K
5
5
MATH-500
5
5
BBH
2.5
2.5
Appendix
Table 4: Steering hyperparameters selected by validation-set sweep. α controls the scale of the adaptive update, and λmax clips the maximum intervention magnitude.
Model
N
SC@ N (%)
Oracle@ N (%)
Gap (pp)
Llama-3.1-8B
1
68.2
68.2
0.0
Llama-3.1-8B
4
72.4
85.0
12.6
Llama-3.1-8B
8
75.0
90.6
15.6
Llama-3.1-8B
16
76.4
95.4
19.0
Llama-3.1-8B
32
75.0
97.4
22.4
Qwen3-4B
1
83.2
83.2
0.0
Appendix
Table 5: Self-consistency and oracle accuracy on the same fixed MMLU candidate sets. Each model uses 500 official-test questions; N is the number of raw draws retained per question. Gap is Oracle minus SC in percentage points.
Reasoning-trained language models often spend more tokens on harder problems, but longer chains of thought do not show whether a model is merely computing for more steps or following a different internal trajectory. We study this distinction through hidden-state trajectories during chain-of-thought generation across competitive programming, mathematics, and Boolean satisfiability. Raw trajectory geometry is strongly shaped by generation length: longer generations mechanically alter path statistics, so difficulty-dependent comparisons are misleading without adjustment. After residualizing trajectory statistics on length, difficulty remains systematically coupled to corrected trajectory geometry across all domains studied. The clearest reasoning-specific separation appears in the code domain, where harder problems show more direct corrected trajectories and less heterogeneous local curvature in reasoning-trained models than in matched instruction-tuned baselines. Corrected difficulty-geometry coupling is weaker, but still present, in mathematics and Boolean satisfiability. Prompt-stage linear probes do not mirror the code-domain separation, and behavioral annotations show that stronger corrected coupling co-occurs with strategy shifts and uncertainty monitoring. Together, these findings establish length correction as a prerequisite for generation-time trajectory analysis and show that reasoning training can be associated with distinct corrected trajectory geometry, with the strength of the effect depending on the domain.
Anders Gjølbye, Lars Kai Hansen, Sanmi Koyejo
Technical University of Denmark · Stanford University
Chain-of-thought reasoning provides a structured computation between a model's input and final answer. Yet it is often evaluated through endpoint accuracy, which ignores the path taken to reach that answer. An emerging line of work addresses this limitation using entropy profiles, which track how uncertainty evolves over the reasoning process but do not reveal which competing hypotheses account for that uncertainty. We introduce answer-distribution trajectories, a stochastic-dynamics-inspired representation that tracks the model's full predictive distribution over answers as reasoning unfolds. As a strictly finer representation than endpoint and entropy summaries, answer-distribution trajectories enable us to characterize a trace through a dynamical reasoning profile spanning exploration, revision, motion, and commitment, and to distinguish different dynamical mechanisms of reasoning success and failure. Across sixteen open-weight language models and four reasoning benchmarks, we show that traces with the same endpoint and similar entropy profiles can exhibit substantially different reasoning dynamics. We further find substantial variation in these dynamics both within and across models and tasks, with different objectives favoring different dynamical profiles. Additionally, we show that training and inference choices systematically reshape these profiles. Our results suggest that answer-distribution trajectories provide a rich framework for analysing and evaluating the dynamics of LLM reasoning.
Mar Gonzàlez I Català, Haitz Sáez de Ocáriz Borde, Davide Murari +3
Multi-path reasoning methods such as self-consistency (SC) sample K reasoning paths and choose the most frequent answer. However, their gains quickly plateau as K increases, and existing methods do not predict when this saturation will occur. We formalize multi-path LLM reasoning as a diversity combining problem from wireless communications: each path is a noisy channel observation, and the pairwise correlation of path correctness caps the design-effect effective sample size of the vote at a finite ceiling. Generalized least squares (GLS) analysis shows that, under exchangeability, the optimal symmetric linear combiner of latent embeddings is uniform, supporting majority vote as the natural default in standard SC while leaving room for weighting or pruning under heterogeneous prompt-template branches. Across 5 models and 12 benchmarks, prompt-template diversity reduces path correlation in 55 of 57 valid cells, with the strongest effect on open-ended QA. We derive an Adaptive-K rule that uses a four-path pilot to select K∗, retaining 96--103% of MV@K=32 accuracy across Math, QA, and NLU.
Guangsheng Yu, Litianyi Zhang, Qin Wang +5
University of Technology Sydney · The University of Sydney · CSIRO +2