Large language models often improve problem-solving performance by generating multi-step reasoning paths, yet how to analyze the hidden states along these paths remains unclear. Existing approaches typically assign each intermediate state the final-answer correctness label and train probes across heterogeneous questions. We argue that this protocol obscures reasoning dynamics in two ways: (1) correctness prediction can exploit question-level variation rather than path quality, and (2) states aligned by absolute step indices may correspond to different functional phases of reasoning. In this work, we propose to view reasoning paths as phase-structured trajectories within fixed questions. We instantiate this view as PAIR, short for Phase-Aligned Intra-question Reasoning. PAIR samples multiple trajectories for each question, maps variable-length paths into shared relative phases based on normalized trajectory progress, and compares successful and unsuccessful trajectories only within the same question and phase. This yields phase-specific path-quality directions that better isolate path-quality signals from question-level variation. Empirically, we find that standard across-question correctness probes lose much of their predictive power under within-question evaluation, suggesting that these probes partly rely on question-level information. PAIR improves within-question trajectory ranking and Best-of-N trajectory selection across models and benchmarks. Phase-wise steering further shows that the learned directions can change generation outcomes, providing causal evidence that they capture trajectory-relevant information.
Figures & tables
Figure 1: Overview of our motivation. (a) Prior probing labels intermediate states by final-answer correctness across different questions. (b) Across-question prediction can exploit question identity rather than path quality. (c) Absolute step indices can misalign reasoning phases across questions. (d) We instead align trajectories by phase and compare paths within the same question.
Figure 2: Overview of PAIR. (a) For each question, PAIR samples multiple reasoning trajectories and aligns variable-length paths into a fixed set of relative phase slots. (b) Within each phase, PAIR constructs successful–unsuccessful trajectory pairs from the same question and learns a phase-specific path-quality direction using a pairwise ranking objective. (c) The learned directions are used for phase-wise activation steering, where hidden states are shifted toward successful trajectory directions during generation.
Model
GSM8K ( Cobbe et al., 2021 )
MATH ( Lightman et al., 2023 )
BBH ( Suzgun et al., 2022 )
AUC
Δ
Balanced Acc.
Δ
AUC
Δ
Balanced Acc.
Δ
AUC
Δ
Balanced Acc.
Δ
LLaMA-3.1-8B-Instruct
0.72 / 0.56
-0.16
0.66 / 0.51
-0.15
0.65 / 0.49
-0.16
0.64 / 0.50
-0.14
0.72 / 0.53
-0.19
0.66 / 0.51
-0.15
Qwen3-4B
0.72 / 0.53
-0.19
0.73 / 0.64
-0.09
0.64 / 0.51
-0.13
0.60 / 0.57
-0.03
0.77 / 0.53
-0.24
0.76 / 0.51
-0.25
Gemma3-4B
0.65 / 0.50
-0.15
0.63 / 0.52
-0.11
0.68 / 0.54
-0.14
0.61 / 0.58
-0.03
0.71 / 0.51
-0.20
0.71 / 0.41
-0.30
DeepSeek-R1-Distill-Llama-8B
0.63 / 0.62
-0.01
0.62 / 0.59
-0.03
0.65 / 0.51
-0.14
0.64 / 0.56
-0.08
0.68 / 0.51
-0.17
0.66 / 0.57
-0.09
Average
0.68 / 0.55
-0.13
0.66 / 0.57
-0.10
0.66 / 0.51
-0.14
0.62 / 0.55
-0.07
0.72 / 0.52
-0.20
0.70 / 0.50
-0.20
Table 1: Across-question and within-question evaluation of correctness probes. For each benchmark, we report across-question (AQ) and within-question (WQ) AUROC and balanced accuracy, together with Δ=WQ−AQ . AQ/WQ entries are formatted as AQ / WQ .
Model
Method
GSM8K ( Cobbe et al., 2021 )
MATH ( Lightman et al., 2023 )
BBH ( Suzgun et al., 2022 )
WQ-AUC ↑
BoN@32 ↑
WQ-AUC ↑
BoN@32 ↑
WQ-AUC ↑
BoN@32 ↑
LLaMA-3.1 -8B-Instruct Meta (2024)
Pre-Ans token
0.71
83.3
0.83
46.6
0.65
84.0
Pooled correctness
0.66
84.6
0.69
36.0
0.51
77.3
Absolute-step
0.63
81.6
0.63
36.6
0.57
81.3
Normalized-step
0.67
85.3
0.71
40.0
0.58
68.0
PAIR
0.76
87.3
0.77
46.6
0.74
92.0
Table 2: Evaluation of trajectory-quality estimation and Best-of-32 trajectory selection. We highlight the best value in green with bold text and the second best value in blue .
Intervention
GSM8K
MATH
BBH
Acc.
W→C
C→W
Tok.
Acc.
W→C
C→W
Tok.
Acc.
W→C
C→W
Tok.
Baseline
79.3
–
–
234 ± 70
42.0
–
–
383 ± 129
74.6
–
–
146 ± 28
Single P1
83.3
32.2
3.3
242 ± 71
47.3
12.6
4.7
399 ± 181
76.0
57.8
17.9
167 ± 25
Single P2
82.6
25.8
2.5
233 ± 67
44.7
6.9
3.2
386 ± 121
70.6
5.2
7.1
187 ± 34
Single P3
80.0
3.2
0
234 ± 70
42.0
2.3
3.2
387 ± 170
69.3
36.8
19.6
205 ± 66
Single P4
79.3
0.0
0.0
234 ± 70
42.7
1.1
0
386 ± 123
76.0
10.5
1.8
198 ± 54
Table 3: Phase-wise steering results across benchmarks. Baseline denotes generation without intervention. Single Pb applies the phase- b direction only at phase Pb , while PAIR applies phase-specific directions progressively along the trajectory. We report final-answer accuracy, wrong-to-correct and correct-to-wrong flip rates, and generated-token length after steering.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
α
λmax
GSM8K
5
5
MATH-500
5
5
BBH
2.5
2.5
Appendix
Table 4: Steering hyperparameters selected by validation-set sweep. α controls the scale of the adaptive update, and λmax clips the maximum intervention magnitude.
Model
N
SC@ N (%)
Oracle@ N (%)
Gap (pp)
Llama-3.1-8B
1
68.2
68.2
0.0
Llama-3.1-8B
4
72.4
85.0
12.6
Llama-3.1-8B
8
75.0
90.6
15.6
Llama-3.1-8B
16
76.4
95.4
19.0
Llama-3.1-8B
32
75.0
97.4
22.4
Qwen3-4B
1
83.2
83.2
0.0
Appendix
Table 5: Self-consistency and oracle accuracy on the same fixed MMLU candidate sets. Each model uses 500 official-test questions; N is the number of raw draws retained per question. Gap is Oracle minus SC in percentage points.