The Error You See Is Not the Error You Made: Progression-aware Reasoning Origin for Reasoning Error Localization
Authors: Yiguo Wang, Ziyuan Yang, Yi Zou, Dan Lin, Rongsheng Li, Yi Zhang
Organizations: Xiaopeng Honors College, Nanchang Hangkong University, China · School of Cyber Science and Engineering, Sichuan University, China · School of Software, Nanchang Hangkong University, China · School of Computer Science and Technology, Harbin Engineering University, China
Verifying multi-step LLM reasoning requires more than determining whether a trace is correct: a useful verifier should identify where the reasoning first goes wrong. However, existing holistic methods provide little positional evidence, while forward sequential verification often treats the first rejected step as the error source. Under error propagation, this assumption can fail, since an earlier mistake may remain locally plausible and become observable only through its downstream consequences. We therefore rethink reasoning verification as a progression-aware error-source localization problem: rather than asking only where a reasoning trace first appears inconsistent, we ask which earlier step best explains how that inconsistency emerges along the trajectory. Based on this view, we propose Progression-aware Reasoning Origin (PRO), a training-free framework for first-error localization. PRO jointly models incoming support from the preceding context and outgoing compatibility with subsequent reasoning, selectively refines regions where these signals disagree, and finally performs detector-conditioned source attribution with intervention-based evidence to distinguish the true error origin from its propagated manifestations. We further formalize the gap between forward rejection and structural exposure, showing why incoming-side evidence alone is insufficient for reliable localization under error propagation. Experiments across open-form, medical, and structured reasoning tasks demonstrate consistent improvements over strong verification baselines, supporting progression-aware source attribution as a more faithful formulation of reasoning verification.
Figures & tables
Figure 1: The overview of the proposed PRO.
Method
Metric
Qwen2.5-32B -Instruct
Qwen2.5-72B -Instruct
DeepSeek -V3.2
Llama3.3-70B -Instruct
Holistic Verification
Correct Acc
98.45
93.26
79.27
94.30
Error Acc
45.89
65.70
80.68
68.60
PARC
Correct Acc
83.94
75.13
82.90
89.12
Error Acc
57.97
68.60
77.78
63.77
GoV
Correct Acc
94.82
97.41
95.34
93.26
Error Acc
68.12
71.01
79.71
68.60
Table 1: ProcessBench. Performance comparison on ProcessBench (%).
Method
Metric
Qwen2.5-32B -Instruct
Qwen2.5-72B -Instruct
DeepSeek -V3.2
Llama3.3-70B -Instruct
Holistic Verification
Correct Acc
98.67
84.89
33.33
95.56
Error Acc
17.78
27.11
44.89
11.56
PARC
Correct Acc
88.00
78.22
68.44
89.33
Error Acc
25.33
35.56
23.56
17.78
GoV
Correct Acc
81.78
80.44
59.11
88.44
Error Acc
27.56
28.89
33.78
12.44
Table 2: MedReason. Performance comparison on MedReason (%).
Problem Size
Llama3.3-70B-Instruct
GoV
PRO
N=2
99.40
98.37
N=4
90.59
94.13
N=6
72.30
79.82
N=8
46.62
56.54
Table 3: The results on Number Triangle Summation (%).
Backward
Refinement
Top- k Rerank
Err. Acc.
F1
×
×
×
56.28
68.64
×
✓
✓
58.94
69.28
✓
×
✓
50.72
64.51
✓
✓
×
44.93
57.78
✓
✓
✓
68.12
79.10
Table 4: The ablation study of PRO.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Figure A1: Representative MedReason case study. Type 1 diabetes reasoning trajectory. The true first error is at Step 2. GoV misattributes the failure to Step 5 where corruption becomes explicit, while PRO correctly identifies the early source error.
School of Electrical and Computer Engineering, Georgia Institute of Technology, USA · Department of Electrical and Computer Engineering, Bogazici University, Turkey