In LLM-based agent systems, failures can originate from early steps whose effects propagate through subsequent interactions, making their origins difficult to identify. To trace such failures back to their origin, failure attribution has been formulated as the task of identifying the earliest step responsible for the failure. Recent methods leverage LLM internal signals for failure attribution, typically using hidden states as step representations. We therefore conduct an empirical study to evaluate how effectively these representations distinguish root-cause steps from other steps and find limited separation. Motivated by this observation, we propose ReCast, a step representation learning method that transforms hidden states from a frozen LLM into attribution-oriented step representations. ReCast first selects attribution-relevant layers, then constructs complementary pattern and deviation features, and finally learns contextualized step representations through an encoder trained with contrastive and ranking objectives. We also introduce ReCast-2K, a training dataset for failure attribution. ReCast achieves the best Hit@1 across four benchmarks, surpassing the strongest baseline by 5.65 and 9.19 pp on Who&When Algorithm and Handcrafted, respectively. Code is available at https://anonymous.4open.science/r/ReCast-5FB6 .
Figures & tables
Figure 1: Comparison of text embeddings and last-layer hidden-state representations on failed trajectories from ReCast-2K and the Handcrafted and Algorithm subsets of Who&When. Lower ARS/GRS and higher SI indicate better separation. See Appendix A for metric definitions.
Figure 2: Overview of ReCast, which learns attribution-oriented step representations from frozen LLM hidden states for failure attribution.
Category
Method
Who&When
TraceElephant
Algorithm
Handcrafted
Captain
Magentic
Heuristic
First
16.13 ± 0.00
1.72 ± 0.00
7.06 ± 0.00
6.67 ± 0.00
Random
16.40 ± 4.85
5.75 ± 0.81
12.55 ± 6.10
5.93 ± 2.62
LLM-based
All-at-Once (Qwen)
26.61 ± 0.00
3.45 ± 0.00
7.06 ± 0.00
4.44 ± 0.00
All-at-Once (DS)
22.04 ± 0.76
10.34 ± 2.44
12.55 ± 1.11
9.63 ± 0.52
All-at-Once (GPT)
27.42 ± 0.66
18.39 ± 5.69
14.90 ± 4.74
8.15 ± 0.52
Table 1: Hit@1 attribution accuracy (%) on Who&When and TraceElephant. Results from three random seeds are summarized as mean ± population standard deviation. The LLMs are Qwen3.5-27B (Qwen), DeepSeek V4 Flash (DS), and GPT-5.6 Luna (GPT). Best results are shown in bold , and second-best results are underlined .
Category
Method
AFTraj-2K
Who&When Pro
Heuristic
First
0.00 ± 0.00
7.35 ± 0.00
Random
11.45 ± 0.77
25.69 ± 0.69
LLM-based
All-at-Once (Qwen)
20.86 ± 0.00
21.10 ± 0.00
All-at-Once (DS)
30.47 ± 5.02
25.95 ± 1.29
All-at-Once (GPT)
19.02 ± 0.87
33.33 ± 1.27
Step-by-Step (Qwen)
51.53 ± 0.00
60.33 ± 0.04
Table 2: Hit@1 attribution accuracy (%) on the two in-domain benchmark settings. Results from three random seeds are summarized as mean ± population standard deviation. Best results are shown in bold , and second-best results are underlined .
Figure 3: Cumulative Hit@ k accuracy on Who&When and TraceElephant. Stacked segments denote ranks 1, 2–3, and 4–5; error bars show the population standard deviation over three seeds.
Component
Variant
Hit@1
Δ Hit@1 (pp)
None
Full model
41.58 ± 1.04
—
Layer selection
All 64 layers
38.83 ± 2.30
−2.75
Uniform 8
38.10 ± 2.55
−3.48
Last 8
37.55 ± 0.93
−4.03
Random 8
37.18 ± 0.69
−4.40
Global probe top-8
38.83 ± 1.13
−2.75
Table 3: Ablation results on Who&When. Hit@1 (%) over all trajectories is reported as mean ± standard deviation across three seeds. Δ Hit@1 reports the change from the full-model mean in percentage points (pp).
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Representation
ARS ↓
GRS ↓
SI ↑
Training
Text embedding
0.5478
0.6173
0.1482
Hidden[-1] (Mean)
0.7254
0.7630
0.0131
Hidden[-1] (Last)
0.7832
0.7686
0.0228
Projected features
0.5763
0.6001
0.0566
ReCast
−0.2590±0.2116
−0.2749±0.1977
0.7032±0.1804
Handcrafted
Text embedding
0.5195
0.5885
0.0959
Appendix
Table 4: Separation of step representations measured by ARS, GRS, and SI. ReCast results show mean ± population standard deviation across seeds 6 , 20 , and 42 ; other representations are fixed. Best results are bolded for each dataset and metric, based on mean values for ReCast.
Task Category
# Trajectories
# Steps
# Tokens
Total
Train
Test
Avg.
Max
Avg.
Max
Embodied
311
249
62
30.0
30
1,444.6
1,636
STEM
1,517
1,213
304
5.6
30
4,199.5
23,786
Data Science
2,446
1,958
488
3.7
17
609.1
6,861
Deep Search
1,641
1,313
328
5.6
31
1,551.9
21,948
Coding
342
273
69
3.0
3
3,127.6
7,490
Appendix
Table 5: Composition of the Who&When Pro text-modality subset used in our experiments. Step counts refer to aligned model-action steps, and token counts are computed over the corresponding inputs.
Benchmark
MAS
Model
Initial
Final Dataset
Runs
Succ.
Fail.
Total
AssistantBench
Captain-Agent
DeepSeek V4 Flash
165
20
53
73
GPT-4o-0806
165
14
38
52
Qwen3-235B
165
20
53
73
Subtotal
495
54
144
198
GAIA
Captain-Agent
DeepSeek V4 Flash
451
257
189
446
Appendix
Table 6: Composition of our source training dataset. Initial runs are the aligned executions before filtering. The final total consists of retained successful trajectories and failed trajectories with root-cause annotations. Additional generation runs are grouped under Others.
Model
Type
# Trajectories
# Steps
# Tokens
Avg.
Max
Avg.
Max
DeepSeek V4 Flash
Successful
454
15.7
61
7,005.5
59,153
Failed
500
21.6
55
10,603.5
74,773
GPT-4o-0806
Successful
225
18.5
60
4,460.0
22,757
Failed
515
32.0
61
7,996.1
29,094
Qwen3-235B
Successful
359
18.5
62
4,568.5
21,378
Appendix
Table 7: Statistics of ReCast-2K. Steps count recorded model responses, while token counts cover the question prefix and serialized trajectory using the Qwen3.5-27B tokenizer. Overall averages are computed over all trajectories.
Component
Hyperparameter
Value
Backbone
Hidden-state dimension dh
5,120
Layer selection
Cross-validation folds
5
Layer budget k
8
Ridge ratio
0.10
Selected layers
{7,15,19,32,33,47,49,57}
Projection
Dimension per branch dp
2,048
Appendix
Table 8: Hyperparameter settings for our method.
Method
LR
Weight decay
Batch
Budget
OAT
10−4
10−5
32
300 epochs
StepFinder
10−3
10−5
16
50 epochs
ASCon
5×10−5
10−5
1
12 epochs
Appendix
Table 9: Optimization settings for trainable feature-based baselines.
Figure 4: MRR@ K across rank cutoffs K=1,…,5 on the four benchmarks. Curves show the cumulative reciprocal-rank score as the cutoff increases; error bars denote the population standard deviation over three seeds.
Figure 5: Receiver operating characteristic (ROC) curves on the four benchmarks. Each curve is computed by pooling the candidate-step predictions within a benchmark. Solid lines show the mean over three seeds, and shaded regions indicate the population standard deviation.
Figure 6: Precision–recall (PR) curves on the four benchmarks. Each curve is computed by pooling the candidate-step predictions within a benchmark. Solid lines show the mean over three seeds, and shaded regions indicate the population standard deviation.
Backbone
# Layers
Hidden size
Selected layers
Qwen3.5-0.8B
24
1,024
3, 5, 8, 11, 13, 18, 19, 22
Qwen3.5-9B
32
4,096
3, 7, 11, 16, 17, 23, 26, 31
Qwen3.5-27B
64
5,120
7, 15, 19, 32, 33, 47, 49, 57
Appendix
Table 10: Hidden-state backbone configurations. Each backbone retains eight layers selected independently using the same source-only procedure.
Evaluation population
Backbone
Hit@1 (%)
Hit@3 (%)
Hit@5 (%)
MRR (%)
AUROC (%)
AUPRC (%)
Who&When Handcrafted
Qwen3.5-0.8B
30.46 ± 2.15
41.95 ± 2.15
55.75 ± 0.81
41.80 ± 1.07
80.00 ± 2.12
19.09 ± 0.84
Qwen3.5-9B
29.89 ± 3.54
44.83 ± 3.72
52.87 ± 4.30
41.61 ± 0.83
79.94 ± 2.41
18.55 ± 2.91
Qwen3.5-27B
33.33 ± 1.63
45.98 ± 0.81
57.47 ± 0.81
44.21 ± 0.77
79.22 ± 0.91
21.88 ± 0.92
Who&When Algorithm
Qwen3.5-0.8B
38.17 ± 2.31
74.73 ± 2.31
91.94 ± 1.14
59.20 ± 2.00
72.00 ± 0.38
33.42 ± 0.43
Qwen3.5-9B
38.98 ± 0.38
74.46 ± 2.01
92.74 ± 2.37
59.85 ± 0.67
73.64 ± 1.58
35.10 ± 0.47
Qwen3.5-27B
45.43 ± 2.01
72.31 ± 1.01
92.47 ± 1.66
63.28 ± 1.42
74.75 ± 1.00
38.52 ± 1.84
Appendix
Table 11: Effect of the hidden-state extraction backbone. Results report mean ± population standard deviation over three encoder seeds. Bold indicates the best mean within each evaluation population, determined before rounding.
Figure 7: Hit@1 accuracy (%) versus measured time per trajectory on Who&When and TraceElephant. Points show three-seed means; time is in seconds on a logarithmic scale. Feature-based methods report attribution time excluding feature extraction, while LLM methods report inference time excluding backend initialization. Upper-left positions indicate higher accuracy at lower measured cost. A and S denote All-at-Once and Step-by-Step.
Figure 8: Step-level attribution scores on three HC trajectories (a–c) and three AG trajectories (d–f). Original zero-based step indices are retained; gaps indicate events without candidate scores.
LLM-based agents increasingly solve complex tasks through long trajectories involving reasoning steps, tool calls, and inter-agent communication. However, when these agents fail, it is often unclear which agent caused the failure and which step introduced the decisive error. This attribution problem is challenging because mistakes can propagate across the trajectory: later actions may appear incorrect, but only because they depend on an earlier corrupted state. Therefore, failure attribution cannot be treated as independent step-level classification. We propose FALAT, a diagnostic framework for failure attribution in LLM agent trajectories. FALAT frames attribution as a dependency-guided search problem. It first constructs an expectation of how the task should be solved and uses this expectation to identify suspicious regions in the trajectory. It then traces dependencies among decisions, tool outputs, and agent messages to distinguish error-introducing steps from steps that merely inherit or propagate prior mistakes. Finally, FALAT evaluates whether correcting a candidate step would be sufficient to recover the expected outcome, allowing it to identify both the responsible agent and the decisive failure step. We evaluate FALAT on the Who&When benchmark, which includes both algorithm-generated and hand-crafted multi-agent failure trajectories. The results show that FALAT consistently improves responsible-agent and decisive-step attribution. Its best configurations achieve 46.0% step-level accuracy on algorithm-generated trajectories and 29.1% on the more challenging hand-crafted trajectories, outperforming specialized attribution baselines and direct prompting with standalone LLMs. These findings suggest that dependency-aware reasoning is essential for reliable failure diagnosis in LLM agent systems.
Md Nakhla Rafi, Md Ahasanuzzaman, Dong Jae Kim +2
SPEAR Lab Concordia University Montreal, Canada · DePaul University Chicago, USA
LLM-based multi-agent systems (MASs) are increasingly used to solve complex tasks through coordinated reasoning, tool use, and interaction with external resources. However, attributing failures in such systems remains challenging because the observed outcome often does not directly reveal the error responsible for the failed execution. In this work, the attribution target is the decisive error, defined as the agent--step pair whose correction would recover the failed execution. Existing approaches largely identify suspicious steps without explicitly modeling how errors propagate across interactions or persist in unresolved loops, making decisive errors difficult to distinguish from downstream failure symptoms. We propose \textbf{E}rror-Propagation \textbf{M}odeling for \textbf{F}ailure \textbf{A}ttribution (\textbf{EMFA}). EMFA constructs a structured representation of the failed trajectory, models both cascading propagation and persistent interaction loops, and uses propagation-aware candidate screening followed by counterfactual verification to identify the decisive agent--step pair. On the Who&When benchmark, EMFA achieves state-of-the-art step-level attribution accuracy and remains competitive at the agent level. It improves the previous best step-level results by 3.45 and 4.40 percentage points on the Hand-Crafted and Algorithm-Generated subsets, respectively.
Jiaqi Liao, Yuanzhao Zhai, Huanxi Liu +5
College of Computer Science and Technology, National University of Defense Technology; State Key Laboratory of Complex & Critical Software Environment, Changsha, Hunan, China · School of Mechanical Engineering, Tianjin University, Tianjin, China · College of Computer Science and Technology, National University of Defense Technology; National Key Laboratory of Parallel and Distributed Computing, Changsha, Hunan, China
LLM-based multi-agent systems exhibit remarkable collaborative capabilities in complex multi-step tasks. However, these systems are highly sensitive to single-step execution errors that can propagate through agent interactions and lead to cascading failures. To understand the causes of failure and improve system reliability, failure attribution has been introduced as a task that aims to automatically identify the root cause step responsible for a failure. Existing failure attribution methods mainly rely on LLMs to reason over original execution trajectories, which not only incur high inference costs and latency, but also suffer from interference caused by redundant and noisy execution logs, causing LLMs to struggle in accurately identifying the true root cause step. To address this, we propose StepFinder, a lightweight failure attribution framework. We use LLMs solely during the feature construction phase to encode execution logs into temporal semantic sequences. Subsequently, a parameter-efficient combination of temporal modeling and attention modules is applied to capture the sequential evolution and cross-step dependencies of the trajectories. Finally, the step-level error score is refined through multi-scale differences and position bias, enabling precise root cause identification. Experimental results on the Who&When benchmark demonstrate that StepFinder outperforms LLM-based methods in step-level failure attribution while achieving substantially higher inference efficiency, reducing inference time by 79% compared with the fastest LLM-based method, with no text generation overhead. Our code is available at https://github.com/taiyu-zhu/StepFinder.