Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures
Organizations: Scale AI · Scale AI Research
Abstract
The increasing deployment of AI agents in long-horizon tasks yields massive execution logs. Diagnosing failures within these records is crucial for reliability, as it transforms outcome-level signals into actionable interventions. The sheer scale of the data renders human review impractical, driving the need for automated root-cause attribution (RCA). However, automated RCA methods using LLMs suffer from low diagnostic accuracy, especially as execution traces grow larger. They struggle because relevant information is often sparse, distributed across distant actions, and disconnected from the visible failure, reducing root-cause attribution to a massive search problem. Existing RCA methods typically rely on one-shot LLM judgments to diagnose failures from execution traces. While effective for shorter trajectories, these judges tend to settle on a plausible diagnosis early, leaving critical evidence in longer traces unexamined. We introduce Continual Search, an iterative framework that nudges the judge, over successive turns, to keep searching for unresolved diagnostic evidence. We evaluate Continual Search across four existing RCA benchmarks. Recognizing the lack of massive execution traces in current benchmarks, we introduce MegaRCA-Mix to evaluate RCA at scale. MegaRCA-Mix provides a challenging testbed of 50 human-annotated failure trials spanning long-horizon, execution-heavy tasks. Across multiple benchmark suites and model families, Continual Search consistently improves attribution performance. On MegaRCA-Mix, for example, it improves Opus-4.8's F1 score by 29%, from to . More interestingly, within the same model family, lower-tier models can even surpass their higher-tier counterparts, demonstrating that effective search supersedes raw model scale.
Figures & tables
| Dataset | Description | Median tokens | Median size | |
| MegaRCA-Mix | First consequential error in Harbor Index evaluation records ( Shi et al., 2026 ) , labeled with the interaction-centric taxonomy ( Raj et al., 2026 ) . The mixture spans GAIA ( Mialon et al., 2024 ) , SWE-Bench ( Jimenez et al., 2024 ; Deng et al., 2025 ) , HLE ( Phan et al., 2025 ) , OpenRCA ( Xu et al., 2025 ) , and 15 additional benchmarks. | 50 | 286K | 1.05 MiB |
| TRAIL ( Deshpande et al., 2025 ) | Error location and category in GAIA ( Mialon et al., 2024 ) and SWE-Bench ( Jimenez et al., 2024 ) trajectories | 148 | 100K | 430 KiB |
| TELBench ( Wang et al., 2026a ) | Earliest harmful error span in deep-research trajectories from GAIA ( Mialon et al., 2024 ) , BrowseComp ( Wei et al., 2025 ) , and xbench ( Chen et al., 2025 ) | 33 | 39K | 144 KiB |
| AgentRx ( Barke et al., 2026 ) | Critical failure step in retail -bench ( Yao et al., 2024 ) trajectories | 29 | 7.2K | 23.9 KiB |
| Who&When ( Zhang et al., 2025 ) | Responsible agent and decisive error step in multi-agent trajectories over GAIA ( Mialon et al., 2024 ) and AssistantBench ( Yoran et al., 2024 ) | 126 | 2.4K | 8.9 KiB |
| GPT-5.5 | Opus-4.8 | |||||
| Dataset | Single-turn | Passive | Continual Search | Single-turn | Passive | Continual Search |
| MegaRCA-Mix | 0.356 | 0.569 | 0.569 | 0.471 | 0.479 | 0.608 |
| TRAIL (Joint Acc.) | 0.143 | 0.170 | 0.207 | 0.133 | 0.135 | 0.152 |
| TRAIL (Weighted F1) | 0.429 | 0.459 | 0.494 | 0.482 | 0.488 | 0.543 |
| TELBench | 0.182 | 0.212 | 0.242 | 0.091 | 0.091 | 0.152 |
| GPT-5.5 | Opus-4.8 | |||||
| Dataset | Single-turn | Passive | Continual Search | Single-turn | Passive | Continual Search |
| AgentRx | 0.310 | 0.345 | 0.241 | 0.414 | 0.379 | 0.379 |
| Who&When | 0.516 | 0.492 | 0.437 | 0.373 | 0.349 | 0.349 |
| Weighted F1 | Cost $ | |||
| Method | Opus-4.8 | GPT-5.5 | Opus-4.8 | GPT-5.5 |
| Single-turn | 0.482 | 0.429 | 1.59 | 2.22 |
| Self-consistency | 0.430 | 0.383 | 6.52 | 8.65 |
| Judge panel | 0.431 | 0.431 | 4.73 | 4.73 |
| Continual Search | 0.543 | 0.494 | 3.43 | 5.83 |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Joint Accuracy | Weighted F1 | |||||
| Run | Single-turn | Passive | Continual Search | Single-turn | Passive | Continual Search |
| 1 | 0.133 | 0.135 | 0.152 | 0.482 | 0.488 | 0.543 |
| 2 | 0.152 | 0.154 | 0.197 | 0.465 | 0.484 | 0.505 |
| 3 | 0.129 | 0.135 | 0.189 | 0.439 | 0.452 | 0.516 |
| 4 | 0.138 | 0.145 | 0.182 | 0.475 | 0.484 | 0.531 |