cs.AISep 11, 2026

Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures

Authors: Harsh Raj, David Lee, Anas Mahmoud, Renxiong Wang, Razvan-Gabriel Dumitru, Chenguang Wang, Tong Zhao, Yunzhong He, +2 more

Organizations: Scale AI · Scale AI Research

Abstract

The increasing deployment of AI agents in long-horizon tasks yields massive execution logs. Diagnosing failures within these records is crucial for reliability, as it transforms outcome-level signals into actionable interventions. The sheer scale of the data renders human review impractical, driving the need for automated root-cause attribution (RCA). However, automated RCA methods using LLMs suffer from low diagnostic accuracy, especially as execution traces grow larger. They struggle because relevant information is often sparse, distributed across distant actions, and disconnected from the visible failure, reducing root-cause attribution to a massive search problem. Existing RCA methods typically rely on one-shot LLM judgments to diagnose failures from execution traces. While effective for shorter trajectories, these judges tend to settle on a plausible diagnosis early, leaving critical evidence in longer traces unexamined. We introduce Continual Search, an iterative framework that nudges the judge, over successive turns, to keep searching for unresolved diagnostic evidence. We evaluate Continual Search across four existing RCA benchmarks. Recognizing the lack of massive execution traces in current benchmarks, we introduce MegaRCA-Mix to evaluate RCA at scale. MegaRCA-Mix provides a challenging testbed of 50 human-annotated failure trials spanning long-horizon, execution-heavy tasks. Across multiple benchmark suites and model families, Continual Search consistently improves attribution performance. On MegaRCA-Mix, for example, it improves Opus-4.8's F1 score by 29%, from 0.4710.471 to 0.6080.608. More interestingly, within the same model family, lower-tier models can even surpass their higher-tier counterparts, demonstrating that effective search supersedes raw model scale.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents

    Aug 5, 2026Zhixiang Liang, Yifei Liu, Yidan Huang +5Search AgentsModel Auditing

  2. OpenRCA 2.0: From Outcome Labels to Causal Process Supervision

    Jun 25, 2026Aoyang Fang, Yifan Yang, Jin'ao Shang +7Root Cause AnalysisReal-World Cloud Fault Injection Dataset

  3. AgentRx: Diagnosing AI Agent Failures from Execution Trajectories

    Feb 2, 2026Shraddha Barke, Arnav Goyal, Alind Khare +3Artificial Intelligence AgentsLatent Failure Patterns