cs.AIJan 30, 2026

Why Your Deep Research Agent Fails? On Hallucination Evaluation in Full Research Trajectory

Authors: Yuhao Zhan, Tianyu Fan, Linxuan Huang, Zirui Guo, Chao Huang

Organizations: Zhejiang University · The University of Hong Kong

Abstract

Diagnosing failure patterns in Deep Research Agents (DRAs) remains a critical challenge. Existing benchmarks predominantly rely on end-to-end evaluation, obscuring intermediate hallucinations that accumulate throughout the research trajectory. To bridge this gap, we propose a shift from outcome-based to process-aware evaluation by auditing hallucinations in the full plan-search-summarize trajectory. We introduce the PING Taxonomy, which categorizes DRA hallucinations into four complementary types: Propagation, Intent, Noise-induced, and Grounding. We further instantiate this taxonomy into a fine-grained evaluation framework that decomposes trajectories into atomic actions, claims, and sub-queries for rigorous verification, and we validate its reliability on standard fact-checking benchmarks and human-reviewed trajectories. Leveraging this framework to isolate 100 hallucination-prone tasks, including adversarial scenarios, we curate DeepHalluBench. Experiments on six representative DRAs show that, on our hallucination-prone stress-test set, all evaluated systems still exhibit non-negligible reliability gaps. Furthermore, our diagnostic analysis traces these failures to systemic deficits, especially hallucination propagation and cognitive biases, providing actionable insights for future architectural optimization. Code and data are available at https://github.com/yuhao-zhan/DeepHalluBench.

Figures & tables

Appendix figures & tables19 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. DR3^{3}-Eval: Towards Realistic and Reproducible Deep Research Evaluation

    Apr 16, 2026Qianqian Xie, Qingheng Xiong, He Zhu +16Deep ResearchMulti-Dimensional Evaluation

  2. Where Do Deep-Research Agents Go Wrong? Span-Level Error Localization in Agent Trajectories

    Jun 1, 2026Jiaming Wang, Ziteng Feng, Jiangtao Wu +8Deep ResearchModel Auditing

  3. DeepResearch Bench II: Diagnosing Deep Research Agents via Rubrics from Expert Reports

    Date pendingRuizhe Li, Mingxuan Du, Benfeng Xu +3Deep ResearchRubrics