cs.AISep 29, 2026

Where Scientific Search Agents Fail: Decision-Checkpoint Auditing of Exposure and Inspection Attempts

Authors: Hongmin Li, Wanli Zhao

Organizations: School of Life Science and Technology, Institute of Science Tokyo Tokyo, Japan · Department of Computational Biology and Medical Sciences Graduate School of Frontier Sciences, The University of Tokyo Kashiwa, Japan · Department of Information Science Graduate School of Advanced Science and Engineering, Hiroshima University Higashihiroshima, Japan

Abstract

Final-answer accuracy does not reveal whether a scientific-search agent failed to encounter a target paper, attempt to inspect it, or return an accepted answer after inspection. We introduce decision checkpoints that record observations and tool actions without benchmark labels during inference, then join target identities and evaluator labels to assign outcome categories from recorded events. Across five conditions on 540 answerable AutoResearchBench Deep questions in a fixed, target-enriched environment, keyword search achieves 24.6% accuracy, compared with 17.8% for raw search. The keyword condition has fewer incorrect answers with neither target exposure nor inspection, but more incorrect answers after the target is exposed and left uninspected. Compared with keyword search, read-first has 27.4% more recorded evidence-search calls. Target inspection attempts occur on 199 questions under read-first and 191 under keyword search; both conditions achieve 24.6% accuracy. The checkpoint protocol makes these question-level differences explicit, distinguishing target exposure and inspection from aggregate accuracy and total tool use.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Diagnosing Search Behavior and Failure Modes in Long-Horizon Search Agents

    Aug 3, 2026Qi Liu, Jiaxin Mao, Fengbin Zhu +1Search AgentsRelevance

  2. SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents

    Aug 5, 2026Zhixiang Liang, Yifei Liu, Yidan Huang +5Search AgentsModel Auditing

  3. Search, Inspect, Fetch: Exploiting Structure-Aware Boolean Retrieval for Deep-Research Agents

    Aug 3, 2026Shuai Wang, Haodong Chen, Yu Yin +3RetrieversFetch