cs.AISep 30, 2026

Search Shapes Conclusions: Auditing Evidence Selection Bias in Deep Research Agents

Authors: Shuyao Xiao, Shengling Wang, Xuan Chen, Ke Chao, Ming Cui, Feifei Qian, Chaoyang Mei, Fanlin Meng, +3 more

Organizations: Beijing Normal University · Ke Holdings

Abstract

Deep Research agents synthesize evidence into cited reports, yet a well-cited report can still reach a misleading conclusion. Citation correctness checks whether cited sources support individual claims. It does not show whether adaptive search exposed a representative view of all documents made available for evaluation, which we call the candidate pool. Early findings redirect later queries, document choices, and stopping, so the documents an agent reads form a selective sample. Existing evaluations rarely account for this selection. We formulate the problem as adaptive evidence sampling and introduce Causal Evidence Selection Correction (CESS). CESS predicts each candidate document's evidence direction and corrects the candidate-pool average using the logged probabilities of selecting each document and reaching each search round. Shrinkage stabilizes short searches, while intervals replace point estimates when some documents cannot be sampled. We also prove that estimating the average evidence direction of a common pool differs from measuring how a change in search policy alters the evidence read. The latter requires intervention. On questions from the MS2 systematic-review benchmark, CESS reduces mean absolute error against the candidate-pool average by 9.2%9.2\% and reduces the estimate's change under opposing document rankings by 39.4%39.4\% relative to averaging the evidence scores of documents read. Across trajectories from a public Open Deep Research agent, the corresponding reductions reach 60.1%60.1\% and 87.2%87.2\%. A further 4,800 trajectories under paired interventions confirm that correcting a pool estimate and measuring a policy effect are different tasks. CESS therefore audits whether the evidence direction underlying a report reflects the documents available for evaluation, while a separate intervention analysis measures the effect of search decisions.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. DRNOISE: Benchmarking Deep Research Agents in Misleading Evidence Environments

    Jul 19, 2026Jun Nie, Zhiqin Yang, Zhenheng Tang +4Deep ResearchFalsehood

  2. Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions

    Jul 23, 2026Pengyu Zhu, Lijun Li, Longju Yang +2Deep ResearchFalsehood

  3. DeepStress: Stress-Testing Deep Search Agents

    Jul 15, 2026Ismael Rousseau, Geraldine Damnati, Frederic BechetSearch AgentsStress Testing