Deep Research agents synthesize evidence into cited reports, yet a well-cited report can still reach a misleading conclusion. Citation correctness checks whether cited sources support individual claims. It does not show whether adaptive search exposed a representative view of all documents made available for evaluation, which we call the candidate pool. Early findings redirect later queries, document choices, and stopping, so the documents an agent reads form a selective sample. Existing evaluations rarely account for this selection. We formulate the problem as adaptive evidence sampling and introduce Causal Evidence Selection Correction (CESS). CESS predicts each candidate document's evidence direction and corrects the candidate-pool average using the logged probabilities of selecting each document and reaching each search round. Shrinkage stabilizes short searches, while intervals replace point estimates when some documents cannot be sampled. We also prove that estimating the average evidence direction of a common pool differs from measuring how a change in search policy alters the evidence read. The latter requires intervention. On questions from the MS2 systematic-review benchmark, CESS reduces mean absolute error against the candidate-pool average by 9.2% and reduces the estimate's change under opposing document rankings by 39.4% relative to averaging the evidence scores of documents read. Across trajectories from a public Open Deep Research agent, the corresponding reductions reach 60.1% and 87.2%. A further 4,800 trajectories under paired interventions confirm that correcting a pool estimate and measuring a policy effect are different tasks. CESS therefore audits whether the evidence direction underlying a report reflects the documents available for evaluation, while a separate intervention analysis measures the effect of search decisions.
Figures & tables
Figure 1: How earlier evidence shapes later queries, document selection (“Inclusion” in the diagram), and stopping. The naive estimate averages the evidence scores of documents read. This is the Opened Mean defined in the text. CESS estimates the candidate-pool target from the logged trajectory, while paired interventions measure how changing search decisions changes the average evidence direction among the documents read.
Figure 2: Ranking sensitivity with the candidate evidence held fixed. (a) In a controlled simulation, the unshrunk sequential correction reduces sensitivity from 0.7478 to 0.0014 before shrinkage and clipping. (b) On MS2 review questions, Task-wise CESS reduces sensitivity for both query models. Lower is better.
Method
MAE
Ranking sensitivity
Worst-ranking MAE
Tuning-set Mean
0.2733
0.0000
0.2733
Outcome Regression
0.2571
0.0000
0.2571
Opened Mean
0.1743
0.2445
0.2522
Sequential IPW
0.1883
0.2573
0.2727
Sequential Doubly Robust
0.1893
0.2626
0.2763
CESS
0.1583
0.1481
0.2165
Table 1: Accuracy and sensitivity to document order on MS2, averaged across the two query models. Shrinkage settings are chosen on separate tuning questions and kept unchanged on the test questions. Worst-ranking MAE is error under the less favorable of the two rankings. Lower is better.
Comparison
CESS lower sensitivity
No detected difference
CESS higher sensitivity
At matched MAE
10
18
0
Table 2: Ranking sensitivity at matched MAE across 28 dataset–model–budget–prediction settings. Each count records the direction of the paired comparison using a 95% question-level bootstrap interval. No multiplicity adjustment is applied.
Figure 3: MAE across search budgets for Opened Mean and Task-wise CESS. Percentages report relative reductions within a round. Gains are largest when only one or two documents can be opened.
Method
MAE
Ranking sensitivity
Direction error
Opened Mean
0.5727
0.5849
0.4907
Outcome Regression
0.2404
0.0000
0.3333
Sequential Doubly Robust
0.3053
0.1492
0.3287
CESS
0.2287
0.0746
0.3102
Table 3: Transfer to a public Open Deep Research agent on 36 PERSPECTRUM tasks, two evidence rankings, and three seeds (216 trajectories). Results use the primary budget K=2 . Lower is better.
Model
Sel. ICC
Stop. ICC
Contrast MAE (adaptive)
Rank corr. (adaptive)
Contrast MAE (fixed)
OLMo3-7B
0.679
0.099
0.151
0.269
0.193
Qwen2.5-32B
0.685
0.102
0.146
0.236
0.197
Table 4: Paired LLM search experiment on 200 questions per model. Intraclass correlation (ICC) measures whether the directly observed effects repeat across three runs. Contrast MAE and rank correlation compare the difference between two CESS estimates with the directly observed selection effect. “Adaptive” lets later queries respond to selected evidence. “Fixed” replays a common query sequence. Higher ICC and rank correlation, and lower MAE, are better.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Selection
Stopping
Full correction
Mechanism removed
No
No
−0.0032(0.0031)
−0.0032(0.0031)
No
Informative
−0.0032(0.0052)
−0.0086(0.0028)
Biased
No
−0.0049(0.0034)
−0.0802(0.0036)
Biased
Informative
−0.0103(0.0091)
−0.0876(0.0096)
Appendix
Table 5: Bias in the unshrunk estimator under controlled selection and stopping. The comparison removes only the probability correction for the decision being tested. Parentheses give standard errors calculated across questions, with repeated runs of each question kept together.
Metric
Opened Mean
CESS
Change
MAE ↓
0.4734
0.4250
−10.2%
RMSE ↓
0.5779
0.5240
−9.3%
Task-level bias ↓
0.4280
0.3638
−15.0%
Ranking sensitivity ↓
1.0651
0.8965
−15.8%
Direction-error rate ↓
0.4190
0.4213
+0.23 pp
Appendix
Table 6: Aggregate results on 36 PERSPECTRUM topics. Bold values indicate the better result for each metric.
Component
Configuration
Opened Mean
CESS
Reduction
Query generator
Qwen
0.4834
0.4344
10.1%
Query generator
Llama
0.4634
0.4157
10.3%
Search controller
Agent
0.4885
0.4364
10.7%
Search controller
MMR
0.4583
0.4137
9.7%
Appendix
Table 7: PERSPECTRUM MAE across query generators and search controllers. Lower is better.
Method
MAE
S
Direction disagreement
Opened Mean
0.4108
0.9574
0.4186
Outcome Regression
0.4200
0.0000
0.6667
Tuning-set Mean
0.2787
0.0000
0.3889
Opened Mean shrunk to Tuning Mean
0.2427
0.2393
0.3893
CESS
0.3103
0.1608
0.5208
Appendix
Table 8: PERSPECTRUM results for searches of up to ten rounds. S measures the difference between rankings after averaging repeated runs. Lower is better.
Variant
Selection correction
MAE
S
D
A (clip, then shrink)
Yes
0.3103
0.1608
0.5208
C (clip, then shrink)
No
0.3071
0.4346
0.5278
B (shrink, then clip)
Yes
0.3279
0.1977
0.5208
D (shrink, then clip)
No
0.3111
0.4650
0.5278
Appendix
Table 9: Four calculations on the same ten-round search runs, varying the order of clipping and shrinkage and whether document-selection weights are used. All retain the same initial prediction, coefficient 0.5, and stopping correction. S follows Eq. 16 . D is the fraction of conclusions with a different direction from the reference.
Round
Opened Mean
CESS
1
0.8424
0.4035
2
0.6189
0.3841
4
0.5039
0.3588
6
0.4631
0.3338
8
0.4305
0.3199
10
0.4108
0.3103
Appendix
Table 10: MAE after each search round, including runs that stopped earlier.
Matched quantity
CESS lower
No detected difference
CESS higher
Matched MAE, compare sensitivity
10
18
0
Matched sensitivity, compare MAE
3
15
10
Appendix
Table 11: Interpolated comparisons across 28 dataset–model–budget–prediction settings. “No detected difference” denotes a 95% interval containing zero. The counts are descriptive and are not adjusted for multiple comparisons.
Selection prob.
Prediction
λ
MAE
S
Min. prob.
Exact
Oracle rule
0.05
0.0453
0.0137
0.00667
Exact
Linear (all features)
0.05
0.0576
0.0172
0.00667
Exact
Linear (reduced)
0.05
0.0768
0.0225
0.00667
Exact
Constant
0.20
0.1854
0.0957
0.00667
Estimated selection prob.
Linear (all features)
0.05
0.0570
0.0137
0.01887
Limited overlap
Linear (all features)
0.00
0.0583
0.0000
0.000006
Appendix
Table 12: Selected results from the controlled simulation. Each row uses 8,000 test questions and a coefficient chosen on separate tuning questions. Lower MAE and ranking sensitivity S are better.
Deep research agents increasingly operate over the open web, where relevant records coexist with redundant summaries, outdated reports, and misleading documents. Existing evaluations offer limited insight into whether agents preserve sound evidential standards when an ordinary-looking false document is deliberately seeded into a searchable environment and offers a direct shortcut to a conflicting answer. We introduce DRNOISE, a 100-task benchmark for answer recovery under misleading evidence. Each task has a unique gold answer supported by two corroborating indirect record chains; the paired noisy condition adds one plausible document that states a conflicting answer directly. The benchmark spans ten families of evidence operations. Across agents with strong clean-task performance, this single intervention causes 66-88 percentage-point accuracy drops. Trace analyses identify verification inertia as the dominant failure mode: agents often retrieve truthful records but stop before completing and reconciling the evidence chain, instead deferring to the answer-like document. Generic verification prompts reduce but do not close this gap. The setting is especially relevant to open-web deployment, where plausible falsehoods arrive through ordinary-looking pages rather than explicit attacks. Reliable deep research therefore requires more than retrieval and citation; it requires active reconciliation of direct claims with record-level evidence.
Jun Nie, Zhiqin Yang, Zhenheng Tang +4
Hong Kong Baptist University · University of Science and Technology of China · The Hong Kong University of Science and Technology
Deep Research agents conduct long-horizon investigations by iteratively planning, retrieving evidence, and generating reports. However, it remains unclear whether they can resist apparently credible but factually false information introduced into these workflows. To study this failure mode, we introduce MisKnow-Agent, a controlled evaluation framework that constructs task-specific documents supporting manually audited false conclusions with controlled authority cues and source styles. Applied to the tasks from DeepResearch Bench, it generates 5,933 misleading documents after filtering. We evaluate DeerFlow and WebThinker with three backbone LLMs, together with Gemini Deep Research, using a report-level false-conclusion adoption rate (FCAR) that counts only reports endorsing the false conclusion. Across the configurations, introducing one misleading document increases the mean FCAR from 0% in the no-injection control to 54.7%. FCAR varies substantially with lifecycle stage and framework design, and also with source authority and presentation style, whereas search-result rank and additional documents beyond the first have limited influence. Although cross-model verification consistently classifies retained instances as misleading, Deep Research agents can still adopt the corresponding false conclusions during long-horizon research. Pre- and post-research defenses reduce FCAR but do not eliminate adoption, motivating continuous verification when evidence enters intermediate research states and final synthesis. To facilitate reproducibility, our code and dataset are publicly available at https://github.com/whfeLingYu/MisKnow-Agent and https://huggingface.co/datasets/whfeLingYu/Misleading_Knowledge, respectively.
Pengyu Zhu, Lijun Li, Longju Yang +2
1Beijing University of Posts and Telecommunications · 2Shanghai Artificial Intelligence Laboratory · 3Chongqing University of Posts and Telecommunications
While search agents demonstrate impressive capabilities in multi-step question answering, their robustness to poor-quality evidence remains under-explored. This phenomenon occurs rarely in realistic benchmarks but can lead to dramatic failure in real life applications. Therefore in this study we propose DeepStress, a stress testing framework that controls the frequency of challenging evidence by replacing the retrieval module of search agents with a controlled synthetic environment. We use this framework to control three dimensions that can affect document reliability: trustworthiness, relevance, and factuality. Testing several search agents on HotpotQA and BrowseCompPlus, we demonstrate that agents exhibit substantial differences in their ability to handle unreliable information and propose new metrics that better document systems outcomes as well as the interactions between conflicting parametric and retrieved knowledge.