Deep Research agents synthesize evidence into cited reports, yet a well-cited report can still reach a misleading conclusion. Citation correctness checks whether cited sources support individual claims. It does not show whether adaptive search exposed a representative view of all documents made available for evaluation, which we call the candidate pool. Early findings redirect later queries, document choices, and stopping, so the documents an agent reads form a selective sample. Existing evaluations rarely account for this selection. We formulate the problem as adaptive evidence sampling and introduce Causal Evidence Selection Correction (CESS). CESS predicts each candidate document's evidence direction and corrects the candidate-pool average using the logged probabilities of selecting each document and reaching each search round. Shrinkage stabilizes short searches, while intervals replace point estimates when some documents cannot be sampled. We also prove that estimating the average evidence direction of a common pool differs from measuring how a change in search policy alters the evidence read. The latter requires intervention. On questions from the MS2 systematic-review benchmark, CESS reduces mean absolute error against the candidate-pool average by 9.2% and reduces the estimate's change under opposing document rankings by 39.4% relative to averaging the evidence scores of documents read. Across trajectories from a public Open Deep Research agent, the corresponding reductions reach 60.1% and 87.2%. A further 4,800 trajectories under paired interventions confirm that correcting a pool estimate and measuring a policy effect are different tasks. CESS therefore audits whether the evidence direction underlying a report reflects the documents available for evaluation, while a separate intervention analysis measures the effect of search decisions.
Figures & tables
Figure 1: How earlier evidence shapes later queries, document selection (“Inclusion” in the diagram), and stopping. The naive estimate averages the evidence scores of documents read. This is the Opened Mean defined in the text. CESS estimates the candidate-pool target from the logged trajectory, while paired interventions measure how changing search decisions changes the average evidence direction among the documents read.
Figure 2: Ranking sensitivity with the candidate evidence held fixed. (a) In a controlled simulation, the unshrunk sequential correction reduces sensitivity from 0.7478 to 0.0014 before shrinkage and clipping. (b) On MS2 review questions, Task-wise CESS reduces sensitivity for both query models. Lower is better.
Method
MAE
Ranking sensitivity
Worst-ranking MAE
Tuning-set Mean
0.2733
0.0000
0.2733
Outcome Regression
0.2571
0.0000
0.2571
Opened Mean
0.1743
0.2445
0.2522
Sequential IPW
0.1883
0.2573
0.2727
Sequential Doubly Robust
0.1893
0.2626
0.2763
CESS
0.1583
0.1481
0.2165
Table 1: Accuracy and sensitivity to document order on MS2, averaged across the two query models. Shrinkage settings are chosen on separate tuning questions and kept unchanged on the test questions. Worst-ranking MAE is error under the less favorable of the two rankings. Lower is better.
Comparison
CESS lower sensitivity
No detected difference
CESS higher sensitivity
At matched MAE
10
18
0
Table 2: Ranking sensitivity at matched MAE across 28 dataset–model–budget–prediction settings. Each count records the direction of the paired comparison using a 95% question-level bootstrap interval. No multiplicity adjustment is applied.
Figure 3: MAE across search budgets for Opened Mean and Task-wise CESS. Percentages report relative reductions within a round. Gains are largest when only one or two documents can be opened.
Method
MAE
Ranking sensitivity
Direction error
Opened Mean
0.5727
0.5849
0.4907
Outcome Regression
0.2404
0.0000
0.3333
Sequential Doubly Robust
0.3053
0.1492
0.3287
CESS
0.2287
0.0746
0.3102
Table 3: Transfer to a public Open Deep Research agent on 36 PERSPECTRUM tasks, two evidence rankings, and three seeds (216 trajectories). Results use the primary budget K=2 . Lower is better.
Model
Sel. ICC
Stop. ICC
Contrast MAE (adaptive)
Rank corr. (adaptive)
Contrast MAE (fixed)
OLMo3-7B
0.679
0.099
0.151
0.269
0.193
Qwen2.5-32B
0.685
0.102
0.146
0.236
0.197
Table 4: Paired LLM search experiment on 200 questions per model. Intraclass correlation (ICC) measures whether the directly observed effects repeat across three runs. Contrast MAE and rank correlation compare the difference between two CESS estimates with the directly observed selection effect. “Adaptive” lets later queries respond to selected evidence. “Fixed” replays a common query sequence. Higher ICC and rank correlation, and lower MAE, are better.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Selection
Stopping
Full correction
Mechanism removed
No
No
−0.0032(0.0031)
−0.0032(0.0031)
No
Informative
−0.0032(0.0052)
−0.0086(0.0028)
Biased
No
−0.0049(0.0034)
−0.0802(0.0036)
Biased
Informative
−0.0103(0.0091)
−0.0876(0.0096)
Appendix
Table 5: Bias in the unshrunk estimator under controlled selection and stopping. The comparison removes only the probability correction for the decision being tested. Parentheses give standard errors calculated across questions, with repeated runs of each question kept together.
Metric
Opened Mean
CESS
Change
MAE ↓
0.4734
0.4250
−10.2%
RMSE ↓
0.5779
0.5240
−9.3%
Task-level bias ↓
0.4280
0.3638
−15.0%
Ranking sensitivity ↓
1.0651
0.8965
−15.8%
Direction-error rate ↓
0.4190
0.4213
+0.23 pp
Appendix
Table 6: Aggregate results on 36 PERSPECTRUM topics. Bold values indicate the better result for each metric.
Component
Configuration
Opened Mean
CESS
Reduction
Query generator
Qwen
0.4834
0.4344
10.1%
Query generator
Llama
0.4634
0.4157
10.3%
Search controller
Agent
0.4885
0.4364
10.7%
Search controller
MMR
0.4583
0.4137
9.7%
Appendix
Table 7: PERSPECTRUM MAE across query generators and search controllers. Lower is better.
Method
MAE
S
Direction disagreement
Opened Mean
0.4108
0.9574
0.4186
Outcome Regression
0.4200
0.0000
0.6667
Tuning-set Mean
0.2787
0.0000
0.3889
Opened Mean shrunk to Tuning Mean
0.2427
0.2393
0.3893
CESS
0.3103
0.1608
0.5208
Appendix
Table 8: PERSPECTRUM results for searches of up to ten rounds. S measures the difference between rankings after averaging repeated runs. Lower is better.
Variant
Selection correction
MAE
S
D
A (clip, then shrink)
Yes
0.3103
0.1608
0.5208
C (clip, then shrink)
No
0.3071
0.4346
0.5278
B (shrink, then clip)
Yes
0.3279
0.1977
0.5208
D (shrink, then clip)
No
0.3111
0.4650
0.5278
Appendix
Table 9: Four calculations on the same ten-round search runs, varying the order of clipping and shrinkage and whether document-selection weights are used. All retain the same initial prediction, coefficient 0.5, and stopping correction. S follows Eq. 16 . D is the fraction of conclusions with a different direction from the reference.
Round
Opened Mean
CESS
1
0.8424
0.4035
2
0.6189
0.3841
4
0.5039
0.3588
6
0.4631
0.3338
8
0.4305
0.3199
10
0.4108
0.3103
Appendix
Table 10: MAE after each search round, including runs that stopped earlier.
Matched quantity
CESS lower
No detected difference
CESS higher
Matched MAE, compare sensitivity
10
18
0
Matched sensitivity, compare MAE
3
15
10
Appendix
Table 11: Interpolated comparisons across 28 dataset–model–budget–prediction settings. “No detected difference” denotes a 95% interval containing zero. The counts are descriptive and are not adjusted for multiple comparisons.
Selection prob.
Prediction
λ
MAE
S
Min. prob.
Exact
Oracle rule
0.05
0.0453
0.0137
0.00667
Exact
Linear (all features)
0.05
0.0576
0.0172
0.00667
Exact
Linear (reduced)
0.05
0.0768
0.0225
0.00667
Exact
Constant
0.20
0.1854
0.0957
0.00667
Estimated selection prob.
Linear (all features)
0.05
0.0570
0.0137
0.01887
Limited overlap
Linear (all features)
0.00
0.0583
0.0000
0.000006
Appendix
Table 12: Selected results from the controlled simulation. Each row uses 8,000 test questions and a coefficient chosen on separate tuning questions. Lower MAE and ranking sensitivity S are better.
1Beijing University of Posts and Telecommunications · 2Shanghai Artificial Intelligence Laboratory · 3Chongqing University of Posts and Telecommunications