In agentic search, an LLM agent searches the web, reads the pages it finds, and decides what to look for next before returning an answer. But can we trust an answer simply because the agent returns it? Often not, and even a correct answer can be a lucky guess: on questions with several constraints, we find that agents conclude the task is complete while a constraint remains unverified in up to 48% of their correct answers. We call this illusory completion. To see how it arises, we introduce the Epistemic Ledger, which tracks at every turn what the retrieved pages establish about each constraint and what the agent claims. Across 13 agents, from 7B RL-trained models to frontier LLMs, training and scale raise accuracy but change the pattern of verification failures rather than eliminating them: constraints may be left unchecked, assumed without support, or retained despite refuting evidence. To measure what agents lose without tracking their constraints, we show them each constraint's state, approximated by LiveLedger, a lightweight 4B tracker. Agents then answer 4.4-16.1 points more questions correctly, suggesting that on their own, they may not track what they have verified and what remains.
Figures & tables
Figure 1: Illusory completion in search agents. Answers from five search agents on multi-constraint questions are split into correct and incorrect . In both, the striped segments are unsubstantiated answers, where the agent shows illusory completion.
Figure 2: Illusory completion through the Epistemic Ledger . The agent answers Drake once C1 is verified , while its other constraints remain unverified: it affirms C2 without evidence ( assumed ), keeps Drake although retrieved evidence contradicts C3 ( refuted ), and neither establishes nor addresses C4 ( unchecked ). The final answer alone reveals none of these.
Figure 3: Constraint outcomes of the Committed candidate. For each question, we compute the share of the candidate’s constraints in LT that are verified , assumed , refuted , or unchecked , and average it over all questions: (a) training recipes at 7–8B, (b) model size, and (c) backbone. Each bar is split by the run result into correct (dark) and incorrect (light) answers. Higher accuracy shifts constraints between failures rather than reducing all of them.
Agent
Acc ( ↑ )
UAR ( ↓ )
Trained
Search-R1 ( Jin et al., 2025 )
18.6
78.9
ASearcher ( Gao et al., 2025 )
18.0
80.4
RAG-R1 ( Tan et al., 2025 )
18.0
79.3
DR-Tulu ( Shao et al., 2025 )
21.7
71.5
WebExplorer ( Liu et al., 2025 )
28.9
66.9
Table 1: Main results on all baselines. Acc (%) and UAR (%) on the 484 questions. Darker cells mean higher values. The best score in each group is in bold .
Figure 4: Constraint outcomes with LiveLedger . For each agent, the left bar of each outcome is without LiveLedger and the hatched right bar is with it, computed and split by the run result as in Figure 3 . With LiveLedger , most agents verify more constraints and leave fewer unchecked (by up to 10.5 points), with far smaller shifts in assumed and refuted ones (at most 2.4 points); more of the verified constraints come from correct answers, raising accuracy.
Figure 5: Stalled runs with and without LiveLedger . Share of questions stalled (§ 5.1 ) at each tool call, without (dotted) and with LiveLedger (solid). Numbers on the right are the shares at the end of the budget. With LiveLedger , fewer questions remain stalled at the end of the budget for all six agents.
Acc ( ↑ )
UAR ( ↓ )
Agent
w/o
w/ LL
w/o
w/ LL
Trained
Agents-A1 ( Bai et al., 2026 )
33.5
41.5
65.7
49.0
TongyiDR ( Tongyi et al., 2025 )
34.7
39.3
64.0
55.9
Prompt-based
gpt-oss-20b ( Agarwal et al., 2025 )
22.3
30.6
70.9
62.2
Table 2: Effect of LiveLedger . Acc and UAR without (w/o) and with LiveLedger (w/ LL). Darker cells mean higher values. † HiPRAG and ReSeek train Qwen2.5-7B with process rewards.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
Source split
#
BrowseComp
test
205
DeepSearchQA
eval, single-answer questions
69
FRAMES
test
33
LiveDRBench
v1-full, test
49
WebWalkerQA
main, English questions
28
BioASQ
rag-mini-bioasq , test (first 100)
100
Appendix
Table 3: Evaluation set. Number of questions kept from each benchmark.
Agent
Cap per run
Search
Page reading
T
top- p
Max tokens
Search-R1
–
Jina Search
none
0.7
–
1,024
ASearcher
30 model calls
Serper, top 10
disabled
0.6
0.95
4,098
RAG-R1
10 retrievals
Jina Search
none
0.7
–
2,048
DR-Tulu
–
Serper, top 10
Jina
1.0
–
32,000
WebExplorer
30 model calls
Serper, top 10
Jina, summarized ∗
0.6
0.95
10,000
TongyiDR
30 model calls
Serper, top 10; Scholar
Jina, summarized ∗
0.85
0.95 †
10,000
Appendix
Table 4: Per-agent settings. An agent call may issue several tool calls. “–” marks a setting that is not set or that we could not confirm from the run logs. ∗ By gpt-4o-mini . † With presence penalty 1.1. ‡ vLLM defaults; no sampling parameters are sent. The last group follows our earlier protocol.
Table 5: Checkpoints. ∗ Also the two Search-o1 backbones. The last group follows our earlier protocol.
Acc ( ↑ )
UAR ( ↓ )
Agent
w/o
w/ LL
w/o
w/ LL
TongyiDR
35.3 ± 0.5
42.1 ± 2.8
64.1 ± 0.5
52.2 ± 3.3
gpt-oss-20b
23.4 ± 1.4
28.5 ± 2.0
70.6 ± 1.1
61.3 ± 1.1
gpt-oss-120b
22.6 ± 1.0
34.2 ± 0.4
71.4 ± 1.0
56.4 ± 0.6
Appendix
Table 6: Results over three runs. Acc (%) and UAR (%) on the 484 questions without (w/o) and with LiveLedger (w/ LL), as mean ± standard deviation over three runs of each agent.
Figure 6: An illustration of LiveLedger integration into the agent.
Epistemic Ledger
LiveLedger
Agent
Calls
USD
Calls
Wait (%)
TongyiDR
48.2
0.093
20.9
7.4
Agents-A1
50.5
0.108
19.4
5.5
gpt-oss-20b
52.3
0.089
19.7
2.0
gpt-oss-120b
57.2
0.095
18.7
1.1
DeepSeek-V4-Pro
46.8
0.089
15.7
8.4
Appendix
Table 7: Cost per question. Epistemic Ledger : evaluator calls and cost at gpt-5-nano list prices, for the agent without LiveLedger . LiveLedger : tracker calls and the share of end-to-end latency the agent spends waiting on the tracker.
Figure 7: Example of an incorrect agent execution, abbreviating the agent’s reasoning text for space.
Candidate : Stephen Lewis Fuchs
Status : Active
Constraint
Evidential Support
Agent Belief
C1: Individual must be a rabbi
✓
✓
Rabbi Stephen Lewis Fuchs
is Rabbi Stephen Lewis Fuchs
C2: Worked for Keneseth Israel
✓
✓
in Philadelphia
Rabbi Stephen Fuchs officiated at
served Reform Congregation
Appendix
Table 8: The Epistemic Ledger after the agent execution in Figure 7 .
Figure 8: Example of an correct agent execution with LiveLedger , abbreviating the agent’s reasoning text for space.
Are LLM-based search agents genuinely searching, or using the web to verify what they already know? We study this question on BrowseComp with three diagnostics. Our analysis reveals Intrinsic Knowledge Dependence (IKD): even with tool access, agents often rely on intrinsic knowledge -- information encoded in the model before retrieval -- rather than on external evidence. Agents answer up to 44.5% of BrowseComp questions without tools, generate more than half of their search queries from internally produced hypotheses rather than retrieved leads, and perform worse than closed-book baselines when answer-supporting evidence is removed. These results suggest that static search benchmarks can reward memory-backed verification rather than evidence-driven discovery, conflating what agents already know with what they can find. We then introduce LiveBrowseComp, a deep-search benchmark designed to evaluate agents beyond intrinsic coverage. It contains 335 human-authored questions whose answers depend on facts published within the 90 days preceding benchmark construction, drawn from six updated sources and filtered to exclude globally salient events. On LiveBrowseComp, all evaluated agents fall below 2% closed-book accuracy, search-augmented scores drop by 25-40 points relative to BrowseComp, and prior model rankings no longer reliably predict performance. LiveBrowseComp is available at https://huggingface.co/datasets/Forival/LiveBrowseComp.
LLM search agents are often evaluated on final-answer accuracy, overlooking the process. Analyzing a search strategy requires understanding how credible evidence is retrieved to address question constraints. This valuable information is buried in raw search trajectories that are long and difficult to parse. We introduce SearchAtlas, a framework that converts search trajectories into structured graphs whose edges represent how evidence is propagated across the reasoning trace, from the query that retrieves it to the final answer. Our automated parsing pipeline achieves a mean edge F1 of 86.0% against human-annotated graphs and remains consistent across repeated runs. We analyze five search agents on three benchmarks, revealing systematic differences in search scale and evidence aggregation. SearchAtlas exposes fragmented answer support, question constraints that do not reach the answer, and unverified parametric knowledge entering the response. These process failures are strongly associated with incorrect answers, even more so than an LLM judge given either the raw trajectory or the ordered query list, suggesting that the constructed graphs provide useful interpretability. Moreover, an audit of cases in which process-diagnostic scores disagree with final-answer correctness shows that they capture information not reducible to answer accuracy.
Large language model (LLM)-based search agents answer questions through multi-step interactions with external environments. However, providing complete execution trajectories to the LLM causes unbounded context growth and introduces noise. Existing compression methods reduce context at the cost of important details and often replace erroneous facts without repairing downstream reasoning derived from them. To address this problem, we propose ReTree, a self-correcting tree-structured memory mechanism for search agents. ReTree constructs a bounded per-step reasoning context while preserving source-linked evidence. It models search as an evidence tree whose nodes store bounded summaries, evidence, and revision histories. When newly retrieved evidence contradicts an earlier claim, ReTree traces back to the node where the claim was introduced, replaces outdated evidence, regenerates summaries, prunes affected branches, and resumes search. Source-grounded evidence provenance supports reliable conflict localization and keeps final claims traceable to retrieved passages. Experiments on four public question-answering and search benchmarks show that ReTree consistently outperforms Full-Trajectory ReAct, improving answer accuracy by up to 25.6 percentage points (pp); the average maximum per-step reasoning context of Full-Trajectory ReAct is 1.27--1.51× that of ReTree. These results establish ReTree as an effective self-correcting memory abstraction for long-horizon search.