In agentic search, an LLM agent searches the web, reads the pages it finds, and decides what to look for next before returning an answer. But can we trust an answer simply because the agent returns it? Often not, and even a correct answer can be a lucky guess: on questions with several constraints, we find that agents conclude the task is complete while a constraint remains unverified in up to 48% of their correct answers. We call this illusory completion. To see how it arises, we introduce the Epistemic Ledger, which tracks at every turn what the retrieved pages establish about each constraint and what the agent claims. Across 13 agents, from 7B RL-trained models to frontier LLMs, training and scale raise accuracy but change the pattern of verification failures rather than eliminating them: constraints may be left unchecked, assumed without support, or retained despite refuting evidence. To measure what agents lose without tracking their constraints, we show them each constraint's state, approximated by LiveLedger, a lightweight 4B tracker. Agents then answer 4.4-16.1 points more questions correctly, suggesting that on their own, they may not track what they have verified and what remains.
Figures & tables
Figure 1: Illusory completion in search agents. Answers from five search agents on multi-constraint questions are split into correct and incorrect . In both, the striped segments are unsubstantiated answers, where the agent shows illusory completion.
Figure 2: Illusory completion through the Epistemic Ledger . The agent answers Drake once C1 is verified , while its other constraints remain unverified: it affirms C2 without evidence ( assumed ), keeps Drake although retrieved evidence contradicts C3 ( refuted ), and neither establishes nor addresses C4 ( unchecked ). The final answer alone reveals none of these.
Figure 3: Constraint outcomes of the Committed candidate. For each question, we compute the share of the candidate’s constraints in LT that are verified , assumed , refuted , or unchecked , and average it over all questions: (a) training recipes at 7–8B, (b) model size, and (c) backbone. Each bar is split by the run result into correct (dark) and incorrect (light) answers. Higher accuracy shifts constraints between failures rather than reducing all of them.
Agent
Acc ( ↑ )
UAR ( ↓ )
Trained
Search-R1 ( Jin et al., 2025 )
18.6
78.9
ASearcher ( Gao et al., 2025 )
18.0
80.4
RAG-R1 ( Tan et al., 2025 )
18.0
79.3
DR-Tulu ( Shao et al., 2025 )
21.7
71.5
WebExplorer ( Liu et al., 2025 )
28.9
66.9
Table 1: Main results on all baselines. Acc (%) and UAR (%) on the 484 questions. Darker cells mean higher values. The best score in each group is in bold .
Figure 4: Constraint outcomes with LiveLedger . For each agent, the left bar of each outcome is without LiveLedger and the hatched right bar is with it, computed and split by the run result as in Figure 3 . With LiveLedger , most agents verify more constraints and leave fewer unchecked (by up to 10.5 points), with far smaller shifts in assumed and refuted ones (at most 2.4 points); more of the verified constraints come from correct answers, raising accuracy.
Figure 5: Stalled runs with and without LiveLedger . Share of questions stalled (§ 5.1 ) at each tool call, without (dotted) and with LiveLedger (solid). Numbers on the right are the shares at the end of the budget. With LiveLedger , fewer questions remain stalled at the end of the budget for all six agents.
Acc ( ↑ )
UAR ( ↓ )
Agent
w/o
w/ LL
w/o
w/ LL
Trained
Agents-A1 ( Bai et al., 2026 )
33.5
41.5
65.7
49.0
TongyiDR ( Tongyi et al., 2025 )
34.7
39.3
64.0
55.9
Prompt-based
gpt-oss-20b ( Agarwal et al., 2025 )
22.3
30.6
70.9
62.2
Table 2: Effect of LiveLedger . Acc and UAR without (w/o) and with LiveLedger (w/ LL). Darker cells mean higher values. † HiPRAG and ReSeek train Qwen2.5-7B with process rewards.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
Source split
#
BrowseComp
test
205
DeepSearchQA
eval, single-answer questions
69
FRAMES
test
33
LiveDRBench
v1-full, test
49
WebWalkerQA
main, English questions
28
BioASQ
rag-mini-bioasq , test (first 100)
100
Appendix
Table 3: Evaluation set. Number of questions kept from each benchmark.
Agent
Cap per run
Search
Page reading
T
top- p
Max tokens
Search-R1
–
Jina Search
none
0.7
–
1,024
ASearcher
30 model calls
Serper, top 10
disabled
0.6
0.95
4,098
RAG-R1
10 retrievals
Jina Search
none
0.7
–
2,048
DR-Tulu
–
Serper, top 10
Jina
1.0
–
32,000
WebExplorer
30 model calls
Serper, top 10
Jina, summarized ∗
0.6
0.95
10,000
TongyiDR
30 model calls
Serper, top 10; Scholar
Jina, summarized ∗
0.85
0.95 †
10,000
Appendix
Table 4: Per-agent settings. An agent call may issue several tool calls. “–” marks a setting that is not set or that we could not confirm from the run logs. ∗ By gpt-4o-mini . † With presence penalty 1.1. ‡ vLLM defaults; no sampling parameters are sent. The last group follows our earlier protocol.
Table 5: Checkpoints. ∗ Also the two Search-o1 backbones. The last group follows our earlier protocol.
Acc ( ↑ )
UAR ( ↓ )
Agent
w/o
w/ LL
w/o
w/ LL
TongyiDR
35.3 ± 0.5
42.1 ± 2.8
64.1 ± 0.5
52.2 ± 3.3
gpt-oss-20b
23.4 ± 1.4
28.5 ± 2.0
70.6 ± 1.1
61.3 ± 1.1
gpt-oss-120b
22.6 ± 1.0
34.2 ± 0.4
71.4 ± 1.0
56.4 ± 0.6
Appendix
Table 6: Results over three runs. Acc (%) and UAR (%) on the 484 questions without (w/o) and with LiveLedger (w/ LL), as mean ± standard deviation over three runs of each agent.
Figure 6: An illustration of LiveLedger integration into the agent.
Epistemic Ledger
LiveLedger
Agent
Calls
USD
Calls
Wait (%)
TongyiDR
48.2
0.093
20.9
7.4
Agents-A1
50.5
0.108
19.4
5.5
gpt-oss-20b
52.3
0.089
19.7
2.0
gpt-oss-120b
57.2
0.095
18.7
1.1
DeepSeek-V4-Pro
46.8
0.089
15.7
8.4
Appendix
Table 7: Cost per question. Epistemic Ledger : evaluator calls and cost at gpt-5-nano list prices, for the agent without LiveLedger . LiveLedger : tracker calls and the share of end-to-end latency the agent spends waiting on the tracker.
Figure 7: Example of an incorrect agent execution, abbreviating the agent’s reasoning text for space.
Candidate : Stephen Lewis Fuchs
Status : Active
Constraint
Evidential Support
Agent Belief
C1: Individual must be a rabbi
✓
✓
Rabbi Stephen Lewis Fuchs
is Rabbi Stephen Lewis Fuchs
C2: Worked for Keneseth Israel
✓
✓
in Philadelphia
Rabbi Stephen Fuchs officiated at
served Reform Congregation
Appendix
Table 8: The Epistemic Ledger after the agent execution in Figure 7 .
Figure 8: Example of an correct agent execution with LiveLedger , abbreviating the agent’s reasoning text for space.