Search agents adapt their queries, yet fixed search interfaces leave candidate processing and evidence presentation outside the agent's direct control. Our trajectory analysis shows that supporting passages can be retrieved yet never delivered to the agent; a same-page oracle intervention shows that changing the returned evidence can reduce subsequent search. We introduce Programmatic Search Agent (PSA), which makes a local executable computation over candidates the unit of a search action. PSA unifies a persistent candidate workspace, flexible primitive composition, and selective evidence presentation. It incrementally generates program cells that reuse candidates, execute dependent operations, and select what the agent inspects next. The runtime resolves specified data dependencies within each cell, while the agent adapts its search strategy across cells as new evidence arrives. We compare PSA with the Query-based Agent and Tool-based Agent on InfoSeek-Eval and BrowseComp-Plus using five policy backbones without task-specific training. All three interfaces share the search substrate, and the Tool-based Agent also shares PSA's primitives and persistent workspace. Relative to the Query-based Agent, PSA improves macro-averaged task success by 4.00 and 7.56 percentage points on the two benchmarks, respectively; within-backbone reductions in final-step tokens average 28.3% and 33.9%. These results support extending agent control beyond query reformulation to the processing and presentation of retrieved evidence. Code will be released subject to approval.
Figures & tables
Figure 1: From query generation to programmatic search. All three agents iteratively act on search observations, but differ in their decision units: queries, tool calls, or a program cell. The Query-based Agent invokes a fixed search pipeline, the Tool-based Agent selects one primitive stage per turn (possibly batching independent calls), and PSA generates executable cells that compose operations over retained and new candidates. The illustrated cell filters retained candidates, conditionally retrieves and merges more candidates, then reranks and extracts evidence within one action.
Figure 2: Evidence-delivery limitations of a fixed search pipeline. (a) Final passage-delivery recall across tasks with verified supporting passages (smoothed density and individual-task marks). (b) Outcomes for supporting passages: immediate delivery, later selection, or no delivery. (c) Judge-based task success under paired control and same-page oracle continuations. Error bars in (b–c) are source-stratified task-bootstrap 95% confidence intervals; (c) shows per-arm uncertainty. (d) Empirical cumulative distributions of subsequent search decisions for complete omissions: further left indicates fewer searches, not higher answer accuracy. Colors follow (c). Complete/partial refers to the target extraction event, not the entire interaction history.
Figure 3: Overview of Programmatic Search Agent (PSA). The policy generates a program cell from the question, interaction history, and workspace description. A persistent workspace retains candidates; executable cells compose operations over them; and the selected output determines the next observation. The policy then chooses another cell or a final answer. Retained objects and returned observations are distinct.
Agent Backbone
Method
InfoSeek-Eval
BrowseComp-Plus
Task success ↑
Avg. tokens (K) ↓
Task success ↑
Recall ↑
Avg. tokens (K) ↓
DeepSeek-V4.1-Flash
Query-based Agent
81.00
33.9
54.94
59.30
115.7
Tool-based Agent
83.33 (+2.33)
31.3 ( ↓ 7.7%)
57.71 (+2.77)
62.07 (+2.77)
118.2 ( ↑ 2.2%)
PSA
84.67 (+3.67)
27.1 ( ↓ 20.1%)
62.29 (+7.35)
64.83 (+5.53)
61.9 ( ↓ 46.5%)
GLM-5.3-Flash
Query-based Agent
80.33
15.6
44.94
40.02
33.2
Tool-based Agent
81.67 (+1.34)
23.2 ( ↑ 49.2%)
47.11 (+2.17)
40.99 (+0.97)
37.0 ( ↑ 11.5%)
Table 1: Overall performance on the complete InfoSeek-Eval and BrowseComp-Plus evaluation sets. Methods within each backbone block share the search substrate and per-call funnel. Task success is semantically judged, and Recall measures supporting-document recall on BrowseComp-Plus. Avg. tokens uses the final-step metric in Section 4 ; values are shown in thousands (K = 1,000), rounded to one decimal place. Parentheses for the Tool-based Agent and PSA give changes from the Query-based Agent: percentage-point differences for task success and Recall, and relative token changes computed before rounding to K ( ↓ reduction; ↑ increase). Bold marks the best value within each backbone and dataset metric.
Figure 4: Evidence feedback on 128 paired tasks, comparing program-selected and full-view feedback. Means and per-arm 95% bootstrap CIs; tokens are final-step tokens.
Figure 5: Decision-budget sensitivity on all 128 sampled tasks. Means over tasks; panel (a) spans 20–60%. Final-step tokens are shown in thousands (K = 1,000).
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Primitive
Evaluation contract
retrieve
Returns at most 20 documents for a query.
filter
Applies field predicates with conjunction, disjunction, and negation; at most eight leaves and depth three.
rerank
Retains at most 10 documents, using their stored source query. Query branches are reranked separately before merging.
extract
Selects one passage per document from at most the first 10 input documents, preserving document order. Takes an evidence instruction and optional output size and schema.
dedupe
Removes duplicates by doc_id or passage_id , preserving first-seen order.
Appendix
Table 2: Shared search-primitive semantics.
Resource
InfoSeek-Eval
BrowseComp-Plus
Pooled
Query-based
PSA
Query-based
PSA
Query-based
PSA
Agent
Agent
Agent
Retrieval requests
9.12
9.79
43.52
30.88
34.39
25.28
Returned documents
182.5
195.8
870.5
617.6
687.8
505.6
Rerank pairs
178.2
195.7
870.5
617.5
686.7
505.5
Extract calls
9.12
9.79
43.52
30.84
34.39
25.25
Appendix
Table 3: Mean per-task resources in the complete DeepSeek-V4.1-Flash runs. Pooled means weight the 300 InfoSeek-Eval and 830 BrowseComp-Plus tasks equally. Rerank pairs count scored documents; extract pairs count passages scored by the passage locator. Final-step tokens are shown in thousands.
Deep search agents answer difficult information-seeking questions by iteratively issuing search queries to gather supporting evidence, but it remains unclear whether and how greater search effort leads to better answers. We study these questions through a trajectory-level diagnosis of long-horizon search agents. Using human-annotated document-level relevance judgments, we evaluate the evidence retrieved at each search step and separate two stages of agent behavior: what evidence an agent retrieves and how effectively it uses that evidence. This distinction further allows us to decompose failures into retrieval gaps, where the necessary evidence is never found, and utilization gaps, where relevant evidence is retrieved but not used correctly. With the retrieval model and evaluation harness held fixed, we compare six agents on BrowseComp-Plus and further validate our findings on BrowseComp with an open-web search API. Across settings, we find that search effort and answer quality are only weakly aligned. Answer accuracy is better correlated with the quality of retrieved evidence, especially cumulative retrieval recall, than with the number of searches or the amount of context consumed. Useful evidence often appears early in the trajectory, yet agents tend to continue searching, producing a long tail of low-yield retrieval steps. At the query level, exploratory reformulations remain useful, but the best-performing agents issue far fewer redundant queries. Overall, by systematically characterizing the search behavior and failure modes of long-horizon search agents, this work points to practical directions for building better deep research systems, including stronger query formulation, more effective evidence selection and context management, and stopping criteria based on whether sufficient supporting evidence has been retrieved.
Qi Liu, Jiaxin Mao, Fengbin Zhu +1
Renmin University of China Beijing, China · National University of Singapore Singapore
Are LLM-based search agents genuinely searching, or using the web to verify what they already know? We study this question on BrowseComp with three diagnostics. Our analysis reveals Intrinsic Knowledge Dependence (IKD): even with tool access, agents often rely on intrinsic knowledge -- information encoded in the model before retrieval -- rather than on external evidence. Agents answer up to 44.5% of BrowseComp questions without tools, generate more than half of their search queries from internally produced hypotheses rather than retrieved leads, and perform worse than closed-book baselines when answer-supporting evidence is removed. These results suggest that static search benchmarks can reward memory-backed verification rather than evidence-driven discovery, conflating what agents already know with what they can find. We then introduce LiveBrowseComp, a deep-search benchmark designed to evaluate agents beyond intrinsic coverage. It contains 335 human-authored questions whose answers depend on facts published within the 90 days preceding benchmark construction, drawn from six updated sources and filtered to exclude globally salient events. On LiveBrowseComp, all evaluated agents fall below 2% closed-book accuracy, search-augmented scores drop by 25-40 points relative to BrowseComp, and prior model rankings no longer reliably predict performance. LiveBrowseComp is available at https://huggingface.co/datasets/Forival/LiveBrowseComp.
LLM search agents are often evaluated on final-answer accuracy, overlooking the process. Analyzing a search strategy requires understanding how credible evidence is retrieved to address question constraints. This valuable information is buried in raw search trajectories that are long and difficult to parse. We introduce SearchAtlas, a framework that converts search trajectories into structured graphs whose edges represent how evidence is propagated across the reasoning trace, from the query that retrieves it to the final answer. Our automated parsing pipeline achieves a mean edge F1 of 86.0% against human-annotated graphs and remains consistent across repeated runs. We analyze five search agents on three benchmarks, revealing systematic differences in search scale and evidence aggregation. SearchAtlas exposes fragmented answer support, question constraints that do not reach the answer, and unverified parametric knowledge entering the response. These process failures are strongly associated with incorrect answers, even more so than an LLM judge given either the raw trajectory or the ordered query list, suggesting that the constructed graphs provide useful interpretability. Moreover, an audit of cases in which process-diagnostic scores disagree with final-answer correctness shows that they capture information not reducible to answer accuracy.