Programmatic Search Agents: Extending Agentic Search Beyond Query Reformulation
Organizations: Zhejiang University · Yuanbao Team, Tencent · Peking University
Abstract
Search agents adapt their queries, yet fixed search interfaces leave candidate processing and evidence presentation outside the agent's direct control. Our trajectory analysis shows that supporting passages can be retrieved yet never delivered to the agent; a same-page oracle intervention shows that changing the returned evidence can reduce subsequent search. We introduce Programmatic Search Agent (PSA), which makes a local executable computation over candidates the unit of a search action. PSA unifies a persistent candidate workspace, flexible primitive composition, and selective evidence presentation. It incrementally generates program cells that reuse candidates, execute dependent operations, and select what the agent inspects next. The runtime resolves specified data dependencies within each cell, while the agent adapts its search strategy across cells as new evidence arrives. We compare PSA with the Query-based Agent and Tool-based Agent on InfoSeek-Eval and BrowseComp-Plus using five policy backbones without task-specific training. All three interfaces share the search substrate, and the Tool-based Agent also shares PSA's primitives and persistent workspace. Relative to the Query-based Agent, PSA improves macro-averaged task success by 4.00 and 7.56 percentage points on the two benchmarks, respectively; within-backbone reductions in final-step tokens average 28.3% and 33.9%. These results support extending agent control beyond query reformulation to the processing and presentation of retrieved evidence. Code will be released subject to approval.
Figures & tables
| Agent Backbone | Method | InfoSeek-Eval | BrowseComp-Plus | |||
|---|---|---|---|---|---|---|
| Task success | Avg. tokens (K) | Task success | Recall | Avg. tokens (K) | ||
| DeepSeek-V4.1-Flash | Query-based Agent | 81.00 | 33.9 | 54.94 | 59.30 | 115.7 |
| Tool-based Agent | 83.33 (+2.33) | 31.3 ( 7.7%) | 57.71 (+2.77) | 62.07 (+2.77) | 118.2 ( 2.2%) | |
| PSA | 84.67 (+3.67) | 27.1 ( 20.1%) | 62.29 (+7.35) | 64.83 (+5.53) | 61.9 ( 46.5%) | |
| GLM-5.3-Flash | Query-based Agent | 80.33 | 15.6 | 44.94 | 40.02 | 33.2 |
| Tool-based Agent | 81.67 (+1.34) | 23.2 ( 49.2%) | 47.11 (+2.17) | 40.99 (+0.97) | 37.0 ( 11.5%) | |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Primitive | Evaluation contract |
|---|---|
| retrieve | Returns at most 20 documents for a query. |
| filter | Applies field predicates with conjunction, disjunction, and negation; at most eight leaves and depth three. |
| rerank | Retains at most 10 documents, using their stored source query. Query branches are reranked separately before merging. |
| extract | Selects one passage per document from at most the first 10 input documents, preserving document order. Takes an evidence instruction and optional output size and schema. |
| dedupe | Removes duplicates by doc_id or passage_id , preserving first-seen order. |
| Resource | InfoSeek-Eval | BrowseComp-Plus | Pooled | |||
|---|---|---|---|---|---|---|
| Query-based | PSA | Query-based | PSA | Query-based | PSA | |
| Agent | Agent | Agent | ||||
| Retrieval requests | 9.12 | 9.79 | 43.52 | 30.88 | 34.39 | 25.28 |
| Returned documents | 182.5 | 195.8 | 870.5 | 617.6 | 687.8 | 505.6 |
| Rerank pairs | 178.2 | 195.7 | 870.5 | 617.5 | 686.7 | 505.5 |
| Extract calls | 9.12 | 9.79 | 43.52 | 30.84 | 34.39 | 25.25 |