Harness-Search: Guiding Long-Horizon Search through Multi-Agent Coordination
Authors: Shanyong Wang, Zhenwen Ji, Lei Jin, Yining Zhao, Yicheng Qian, Chengqiang Lu, Yi Wu, Yao Hu, +2 more
Organizations: Xiaohongshu Inc. · the Joint SDU-NTU Centre for Artificial Intelligence Research (C-FAIR), Shandong University · University of Illinois at Urbana-Champaign
Long-horizon search requires agents to gather evidence across multiple steps and synthesize it into well-supported answers. The recent agent harnesses provide a natural and promising framework to support such long-running search processes. As interaction histories grow, one single agent in harnesses might get stuck and cause the policy to lose track of unresolved questions, overlook useful evidence, or terminate before sufficient support has been collected. One of promising way is to decouple three distinct responsibilities of proposing retrieval actions, updating persistent state, and deciding when to stop rather than concentrating them within a single policy. Targeted at it, we introduce Harness-Search, a multi-agent search harness to reduce the local errors propagating across subsequent exploration, evidence curation, and termination decisions. In particular, Harness-Search assigns these responsibilities to three permission-bounded authorities: a Retrieval Policy that proposes search operations, a Memory Operator that validates and commits persistent-state updates, and a Summary Auditor that accepts or rejects termination based on the sufficiency of the curated evidence. Together, these roles form a Propose-Commit-Audit loop in which actions are proposed, persistent evidence is selectively committed, and stopping decisions are subjected to an explicit sufficiency check. Across seven long-horizon search benchmarks, Harness-Search improves both retrieval and answer generation under the same policy backbone, increasing Recall by 4.60-27.92 points and Final-Answer Recall by 12.34-30.13 points over the strongest harness-based baseline on each evidence-retrieval benchmark. Moreover, trajectory-level analyses show that Harness-Search continues to accumulate useful evidence and expand evidence coverage with less redundant retrieval as the search history grows.
Figures & tables
Figure 1: Harness-Search turns open-ended search into a guided, evidence-driven process. Naive iterative search often gets stuck exploring the same direction, leading to redundant retrieval and missing evidence. Harness-Search explicitly alternates between planning, searching, curating, and auditing, and redirects the search when the auditor identifies an evidence gap. As a result, it continually expands search coverage instead of saturating early ( Bottom Right ), yielding substantial improvements across diverse search-intensive benchmarks ( Bottom Left ).
Figure 2: Overview of Harness-Search . Harness-Search addresses long-horizon search with three permission-bounded agents: a Retrieval Policy proposes actions, a Memory Operator validates and commits persistent state, and a Summary Auditor audits evidence sufficiency before termination.
Table 3Figure 4
Figure 4: Behavior following a Summary Auditor rejection.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Tool signature
Effect and output
fan_out_search(queries)
Executes up to five complementary hybrid searches and returns the combined ranked document IDs and text snippets.
search(query)
Performs a single hybrid search using sparse and dense retrieval, followed by fusion and reranking; returns ranked document IDs and text snippets.
grep_corpus(pattern)
Matches an exact or regular-expression pattern against the corpus and returns matching document IDs and text excerpts.
read(doc_id)
Retrieves the full content of the specified document and adds it to the document memory.
review_docs(doc_ids)
Re-renders up to five previously retrieved documents from outer memory without issuing an additional corpus call.
redirect(doc_ids, reasoning)
Sends up to five newly retrieved candidate documents and the policy’s rationale to the intent model; returns an updated active retrieval intent.
Table 7: Repeated experiments across same and random seeds
Category
System or intervention
Rablation
ΔR (95% CI)
FA-R
TR
Full
–
57.81
—
61.51
66.86
Component
Without Memory Operator
44.74
+13.07[10.71,15.43]
49.89
59.08
Without Summary Auditor
52.07
+5.74[3.52,7.90]
56.67
65.34
Without Document Relevance Judge
52.98
+4.83[2.78,6.86]
55.99
64.26
Without Direction Planner
52.58
+5.23[3.52,6.94]
54.43
63.12
Monolithic
True single agent + structured tools
54.40
+3.41[1.29,5.55]
55.93
63.34
Appendix
Table 8: Full Multi-agent decomposition study. Rctrl is the Recall of each control. ΔR=Rfull−Rablation ; positive values favor the complete system. FA-R denotes final-answer Recall, TR denotes trajectory Recall. All Recall values are percentages.
Figure 5: Sensitivity to inference-time interaction.
Max Turn
R
FA-R
TR
Mean latency(s)
P50
P95
P99
10
44.74
49.12
51.90
27.03
25.13
47.78
61.23
20
54.69
57.74
61.79
39.42
37.93
63.96
88.94
30
56.55
59.36
63.81
47.95
46.68
83.89
99.88
40
57.81
61.51
66.86
52.15
52.50
90.79
114.96
50
57.77
63.11
68.59
60.34
59.76
112.19
138.69
60
56.40
60.88
66.12
67.91
68.95
129.04
166.30
Appendix
Table 9: Sensitivity to the total interaction budget and the search-before-curation budget on BrowseComp-Plus using GPT-OSS-20B. In the upper block, we vary the maximum number of total interaction turns Tmax while fixing Smax=5 . In the lower block, we vary the maximum number of consecutive search actions before a required curate operation, Smax , while fixing Tmax=40 . R, FA-R, and TR denote final Recall, final-answer Recall, and trajectory Recall, respectively. We select Tmax=40 and Smax=5 for the main experiments as a quality–latency trade-off.
Figure 6: Role-specific model routing on BrowseComp-Plus. The Retrieve Policy is fixed to GPT-OSS-20B, while the Memory Operator (columns) and Summary Auditor (rows) are independently assigned GPT-OSS-20B or Qwen3 models of different sizes. Panels report (a) Recall of the final curated set, (b) final-answer Recall over gold documents, and (c) trajectory Recall over documents encountered during search. Values are percentages, with darker colors indicating higher scores.
Figure 7: Training loss figures of three agents.
Figure 8: Case Study 1. Compared to other method(Left), the improvement mainly comes from Redirect by Retrieve Policy and search_more by Summary Auditor.
Figure 9: Case Study 2. The step-by-step results show that most of the improvement comes from Redirect decisions made by Retrieval Policy.
Search agents are often trained as policies over growing transcripts: the model must decide how to search while also remembering what it has seen, which evidence is useful, which constraints remain open, and which claims have actually been checked. We argue that this formulation puts too much routine state management inside the policy: reinforcement learning is forced to optimize both semantic search decisions and recoverable bookkeeping that the environment can maintain more reliably. We introduce Harness-1, a 20B search agent (retrieval subagent) trained with reinforcement learning inside a stateful search harness. The harness maintains environment-side working memory, including a candidate pool, an importance-tagged curated set, compact evidence links, verification records, compressed and deduplicated observations, and budget-aware context rendering. The policy retains the semantic decisions: what to search, which documents to keep or discard, what to verify, and when to stop. Across eight retrieval benchmarks spanning web, finance, patents, and multi-hop QA, Harness-1 achieves 0.730 average curated recall, outperforming the next strongest open search subagent by +11.4 points and remaining competitive with much larger frontier-model searchers. Its gains are especially strong on held-out transfer benchmarks, suggesting that reinforcement learning over explicit search state can produce retrieval behaviors that generalize beyond the training domains. Our code is available at https://github.com/pat-jj/harness-1.
Pengcheng Jiang, Zhiyi Shi, Kelly Hong +5
University of Illinois at Urbana-Champaign · †UC Berkeley · ✧Chroma
Deep search agents tackle complex knowledge tasks through iterative retrieval, multi-hop reasoning, and evidence synthesis across multiple sources. Existing approaches typically assume relatively stable retrieval systems and operate over short-horizon tool interaction. However, when retrieval is sensitive to query formulation, even a semantically appropriate query may fail to surface critical evidence because of mismatched entity names, aliases, or keyword combinations. Recovering from such failures requires repeated query reformulation and longer interaction trajectories. This setting poses a distinct training challenge, as the policy must sustain long-horizon query exploration while managing an expanding volume of retrieved content. We propose LexiHorizon, a framework for training search agents over long horizons that expands the trajectory context budget, manages accumulated retrieval content using a window over recent tool observations while preserving the reasoning history, and introduces an outcome-gated search-effort reward that provides a bounded bonus for tool invocations to trajectories with nonzero answer reward. Experiments on XBench, WebWalkerQA, and BrowseComp-ZH show that the resulting 9B model consistently outperforms both its base model and MiroThinker-1.7-mini, with maximum absolute gains of 8.7 and 23.8 percentage points, respectively. These results suggest that combining an extended context budget with reasoning-preserving context management benefits long-horizon deep search agents.
Reinforcement learning (RL) search agents commonly model retrieval as free-form natural-language query generation and optimize multi-turn interactions using final-answer rewards. Current studies mainly improve training with denser or more structured credit signals, but rarely examine whether retrieval is properly formulated at the policy-environment interface. We observe pronounced retrieval aliasing during Search-R1 training: rollouts for the same question continue to generate distinct query strings, yet their accumulated evidence sets increasingly overlap. We call this phenomenon retrieval-equivalence collapse; in this regime, trajectories approach utility equivalence with respect to retrieval decisions, leaving within-group returns with little effective retrieval contrast. To address this problem, we propose Harness-G, a graph-structured retrieval framework that redesigns this interface. It reformulates free-form query generation as finite action selection: the policy selects an evidence sentence or entity, or chooses to answer, while the environment constructs the menu, tracks retrieval state, and validates and executes each choice. This interface reduces linguistic aliasing and makes same-state alternatives directly comparable. Building on this interface, we introduce Structured Non-myopic Credit (SNC), which uses a frozen answer scorer to compare the selected action with its alternatives and assigns downstream gains to the earlier actions that enabled them. Across six QA benchmarks, Harness-G achieves the highest average F1 at both evaluated model scales, outperforming the strongest baseline, Graph-R1, by 10.74 points at 1.5B and 3.98 points at 3B.
Yanning Hou, Haoyuan Chen, Sihang Zhou +7
National University of Defense Technology · Changsha, China