Harness-Search: Guiding Long-Horizon Search through Multi-Agent Coordination
Organizations: Xiaohongshu Inc. · the Joint SDU-NTU Centre for Artificial Intelligence Research (C-FAIR), Shandong University · University of Illinois at Urbana-Champaign
Abstract
Long-horizon search requires agents to gather evidence across multiple steps and synthesize it into well-supported answers. The recent agent harnesses provide a natural and promising framework to support such long-running search processes. As interaction histories grow, one single agent in harnesses might get stuck and cause the policy to lose track of unresolved questions, overlook useful evidence, or terminate before sufficient support has been collected. One of promising way is to decouple three distinct responsibilities of proposing retrieval actions, updating persistent state, and deciding when to stop rather than concentrating them within a single policy. Targeted at it, we introduce Harness-Search, a multi-agent search harness to reduce the local errors propagating across subsequent exploration, evidence curation, and termination decisions. In particular, Harness-Search assigns these responsibilities to three permission-bounded authorities: a Retrieval Policy that proposes search operations, a Memory Operator that validates and commits persistent-state updates, and a Summary Auditor that accepts or rejects termination based on the sufficiency of the curated evidence. Together, these roles form a Propose-Commit-Audit loop in which actions are proposed, persistent evidence is selectively committed, and stopping decisions are subjected to an explicit sufficiency check. Across seven long-horizon search benchmarks, Harness-Search improves both retrieval and answer generation under the same policy backbone, increasing Recall by 4.60-27.92 points and Final-Answer Recall by 12.34-30.13 points over the strongest harness-based baseline on each evidence-retrieval benchmark. Moreover, trajectory-level analyses show that Harness-Search continues to accumulate useful evidence and expand evidence coverage with less redundant retrieval as the search history grows.
Figures & tables
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Tool signature | Effect and output |
|---|---|
| fan_out_search(queries) | Executes up to five complementary hybrid searches and returns the combined ranked document IDs and text snippets. |
| search(query) | Performs a single hybrid search using sparse and dense retrieval, followed by fusion and reranking; returns ranked document IDs and text snippets. |
| grep_corpus(pattern) | Matches an exact or regular-expression pattern against the corpus and returns matching document IDs and text excerpts. |
| read(doc_id) | Retrieves the full content of the specified document and adds it to the document memory. |
| review_docs(doc_ids) | Re-renders up to five previously retrieved documents from outer memory without issuing an additional corpus call. |
| redirect(doc_ids, reasoning) | Sends up to five newly retrieved candidate documents and the policy’s rationale to the intent model; returns an updated active retrieval intent. |
| Seed | BC+ | Web | Sec | LongSeal |
|---|---|---|---|---|
| 20 | 55.00 | 59.20 | 35.20 | 86.00 |
| 42 | 57.81 | 62.19 | 37.97 | 88.98 |
| 314 | 61.50 | 65.50 | 41.00 | 92.20 |
| Avg. | 58.10{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 3.26} | 62.30{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 3.15} | 38.06{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 2.90} | 89.06{\color[rgb]{0.5,0.5,0.5}\scriptstyle\pm 3.10} |
| 42 | 56.32 | 61.05 | 36.41 | 87.52 |
| 42 | 58.27 | 62.48 | 37.83 | 89.17 |
| Category | System or intervention | (95% CI) | FA-R | TR | |
| Full | – | 57.81 | — | 61.51 | 66.86 |
| Component | Without Memory Operator | 44.74 | 49.89 | 59.08 | |
| Without Summary Auditor | 52.07 | 56.67 | 65.34 | ||
| Without Document Relevance Judge | 52.98 | 55.99 | 64.26 | ||
| Without Direction Planner | 52.58 | 54.43 | 63.12 | ||
| Monolithic | True single agent + structured tools | 54.40 | 55.93 | 63.34 |
| Max Turn | R | FA-R | TR | Mean latency(s) | P50 | P95 | P99 |
|---|---|---|---|---|---|---|---|
| 10 | 44.74 | 49.12 | 51.90 | 27.03 | 25.13 | 47.78 | 61.23 |
| 20 | 54.69 | 57.74 | 61.79 | 39.42 | 37.93 | 63.96 | 88.94 |
| 30 | 56.55 | 59.36 | 63.81 | 47.95 | 46.68 | 83.89 | 99.88 |
| 40 | 57.81 | 61.51 | 66.86 | 52.15 | 52.50 | 90.79 | 114.96 |
| 50 | 57.77 | 63.11 | 68.59 | 60.34 | 59.76 | 112.19 | 138.69 |
| 60 | 56.40 | 60.88 | 66.12 | 67.91 | 68.95 | 129.04 | 166.30 |