Long-horizon information-seeking agents often accumulate noisy or misleading context, causing early mistakes to persist and making recovery increasingly difficult. We introduce an autonomous search harness in which the agent manages its own search process through three states: Rubric, Answer, and Verify. The agent first defines criteria for a valid answer, searches under these criteria, and then independently verifies the result before deciding whether to terminate or continue searching. It is further equipped with a Seal Memory tool that enables active context management. Training this behavior with reinforcement learning, however, can induce Seal Collapse, resulting in unstable training and preventing the agent from reliably learning when and how to use its memory tools. We solve this with a simple strategy that trains only the final segment after context management. Our 35B model achieves 72.83 on BrowseComp, outperforming comparable open-source systems, and consistently improves over the base model across BrowseComp-ZH, xbench, DeepSearchQA, WideSearch, financial investigation, and product search. Ablations show that autonomous compression outperforms automatic compaction and validate our RL design.
Figures & tables
Figure 1: Overview of Traverse. Top: Traverse integrates rubric-guided exploration, autonomous context management, and self-verification into an iterative search workflow. Bottom: During RL, trajectories are segmented by Seal Memory operations, and only the final segment is optimized using group-relative advantages.
Model
Browse Comp
BrowseComp- ZH
xbench- 2510
DeepSearch QA
Wide Search
Closed-source Models
GPT-5.5
84.4
–
–
–
–
Gemini 3.1 Pro
85.9
–
53.0
93.3
66.4
Claude Opus 5
90.8
–
–
95.0
–
General-purpose Models
Qwen3.5-35B-A3B
61.0
69.5
50.3
68.5
57.1
Table 1: Main results on five search-agent benchmarks. We report accuracy on BrowseComp, BrowseComp-ZH, and xbench-DeepSearch-2510, Macro-F1 on DeepSearchQA, and Item-F1 on WideSearch. “–” denotes unavailable results.
GAIA
FinSearchComp
ShoppingComp
Model
Acc. (%)
Acc. (%)
VPR
SoP
Base
61.81
38.36
0.6660
0.2245
Ours
80.26
58.06
0.6680
0.3576
Table 2: Generalization results on GAIA, FinSearchComp, and ShoppingComp. VPR measures the valid-response rate, while SoP denotes Satisfaction of Products.
Figure 2: Comparison of SFT, final-segment RL, all-segment RL, and all-segment RL with 1/Ki weighting on BC185 and WideSearch. In the weighted variant, each of the Ki segments in trajectory i is assigned weight 1/Ki , so every trajectory has the same total weight. Bars report task performance on the left axis, and lines report average answer turns on the right axis, where lower is better. All evaluations are conducted with answer verification disabled.
Figure 3: Training dynamics of equal-weight all-segment RL, 1/Ki -weighted all-segment RL, and final-segment-only RL. The panels report trajectory accuracy, hard context-budget termination, entry into the turn-penalty zone, Search/LinkSummary tool calls, Seal Memory usage and timing, repeated Search/LinkSummary loops, and the number of segments used for training.
Strategy
Avg@3
Pass@3
Maj@3
First Turn
First Tokens
Avg. Turns
Auto Compaction
54.95
68.11
56.22
217
231.4K
112.2
Active Seal
65.23
76.22
68.11
67
67.5K
217.7
Table 3: Active sealing versus automatic context compaction on BC185. First Turn and First Tokens denote the median turn index and token usage at the first memory operation, respectively, computed over trajectories containing a memory operation. Finally, Avg. Turns denotes the average number of turns across all trajectories.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Case
Before Seal
Sealed State
Strategy after Reset
Outcome
Hipolit Łossowski
The model repeatedly examines Polish officers, but each candidate violates at least one constraint, such as military branch, award rank, death place, or burial site.
It records the rejected candidates together with their exclusion reasons, while retaining the unresolved constraints and verified historical facts.
The model replaces person-by-person enumeration with a structured query combining war participation, military rank, award, death place, and burial location.
1881 ✓
Hiroaki Yamamoto
The model searches the current Rakuten Eagles roster but cannot find a player satisfying both the birth-year and weight constraints, eventually noting that it is “running in circles.”
It preserves the likely team, the failure of the current-roster hypothesis, conflicting weight evidence, and the possibility that the target is a former player.
The model re-verifies the weight constraint, expands the search from the current roster to historical players, and corrects its structured-query formulation.
Hiroaki Yamamoto ✓
Appendix
Table 4: Representative trajectories exhibiting autonomous context recovery. In both cases, the model invokes Seal Memory without an externally imposed compaction trigger, preserves the useful search state, and changes its strategy after entering a refreshed context.
# Runs
Deviation tolerance (percentage points)
Setting
1
2
3
4
5
Best-of-1
5
345
230
185
–
–
Best-of-5
4
785
695
650
190
185
Appendix
Table 5: Minimum subset size K required to satisfy different subset-to-full-set deviation tolerances on the historical calibration runs.
Figure 4: Mechanism and qualitative comparison of automatic compaction and active Seal . (a) Automatic compaction is triggered by context length and can miss an earlier point of semantic stagnation; when triggered late, it may summarize an already entrenched search branch. Active Seal instead couples the agent’s recognition of a dead end with an immediate, structured state transition. (b) Three anonymized cases retain only the behavioral differences between the two strategies: multi-constraint entity joining, rare-constraint prioritization, and entity-centered redirection. To avoid potential benchmark contamination, we anonymize the cases by removing task-specific entities, sources, dates, topics, quotations, clues, and final answers while retaining only behavior-level differences.
Benchmark
Judge Model
Evaluation Protocol
Metric
BC
DeepSeek-V4-Flash
Semantic equivalence between the predicted and reference answers
Accuracy
BC-ZH
DeepSeek-V4-Flash
The same answer-equivalence protocol as BC, with Chinese questions and answers
Accuracy
GAIA
DeepSeek-V4-Flash
Benchmark-specific answer matching covering numerical, textual, list, and date equivalence
Accuracy
xbench
DeepSeek-V4-Flash
Chinese answer-extraction and equivalence prompt following the benchmark grading format
Accuracy
DeepSearchQA
Gemini-2.5-Flash
Official component-level evaluation for single- and set-valued answers ( Gupta et al., 2026 )
Macro-F1
WideSearch
GPT-4.1-2025-04-14
Official table parsing, normalization, semantic alignment, and column-level grading ( Wong et al., 2025 )
Item-F1
Appendix
Table 6: Judge models and evaluation protocols used in our experiments.
Long-horizon search agents must manage a rapidly growing working context as they reason, call tools, and observe information. Naively accumulating all intermediate content can overwhelm the agent, increasing costs and the risk of errors. We propose that effective context management should be adaptive: parts of the agent's trajectory are maintained at different levels of detail depending on their current relevance to the task. To operationalize this principle, we introduce Context-ReAct, a general agentic paradigm for elastic context orchestration that integrates reasoning, context management, and tool use in a unified loop. Context-ReAct provides five atomic operations: Skip, Compress, Rollback, Snippet and Delete, which allow the agent to dynamically reshape its working context, preserving important evidence, summarizing resolved information, discarding unhelpful branches, and controlling context size. We prove that the Compress operator is expressively complete, while the other specialized operators provide efficiency and fidelity guarantees that reduce generation cost and hallucination risk. Building on this paradigm, we develop LongSeeker, a long-horizon search agent fine-tuned from Qwen3-30B-A3B on 10k synthesized trajectories. Across four representative search benchmarks, LongSeeker achieves 61.5% on BrowseComp and 62.5% on BrowseComp-ZH, substantially outperforming Tongyi DeepResearch (43.2% and 46.7%) and AgentFold (36.2% and 47.3%). These results highlight the potential of adaptive context management, showing that agents can achieve more reliable and efficient long-horizon reasoning by actively shaping their working memory.
Search agents are often trained as policies over growing transcripts: the model must decide how to search while also remembering what it has seen, which evidence is useful, which constraints remain open, and which claims have actually been checked. We argue that this formulation puts too much routine state management inside the policy: reinforcement learning is forced to optimize both semantic search decisions and recoverable bookkeeping that the environment can maintain more reliably. We introduce Harness-1, a 20B search agent (retrieval subagent) trained with reinforcement learning inside a stateful search harness. The harness maintains environment-side working memory, including a candidate pool, an importance-tagged curated set, compact evidence links, verification records, compressed and deduplicated observations, and budget-aware context rendering. The policy retains the semantic decisions: what to search, which documents to keep or discard, what to verify, and when to stop. Across eight retrieval benchmarks spanning web, finance, patents, and multi-hop QA, Harness-1 achieves 0.730 average curated recall, outperforming the next strongest open search subagent by +11.4 points and remaining competitive with much larger frontier-model searchers. Its gains are especially strong on held-out transfer benchmarks, suggesting that reinforcement learning over explicit search state can produce retrieval behaviors that generalize beyond the training domains. Our code is available at https://github.com/pat-jj/harness-1.
Pengcheng Jiang, Zhiyi Shi, Kelly Hong +5
University of Illinois at Urbana-Champaign · †UC Berkeley · ✧Chroma
Large language model (LLM)-based search agents answer questions through multi-step interactions with external environments. However, providing complete execution trajectories to the LLM causes unbounded context growth and introduces noise. Existing compression methods reduce context at the cost of important details and often replace erroneous facts without repairing downstream reasoning derived from them. To address this problem, we propose ReTree, a self-correcting tree-structured memory mechanism for search agents. ReTree constructs a bounded per-step reasoning context while preserving source-linked evidence. It models search as an evidence tree whose nodes store bounded summaries, evidence, and revision histories. When newly retrieved evidence contradicts an earlier claim, ReTree traces back to the node where the claim was introduced, replaces outdated evidence, regenerates summaries, prunes affected branches, and resumes search. Source-grounded evidence provenance supports reliable conflict localization and keeps final claims traceable to retrieved passages. Experiments on four public question-answering and search benchmarks show that ReTree consistently outperforms Full-Trajectory ReAct, improving answer accuracy by up to 25.6 percentage points (pp); the average maximum per-step reasoning context of Full-Trajectory ReAct is 1.27--1.51× that of ReTree. These results establish ReTree as an effective self-correcting memory abstraction for long-horizon search.