Traverse: Learning When to Remember, Reset, and Redirect for Long-Horizon Web Search
Organizations: State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University · ByteDance
Abstract
Long-horizon information-seeking agents often accumulate noisy or misleading context, causing early mistakes to persist and making recovery increasingly difficult. We introduce an autonomous search harness in which the agent manages its own search process through three states: Rubric, Answer, and Verify. The agent first defines criteria for a valid answer, searches under these criteria, and then independently verifies the result before deciding whether to terminate or continue searching. It is further equipped with a Seal Memory tool that enables active context management. Training this behavior with reinforcement learning, however, can induce Seal Collapse, resulting in unstable training and preventing the agent from reliably learning when and how to use its memory tools. We solve this with a simple strategy that trains only the final segment after context management. Our 35B model achieves 72.83 on BrowseComp, outperforming comparable open-source systems, and consistently improves over the base model across BrowseComp-ZH, xbench, DeepSearchQA, WideSearch, financial investigation, and product search. Ablations show that autonomous compression outperforms automatic compaction and validate our RL design.
Figures & tables
| Model | Browse Comp | BrowseComp- ZH | xbench- 2510 | DeepSearch QA | Wide Search |
| Closed-source Models | |||||
| GPT-5.5 | 84.4 | – | – | – | – |
| Gemini 3.1 Pro | 85.9 | – | 53.0 | 93.3 | 66.4 |
| Claude Opus 5 | 90.8 | – | – | 95.0 | – |
| General-purpose Models | |||||
| Qwen3.5-35B-A3B | 61.0 | 69.5 | 50.3 | 68.5 | 57.1 |
| GAIA | FinSearchComp | ShoppingComp | ||
|---|---|---|---|---|
| Model | Acc. (%) | Acc. (%) | VPR | SoP |
| Base | 61.81 | 38.36 | 0.6660 | 0.2245 |
| Ours | 80.26 | 58.06 | 0.6680 | 0.3576 |
| Strategy | Avg@3 | Pass@3 | Maj@3 | First Turn | First Tokens | Avg. Turns |
|---|---|---|---|---|---|---|
| Auto Compaction | 54.95 | 68.11 | 56.22 | 217 | 231.4K | 112.2 |
| Active Seal | 65.23 | 76.22 | 68.11 | 67 | 67.5K | 217.7 |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Case | Before Seal | Sealed State | Strategy after Reset | Outcome |
|---|---|---|---|---|
| Hipolit Łossowski | The model repeatedly examines Polish officers, but each candidate violates at least one constraint, such as military branch, award rank, death place, or burial site. | It records the rejected candidates together with their exclusion reasons, while retaining the unresolved constraints and verified historical facts. | The model replaces person-by-person enumeration with a structured query combining war participation, military rank, award, death place, and burial location. | 1881 ✓ |
| Hiroaki Yamamoto | The model searches the current Rakuten Eagles roster but cannot find a player satisfying both the birth-year and weight constraints, eventually noting that it is “running in circles.” | It preserves the likely team, the failure of the current-roster hypothesis, conflicting weight evidence, and the possibility that the target is a former player. | The model re-verifies the weight constraint, expands the search from the current roster to historical players, and corrects its structured-query formulation. | Hiroaki Yamamoto ✓ |
| # Runs | Deviation tolerance (percentage points) | |||||
| Setting | ||||||
| Best-of-1 | 5 | 345 | 230 | 185 | – | – |
| Best-of-5 | 4 | 785 | 695 | 650 | 190 | 185 |
| Benchmark | Judge Model | Evaluation Protocol | Metric |
|---|---|---|---|
| BC | DeepSeek-V4-Flash | Semantic equivalence between the predicted and reference answers | Accuracy |
| BC-ZH | DeepSeek-V4-Flash | The same answer-equivalence protocol as BC, with Chinese questions and answers | Accuracy |
| GAIA | DeepSeek-V4-Flash | Benchmark-specific answer matching covering numerical, textual, list, and date equivalence | Accuracy |
| xbench | DeepSeek-V4-Flash | Chinese answer-extraction and equivalence prompt following the benchmark grading format | Accuracy |
| DeepSearchQA | Gemini-2.5-Flash | Official component-level evaluation for single- and set-valued answers ( Gupta et al., 2026 ) | Macro-F1 |
| WideSearch | GPT-4.1-2025-04-14 | Official table parsing, normalization, semantic alignment, and column-level grading ( Wong et al., 2025 ) | Item-F1 |