cs.CLSep 29, 2026

Traverse: Learning When to Remember, Reset, and Redirect for Long-Horizon Web Search

Authors: Jingyuan Ma, Lynx Aster, He Zhang, Siyao Song, Weijie Yuan, Zhe Zhang, Kai Jia, Zhifang Sui

Organizations: State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University · ByteDance

Abstract

Long-horizon information-seeking agents often accumulate noisy or misleading context, causing early mistakes to persist and making recovery increasingly difficult. We introduce an autonomous search harness in which the agent manages its own search process through three states: Rubric, Answer, and Verify. The agent first defines criteria for a valid answer, searches under these criteria, and then independently verifies the result before deciding whether to terminate or continue searching. It is further equipped with a Seal Memory tool that enables active context management. Training this behavior with reinforcement learning, however, can induce Seal Collapse, resulting in unstable training and preventing the agent from reliably learning when and how to use its memory tools. We solve this with a simple strategy that trains only the final segment after context management. Our 35B model achieves 72.83 on BrowseComp, outperforming comparable open-source systems, and consistently improves over the base model across BrowseComp-ZH, xbench, DeepSearchQA, WideSearch, financial investigation, and product search. Ablations show that autonomous compression outperforms automatic compaction and validate our RL design.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. LongSeeker: Elastic Context Orchestration for Long-Horizon Search Agents

    May 6, 2026Yijun Lu, Rui Ye, Yuwen Du +3Long-Horizon AgentsOrchestration

  2. Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses

    Jun 1, 2026Pengcheng Jiang, Zhiyi Shi, Kelly Hong +5Search AgentsRetrievers

  3. Self-Correcting Long-Horizon Search Agents via Tree-Structured Memory

    Aug 11, 2026Aijun Yang, Qianxue Guo, Ziyi Huang +3Search Agents