cs.AIOct 4, 2026

LexiHorizon: Stabilizing Reinforcement Learning for Long-Horizon Deep Search

Authors: Zhiqing Nong, Liang Wen, Chao-Hsuan Liu

Abstract

Deep search agents tackle complex knowledge tasks through iterative retrieval, multi-hop reasoning, and evidence synthesis across multiple sources. Existing approaches typically assume relatively stable retrieval systems and operate over short-horizon tool interaction. However, when retrieval is sensitive to query formulation, even a semantically appropriate query may fail to surface critical evidence because of mismatched entity names, aliases, or keyword combinations. Recovering from such failures requires repeated query reformulation and longer interaction trajectories. This setting poses a distinct training challenge, as the policy must sustain long-horizon query exploration while managing an expanding volume of retrieved content. We propose LexiHorizon, a framework for training search agents over long horizons that expands the trajectory context budget, manages accumulated retrieval content using a window over recent tool observations while preserving the reasoning history, and introduces an outcome-gated search-effort reward that provides a bounded bonus for tool invocations to trajectories with nonzero answer reward. Experiments on XBench, WebWalkerQA, and BrowseComp-ZH show that the resulting 9B model consistently outperforms both its base model and MiroThinker-1.7-mini, with maximum absolute gains of 8.7 and 23.8 percentage points, respectively. These results suggest that combining an extended context budget with reasoning-preserving context management benefits long-horizon deep search agents.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SearchArt: Training Long-Horizon Search Agent with Scalable Synthetic and Verified Task

    Jul 25, 2026Lang Mei, Xiaohan Yu, Chong Chen +27Synthetic TaskSynthetic Data

  2. Harness-Search: Guiding Long-Horizon Search through Multi-Agent Coordination

    Oct 4, 2026Shanyong Wang, Zhenwen Ji, Lei Jin +7Search AgentsLong-Horizon Agents

  3. Lost in the Maze: Overcoming Context Limitations in Long-Horizon Agentic Search

    Oct 21, 2025Howard Yen, Yoonsang Lee, Ashwin Paranjape +5Long-Horizon AgentsDeep Research