cs.CLSep 24, 2026

IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis

Authors: Xingyu Wu, Yuchen Yan, Zhengxi Lu, Siqi Chen, Xin ZHANG, Aiting Liu, Chao Deng, Jie Liu, +4 more

Organizations: Zhejiang University · Tencent

Abstract

Deep search requires LLM agents to decompose complex queries, search for evidence, and synthesize grounded answers, yet existing ReAct-style agents suffer from two limitations: role coupling, where one policy must handle planning, evidence use, and synthesis; and context accumulation, where growing search histories introduce noise and obscure useful information. To address these issues, we propose IterSynth, a role-decoupled and summary-based paradigm that alternates between a Planner for identifying information needs and a Synthesizer for integrating evidence into an evolving summary state. This design separates planning from synthesis while using the summary as the persistent state of search, reducing both capability coupling and context noise. To train IterSynth effectively, we further introduce Role-Decoupled Policy Optimization (RDPO) for reinforcement learning, which combines terminal outcome rewards with turn-level rubric evaluations and computes role-specific advantages for more precise credit assignment. Experiments on five long-horizon deep-search benchmarks such as BrowseComp and Xbench-DS show that IterSynth-8B achieves an average score of 50.7, surpassing the strongest prior ≤\leq8B agent by +4.2%. Moreover, IterSynth serves as a model-agnostic prompting paradigm, delivering substantial zero-shot gains over ReAct and similar prompting paradigms on frontier proprietary models.

Figures & tables

Explore similar work

CardsList
  1. SearchArt: Training Long-Horizon Search Agent with Scalable Synthetic and Verified Task

    Jul 25, 2026Lang Mei, Xiaohan Yu, Chong Chen +27Synthetic TaskSynthetic Data

  2. QUEST: Training Frontier Deep Research Agents with Fully Synthetic Tasks

    May 22, 2026Jian Xie, Tianhe Lin, Zilu Wang +16Deep ResearchSynthetic Task

  3. G-ReAct: Graph-Guided Deep Search via Structure-State Co-Evolution

    Aug 2, 2026Shaoxiong Yang, Mengyuan Zhang, Shaojun Lin +4DeepseekReactive