DAGent: Evaluate-then-Grow Planning for Deep Research Agents
Organizations: New York University · New York University Shanghai
Abstract
Deep research tasks require agents to navigate large knowledge spaces, synthesize evidence across many sources, and adapt their plans as findings emerge. Directed acyclic graph (DAG)-based multi-agent systems suit this setting because they support parallel execution and isolate each sub-task within a focused dependency context. Yet existing DAG-based agents instantiate a task-level plan before execution and repair the graph only after failures or missing evidence are observed. This Plan-then-Patch strategy is brittle for deep research: the system commits most strongly when its evidence is weakest, and later revisions waste computation on branches that should not have been planned. We propose DAGent, a DAG-based multi-agent framework with Evaluate-then-Grow incremental planning: an Orchestrator grows the task graph one batch at a time, conditioning each expansion on confidence and uncertainty signals from completed nodes. A hierarchical context layer propagates compact QueryDocs by default while preserving full execution traces for on-demand recall. The recorded DAG topology admits structural RL signals that outcome-only recipes cannot define; DAGRPO, a GRPO adaptation, injects topology-conditioned credit on Executor rollouts and a structural compliance regularization on Orchestrator plans. Across BrowseComp-Plus, GAIA, and xbench-DeepSearch, DAGent surpasses the strongest open-source baseline by 5.3 / 5.8 / 2.0 points at the Qwen3-235B-A22B scale, and the lead replicates across four open-source backbones and extends to GPT-5 at 327K context. At the Qwen3-8B scale, DAGRPO improves over a same-budget outcome-only GRPO baseline by 3.0 average Pass@1 points. A same-architecture comparison shows that evidence-conditioned planning reaches higher accuracy at lower per-task token, tool-call, and step footprints than its Plan-then-Patch counterpart. Code: https://github.com/hanwenliu6825/DAGent
Figures & tables
| Backbone | Max #Token | Agent Paradigm | BrowseComp-Plus | GAIA | xbench-DS | ||||||
| Easy | Med. | Hard | Avg. | L1 | L2 | L3 | Avg. | Avg. | |||
| Training-free | |||||||||||
| Qwen3-8B | 32K | ReAct Agent | 46.0 | 14.0 | 0.0 | 20.0 | 38.5 | 23.1 | 25.0 | 29.1 | 39.0 |
| 109K | ReAct Agent | 64.0 | 18.0 | 2.0 | 28.0 | 46.2 | 26.9 | 25.0 | 34.0 | 44.0 | |
| 32K N | Summary Agent | 76.0 | 20.0 | 2.0 | 32.7 | 51.3 | 26.9 | 16.7 | 35.0 | 54.0 | |
| 32K N | Fold Agent | 78.0 | 24.0 | 2.0 | 34.7 | 48.7 | 28.8 | 16.7 | 35.0 | 52.0 | |
| Method | BrowseComp-Plus | GAIA | xbench-DS | Overall | ||||
| Avg. | Avg. | Avg. | Avg. | |||||
| DAGent (Full) | 47.3 | — | 55.3 | — | 65.0 | — | 55.9 | — |
| Plan-then-Patch Variant | 42.0 | -5.3 | 50.5 | -4.8 | 61.0 | -4.0 | 51.2 | -4.7 |
| w/o Evaluate-then-Grow Planning | 33.3 | -14.0 | 42.7 | -12.6 | 55.0 | -10.0 | 43.7 | -12.2 |
| w/o Selective Propagation | 34.7 | -12.6 | 44.7 | -10.6 | 57.0 | -8.0 | 45.5 | -10.4 |
| w/o QueryDoc | 39.3 | -8.0 | 47.6 | -7.7 | 60.0 | -5.0 | 49.0 | -6.9 |
| Regularization | Configuration | BrowseComp-Plus | GAIA | xbench-DS | Overall | |||||
| Avg. | Avg. | Avg. | Avg. | |||||||
| 0.00 | on | — | 44.0 | -5.6 | 49.5 | -3.9 | 61.0 | -4.7 | 51.5 | -4.7 |
| 0.25 | on | — | 48.7 | -0.9 | 52.4 | -1.0 | 64.0 | -1.7 | 55.0 | -1.2 |
| 0.50 | on | DAGRPO | 49.6 | — | 53.4 | — | 65.7 | — | 56.2 | — |
| 0.75 | on | — | 48.0 | -1.6 | 52.4 | -1.0 | 64.0 | -1.7 | 54.8 | -1.4 |
| 1.00 | on | w/o topology credit | 47.3 | -2.3 | 51.5 | -1.9 | 64.0 | -1.7 | 54.3 | -2.0 |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| BrowseComp-Plus | GAIA | xbench-DeepSearch | |
| Tasks | 150 | 103 (text-only validation) | 100 (2505 release) |
| Difficulty splits | 50 easy / 50 medium / 50 hard | 39 L1 / 52 L2 / 12 L3 | none |
| Language | English | English | Chinese |
| Search | Local dense retrieval | Google search via Serper | Google search via Serper |
| Page access | Verified corpus | Jina extraction | Jina extraction |
| Judge prompt | Benchmark rubric | Equivalence prompt | Benchmark rubric (Chinese) |
| Parameter | Scope | Value |
| Prompt length | per context | tokens |
| Response length, including tool outputs | per context | tokens |
| Orchestrator iterations | per task | |
| Session timeout | per task | s |
| Executor turns | per sub-task | |
| Executor timeout, evaluation / training | per sub-task | s / s |
| Workflow | Setting | Cap |
| ReAct Agent | Context window | K or K tokens |
| Summary Agent / Fold Agent | Sessions per task | |
| Flash-Searcher | Action steps per task | |
| Summary interval | (release default ) | |
| FlowSearch | Main iterations per task | |
| Planner iterations / nodes | / |
| Parameter | Value |
| Hardware | H200, tensor parallel |
| LoRA rank / alpha | / , all linear layers |
| Optimizer, learning rate | AdamW, |
| Tasks per batch, rollouts per task | , |
| PPO mini-batch (tasks / trajectories) | / |
| Asymmetric clip / | / |
| Subset | N | Fleiss’ (3H) | Cohen’s (H vs J) | Raw % |
| BrowseComp-Plus | 50 | 0.98 | 0.96 | 98.0 |
| GAIA | 50 | 0.98 | 0.96 | 98.0 |
| xbench-DeepSearch | 50 | 0.96 | 0.92 | 96.0 |
| Aggregate | 150 | 0.97 | 0.95 | 97.3 |
| Backbone | Vendor | Agent Paradigm | BrowseComp-Plus | GAIA | xbench-DS | ||||||
| Easy | Med. | Hard | Avg. | L1 | L2 | L3 | Avg. | Avg. | |||
| Qwen3-32B | Alibaba | ReAct Agent (32K) | 50.0 | 16.0 | 2.0 | 22.7 | 46.2 | 28.8 | 8.3 | 33.0 | 54.0 |
| ReAct Agent (109K) | 66.0 | 24.0 | 4.0 | 31.3 | 51.3 | 36.5 | 16.7 | 39.8 | 58.0 | ||
| Summary Agent | 80.0 | 36.0 | 0.0 | 38.7 | 51.3 | 28.8 | 25.0 | 36.9 | 55.0 | ||
| Fold Agent | 82.0 | 40.0 | 2.0 | 41.3 | 56.4 | 38.5 | 16.7 | 42.7 | 55.0 | ||
| Flash-Searcher | 80.0 | 34.0 | 2.0 | 38.7 | 56.4 | 40.4 | 25.0 | 44.7 | 58.0 | ||
| Benchmark | Workflow | Pass@1 | Tool calls | External calls | Steps | Calls / step | Tokens (M) | Time (s) | Nodes | Off-chain |
| BrowseComp-Plus | DAGent (Full) | 47.3 | 67.6 | 60.8 | 42.7 | 1.58 | 1.20 | 541.6 | 11.8 | 0.25 |
| Plan-then-Patch variant | 42.0 | 87.9 | 82.3 | 54.8 | 1.60 | 1.68 | 731.9 | 17.6 | 0.40 | |
| Flash-Searcher | 38.7 | 51.3 | 42.1 | 35.8 | 1.43 | 0.91 | 388.7 | – | – | |
| FlowSearch | 43.3 | 83.5 | 76.7 | 47.6 | 1.75 | 1.57 | 714.5 | – | – | |
| GAIA | DAGent (Full) | 55.3 | 36.8 | 33.1 | 31.2 | 1.18 | 0.66 | 461.8 | 5.4 | 0.20 |
| Plan-then-Patch variant | 50.5 | 45.3 | 42.0 | 37.6 | 1.20 | 0.85 | 584.6 | 7.3 | 0.30 |
| Benchmark | Setting | Pass@1 | Tokens | Tool calls | Steps | Nodes | Off-chain |
| BrowseComp-Plus | DAGent (training-free) | 40.0 | 1.14M | 65.5 | 41.5 | 12.2 | 0.25 |
| + RL (GRPO) | 46.0 | 0.96M | 59.6 | 36.5 | 11.3 | 0.21 | |
| + RL (DAGRPO) | 49.6 | 0.90M | 56.3 | 34.5 | 10.3 | 0.16 | |
| GAIA | DAGent (training-free) | 46.6 | 0.62M | 35.5 | 29.5 | 5.3 | 0.20 |
| + RL (GRPO) | 50.8 | 0.53M | 33.0 | 26.8 | 4.6 | 0.17 | |
| + RL (DAGRPO) | 53.4 | 0.51M | 31.6 | 25.7 | 4.0 | 0.13 |