cs.CLJan 10, 2026

IDRBench: Benchmarking the Interactive Capabilities of Deep Research Agents

Authors: Yingchaojie Feng, Qiang Huang, Xiaoya Xie, Zhaorui Yang, Jun Yu, Wei Chen, Anthony K. H. Tung

Organizations: School of Computing, National University of Singapore · School of Intelligence Science and Engineering, Harbin Institute of Technology (Shenzhen) · Zhejiang University · State Key Lab of CAD&CG, Zhejiang University

Abstract

Large Language Model (LLM)-based deep research agents perform multi-step reasoning, web exploration, and long-form report generation. In these long-horizon workflows, early deviations from user intent can misdirect research and propagate through planning, search, and synthesis, making timely interaction essential. However, existing benchmarks primarily treat deep research as a static input-output task, overlooking agents' ability to elicit and use user feedback. We introduce IDRBench, a benchmark for evaluating interactive deep research with controlled opportunities for clarification. Within a common workflow and stage-wise interaction budget, IDRBench compares autonomous and interactive trajectories, measuring interaction benefit through changes in task-specific report alignment and interaction cost through turns and tokens. Comprehensive experiments on 100 tasks with seven proprietary and open-weight LLMs show that interaction improves all five alignment measures for every model, yielding an average gain of 6.39 points, while revealing distinct trade-offs among autonomous performance, alignment gain, and communication cost. At the task level, interaction improves performance in 74.4% of cases but degrades it in 19.9%, demonstrating that access to clarification alone does not guarantee better outcomes: success depends on what agents ask and how effectively they incorporate the resulting feedback.

Explore similar work

CardsList
  1. An Interactive Paradigm for Deep Research

    May 22, 2026Lin Ai, Victor S. Bursztyn, Xiang Chen +2Deep ResearchLong-Horizon Tasks