cs.AISep 27, 2026

What Happens During Autonomous Deep Research After the User Steps Away?

Authors: Yimin Liu, Yijia Zhang, Yanmin Li, Tangwen Luo, Yuze Li, Ziling Yao, Zhi Yang

Organizations: University of Southern California · University of Michigan, Ann Arbor · Institute of Automation, Chinese Academy of Sciences · Zhejiang University · Beijing University of Posts and Telecommunications · Shanghai University of Finance and Economics

Abstract

In autonomous deep research, a user provides a task and relevant background, then leaves the agent to conduct an extended investigation without further human intervention. We study how this initial user information is reflected in intermediate actions and how these actions relate to final recommendations. We introduce DRaligned, a counterfactual behavioral evaluation framework built on PDR-Bench. By varying one task-relevant user factor while keeping the remaining context fixed, we compare acquisition requests, working drafts, and final reports. Source-grounded extraction, blinded local judgments, and deterministic aggregation yield coarse directional measurements while leaving ambiguous cases unresolved. Our experiments show that strong user-specific delivery can emerge from a largely shared research process: agents investigate similar broad questions but allocate requests differently, and final recommendations distinguish user conditions more clearly than explicit requests do. Reports can also integrate user factors that were not jointly visible during acquisition. In readable draft-to-report comparisons, recommendations often retain their coarse user-specific direction despite substantial rewriting. Final directional differences recur across tested agent models, execution harnesses, and evaluator models, even as execution paths vary. These findings describe how initial user information shapes autonomous research and clarify the relationship between the process an agent follows and the recommendations it delivers.

Figures & tables

Appendix figures & tables26 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. DRACULA: Hunting for the Actions Users Want Deep Research Agents to Execute

    Apr 26, 2026Nishant Balepur, Malachi Hamada, Varsha Kishore +9Action SelectionLong-Horizon Agents

  2. DeepResearch Bench II: Diagnosing Deep Research Agents via Rubrics from Expert Reports

    Date pendingRuizhe Li, Mingxuan Du, Benfeng Xu +3Deep ResearchRubrics