cs.CLSep 29, 2026

Beyond Oracle Communication: Benchmarking Interactive Intent Alignment Under Miscommunication and Evolving User Intent

Authors: Zheyuan Zhang, Mengyuan Chao, Ke Xiao, Ziyi Chen, Daoan Zhang, Yan Zhang, Yanfang Ye, Wei Xu

Organizations: ByteDance Inc., USA · University of Notre Dame

Abstract

Modern LLM agents increasingly tackle complex tasks through interactive, long-horizon exchanges with users, while existing benchmarks generally assume that users always accurately and sufficiently communicate a fixed intent. However, this oracle communication assumption rarely holds in practice: users may miscommunicate, change their goals, and run out of patience. We define this task setting as Interactive Intent Alignment, where agents must recover and continuously track the user's current intent despite imperfect communication and evolving goals. To study this setting, we introduce Drift-Bench++, a principled benchmark construction pipeline for verified executable tasks with controlled misalignment and intent shifts, along with an interaction protocol featuring finite patience, diverse simulated users, and silent interaction-conditioned shifts. We further develop GRIP, a comprehensive evaluation protocol covering task grounding, user realism, inquiry effectiveness, and adaptation to evolving intent. Across diverse environments, models, and interaction conditions, stronger interaction consistently helps but remains far from oracle performance; Validation on deployed ProdAgent sessions further shows that the modeled failures are prevalent and consequential in deployment. By providing a unified, executable benchmark for interactive intent alignment, Drift-Bench++ offers a foundation for evaluating and advancing agents under realistic communication and evolving intent.

Figures & tables

Appendix figures & tables18 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. LLMs Get Lost in Evolving User Intent

    Jul 22, 2026Jihoon Tack, Philippe Laban, Jennifer NevilleMulti-Turn InteractionsConversational Context

  2. IDRBench: Benchmarking the Interactive Capabilities of Deep Research Agents

    Jan 10, 2026Yingchaojie Feng, Qiang Huang, Xiaoya Xie +4Deep ResearchAgentic Benchmarks

  3. Interactive Task Alignment as a POMDP

    Jul 17, 2026Andy Dai, Zexue He, Zhenyu Zhang +2Partially Observable Markov Decision Process