Modern LLM agents increasingly tackle complex tasks through interactive, long-horizon exchanges with users, while existing benchmarks generally assume that users always accurately and sufficiently communicate a fixed intent. However, this oracle communication assumption rarely holds in practice: users may miscommunicate, change their goals, and run out of patience. We define this task setting as Interactive Intent Alignment, where agents must recover and continuously track the user's current intent despite imperfect communication and evolving goals. To study this setting, we introduce Drift-Bench++, a principled benchmark construction pipeline for verified executable tasks with controlled misalignment and intent shifts, along with an interaction protocol featuring finite patience, diverse simulated users, and silent interaction-conditioned shifts. We further develop GRIP, a comprehensive evaluation protocol covering task grounding, user realism, inquiry effectiveness, and adaptation to evolving intent. Across diverse environments, models, and interaction conditions, stronger interaction consistently helps but remains far from oracle performance; Validation on deployed ProdAgent sessions further shows that the modeled failures are prevalent and consequential in deployment. By providing a unified, executable benchmark for interactive intent alignment, Drift-Bench++ offers a foundation for evaluating and advancing agents under realistic communication and evolving intent.
Figures & tables
Figure 1 : Overview of Drift-Bench++. a) A user query may misalign with latent intent, and the agent’s response may trigger a silent intent shift. b) Persona-conditioned behavior shapes the interaction beyond prompting style, and c) GRIP evaluates the resulting alignment process.
Figure 2 : Overview of the Drift-Bench++ data synthesis pipeline. Host tasks are mapped into intent graphs, and converted into verified misaligned instances. Lightweight benchmark-specific adapters and native oracles allow the remaining construction to be shared across domains.
Table 2 : Main comparison under the default evaluation condition (DeepSeek-V4-Pro backbone, Rational persona, single-fault requests, and intent shifting enabled). Results report mean ± sd over three seeded runs. Best results in each column are bolded and second-best results are underlined.
Figure 3 : Comparison between Various LLM Backbones. The benefit of interaction is backbone-robust. We show success rate of three agents across five LLM backbones. Clarification beats No-Ask baseline on every backbone and benchmark, and Self-evolving leads in 14 of 15 cells.
Figure 5
Figure 6 : Persona profiles of the Self-evolving agent on WebShop. Fans: one metric each, one wedge per persona, avg/max below. Bars: persona identification and consistency, judged from user messages alone with user profile prompt hidden on purpose.
Figure 7 : ProdAgent validation on 16,596 real production agent sessions over a 10-week deployment. Left: Comparison of misaligned and matched clean sessions in terms of platform evaluation scores and average tool calls per session; and Right: fault-type and shift-type distributions under strict two-judge agreement against the extracted intents provided by the platform.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmarks
Tasks
Graphs
Rejection Rate
Samples
WebShop
500
257
31%
1,780
τ2 -Retail
114
98
47%
555
τ2 -Airline
50
40
48%
208
Appendix
Table 4: Key statistics of the construction pipeline with the rejection rate of the three-step verifier and the resulting sample numbers.
Figure 8 : Communication-fault taxonomy for interactive intent misalignment. Drift-Bench++ organizes communication failures into four families: flaws of intention, premise, parameter, and expression , which are further instantiated as eleven operational fault types used in verified query perturbation. The hierarchy highlights where cooperative breakdown occurs, ranging from obscured goals and false assumptions to missing task parameters and ambiguous language.
Disclosure
Patience & Feedback
Silent Intent Shifts
Persona
Slots/Answer
Volunteer
Refuse
Patience ×
Hint
Placements
Suggestibility
Rational
1
10%
0%
1.0
30%
100%
20%
Dependent
2
35%
0%
1.2
60%
66%, 33%
80%
Avoidant
1
0%
50%
0.8
10%
100%
10%
Intuitive
1
20%
10%
0.9
30%
100%, 50%
50%
Spontaneous
1
30%
5%
0.7
40%
100%, 50%
60%
Appendix
Table 5: Persona configurations. Disclosure : slots revealed per answered question, the probability of volunteering an unasked requirement, and the probability of refusing an otherwise-valid answer. Patience & feedback : the multiplier scales the base patience budget B=10 ; hint is the probability that a rejection names what is wrong rather than a bare refusal. Silent intent shifts : placements position each scheduled shift as a percentage of the persona’s own initial budget, crossed as patience falls to or below that point; each crossing fires with probability 60% (the scheduled fire coin), and an on-topic clarification question can instead fire a shift early with the persona’s suggestibility, consuming the last remaining placement so the shift count never grows (Appendix B.5 ).
Figure 9 : Left: Aim-Recovery changes from Just Ask (circle) to Self-Evolving (arrowhead) across five backbones. Right: successful hidden-intent episodes decomposed into Earned (solid) and Inferred (hatched) outcomes. Together, the panels show that stronger interaction primarily improves intent recovery and often shifts success toward explicitly recovered requirements, while Aim alone does not guarantee better recovery. Results are averaged over three seeded runs.
Figure 10 : Reaction under composed faults. Median turns from the last relevant intent shift to acceptance as fault complexity k increases. Successful interactive trajectories generally require more post-shift turns at higher complexity, revealing the interaction cost of recovering coupled intent errors. Mean ± sd over three seeds; axes are scaled per panel.
Figure 11 : Patience consumed by intent shifting. Each arrow shows the change in remaining patience from shifts disabled (open circle) to enabled (filled marker). Shifts consume additional interaction budget in every method and benchmark, with the largest drain on τ2 -Retail. Values are from Table 8 ; axes are scaled per panel.
Figure 12 : Persona rankings across GRIP dimensions. Each cell ranks the five personas within one benchmark and metric for Self-Evolving (1 = best; darker = higher rank). Rankings show a stable broad difficulty structure but metric-dependent reversals, particularly between Rational and Dependent users. Just Ask exhibits a similar overall pattern. Values are from Table 9 .
Figure 13 : Generation prompt for a single communication fault, instantiated with a real task from the released corpus. Each of the eleven fault types has its own card and a type-specific final rule; the writer never sees the target product or the ground-truth answer set. Every output must pass dual extraction and executable admission before it ships. Section labels are added only for presentation.
Figure 14 : Extraction prompts that recover the structured edit a generated query actually expresses. Each questionnaire is answered independently by two models from different families at temperature zero, and a candidate proceeds only when both passes agree on the fault family and its target slot. Section labels are added only for presentation.
Figure 15 : Composite generation prompt, shown for a drawn pair of faults. Every drawn fault is pre-assigned its own target requirement before writing, and all cards are issued in one pass so no later rewrite can silently repair an already-checked fault; when three or more faults include a whole-message fault, that fault is applied as a second transformation-only pass. The final text is verified per component. Section labels are added only for presentation.
Figure 16 : The simulated user. The model only voices replies: what may be revealed is computed by code (the [PRIVATE] tags), acceptance of a proposal is decided by the executable verifier rather than by the model, and whether a rejection carries a hint is a persona trait rolled per rejection. A spoken reaction that contradicts the executable verdict is regenerated once with the verdict as a constraint. Section labels are added only for presentation.
Figure 17 : The five persona bios conditioning the simulated user (abridged; inherited verbatim from Drift-Bench so persona-conditioned results remain comparable), and the behavioral parameters Drift-Bench++ adds: slots revealed per answered question, the probability of volunteering an unasked requirement, the probability of declining a valid reveal, the patience multiplier, the probability a rejection carries a hint, intent-shift placements on the patience axis, and suggestibility to agent-prompted shifts. A persona changes what is said, never what is true: acceptance is executable and persona-independent.
Figure 18 : The prompt base every evaluated method shares: a deliberately strategy-free action interface and a store manual describing the environment mechanics. The No-Ask baseline receives the same prompt with the user-message action removed from the interface entirely; the patience meter is shown only to methods whose design includes budget planning. Section labels are added only for presentation.
Figure 19 : Method-specific guidance of the two clarification-capable reference methods. Just Ask adds a single ask-early instruction to the shared base. Self-Evolving extends the budget-planner method with lessons it distills from its own finished episodes (Figure 20 ); every run starts with an empty book, nothing crosses runs, personas, seeds, or benchmarks, and once the playbook becomes non-empty, its lessons are available from the beginning of subsequent episodes. Section labels are added only for presentation.
Figure 20 : The reflection prompt that collects Self-Evolving’s lessons between generations of a run, shown with the shopping-domain nouns (a retail run reads "customer" for "shopper"; the instruction is otherwise identical). Episode digests are projected onto agent-observable signals and mechanically redacted before reflection, and the credit labels are computed by code, so the reflection judges strategy rather than reconstructing what happened. Section labels are added only for presentation.
Figure 21 : An illustrative τ2 -Airline episode, condensed; quotes are verbatim. The plural “reservations” is lexically two-way, hiding a second booking; a scheduled relaxation resolves it silently before the first exchange. The deeper difficulty is structural: under airline policy a basic-economy booking cannot simply be modified, so the true intent’s golden sequence is update-then-cancel — the user’s own answer (“changing, not canceling”) describes their goal, not the policy’s implementation, and the missing cancellation surfaces only through rejection feedback. The agent asks the policy-required cancellation reason, the user volunteers the rebooking expectation, and the resubmission is accepted. Thirteen turns, patience 12 → 4.
Figure 22 : An illustrative WebShop episode, condensed from the released run; quotes are verbatim. The vague opening hides the $120 ceiling; the agent grounds the task in the catalogue before asking one two-part question that both disambiguates “natural hair” and elicits the budget — and that question triggers a silent refinement. A mechanically rejected purchase (Buy Now pressed on a details page) costs 4 patience and draws a natural complaint; the agent recovers, a second scheduled shift relaxes the intent, and the resubmitted purchase is accepted against the intent as it stands at commitment.
Figure 23 : An illustrative τ2 -Retail interaction trace showing recovery under repeated silent intent shifts.
As LLMs become more capable, they are increasingly deployed as collaborative agents, taking on user-delegated tasks through iterative interaction. Yet genuine interaction is inherently dynamic: users rarely specify their intent upfront, instead disclosing, revising, and reshaping it as the conversation unfolds. Despite this, LLMs are still predominantly evaluated or trained in single-turn, fully-specified settings, leaving open a fundamental question: how well do LLMs track and act on user intent as it evolves over the course of a conversation? To study this, we introduce a framework that transforms static, single-turn tasks into dynamic multi-turn conversations in which the user's intent evolves across turns--incrementally revealed, revised, and at times redirected mid-conversation--while preserving each task's original evaluation protocol, enabling existing benchmarks to be reused as controlled testbeds without new annotation. Across multiple tasks, we surface a consistent phenomenon: strong static-setting performance does not transfer to the evolving-intent setting, with substantial drops across model families. Our findings point to a fundamental gap: today's LLMs do not yet faithfully track and act on the user's evolving intent, a capability invisible to static evaluation yet critical for future collaborative agents.
Large Language Model (LLM)-based deep research agents perform multi-step reasoning, web exploration, and long-form report generation. In these long-horizon workflows, early deviations from user intent can misdirect research and propagate through planning, search, and synthesis, making timely interaction essential. However, existing benchmarks primarily treat deep research as a static input-output task, overlooking agents' ability to elicit and use user feedback. We introduce IDRBench, a benchmark for evaluating interactive deep research with controlled opportunities for clarification. Within a common workflow and stage-wise interaction budget, IDRBench compares autonomous and interactive trajectories, measuring interaction benefit through changes in task-specific report alignment and interaction cost through turns and tokens. Comprehensive experiments on 100 tasks with seven proprietary and open-weight LLMs show that interaction improves all five alignment measures for every model, yielding an average gain of 6.39 points, while revealing distinct trade-offs among autonomous performance, alignment gain, and communication cost. At the task level, interaction improves performance in 74.4% of cases but degrades it in 19.9%, demonstrating that access to clarification alone does not guarantee better outcomes: success depends on what agents ask and how effectively they incorporate the resulting feedback.
Yingchaojie Feng, Qiang Huang, Xiaoya Xie +4
School of Computing, National University of Singapore · School of Intelligence Science and Engineering, Harbin Institute of Technology (Shenzhen) · Zhejiang University +1
Current benchmarks for language models primarily evaluate execution on fully specified tasks. However, real user tasks are often ambiguous. Users arrive with incomplete, exploratory, or even inconsistent goals, requiring the assistant to first determine the intended task before carrying it out. We study this problem as task alignment: the ability to align with a user on their intended task. We introduce a general framework for converting specified tasks into underspecified interactions, formalized as a POMDP in which the model must infer a latent task from partial and evolving user intent. We validate our user simulator post hoc with a human user study. Across shopping, coding, and professional work settings, we find that while models often perform well once the task is specified, models still struggle with task alignment: current models act prematurely, interact ineffectively, and fail to resolve ambiguous requests. Models on average recover the user's intended task only 22-32% of the time under ambiguity. In a human study in the same setting, humans reach 48%, outperforming all evaluated models. We show that post-training with supervised fine-tuning and reinforcement learning improves task alignment, but models still lag behind humans in resolving uncertainty through interaction. Together, our results suggest that current models still lack key interaction abilities required for reliable agency.