When a large language model handles a multi-turn task and a user proposes a change but ultimately rejects it, the model should continue as if nothing changed. We find a surprising failure: merely mentioning a rejected change can derail task execution, even when the user's final intent remains unchanged. To systematically study language model behavior under evolving user intent, we introduce Intent-Eval, a controlled benchmark spanning tool actions, code, databases, and mathematics. Across diverse tasks, models are vulnerable to both rejected proposals and superseded requirements, consistent with mentioned-as-in-effect confusion: conversational content is treated as active requirements even after it has been rejected or replaced. Accuracy degradation can deepen or persist as interaction continues, highlighting the need to distinguish what has been mentioned from what remains in effect. Building on this insight, we propose Intent-OPSD, a decision-conditioned on-policy self-distillation framework with Teacher and Student initialized from the same model. The frozen Teacher provides active-intent supervision from the complete task matching the user's decision, training the Student on the full dialogue to follow active requirements reflecting user intent.
Figures & tables
Figure 1 : Illustration of active intent. (a) Single-turn controls and targets. (b) The Original task distributed across turns; U3–U4 mark the insertion point for the other conditions. (c) Neutral clarifies; Retained rejects and Revised accepts the same proposal. Only Revised changes the target.
Figure 2 : Overview and statistics of Intent-Eval . The benchmark comprises 414 source tasks across four domains. Each source task has four multi-turn conditions and four single-turn controls, yielding 3,312 evaluation instances per model.
Figure 3 : Intent-OPSD : the frozen Teacher sees the active task; the Student sees the dialogue.
Model
Domain
Single
Multi
Original
Direct Rev.
Revised
Retained
Original
Neutral
Revised
Retained
3.6-27B
Math
98.06
97.09
87.38
93.20
73.79
71.84
66.99
61.17
Code
95.00
93.00
90.00
88.00
81.00
76.00
80.00
72.00
DB
92.52
91.59
91.59
87.85
42.99
43.93
42.06
38.32
Actions
96.15
95.19
96.15
85.58
49.04
44.23
47.12
48.08
3-8B
Math
92.23
90.29
75.73
67.96
64.08
64.08
54.37
37.86
Table 1 : Model-by-domain accuracy (%) across single-turn controls and the four multi-turn conditions. In Single, Direct Rev. states the revised task directly, while Revised includes the proposal and acceptance. Overall aggregates all 414 task sources across the eight models shown here; two additional models appear in Appendix Table 8 . Cell backgrounds indicate changes relative to Single Original: for declines, for gains, and for no change.
Figure 4 : Inactive content in errors. Bars show inclusion rates among new Retained (left) and Revised (right) errors, pooled over ten models; headers aggregate all domains.
Figure 5 : Accuracy change with depth relative to each curve’s k=0 baseline. Y-ranges are shared within domains. Results for six additional models are in Appendix E.3 (Figure 7 ).
Model
Domain
Base
SFT
Intent-OPSD
Orig.
Neut.
Revised
Retained
Orig.
Neut.
Revised
Retained
Orig.
Neut.
Revised
Retained
3-8B
Math
64.08
64.08
54.37
37.86
61.17
61.17
59.22
28.16
66.02
67.96
64.08
23.30
Code
41.00
43.00
43.00
26.00
35.00
36.00
37.00
24.00
39.00
42.00
45.00
28.00
DB
38.32
26.17
32.71
17.76
43.93
38.32
43.93
20.56
45.79
44.86
42.99
27.10
Actions
51.92
50.00
46.15
39.42
68.27
71.15
31.73
57.69
66.35
60.58
66.35
55.77
4-26B- A4B
Math
69.90
65.05
66.02
58.25
70.87
71.84
66.02
62.14
81.55
77.67
72.82
60.19
Table 2 : Accuracy (%) and Intent-OPSD depth gains (pp). For accuracy, conditions and model–domain pairs have equal weight. Colors show changes from Base ( lower, higher, unchanged); shading reflects magnitude. Gains use overall accuracy across the four interaction paths; k counts added event blocks. Appendix E.4 provides detailed results by path and depth.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Table 3: Domain-specific revisions and reference updates.
Condition
Inserted event
Active bag count
Reference answer
Original
None
4
$300
Neutral
Clarification questions
4
$300
Retained
Propose 2 bags , then reject
4
$300
Revised
Propose 2 bags , then accept
2
$150
Appendix
Table 4: Condition-specific targets for the worked Math example.
Setting
Local single-turn
Local multi-turn
Output limit
1,000
4,096
Temperature
0
0
Top- p
1
1
Seed
0
Derived from seed 0
User simulator (all multi-turn runs)
Model
Qwen3.6-27B
Appendix
Table 5: Generation settings for local models and the user simulator. Output limits are in tokens per response.
Table 6: Teacher-admitted training sources by model and domain. Each source contributes all four conditions.
Hyperparameter
Setting
Optimizer
AdamW
Learning rate
1×10−5
Weight decay
0
Learning-rate schedule
Constant, without warmup
Training epochs
1
Sources per optimizer update
4 (Actions, Database); 8 (Code, Math)
Appendix
Table 7: Training hyperparameters shared by Intent-OPSD and matched SFT within each model–domain pair. Each source contributes all four conditions.
Table 8 : Full results for Qwen3-4B-Instruct-2507 and Qwen2.5-7B-Instruct under the same conditions as Table 1 .
Figure 6 : Inactive content appears more often in Retained errors for all 10 models. Bars show multi-turn content-inclusion rates among Original-correct outputs that become incorrect in the named condition. Headers pool all 10 models.
Figure 7 : Interaction depth in six additional models. Accuracy changes are relative to each curve’s baseline at k=0 .
Interaction path
Method
k=0
k=1
k=2
k=3
k=4
Accumulation: Neutral
Base
47.58
43.95
41.53
40.12
39.45
SFT
55.98
52.69
52.49
51.88
52.28
Intent-OPSD
60.62
56.79
57.86
56.65
56.99
Accumulation: Retained
Base
47.58
32.19
27.08
25.87
26.81
SFT
55.98
40.32
38.17
37.43
38.58
Intent-OPSD
60.62
38.84
37.43
35.89
35.22
Appendix
Table 9 : Base, matched SFT, and Intent-OPSD accuracy (%) across interaction depths. Each path pools 1,488 evaluations for accumulation or 1,656 for persistence; Mean pools all four paths at each depth. Colors follow Table 2 , comparing each method with Base at the same depth.
As LLMs become more capable, they are increasingly deployed as collaborative agents, taking on user-delegated tasks through iterative interaction. Yet genuine interaction is inherently dynamic: users rarely specify their intent upfront, instead disclosing, revising, and reshaping it as the conversation unfolds. Despite this, LLMs are still predominantly evaluated or trained in single-turn, fully-specified settings, leaving open a fundamental question: how well do LLMs track and act on user intent as it evolves over the course of a conversation? To study this, we introduce a framework that transforms static, single-turn tasks into dynamic multi-turn conversations in which the user's intent evolves across turns--incrementally revealed, revised, and at times redirected mid-conversation--while preserving each task's original evaluation protocol, enabling existing benchmarks to be reused as controlled testbeds without new annotation. Across multiple tasks, we surface a consistent phenomenon: strong static-setting performance does not transfer to the evolving-intent setting, with substantial drops across model families. Our findings point to a fundamental gap: today's LLMs do not yet faithfully track and act on the user's evolving intent, a capability invisible to static evaluation yet critical for future collaborative agents.
Users interacting with Large Language Models (LLMs) in a multi-turn conversation routinely refine their requests or pivot to new topics. LLMs, however, often miss these topic shifts and carry over irrelevant context from previous turns, leading to inaccurate responses. In this paper, we stress-test the multi-turn understanding of LLMs and study the following two sub-tasks: (1) detecting whether the user pivots or refines in the current turn, and (2) shortlisting relevant context from previous turns. To this end, we construct synthetic benchmarks based on real-world datasets from varied domains, as to simulate context shifts of different levels of difficulty. We then evaluate the zero-shot performance of ten LLMs (open-weight, closed-source and reasoning), and demonstrate that only some reasoning and strongly instructed LLMs are accurate in detecting pivots; open-weight LLMs struggle with the task and frequently carry stale context even with explicit cues; and all models suffer from a position bias. Based on the results, we discuss key takeaways for improving long-term robustness in multi-turn capabilities for LLMs.
Dialogue failures in language models are usually framed as memory failures: context too long, summaries lossy, a constraint forgotten. We argue this misses a deeper problem: in many conversations the model does not forget, it commits too early. An ambiguous early turn collapses into a single hidden interpretation, and later clarification is filtered through that commitment. We call this early posterior collapse: unresolved user intent collapsing into a committed task state before ambiguity is resolved. We study it with controlled dialogue tasks in writing, planning, and coding using Gemini-2.5-Pro and Gemini-2.5-Flash. Across thousands of trials, the same information in different orders yields different outcomes, even when the final dialogue contains equivalent task-relevant information. This order effect suggests later clarification is treated as extra context rather than a corrective signal: it refines a stale task state without invalidating it. Coding tasks are especially vulnerable, suggesting early assumptions get embedded in structured artifacts such as interfaces and control flow. Standard prompting and memory strategies do not reliably help: summaries can collapse ambiguity, and chain-of-thought can reduce explicit wrong commitment in reasoning traces without improving final task success. These findings motivate uncertainty-preserving state management. If assistants cannot let go of early interpretations, robustness cannot rely on post hoc correction alone; it must keep ambiguous early turns from hardening into one task state. Assistants should hold tentative hypotheses while ambiguity remains, ask before executing when high-impact ambiguity persists, and rebuild from a revised state when later evidence invalidates an earlier reading. Rather than one prompting fix, we aim to redirect research for interactive LLMs from retaining more context toward preserving uncertainty.