When a large language model handles a multi-turn task and a user proposes a change but ultimately rejects it, the model should continue as if nothing changed. We find a surprising failure: merely mentioning a rejected change can derail task execution, even when the user's final intent remains unchanged. To systematically study language model behavior under evolving user intent, we introduce Intent-Eval, a controlled benchmark spanning tool actions, code, databases, and mathematics. Across diverse tasks, models are vulnerable to both rejected proposals and superseded requirements, consistent with mentioned-as-in-effect confusion: conversational content is treated as active requirements even after it has been rejected or replaced. Accuracy degradation can deepen or persist as interaction continues, highlighting the need to distinguish what has been mentioned from what remains in effect. Building on this insight, we propose Intent-OPSD, a decision-conditioned on-policy self-distillation framework with Teacher and Student initialized from the same model. The frozen Teacher provides active-intent supervision from the complete task matching the user's decision, training the Student on the full dialogue to follow active requirements reflecting user intent.
Figures & tables
Figure 1 : Illustration of active intent. (a) Single-turn controls and targets. (b) The Original task distributed across turns; U3–U4 mark the insertion point for the other conditions. (c) Neutral clarifies; Retained rejects and Revised accepts the same proposal. Only Revised changes the target.
Figure 2 : Overview and statistics of Intent-Eval . The benchmark comprises 414 source tasks across four domains. Each source task has four multi-turn conditions and four single-turn controls, yielding 3,312 evaluation instances per model.
Figure 3 : Intent-OPSD : the frozen Teacher sees the active task; the Student sees the dialogue.
Model
Domain
Single
Multi
Original
Direct Rev.
Revised
Retained
Original
Neutral
Revised
Retained
3.6-27B
Math
98.06
97.09
87.38
93.20
73.79
71.84
66.99
61.17
Code
95.00
93.00
90.00
88.00
81.00
76.00
80.00
72.00
DB
92.52
91.59
91.59
87.85
42.99
43.93
42.06
38.32
Actions
96.15
95.19
96.15
85.58
49.04
44.23
47.12
48.08
3-8B
Math
92.23
90.29
75.73
67.96
64.08
64.08
54.37
37.86
Table 1 : Model-by-domain accuracy (%) across single-turn controls and the four multi-turn conditions. In Single, Direct Rev. states the revised task directly, while Revised includes the proposal and acceptance. Overall aggregates all 414 task sources across the eight models shown here; two additional models appear in Appendix Table 8 . Cell backgrounds indicate changes relative to Single Original: for declines, for gains, and for no change.
Figure 4 : Inactive content in errors. Bars show inclusion rates among new Retained (left) and Revised (right) errors, pooled over ten models; headers aggregate all domains.
Figure 5 : Accuracy change with depth relative to each curve’s k=0 baseline. Y-ranges are shared within domains. Results for six additional models are in Appendix E.3 (Figure 7 ).
Model
Domain
Base
SFT
Intent-OPSD
Orig.
Neut.
Revised
Retained
Orig.
Neut.
Revised
Retained
Orig.
Neut.
Revised
Retained
3-8B
Math
64.08
64.08
54.37
37.86
61.17
61.17
59.22
28.16
66.02
67.96
64.08
23.30
Code
41.00
43.00
43.00
26.00
35.00
36.00
37.00
24.00
39.00
42.00
45.00
28.00
DB
38.32
26.17
32.71
17.76
43.93
38.32
43.93
20.56
45.79
44.86
42.99
27.10
Actions
51.92
50.00
46.15
39.42
68.27
71.15
31.73
57.69
66.35
60.58
66.35
55.77
4-26B- A4B
Math
69.90
65.05
66.02
58.25
70.87
71.84
66.02
62.14
81.55
77.67
72.82
60.19
Table 2 : Accuracy (%) and Intent-OPSD depth gains (pp). For accuracy, conditions and model–domain pairs have equal weight. Colors show changes from Base ( lower, higher, unchanged); shading reflects magnitude. Gains use overall accuracy across the four interaction paths; k counts added event blocks. Appendix E.4 provides detailed results by path and depth.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Table 3: Domain-specific revisions and reference updates.
Condition
Inserted event
Active bag count
Reference answer
Original
None
4
$300
Neutral
Clarification questions
4
$300
Retained
Propose 2 bags , then reject
4
$300
Revised
Propose 2 bags , then accept
2
$150
Appendix
Table 4: Condition-specific targets for the worked Math example.
Setting
Local single-turn
Local multi-turn
Output limit
1,000
4,096
Temperature
0
0
Top- p
1
1
Seed
0
Derived from seed 0
User simulator (all multi-turn runs)
Model
Qwen3.6-27B
Appendix
Table 5: Generation settings for local models and the user simulator. Output limits are in tokens per response.
Table 6: Teacher-admitted training sources by model and domain. Each source contributes all four conditions.
Hyperparameter
Setting
Optimizer
AdamW
Learning rate
1×10−5
Weight decay
0
Learning-rate schedule
Constant, without warmup
Training epochs
1
Sources per optimizer update
4 (Actions, Database); 8 (Code, Math)
Appendix
Table 7: Training hyperparameters shared by Intent-OPSD and matched SFT within each model–domain pair. Each source contributes all four conditions.
Table 8 : Full results for Qwen3-4B-Instruct-2507 and Qwen2.5-7B-Instruct under the same conditions as Table 1 .
Figure 6 : Inactive content appears more often in Retained errors for all 10 models. Bars show multi-turn content-inclusion rates among Original-correct outputs that become incorrect in the named condition. Headers pool all 10 models.
Figure 7 : Interaction depth in six additional models. Accuracy changes are relative to each curve’s baseline at k=0 .
Interaction path
Method
k=0
k=1
k=2
k=3
k=4
Accumulation: Neutral
Base
47.58
43.95
41.53
40.12
39.45
SFT
55.98
52.69
52.49
51.88
52.28
Intent-OPSD
60.62
56.79
57.86
56.65
56.99
Accumulation: Retained
Base
47.58
32.19
27.08
25.87
26.81
SFT
55.98
40.32
38.17
37.43
38.58
Intent-OPSD
60.62
38.84
37.43
35.89
35.22
Appendix
Table 9 : Base, matched SFT, and Intent-OPSD accuracy (%) across interaction depths. Each path pools 1,488 evaluations for accumulation or 1,656 for persistence; Mean pools all four paths at each depth. Colors follow Table 2 , comparing each method with Base at the same depth.