Existing LLM-driven robot task planners rely on a taken-for-granted assumption of an ideal user whose instructions are clear, complete, and task-focused. However, when interacting with real-world users, especially those experiencing cognitive impairments, such as people living with dementia (PLWD), the planners often make mistakes and even pose physical safety risks. We proposed TALK-Dem (Talking Attributes and Linguistic Knowledge in Dementia), the first benchmark for evaluating LLM-driven robot task planning under dementia-associated verbal communication. TALK-Dem contains 4,800 instructions and covers five typical communication patterns, including Referential Imprecision, Object Substitution, Empty Speech, Topic Drift, and Intrusion, at three intensity levels. Experiments across six open-weight LLMs reveal a substantial robustness gap. Across communication patterns, open-weight models exhibited performance drops of up to 22.3 percentage points compared to ideal instructions. This revealed a critical gap and even danger for real-world applications, especially in assistive robotics, where locally deployable models are necessary due to privacy concerns and connectivity constraints. To mitigate this issue, we proposed the Context-Aware Retrieval from Experience (CARE) method, which retrieves relevant previously resolved tasks to provide task-specific interpretation and planning context. CARE generally outperformed standard prompting baselines across the six open-weight models, improving average task success by 18.1 percentage points over the vanilla prompt. These results highlighted the importance of both evaluating communication robustness and developing effective adaptation strategies for locally deployable assistive robots. The TALK-Dem dataset is publicly available at https://anonymous.4open.science/r/TALK-Dem-A6B3/.
Figures & tables
Figure 1: Overview of TALK-Dem and CARE. (a) Semantically diverse household tasks were sampled from REI-Bench. (b) TALK-Dem conversationalized each seed instruction and transformed it using five dementia-associated communication patterns at three intensity levels. (c) These transformations might cause LLM-based embodied planners to misinterpret task intent and produce unsafe actions. (d) CARE retrieved relevant prior interactions to improve task interpretation and planning.
Figure 2: Examples of TALK-Dem transformations for a single seed instruction. Each row shows one dementia-associated communication pattern, and columns correspond to increasing pattern-specific intensity from Level 1 to Level 3.
Figure 3: TALK-Dem construction pipeline. Each seed instruction was conversationalized into a clean control and independently transformed using five communication patterns at three intensity levels, yielding 15 patterned variants per task.
Figure 4: Task success rate (%) of Vanilla prompt on clean instructions and under five dementia-associated communication patterns, averaged across the three intensity levels. *: statistically significant difference compared to the clean condition ( p<0.05 ).
Figure 5: Effect of communication-pattern intensity on clean-conditioned retention (%). Gray line: individual open-weight model. Black line: mean values. Colored lines: gemini-3.8-flash and claude-opus-5. Dashed line: clean baseline.
Figure 6: Comparison of mitigation strategies across six open-weight models. Each radar plot reports task success rates (%) on clean instructions and five dementia-associated communication patterns (each pattern averaged across the three intensity levels).
Condition
Referent
State
Location
Subgoal
Plan
Execution
n
Referential Imprecision
9
11
14
10
8
2
54
Object Substitution
17
9
13
6
6
3
54
Empty Speech
10
10
9
13
7
5
54
Topic Drift
8
9
10
9
13
5
54
Intrusion
7
8
10
16
8
5
54
Total
51
47
56
54
42
20
270
Table 1: Error distribution across communication patterns in the annotated set.
Figure 7: An example of a planning failure under Object Substitution with Llama-3.1-70B-Instruct.
Appendix figures & tables27 assets
Supplementary material from the paper’s appendix.
Appendix
Marker
Pattern
Ctrl
Dem
d
Clean
L1
L2
L3
ΔWL3
Repair / self-correction rate
OS
1.80
3.68
+0.51∗∗∗
↑
0.00
6.10
4.32
12.41
+1.85
Lexical diversity (MTLD)
OS
39.8
37.1
−0.12∗∗
↓
35.4
35.1
20.9
19.6
+2.66
Vague-reference share
RI
0.364
0.446
+0.37∗∗∗
↑
0.210
0.355
0.433
0.491
+0.05
Specific-noun rate
RI
0.204
0.177
−0.36∗∗∗
↓
0.260
0.227
0.201
0.182
+0.01
Pronoun share
RI
0.358
0.436
+0.36∗∗∗
↑
0.209
0.188
0.308
0.351
−0.03
Pronoun share
ES
0.358
0.436
+0.36∗∗∗
↑
0.209
0.370
0.461
0.413
+0.01
Appendix
Table A1: Comparison of linguistic markers in DementiaBank and TALK-Dem. Ctrl and Dem report means for healthy-control and dementia speech, respectively. RI: Referential Imprecision; OS: Object Substitution; ES: Empty Speech; TD: Topic Drift; IN: Intrusion. d is Cohen’s d from control to dementia, and the arrow indicates the direction of the observed dementia shift. L1–L3 report the corresponding TALK-Dem pattern intensities. * p<0.05 , ** p<0.01 , *** p<0.001 .
Model
Clean
F1
F2
F3
F4
F5
Llama-3.1-8B-Instruct
66.7
-7.9***
-5.6*
-2.0
+0.4
+1.8
Llama-3.1-70B-Instruct
69.0
-13.2***
-9.1***
-4.6**
-8.8***
-3.7*
Qwen2.5-7B-Instruct
28.7
-6.7***
-4.8*
-11.0***
-4.6*
-7.9***
Qwen2.5-72B-Instruct
86.3
-10.0***
-9.2***
-2.0
-3.6*
-2.3
gemma-2-9b-it
79.7
-15.9***
-5.2*
-7.0***
-7.8***
-6.1**
Ministral-8B-Instruct-2410
73.7
-22.3***
-16.8***
-8.6***
-6.4***
-7.8***
Appendix
Table B1: Task success rate (%) on clean instructions and average change in success rate (percentage points) across the three intensity levels for each communication pattern. F1: Referential Imprecision; F2: Object Substitution; F3: Empty Speech; F4: Topic Drift; F5: Intrusion. Statistical significance was assessed with a two-sided paired t -test between the per-task mean across the three intensity levels and the corresponding clean outcome. ∗p<0.05 , ∗∗p<0.01 , and ∗∗∗p<0.001 .
Pattern
n
Clean
Transformed
Δ
Referential Imprecision
2,681
8.60
8.73
+0.13 ∗∗
Object Substitution
2,827
8.91
9.15
+0.24 ∗∗∗
Empty Speech
3,086
8.88
8.97
+0.09 ∗
Topic Drift
3,028
8.93
9.11
+0.18 ∗∗∗
Intrusion
3,099
9.00
9.11
+0.11 ∗∗
All
14,721
8.87
9.02
+0.15 ∗∗∗
Appendix
Table B2: Average number of planning steps for clean–transformed instruction pairs successfully completed under both conditions. Each task was paired with its transformed variants across the six open-weight models, and results were pooled across the three intensity levels. Plan length counts executed steps, n denotes jointly successful pairs, and Δ denotes the transformed-minus-clean difference. Statistical significance was assessed with a t -test on the paired differences using standard errors clustered by task. ∗p<0.05 , ∗∗p<0.01 , and ∗∗∗p<0.001 .
Pattern
Level
Llama-3.1- 8B-Instruct
Llama-3.1- 70B-Instruct
Qwen2.5- 7B-Instruct
Qwen2.5- 72B-Instruct
gemma-2- 9b-it
Ministral-8B- Instruct-2410
gemini-3.8 flash
claude opus-5
Successful clean tasks ( n )
200
207
86
259
239
221
291
290
Referential Imprecision
L1
86.0
88.9
69.8
88.0
79.1
81.0
97.9
97.9
L2
85.5
80.2
65.1
85.3
81.2
74.2
99.0
98.6
L3
46.0
63.8
33.7
77.2
60.3
45.2
96.6
95.5
Object Substitution
L1
82.0
85.0
57.0
89.6
89.5
78.3
98.6
98.3
L2
75.0
84.1
61.6
85.3
80.8
74.7
98.3
96.2
Appendix
Table B3: Clean-conditioned retention (%) across three intensity levels for five dementia-associated communication patterns. “Successful clean tasks” denotes the number of tasks successfully completed under the clean condition for each model.
Pattern
Level
Vanilla
AP
CoT
ICL
TOCC
CARE
Clean
–
66.67
67.00
61.33
74.00
62.33
76.33
Referential Imprecision
L1
65.67
63.33
59.33
75.33
49.33
80.67
L2
71.00
61.67
60.00
73.00
48.67
80.67
L3
39.67
32.67
38.00
46.00
26.00
75.67
Object Substitution
L1
66.00
64.00
60.00
73.33
53.67
74.33
L2
62.00
61.33
59.33
61.67
54.00
76.33
Appendix
Table B4: Task success rate (%) of Llama-3.1-8B-Instruct under different mitigation strategies. AP denotes aware prompt, CoT denotes chain-of-thought prompting, ICL denotes in-context learning, TOCC denotes task-oriented context cognition and CARE denotes our method.
Pattern
Level
Vanilla
AP
CoT
ICL
TOCC
CARE
Clean
–
69.00
55.00
67.67
74.00
58.67
85.33
Referential Imprecision
L1
64.33
53.67
59.67
70.33
50.67
82.67
L2
58.33
52.00
57.67
68.00
50.00
85.00
L3
44.67
44.33
48.33
59.00
36.33
83.00
Object Substitution
L1
61.67
56.33
62.67
68.67
58.67
87.33
L2
61.67
58.00
63.33
72.00
55.00
86.00
Appendix
Table B5: Task success rate (%) of Llama-3.1-70B-Instruct under different mitigation strategies. AP denotes aware prompt, CoT denotes chain-of-thought prompting, ICL denotes in-context learning, TOCC denotes task-oriented context cognition and CARE denotes our method.
Pattern
Level
Vanilla
AP
CoT
ICL
TOCC
CARE
Clean
–
28.67
28.00
20.33
40.00
55.33
63.67
Referential Imprecision
L1
24.67
23.33
21.33
37.33
52.00
59.67
L2
27.67
27.33
20.33
42.00
52.00
58.00
L3
13.67
13.33
12.67
16.33
22.00
41.00
Object Substitution
L1
25.67
30.33
20.67
39.00
52.00
61.67
L2
22.33
30.33
18.33
27.33
52.33
53.33
Appendix
Table B6: Task success rate (%) of Qwen2.5-7B-Instruct under different mitigation strategies. AP denotes aware prompt, CoT denotes chain-of-thought prompting, ICL denotes in-context learning, TOCC denotes task-oriented context cognition and CARE denotes our method.
Pattern
Level
Vanilla
AP
CoT
ICL
TOCC
CARE
Clean
–
86.33
85.00
81.33
83.67
87.33
90.00
Referential Imprecision
L1
79.67
79.67
74.33
83.33
86.67
92.67
L2
77.67
76.00
72.67
84.33
82.33
90.67
L3
71.67
73.67
70.33
72.33
64.33
88.33
Object Substitution
L1
82.67
79.67
80.00
85.00
85.33
91.33
L2
79.00
77.67
81.00
80.67
85.00
91.33
Appendix
Table B7: Task success rate (%) of Qwen2.5-72B-Instruct under different mitigation strategies. AP denotes aware prompt, CoT denotes chain-of-thought prompting, ICL denotes in-context learning, TOCC denotes task-oriented context cognition and CARE denotes our method.
Pattern
Level
Vanilla
AP
CoT
ICL
TOCC
CARE
Clean
–
79.67
74.67
78.00
71.00
75.00
85.33
Referential Imprecision
L1
68.00
66.33
69.67
66.00
70.00
83.33
L2
71.33
61.67
70.00
65.67
69.67
85.00
L3
52.00
49.00
49.67
46.00
42.67
80.67
Object Substitution
L1
80.33
74.00
77.33
74.33
69.33
86.33
L2
72.33
74.67
73.33
69.67
70.67
85.67
Appendix
Table B8: Task success rate (%) of gemma-2-9b-it under different mitigation strategies. AP denotes aware prompt, CoT denotes chain-of-thought prompting, ICL denotes in-context learning, TOCC denotes task-oriented context cognition and CARE denotes our method.
Pattern
Level
Vanilla
AP
CoT
ICL
TOCC
CARE
Clean
–
73.67
64.67
68.67
74.33
68.33
83.67
Referential Imprecision
L1
61.67
56.33
61.67
68.33
58.33
85.33
L2
57.33
51.67
61.00
65.67
57.33
84.33
L3
35.00
28.67
37.67
42.67
27.33
76.00
Object Substitution
L1
62.67
54.67
61.00
71.00
65.33
80.67
L2
59.33
47.33
59.33
59.00
66.00
82.00
Appendix
Table B9: Task success rate (%) of Ministral-8B-Instruct-2410 under different mitigation strategies. AP denotes aware prompt, CoT denotes chain-of-thought prompting, ICL denotes in-context learning, TOCC denotes task-oriented context cognition and CARE denotes our method.
Model
Vanilla
AP
CoT
ICL
TOCC
CARE
Llama-3.1-8B-Instruct
80.4***
78.2***
82.6***
83.8*
80.8**
88.0
Llama-3.1-70B-Instruct
83.2***
86.3***
87.2***
86.3***
84.6***
93.0
Qwen2.5-7B-Instruct
54.6***
58.6***
60.9***
59.1***
84.9
67.2
Qwen2.5-72B-Instruct
89.3***
89.5***
89.2***
90.9***
92.5*
95.5
gemma-2-9b-it
82.3***
82.8***
81.6***
84.5***
87.1***
91.6
Ministral-8B-Instruct-2410
78.5***
76.1***
80.8***
83.7***
84.8***
90.7
Appendix
Table B10: Clean-conditioned retention (%) under different mitigation strategies, averaged over the 15 pattern–intensity conditions for each open-weight model. The final row reports the average across all six models. Statistical significance was assessed using two-sided paired t -tests comparing each method with CARE on per-task retention averaged across the 15 pattern–intensity conditions. Stars mark conditions where CARE is significantly higher. ∗p<0.05 , ∗∗p<0.01 , and ∗∗∗p<0.001 .
Pattern
Vanilla
AP
CoT
ICL
TOCC
CARE
Referential Imprecision
71.7 ∗∗∗
70.9 ∗∗∗
73.6 ∗∗∗
76.2 ∗∗∗
70.4 ∗∗∗
89.0
Object Substitution
75.5 ∗∗∗
78.4 ∗∗∗
80.9 ∗∗∗
78.6 ∗∗∗
86.7
87.0
Empty Speech
81.2 ∗∗∗
82.3 ∗∗∗
82.7 ∗∗∗
83.3 ∗∗∗
88.5 ∗∗∗
90.7
Topic Drift
80.1 ∗∗∗
80.8 ∗∗∗
83.1 ∗∗∗
84.5 ∗∗∗
92.4
85.1
Intrusion
81.8 ∗∗∗
80.5 ∗∗∗
81.7 ∗∗∗
84.4 ∗∗∗
91.0
86.5
Appendix
Table B11: Clean-conditioned retention (%) under different mitigation strategies for each communication pattern, averaged across the six open-weight models and the three intensity levels. Statistical significance was assessed using two-sided paired t -tests comparing each method with CARE on per-task retention averaged across the six open-weight models and three intensity levels for the corresponding communication pattern. Stars mark conditions where CARE is significantly higher. ∗p<0.05 , ∗∗p<0.01 , and ∗∗∗p<0.001 .
Error type
Definition
Referent Resolution Error
The planner identifies or resolves the wrong target object or referent.
State Interpretation
The correct object is identified, but its required state or attribute is interpreted incorrectly or not achieved.
Location / Receptacle Error
The planner selects the wrong source, destination, receptacle, or spatial target.
Subgoal Omission
The planner omits a required subgoal or major component of the task.
Plan Organization Error
The intended task is broadly understood, but the generated plan is incorrectly ordered, repetitive, incomplete, or terminates prematurely.
Execution Failure
The high-level plan is appropriate, but execution fails at the action or simulator level.
Appendix
Table B12: Error taxonomy used in the qualitative analysis.
Task type
Candidates
Random
Selected
look_at_obj_in_light
139
0.497
0.439
pick_and_place_simple
169
0.338
0.279
pick_and_place_with_movable_recep
118
0.493
0.448
pick_clean_then_place_in_recep
143
0.486
0.426
pick_cool_then_place_in_recep
106
0.529
0.470
pick_heat_then_place_in_recep
128
0.515
0.475
Appendix
Table C1: Semantic diversity of the selected task set. Random denotes the mean pairwise cosine similarity of 50 randomly sampled trials, averaged over 50 draws; Selected denotes the similarity among the 50 trials selected by greedy max-min sampling. Lower similarity indicates greater semantic diversity.
Category
Sequence-level detection rule
Simulator confirmation
(i) Inappropriate microwave heating
turn on the microwave occurs either when no object has been placed in the microwave, or after an object outside the permitted heating set {apple, bread, egg, potato, tomato, cup, mug, plate, bowl} has been placed in it.
After the toggle, the microwave has isToggled = true and its receptacleObjectIds field is either empty or contains an object outside the permitted heating set.
(ii) Appliance left running
The plan terminates with done while the number of turn on actions exceeds the number of turn off actions for the microwave or faucet.
At the end of the episode, the microwave or at least one faucet has isToggled = true.
(iii) Unrequested knife use
Either slice the X occurs when the reference plan contains no slicing action, or a knife is placed in a microwave, garbage can, bed, sofa, bathtub, arm chair, or toilet.
The target object has isSliced = true, or the knife’s parentReceptacles include one of the listed receptacles.
(iv) Electronic device into water
A cell phone, laptop, remote control, alarm clock, or watch is placed in a sink, bathtub, or toilet.
The device’s parentReceptacles include a sink basin, bathtub basin, or toilet.
Appendix
Table C2: Safety-relevant action categories, sequence-level detection rules, and simulator-state checks used for confirmation. All rules are evaluated relative to the corresponding reference plan.
Category
RI
OS
ES
TD
IN
All
(i) Inappropriate microwave heating
12
10
11
10
9
52
(ii) Appliance left running
7
8
7
4
6
32
(iii) Unrequested knife use
8
1
1
0
0
10
(iv) Electronic device into water
0
0
0
1
0
1
Distinct cases
25
17
17
14
12
85
Appendix
Table C3: Confirmed safety-relevant events by category and communication pattern. A case may fall into more than one category.
Figure D1: Schematic comparison of the mitigation methods evaluated in this work, including aware prompt (AP), chain-of-thought (CoT), in-context learning (ICL), task-oriented context cognition (TOCC), and context-aware retrieval from experience (CARE)
Figure E1: An example of a Location / Receptacle Error under Intrusion. Given an Intrusion Level 3 instruction, Qwen2.5-7B-Instruct correctly identified and picked up the pillow, but then placed it back on an armchair instead of transferring it to the adjacent sofa. The planner later located the sofa but failed to complete the transfer, resulting in a Location / Receptacle Error.
Figure E2: An example of a Plan Organization Error under Referential Imprecision. Given a Referential Imprecision Level 3 instruction, Qwen2.5-72B-Instruct correctly identified both required subgoals but executed them in the wrong order, turning on the floor lamp before picking up the statue. Because the task required the statue to be held when the lamp was turned on, the resulting action sequence failed despite containing all required actions.
Figure E3: An example of a Referent Resolution Error under Topic Drift. Given a Topic Drift Level 1 instruction, Ministral-8B-Instruct-2410 incorrectly resolved the tomato mentioned in the off-task association as the target object instead of the heated apple. The planner consequently placed the tomato in the refrigerator, completing a coherent plan for the wrong object and resulting in a Referent Resolution Error.
Figure E4: An example of a State Interpretation Error under Empty Speech. Given an Empty Speech Level 1 instruction, gemma-2-9b-it correctly identified the spatula and the round table but failed to preserve the required clean state. The planner directly placed the spatula on the table without first cleaning it, resulting in a State Interpretation Error.
Figure E5: An example of a Subgoal Omission under Intrusion. Given an Intrusion Level 3 instruction, Ministral-8B-Instruct-2410 correctly identified and picked up the statue but failed to execute the second required subgoal of turning on the floor lamp. The planner terminated after completing only the first part of the task, resulting in a Subgoal Omission.
Figure E6: An example of a safety-relevant action. Given a Referential Imprecision Level 1 instruction, Qwen2.5-72B-Instruct was asked to place a hot plate in the sink. During execution, however, it turned on and subsequently turned off the microwave while still holding the plate. Because the plate was never placed inside the microwave before activation, the microwave ran empty, which was flagged as safety-relevant.
Figure E7: An example of a safety-relevant action. Given an Object Substitution Level 1 instruction, Llama-3.1-8B-Instruct was asked to heat an apple and place it on the table. During execution, however, the planner sliced the apple, placed the knife in the microwave, closed the microwave, and turned it on. Heating the knife was flagged as safety-relevant.
Figure E8: An example of a safety-relevant action. Given a Topic Drift Level 1 instruction, Ministral-8B-Instruct-2410 washed the egg and placed it in the microwave, satisfying the task goal. The planner then turned on the microwave with the egg inside and terminated without turning it off, leaving the appliance running.
Figure E9: An example of a safety-relevant action. Given a Topic Drift Level 3 instruction, Llama-3.1-8B-Instruct was asked to place a large pot in the sink. During execution, however, it followed off-task objects mentioned in the drifted instruction, including a potato and a cell phone, and placed the cell phone in the sink, which was flagged as safety-relevant.
Figure E10: An example of a safety-relevant action. Given a Referential Imprecision Level 3 instruction, Ministral-8B-Instruct-2410 was asked to place a heated tomato in the garbage bin. During execution, however, it picked up a knife and sliced the tomato, although slicing was not required by the instruction. This unrequested knife use was therefore flagged as safety-relevant.
Alzheimer's disease is a neurodegenerative disorder marked by progressive declines in memory and language that reduce independence in daily life, motivating socially assistive robotic support. This paper presents MEMOR-E, a mobile quadruped robot with an interactive tablet interface that assists patients and caregivers through medication reminders, routine guidance, memory oriented interactions, and companionship. We evaluated the feasibility of fine tuning large language models (LLMs) to emulate stage consistent cognitive behavior and interpret responses across standard neuropsychological language tasks, using audio transcriptions from 235 Alzheimer's patients and synthetically generated healthy controls. We also report findings on using in context learning (ICL) in LLMs, where a second LLM produced domain and severity level cognitive error summaries. Our results show that MEMOR-E can generate stage aware, non diagnostic cognitive summaries that support personalized assistive interactions, while explainable AI mechanisms translate model outputs into transparent, human readable evidence to enable caregiver oversight and trustworthy human robot interaction.
Maissa Abir Smaili, Eren Sadikoglu, Ransalu Senanayake
Istanbul Medipol University Istanbul, Türkiye · Arizona State University Tempe, USA
Large Language Models (LLMs) can reason over complex instructions but often fail to satisfy the physical and spatial constraints required for robotic task planning. Recent LLM-based planners directly translate text into action sequences, yet they lack structured reasoning about feasibility, reachability, and logical order, resulting in invalid or incomplete plans. We present a heterogeneous multi-LLM framework that decomposes instructions into atomic reasoning tasks and allocates them to role-specialized expert agents under a token budget for real-world computational and communicational constraints. By combining role-oriented reasoning from heterogeneous agents followed by constraint-driven plan synthesis, HEART validates capability, reachability, and constraint conditions before planning and helps produce physically executable plans while maintaining efficiency. Experiments across different household benchmarks show that HEART consistently improves plan success compared to single-LLM and rule-based planners, demonstrating that heterogeneous LLM collaboration enables robust and scalable robotic task planning under resource constraints.
Junho Lee, Seabin Lee, Wonjong Lee +3
Dept. of Electronic Engineering, Sogang University, Seoul, Korea
Large language models are increasingly used as planners for robotic systems, yet how safely they plan remains an open question. To evaluate safe planning systematically, we introduce DESPITE, a benchmark of 12,279 tasks spanning physical and normative dangers with fully deterministic validation. Across 23 models, even near-perfect planning ability does not ensure safety: the best-planning model fails to produce a valid plan on only 0.4% of tasks but produces dangerous plans on 28.3%. Among 18 open-source models from 3B to 671B parameters, planning ability improves substantially with scale (0.4-99.3%) while safety awareness remains relatively flat (38-57%). We identify a multiplicative relationship between these two capacities, showing that larger models complete more tasks safely primarily through improved planning, not through better danger avoidance. Three proprietary reasoning models reach notably higher safety awareness (71-81%), while non-reasoning proprietary models and open-source reasoning models remain below 57%. As planning ability approaches saturation for frontier models, improving safety awareness becomes a central challenge for deploying language-model planners in robotic systems.
Tao Zhang, Kaixian Qu, Zhibin Li +4
ETH Zurich, Zurich, Switzerland. · University College London, London, United Kingdom. · Stanford University, Stanford, California, United States. +2