WSM-Aware HRI: An IoT-Enhanced Framework for Early Detection and Norm-Guided Repair of Failures with LLM Guidance
Organizations: The Chinese University of Hong Kong, Shenzhen, China · X SQUARE Robot, China · Shenzhen Institute of Artificial Intelligence and Robotics for Society, China
Abstract
Human-robot interaction (HRI) failures remain a major barrier to deploying robots in real-world environments. Prior work often treats failures as isolated technical faults or focuses on post-hoc recovery behaviors. In practice, many breakdowns arise because humans and robots operate under inconsistent assumptions about the current world state. We propose WSM-Aware HRI, an IoT-enhanced modular framework that unifies diverse HRI breakdowns as World-State Mismatches (WSMs) between a human's instruction-implied assumptions and a robot's grounded world model built from multimodal perception and digital augmentation. A Large Language Model (LLM) is used to make implicit assumptions explicit, map them to a small set of mismatch types, and specify the evidence needed for verification against the robot's world state. WSM-Aware HRI shifts failure handling from execution-time recovery to proactive mismatch detection during intention formation, enabling interventions guided by safety, norm compliance, and multi-user coordination with transparent explanations. We evaluate mismatch identification in ten everyday cases spanning both visual and latent-state mismatches. The system can accurately produce the expected output results, and ablations show that reliable identification depends on appropriate grounding representations and verification-oriented refinement. These results indicate that treating interaction breakdowns as explicit world-state mismatches enables earlier detection of impending failures and offers a principled mechanism for integrating external evidence and social constraints into human-robot interaction.
Figures & tables
| WSM Type | Definition and Representative Example |
|---|---|
| Instruction Logic Mismatch | Definition: The instruction is incompatible with task-level logic or hazard/causal constraints under the current context. Example: Ask the robot to extinguish a fire with water while the robot detects an active electrical source. |
| Perceptual / Affordance Mismatch | Definition: The user’s instruction assumes physical feasibility or safe affordances that are contradicted by grounded sensing or verification. Example: Request the robot to pick up a cup with cold water while the robot detects hot liquid inside. |
| Temporal / Sequence Mismatch | Definition: The instruction conflicts with temporal ordering or state-machine prerequisites. Example: Ask the robot to place an object before it has been grasped. |
| Social / Ethical (Norm) Mismatch | Definition: The instruction violates applicable social norms, policies, or role-based permissions, regardless of physical feasibility. Example: Request the robot to enter a restricted or private area. |
| Multi-user Conflict | Definition: The instruction is incompatible with constraints or objectives posed by multiple stakeholders, requiring coordination or negotiation. Example: The teacher requests the robot to maintain classroom order while students request entertainment. |
| Prior failure mechanism | Instruction-implied constraint | Evidence needed for verification | Repair implication | Resulting WSM category |
|---|---|---|---|---|
| Perception/action mismatch and physical execution failure ( Honig and Oron-Gilad, 2018 ; Cameron et al., 2024 ) | The instruction assumes that the target object or action is physically feasible, compatible, and safe under the current context. | Grounded perception, object state, affordance checks, force/temperature/weight sensing, or other physical verification. | Verify the relevant property, adapt the plan, avoid unsafe manipulation, or suggest a safer alternative. | Perceptual/Affordance Mismatch |
| Mismatched expectations and invalid causal/task assumptions ( Honig and Oron-Gilad, 2018 ; Salem et al., 2015 ) | The instruction presupposes a causal, task-level, or goal-logic condition that may not hold in the current world state. | Task model, hazard relation, commonsense constraint, causal dependency, or context-specific safety rule. | Clarify the intended goal, reject invalid causal assumptions, or propose a goal-preserving alternative. | Instruction Logic Mismatch |
| Expectation-related timing, prerequisite, or ordering failure ( Tolmeijer et al., 2020 ; Cameron et al., 2024 ) | The instruction assumes that a required prior state has already been achieved or that the requested action is temporally admissible. | Execution history, task-state record, state-machine status, prerequisite checks, or temporal ordering constraints. | Insert missing prerequisites, reorder the plan, delay execution, or request confirmation before proceeding. | Temporal/Sequence Mismatch |
| Social error, privacy violation, permission failure, or role/policy violation ( Tian and Oviatt, 2021 ; Nogueira et al., 2026 ) | The instruction assumes that the requested action is socially, ethically, or institutionally permissible. | Role information, permission status, privacy policy, institutional rule, safety norm, or social constraint. | Request consent or authorization, refuse the action with an explanation, or provide a norm-compliant alternative. | Social/Ethical Norm Mismatch |
| Intertwined social causes and incompatible stakeholder goals ( Civit et al., 2025 ; Sebo et al., 2020 ) | The instruction assumes that the goals or constraints of multiple users are mutually compatible. | Active user requests, stakeholder roles, priority rules, conflict status, or multi-user interaction history. | Initiate negotiation, arbitrate according to priority rules, escalate to a human authority, or propose a compromise. | Multi-user Conflict |
| System variant | WSM type match | Evidence coverage | Pre-execution intervention | Repair match |
|---|---|---|---|---|
| Post-hoc recovery | 28% | N/A | 10% | 32% |
| Affordance/action-grounding | 54% | 48% | 48% | 52% |
| LLM-only | 46% | N/A | 42% | 44% |
| LLM-Hypothesize | 64% | 68% | 60% | 62% |
| No-IoT / no digital evidence | 74% | 70% | 72% | 72% |
| Full WSM-Aware HRI | 88% | 86% | 86% | 84% |
| ID | Case and instruction | WSM type | Required evidence | Expected system response | Correct |
|---|---|---|---|---|---|
| 1 | Toy + cup: “Put the toy into the cup.” | Perceptual / Affordance | L : object and container size from local vision | Detect size/fit incompatibility; explain that the toy cannot fit and ask for an alternative placement. | 5/5 |
| 2 | Water + holed container: “Pour the water into the container.” | Perceptual / Affordance | L : visual defect or leakage cue | Detect that the container cannot hold liquid; refuse direct pouring and suggest a watertight container. | 5/5 |
| 3 | Light state: “Turn off the light.” | Temporal / Sequence | IoT : smart-light on/off state | Check whether the light is already off; avoid redundant action and inform the user of the current state. | 5/5 |
| 4 | Placement before grasp: “Put the toy on the shelf.” | Temporal / Sequence | L : gripper and execution state | Detect the missing prerequisite; insert grasping before placement or explain the required sequence. | 5/5 |
| 5 | Fire + live electricity: “Use water to extinguish the fire.” | Instruction Logic | L : fire observation; IoT : electrical-source status | Detect the causal hazard; refuse water-based extinguishing and suggest a safer alternative. | 5/5 |
| 6 | Opaque cup: “I’m thirsty. Pass me the water.” | Instruction Logic | L : cup localization; IoT : smart cup or coaster state | Verify the hidden liquid state; if insufficient, inform the user and offer to add water or provide another drink. | 4/5 |
| Scene | With-IoT prompt evidence | IoT evidence removed | Accuracy with IoT | Accuracy without IoT |
| Light status | Smart light state indicates that the light is already off. | Smart light state is removed; only the user instruction is provided. | 10/10 | 0/10 |
| Electrical fire | Smart outlet/electrical panel indicates an active electrical source near the fire. | Smart outlet and electrical-panel evidence are removed. | 10/10 | 0/10 |
| Opaque cup | Smart cup/coaster indicates that the cup is nearly empty. | Smart cup and smart-coaster evidence are removed. | 10/10 | 1/10 |
| Do-not-disturb room | Smart door sign, smart lock, and calendar system indicate a privacy-sensitive state. | Door-sign, lock, and calendar evidence are removed. | 10/10 | 0/10 |
| Family TV conflict | Smart TV and family-identification system indicate conflicting parent–child requests. | Smart TV state and family-identification evidence are removed. | 9/10 | 5/10 |
| Medical-treatment conflict | Healthcare platform and identity-verification system indicate conflicting stakeholder roles. | Healthcare-platform and identity-verification evidence are removed. | 9/10 | 4/10 |
| Ablated Output | ||||||
|---|---|---|---|---|---|---|
| Scene ID | 2 | 3 | 2 | 3 | 2 | 3 |
| WSMs identified (X/5) | 5 | 2 | 0 | 5 | 5 | 5 |
| WSMs identified No Ablation (X/5) | 5 | 5 | 5 | 5 | 5 | 5 |
| Ablated module | Refinement phase in Algorithm 1 |
|---|---|
| Scene ID | 4 |
| WSMs identified (X/5) | 3 |
| WSMs identified no ablation (X/5) | 5 |