MRPilot: Supervising and Intervening LLM-Based Multi-Robot Teams through Mixed Reality
Organizations: North Carolina State University Raleigh, North Carolina, USA · Google San Jose, California, USA · Purdue University West Lafayette, Indiana, USA
Abstract
Large language models (LLMs) let users direct heterogeneous multi-robot systems (MRS) through natural language, but make task interpretation, robot assignment, and coordination difficult to inspect and change. Based on a formative study with 12 non-expert users, we developed MRPilot, a mixed reality system organized around four stages of supervision and intervention. MRPilot represents robot-team plans and execution states as structured commitments shared across synchronized situated and overview views. Across four stages, it helps users resolve ambiguous references (Forming), review plans before execution (Reviewing), monitor distributed execution (Following), and make robot-level or team-level changes when problems arise (Repairing). In a within-subjects study with 20 participants in a virtual reality-simulated home, MRPilot reduced workload, increased situational awareness, transparency, trust, and perceived control compared with a conventional LLM-based conversational interface using the same LLM planner and robot capabilities. We provide design implications for multi-scale intervention, adaptive supervision, and calibrated reliance in LLM-based MRS.
Figures & tables
| # | Intervention Event | Message Presented to Participant | Robot | Intervention Type | Initiator |
|---|---|---|---|---|---|
| 1 | Ambiguous object reference | Which item do you want to choose? Select one highlighted option in the space [Glasses 1] [Glasses 2] | Robot Arm | Information Clarification | Agent |
| 2 | Route preference | Choose a route preference: Default route is safer but slower; the Fast route is quicker but riskier. | Drone | Plan Selection or Optimization | Agent |
| 3 | Potentially unstable grasp | Oops! I noticed that the robotic arm is pinching the glasses by the temple, so they may slip. I need to adjust the grasp so it holds the frame securely. | Robot Team | Pre-Failure Action Correction | Human |
| 4 | Delivery robot blocked | Execution failure: The robot is blocked. Assign another robot to take over or roll the delivery robot back to its previous checkpoint and replan a route. | Delivery Robot | Failure Recovery | Agent |
| 5 | Private-room access | The spaceguest room is marked PRIVATE. Allow the drone to enter? | Drone | Authorization or Access | Agent |
| 6 | Recipient relocation | I just realized that the Jim has moved to the studio. I have to pick the glasses back up from the bedside table and deliver them to the desk in the studio. | Robot Team | Task or Goal Update | Human |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| 1 | 2 | 3 | 4 | 5 | 6 | 7 | |
|---|---|---|---|---|---|---|---|
| Strongly disagree | Disagree | Somewhat disagree | Neither agree nor disagree | Somewhat agree | Agree | Strongly agree | |
| I was able to make the robot team do what I intended. | |||||||
| When I wanted to change the robot team’s plan, I was able to change it. | |||||||
| I was in control of how the task was carried out. |
| # | Scenario / Message | Robot | Type | Intervention | Initiator |
|---|---|---|---|---|---|
| 1 | Which room should the drone begin in? Select one highlighted option in the space [Room1] [Room2] | Drone | T1 | Information Clarification | Agent |
| 2 | Choose a route preference: Default route is safer but slower; the Fast route is quicker but riskier. | Drone | T2 | Plan Selection or Optimization | Agent |
| 3 | The room is passcode-protected. Enter the passcode. | Drone | T3 | Authorization or Access | Agent |
| 4 | Oops! The drone is unstable and the bottle may fall during the flight. I need to regrab the bottel before it continues. | T5 | Pre-Failure Action Correction | Human | |
| 5 | The drone is out of battery and has landed. Delete the drone and use other robots to continue the task. | Drone | T6 | Failure Recovery | Agent |
| 6 | Please place the medicine bottle on the delivery robot. | Delivery Robot | T7 | Capability-Gap Assistance | Agent |
| Intervention Type | Description | Include |
|---|---|---|
| Task or Goal Update | The human changes what the agent is trying to achieve after execution has begun—adding, revising, narrowing, or withdrawing requirements—and the agent must re-plan from its current state rather than restart. Distinct from plan editing (the goal is stable) and from norm adjustment (the goal is stable but the acceptable way of reaching it changes). | ( Mozannar et al., 2025 ; Huq et al., 2026 ; Feng et al., 2026 ; Zou et al., 2026 ; Gao et al., 2025 ; Liu et al., 2025b ; Sharma et al., 2022 ; Thierauf et al., 2024 ) |
| Plan Selection or Optimization | Before execution, or between steps, the agent exposes its intended plan or a set of candidate next steps and the human approves, chooses among alternatives, or edits, re-orders, or optimises steps. The goal is unchanged; the human shapes the how . Typically agent-proposed, human-decided. | ( Mozannar et al., 2025 ; Feng et al., 2026 ; Dhanorkar et al., 2026 ; Zhang et al., 2024b ; Ren et al., 2023 ; Takerngsaksiri et al., 2025 ; He et al., 2025 ; Kim et al., 2025 ; He et al., 2026 ; Peng et al., 2025 ) |
| Information Clarification | The agent (occasionally the human) resolves ambiguity, underspecification, or missing information in the task input—such as referents, parameters, facts, or intent—by asking and answering questions before or during execution. The human supplies information the agent lacks ; the goal, plan, permissions, and norms are otherwise unchanged. The intervention is almost always agent-initiated; the central design problem is whether, what, and when to ask while keeping human burden low. | ( Vijayvargiya et al., 2026b ; Su and Cardie, 2026 ; Wu et al., 2026 ; Wang et al., 2025 ; Qian et al., 2024 ; Chen et al., 2024a ; Huq et al., 2026 ; Suri et al., 2026 ; Lin et al., 2025 ; Doğan et al., 2022 ; Ramrakhya et al., 2025 ; Zhang et al., 2024b ; Fang et al., 2025b ; Edwards and Schuster, 2026 ; Shi et al., 2022 ; Zhang et al., 2023 ; Testoni and Fernández, 2024 ; Ren et al., 2023 ; Vijayvargiya et al., 2026a ; Deng et al., 2026 ; Gulati et al., 2026 ; Wu et al., 2025 ; Sun et al., 2025 ; Saracay et al., 2026 ) |
| Pre-Failure Action Correction | While the agent is acting, the human (or a pre-emptive monitor that alerts the human) sees that the next or current action is wrong, premature, or risky and corrects or blocks it before it commits or before error compounds. Fine-grained, real-time, and usually human-initiated; control is normally handed back. | ( Mozannar et al., 2025 ; Huq et al., 2026 ; Huq et al., 2025 ; Dhanorkar et al., 2026 ; Gao et al., 2025 ; Liu et al., 2025b ; Fang et al., 2025a ; Shi et al., 2024 ; Liu et al., 2023b ; Kelly et al., 2019 ; Spencer et al., 2020 ; Saunders et al., 2017 ; He et al., 2025 ; Epperson et al., 2025 ; Zhang et al., 2025a ) |
| Table 3 continued | ||
|---|---|---|
| Intervention Type | Description | Include |
| Failure Recovery | After the agent has failed, stalled, or looped, the human diagnoses, un-sticks, rolls back, or completes the failed step so the agent can resume or the task can be salvaged. Often agent-initiated through a targeted help request, but also human-initiated when the agent does not notice its own failure. | ( Huq et al., 2026 ; Dhanorkar et al., 2026 ; Gao et al., 2025 ; Liu et al., 2025b ; Zhou et al., 2026 ; Shi et al., 2024 ; Liu et al., 2023b ; Sharma et al., 2022 ; Knepper et al., 2015 ; Ahn et al., 2024 ; Das et al., 2021 ; Epperson et al., 2025 ; Zhang et al., 2025a ; Lu et al., 2026 ; Piao et al., 2025 ) |
| Authorization or Access | The agent (or a policy layer) stops at an action boundary and the human grants or denies permission for a consequential action, supplies credentials or access, or sets permission scope in advance. The human decision is binary or scoped; the plan and goal are otherwise unchanged. | ( Mozannar et al., 2025 ; Huq et al., 2025 ; Dhanorkar et al., 2026 ; Ruan et al., 2024 ; Zhang et al., 2025c ; Weng, 2026 ; Luo et al., 2026 ; Jia et al., 2026 ; Korbak et al., 2025 ; Zhou et al., 2026 ; Takerngsaksiri et al., 2025 ; Peng et al., 2025 ) |
| Capability-Gap Assistance | The agent reaches a step it cannot perform—physically, perceptually, or because of a technical or policy barrier—and the human performs that step or supplies the missing resource before returning control. Unlike failure recovery, nothing has necessarily gone wrong; the step lies outside the agent’s capability envelope. | ( Mozannar et al., 2025 ; Huq et al., 2026 ; Huq et al., 2025 ; Feng et al., 2024 ; Knepper et al., 2015 ; Rosenthal et al., 2010 ; Ahn et al., 2024 ; Peng et al., 2025 ; Piao et al., 2025 ) |
| Contextual or Norm Adjustment | The human supplies or changes situational context, preferences, constraints, or norms that govern how the agent should act without changing the goal or a specific action, so that behaviour is appropriate to the setting or person. Such information may be agent-elicited or human-volunteered and may persist beyond the current task. | ( Mozannar et al., 2025 ; Huq et al., 2026 ; Shao et al., 2026 ; Wang et al., 2025 ; Zhang et al., 2024b ; Ramrakhya et al., 2025 ; Ren et al., 2023 ; Glaze and Inclezan, 2025 ; Wu et al., 2023 ; Shankar et al., 2026 ; Jacniacki and Bilski, 2026 ) |