We present iTeach, a deployable system that lets any co-located human fix a robot's perception failures on the spot without expertise, a workstation, or offline retraining. The operator wears a mixed reality (MR) headset, sees the robot's segmentation predictions overlaid on the real scene, and corrects failures hands-free: rearranging objects (HumanPlay), annotating via gaze and voice, and triggering SAM2 backward mask propagation. Each ~20 s interaction yields 150-300 densely labeled training frames; the system fine-tunes the perception model onboard, keeps the better model, and redeploys, all without leaving the deployment site. The full loop requires only an RGB-D camera, onboard GPU, and an MR headset: any mobile robot, any environment. Starting from 26.1 on cluttered real-world scenes, 45 teaching interactions (13K frames) raise segmentation to 80.7 with no catastrophic forgetting; on three standard benchmarks the model never trained on, performance improves as well. Downstream pick-and-place on SceneReplica reaches 72/100, surpassing a model-based pipeline requiring CAD models. A 12-participant user study confirms non-experts match experts on annotation accuracy (~95% box IoU), speed, and task load (NASA-TLX 21/100). The framework is architecture-agnostic: any fine-tunable perception model can serve as backbone.
Figures & tables
Fig. 2 : System setup. A Fetch robot streams RGB-D to an onboard GPU (RTX 4090 in our prototype; replaceable by embeddable compute such as Jetson Thor). A HoloLens 2 connected via Wi-Fi overlays predictions on the real scene and captures gaze-and-voice annotations. All inference, SAM2 propagation, and fine-tuning run onboard.
Fig. 3 : The human navigates the robot through cluttered environments. When the MR overlay reveals a segmentation failure (missed objects, merged instances, or spurious masks), the human initiates data collection.
Fig. 4 : Annotation via MR. The orange cursor tracks eye gaze; green dots are point prompts placed by voice command. SAM2 [ 8 ] converts point prompts to bounding boxes. HumanPlay transitions the scene from cluttered to clean, enabling accurate final-frame annotation.
Fig. 5 : SAM2 backward propagation. Masks generated from the annotated final frame are propagated to earlier, more cluttered frames, converting one annotation into dense video-level supervision.
Fig. 6 : Annotated data collected via MR across diverse environments.
Round
#Scenes
Fo
Fb
I0.75
C
Pre.
–
28.2
21.7
30.9
26.1
1
3
69.0
64.5
75.0
68.4
2
6
70.5
66.2
74.1
69.5
3
9
78.3
72.3
77.4
75.7
4
12
82.2
74.4
79.5
78.5
5
15
74.4
68.0
78.9
72.7
TABLE I : Iterative fine-tuning on D40 (902 test images). The pretrained model ( C=26.1 ) is trained on synthetic data only. Performance saturates at ∼ 20 scenes. D5 (not shown) reaches C=76.6 after 5 single-scene rounds.
Fig. 7 : Qualitative UOIS results across fine-tuning stages on tabletop and beyond-tabletop scenes.
Model
OCID [ 39 ]
OSD [ 40 ]
RobotPush [ 22 ]
iTeach-HP
C
C
C
C
MSMFormer [ 11 ]
86.7
76.4
83.4
29.3
Lu et al. (OCID) [ 22 ]
87.1
74.6
85.6
26.9
Lu et al. (OSD) [ 22 ]
87.4
76.2
89.0
39.9
iTeach-UOIS † (scratch)
76.5
69.7
80.6
58.3
iTeach-UOIS ‡ (pretrained init)
87.6
77.5
87.8
67.8
TABLE II : Longer training (50K iterations) on aggregated HumanPlay data. iTeach-UOIS † : HumanPlay data only, no pretraining. iTeach-UOIS ‡ : pretrained MSMFormer + HumanPlay fine-tuning. The failure-driven data improves the target domain (+38.5 on iTeach-HP) while maintaining OCID (+0.9), OSD (+1.1), and RobotPushing (+4.4).
Fig. 8 : Pick-and-place execution using Pipeline 4. iTeach -UOIS produces clean instance masks, enabling accurate grasp planning and collision-free motion.
#
Perception
Motion
Grasp ↑
P&P ↑
1
GDRNPP [ 44 ] (model-based)
RFP+OMPL [ 45 ]
73
70
2
MSMFormer [ 11 ]
CGNet+OMPL
65
57
3
MSMFormer [ 11 ]
CGNet+GTO
71
65
4
iTeach -UOIS
CGNet+GTO
74
72
TABLE III : SceneReplica manipulation results (out of 100 trials). Pipeline 4 differs from Pipeline 3 only in perception. Pipeline 1 uses model-based perception (GDRNPP) with known CAD models. iTeach -UOIS achieves the best result on this benchmark, surpassing model-based perception (GDRNPP) without requiring CAD models.
Stage
Duration
Details
HumanPlay
5–10 s
Rearrange objects
Annotation
5–10 s/obj
Eye-gaze + voice
SAM2 propagation †
∼ 6 FPS
RTX 4090
Dataset update †
< 2 min
Aggregate frames
Fine-tuning †
10–15 min
2K iters, mixed dataset
Total (incl. model)
15–25 min
Failure → redeployment
TABLE IV : Timing breakdown. Green : system-constant ( ∼ 20 s human effort per scene). Orange † : model-dependent (SAM2 + MSMFormer on RTX 4090); reducible with faster models or LoRA.
Capability
MR (ours)
Desktop
Tablet
Works across mobile robots
✓
×
×
Co-located with robot
✓
×
✓
In-situ prediction overlay
✓
×
×
Hands-free object interaction *
✓
✓
×
Hands-free annotation
✓
×
×
All five simultaneously
✓
×
×
TABLE V : Interface comparison. A desktop cannot travel with the robot; a tablet can, but must be held and shows the scene on a screen rather than in it. Only a worn headset delivers all five capabilities the iTeach loop needs at once. *A held tablet occupies one hand, so neither guiding the robot nor rearranging objects stays hands-free.
Fig. 9 : User study results by expertise ( n=6 per group). Intervals overlap on all three measures.
Fig. 10 : NASA-TLX by subscale. The two profiles overlap on every subscale, so the agreement in Fig. 9 (c) is not an averaging artifact.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Step
Operation
1. Failure observation
Human identifies incorrect masks in robot prediction
SAM2 [ 8 ] video mode propagates masks to earlier frames using Robokit [ 48 ]
Appendix
TABLE VI : iTeach -HumanPlay data collection pipeline.
Fig. 11 : Five scenes from D5 with ground-truth masks from SAM2 [ 8 ] .
Fig. 12 : All 40 scenes from D40. Masks are propagated backward from the annotated final frame. Scenes in red (2, 8, 16, 21, 33) are propagation failures due to occlusion or large object motion.
Method
Interaction
Data Collected
Annotation Effort
Supervision
Robot pushing [ 22 ]
Robot motion
Single frame
None
Self-supervised
Offline labeling
None
Single frame
Full mask
Dense manual
iTeach (ours)
5–10 s HumanPlay
Video sequence
Last-frame only
FS3 Propagated dense
Appendix
TABLE VII : Comparison of data collection strategies.
Model
OCID [ 39 ]
OSD [ 40 ]
RobotPushing [ 22 ]
iTeach -HumanPlay
Fo
Fb
I0.75
C
ΔC
Fo
Fb
I0.75
C
ΔC
Fo
Fb
I0.75
C
ΔC
Fo
Fb
I0.75
C
ΔC
Stage-1 (Stage-1 fine-tuned)
MSMFormer [ 11 ]
88.2
82.9
83.2
84.8
0
82.6
60.9
77.1
73.5
0
87.0
79.5
77.3
81.3
0
28.2
21.7
30.9
26.9
0
Lu et al. [ 22 ] (OCID)
89.9
85.4
82.9
86.7
+1.9
84.6
69.0
76.2
76.7
+3.2
87.8
81.2
75.1
82.6
+1.3
24.3
21.9
28.4
24.1
-2.8
Lu et al. [ 22 ] (OSD)
89.8
84.5
83.0
86.3
+1.5
82.1
61.6
76.4
72.8
-0.7
91.0
85.6
83.8
87.4
+6.1
29.5
26.3
49.8
32.3
+5.4
iTeach -UOIS †
75.2
62.4
68.8
68.8
-16.0
74.8
54.1
68.0
65.6
-7.9
81.1
71.0
72.7
74.9
-6.4
65.3
56.9
55.9
59.4
+32.5
Appendix
TABLE VIII : Longer training (50K iterations) on aggregated HumanPlay data. The main paper uses iterative 2K-iteration fine-tuning. iTeach -UOIS † : trained from scratch on HumanPlay data. iTeach -UOIS ‡ : pretrained MSMFormer fine-tuned on HumanPlay data. Stage-2 is frozen; its gains come solely from improved Stage-1.
Component
Description
Hardware
Robot
Fetch mobile manipulator with RGB-D sensing
Onboard compute
Lenovo Legion laptop with RTX 4090 GPU (16 GB VRAM)
MR interface
Microsoft HoloLens 2 for gaze and voice based annotation and visualization
Networking
Robot–Laptop link
Wired Ethernet for RGB-D data and perception outputs
Appendix
TABLE IX : System configuration.
Group
Pipeline Stage
Command
Action
Visualization
Failure observation
Stream
Start live RGB-D streaming with predicted masks and bounding boxes overlaid on the HoloLens display
Capture
HumanPlay recording
Begin Capture
Start recording an RGB-D video sequence
Stop Capture
End the current recording
Annotation
Final-frame labeling
True Label
Register gaze point as a positive prompt
False Label
Register gaze point as a negative prompt
Erase Label
Remove all point prompts for the current object
Appendix
TABLE X : Voice commands in the iTeach MR interface, mapped to pipeline stages. The user’s hands remain free throughout.
Category
Challenge
Mitigation
Power
Laptop battery drain during GPU training
External power / periodic charging
Power
HoloLens battery limitation
External power bank during long sessions
Thermal
HoloLens overheating
Short pauses between sessions
Deployment
Unity / UWP build compatibility
Fixed dev environment or pre-built app
Networking
Wi-Fi hotspot + ROS communication
Onboard laptop as centralized hub
Compute
Concurrent inference + training
Onboard RTX 4090 GPU
Appendix
TABLE XI : Practical deployment challenges and mitigations.
Fig. 13 : Practical challenges encountered during deployment.
Fig. 14 : Four scenes annotated by P7, two per table surface.
Measure
All ( n=12 )
Expert ( n=6 )
Non-expert ( n=6 )
Box IoU (%)
94.87 (0.49)
94.95 (0.40)
94.78 (0.59)
Completion time (s)
351.19 (36.74)
354.22 (36.86)
348.16 (39.85)
NASA-TLX (/100)
21.04 (12.56)
20.42 (9.88)
21.67 (15.77)
HumanPlay time (s)
14.17 (3.15)
Prompts / object
1.03 (median 1, max 3)
Appendix
TABLE XII : User study results. Values are mean (SD). HumanPlay time and prompts per object are reported as aggregates across all participants.
Fig. 15 : Per-participant annotations, P1–P6 (continued on next page).
Fig. 16 : Per-participant annotations, P7–P12 (continued from previous page). All 48 annotated frames with bounding boxes and instance masks numbered in labeling order. Odd participants (P1, P3, P5, P7, P9, P11) follow W–W–B–B order; even participants (P2, P4, P6, P8, P10, P12) follow B–B–W–W. Corner tags indicate table color.
Long-horizon robot manipulation reuses skills across many task compositions, but improving these compositions with additional end-to-end demonstrations is costly. A practical self-improving system must decide both what to teach next and where to apply that supervision. We present ROBOCOACH, a world-model-guided coaching framework that uses imagined failures to guide demonstration requests and expert updates. Its Route-Imagine-Diagnose-Improve (RIDI) loop executes reusable skill experts inside COACHWORLD, our shared action-conditioned world model, and uses a progress judge to record the first subtask that fails to complete. Aggregated records select which subtask demonstrations to acquire and which expert adapters to update. Across two simulation suites and two real-robot platforms, imagined and deployed success correlate over 22 task-policy pairs (rho = 0.840). Controlled comparisons show that our coaching method outperforms matched baselines under matched data budgets and update schedules. With only 150 additional subtask demonstrations, success rises from 13.3% to 75.0% on Franka and from 40.0% to 83.8% on AgileX. The coached experts also transfer to four held-out compositions, achieving an average success of 35.0%, compared with 0% for a shared-policy baseline updated with uniformly acquired demonstrations. Together, these results show that world models can serve as active coaches, turning imagined failures into targeted supervision for modular policy improvement. Project Page: https://robocoach-ai.github.io/
Jiajun Liu, Yifan Chen, Yichao Liu +7
Renmin University of China · Tsinghua University · Shanghai Qizhi Institute +2
Expert demonstrations are widely assumed to be the gold standard for robot imitation learning. Yet for fine-grained manipulation such as insertion, stacking, and alignment, we uncover a counterintuitive failure mode: fluent demonstrations can be poor teachers. A skilled teleoperator compresses the decisive moments of alignment and recovery into a brief temporal window, leaving the policy flooded with redundant free-space motion and starved of supervision exactly where precision determines success. We address this bottleneck at two levels. At the data level, slowing down near alignment and resampling critical segments both help, yet the gain comes mainly from broadening the coverage of recovery states the policy must learn, not from reweighting frames it already has. Such data-side fixes, however, leave the policy's per-frame view untouched: a single image still maps directly to an action, and the local motion that governs correction stays implicit. We therefore turn to the representation level and introduce STAIR (\textbf{S}patio-\textbf{T}emporal feature \textbf{A}s an \textbf{I}nterface for \textbf{R}obot learning), a compact dynamic feature that bridges the vision-language model and the action expert, distilling the short-horizon motion already recorded in each trajectory into dense, motion-aware supervision. Trained on fluent data alone, STAIR recovers most of the deliberate-demonstration gain (50.0 to 62.2% overall, approaching the 64.4% of deliberate demonstrations). These results call for a more pedagogical view of robot data, optimized for machine learnability rather than human efficiency alone.
Mingyu Liu, Zeju Li, Jiuhe Shu +4
Zhejiang University · Hong Kong University of Science and Technology (GZ) · Shanghai Innovation Institute
Understanding how users perceive and respond to robot failures is essential for building robust and trustworthy robot systems. Prior work, however, (i) often treats failures as independent events, (ii) emphasizes binary failure detection, (iii) with rule-based recovery modeling. We present REPAIR-Bench, built on 214 interaction trials from 41 participants, the benchmark spans four induced failure types and provides synchronized facial action units, head pose, speech transcripts, and post-interaction affect and recovery reports. The benchmark spans three novel evaluation tasks that jointly capture the lifecycle of failure in human-robot interaction (HRI): (i) failure detection over inter-dependent interaction sessions, modeling longitudinal user adaptation across repeated failures; (ii) visual failure-type classification beyond binary success/failure formulations; and (iii) user-centered recovery prediction, inferring users' preferred recovery strategies from interaction context rather than relying on manually designed or rule-based strategies. In baseline experiments, hierarchical recurrent modeling improved failure detection over a single-session model (strict F1: 0.80 vs. 0.68), achieved a failure localization mean signed error of -0.51 s, median absolute error of 2.97 s and, for recovery prediction, a QLoRA-tuned Mistral-7B reached Hit@5=0.76 and F1@5=0.32. REPAIR-Bench provides both the HRI and Medical HRI communities with a standardized framework for (1) evaluating robot failures and (2) building transparent, adaptive, and trustworthy recovery systems.