We present iTeach, a deployable system that lets any co-located human fix a robot's perception failures on the spot without expertise, a workstation, or offline retraining. The operator wears a mixed reality (MR) headset, sees the robot's segmentation predictions overlaid on the real scene, and corrects failures hands-free: rearranging objects (HumanPlay), annotating via gaze and voice, and triggering SAM2 backward mask propagation. Each ~20 s interaction yields 150-300 densely labeled training frames; the system fine-tunes the perception model onboard, keeps the better model, and redeploys, all without leaving the deployment site. The full loop requires only an RGB-D camera, onboard GPU, and an MR headset: any mobile robot, any environment. Starting from 26.1 on cluttered real-world scenes, 45 teaching interactions (13K frames) raise segmentation to 80.7 with no catastrophic forgetting; on three standard benchmarks the model never trained on, performance improves as well. Downstream pick-and-place on SceneReplica reaches 72/100, surpassing a model-based pipeline requiring CAD models. A 12-participant user study confirms non-experts match experts on annotation accuracy (~95% box IoU), speed, and task load (NASA-TLX 21/100). The framework is architecture-agnostic: any fine-tunable perception model can serve as backbone.
Figures & tables
Fig. 2 : System setup. A Fetch robot streams RGB-D to an onboard GPU (RTX 4090 in our prototype; replaceable by embeddable compute such as Jetson Thor). A HoloLens 2 connected via Wi-Fi overlays predictions on the real scene and captures gaze-and-voice annotations. All inference, SAM2 propagation, and fine-tuning run onboard.
Fig. 3 : The human navigates the robot through cluttered environments. When the MR overlay reveals a segmentation failure (missed objects, merged instances, or spurious masks), the human initiates data collection.
Fig. 4 : Annotation via MR. The orange cursor tracks eye gaze; green dots are point prompts placed by voice command. SAM2 [ 8 ] converts point prompts to bounding boxes. HumanPlay transitions the scene from cluttered to clean, enabling accurate final-frame annotation.
Fig. 5 : SAM2 backward propagation. Masks generated from the annotated final frame are propagated to earlier, more cluttered frames, converting one annotation into dense video-level supervision.
Fig. 6 : Annotated data collected via MR across diverse environments.
Round
#Scenes
Fo
Fb
I0.75
C
Pre.
–
28.2
21.7
30.9
26.1
1
3
69.0
64.5
75.0
68.4
2
6
70.5
66.2
74.1
69.5
3
9
78.3
72.3
77.4
75.7
4
12
82.2
74.4
79.5
78.5
5
15
74.4
68.0
78.9
72.7
TABLE I : Iterative fine-tuning on D40 (902 test images). The pretrained model ( C=26.1 ) is trained on synthetic data only. Performance saturates at ∼ 20 scenes. D5 (not shown) reaches C=76.6 after 5 single-scene rounds.
Fig. 7 : Qualitative UOIS results across fine-tuning stages on tabletop and beyond-tabletop scenes.
Model
OCID [ 39 ]
OSD [ 40 ]
RobotPush [ 22 ]
iTeach-HP
C
C
C
C
MSMFormer [ 11 ]
86.7
76.4
83.4
29.3
Lu et al. (OCID) [ 22 ]
87.1
74.6
85.6
26.9
Lu et al. (OSD) [ 22 ]
87.4
76.2
89.0
39.9
iTeach-UOIS † (scratch)
76.5
69.7
80.6
58.3
iTeach-UOIS ‡ (pretrained init)
87.6
77.5
87.8
67.8
TABLE II : Longer training (50K iterations) on aggregated HumanPlay data. iTeach-UOIS † : HumanPlay data only, no pretraining. iTeach-UOIS ‡ : pretrained MSMFormer + HumanPlay fine-tuning. The failure-driven data improves the target domain (+38.5 on iTeach-HP) while maintaining OCID (+0.9), OSD (+1.1), and RobotPushing (+4.4).
Fig. 8 : Pick-and-place execution using Pipeline 4. iTeach -UOIS produces clean instance masks, enabling accurate grasp planning and collision-free motion.
#
Perception
Motion
Grasp ↑
P&P ↑
1
GDRNPP [ 44 ] (model-based)
RFP+OMPL [ 45 ]
73
70
2
MSMFormer [ 11 ]
CGNet+OMPL
65
57
3
MSMFormer [ 11 ]
CGNet+GTO
71
65
4
iTeach -UOIS
CGNet+GTO
74
72
TABLE III : SceneReplica manipulation results (out of 100 trials). Pipeline 4 differs from Pipeline 3 only in perception. Pipeline 1 uses model-based perception (GDRNPP) with known CAD models. iTeach -UOIS achieves the best result on this benchmark, surpassing model-based perception (GDRNPP) without requiring CAD models.
Stage
Duration
Details
HumanPlay
5–10 s
Rearrange objects
Annotation
5–10 s/obj
Eye-gaze + voice
SAM2 propagation †
∼ 6 FPS
RTX 4090
Dataset update †
< 2 min
Aggregate frames
Fine-tuning †
10–15 min
2K iters, mixed dataset
Total (incl. model)
15–25 min
Failure → redeployment
TABLE IV : Timing breakdown. Green : system-constant ( ∼ 20 s human effort per scene). Orange † : model-dependent (SAM2 + MSMFormer on RTX 4090); reducible with faster models or LoRA.
Capability
MR (ours)
Desktop
Tablet
Works across mobile robots
✓
×
×
Co-located with robot
✓
×
✓
In-situ prediction overlay
✓
×
×
Hands-free object interaction *
✓
✓
×
Hands-free annotation
✓
×
×
All five simultaneously
✓
×
×
TABLE V : Interface comparison. A desktop cannot travel with the robot; a tablet can, but must be held and shows the scene on a screen rather than in it. Only a worn headset delivers all five capabilities the iTeach loop needs at once. *A held tablet occupies one hand, so neither guiding the robot nor rearranging objects stays hands-free.
Fig. 9 : User study results by expertise ( n=6 per group). Intervals overlap on all three measures.
Fig. 10 : NASA-TLX by subscale. The two profiles overlap on every subscale, so the agreement in Fig. 9 (c) is not an averaging artifact.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Step
Operation
1. Failure observation
Human identifies incorrect masks in robot prediction
SAM2 [ 8 ] video mode propagates masks to earlier frames using Robokit [ 48 ]
Appendix
TABLE VI : iTeach -HumanPlay data collection pipeline.
Fig. 11 : Five scenes from D5 with ground-truth masks from SAM2 [ 8 ] .
Fig. 12 : All 40 scenes from D40. Masks are propagated backward from the annotated final frame. Scenes in red (2, 8, 16, 21, 33) are propagation failures due to occlusion or large object motion.
Method
Interaction
Data Collected
Annotation Effort
Supervision
Robot pushing [ 22 ]
Robot motion
Single frame
None
Self-supervised
Offline labeling
None
Single frame
Full mask
Dense manual
iTeach (ours)
5–10 s HumanPlay
Video sequence
Last-frame only
FS3 Propagated dense
Appendix
TABLE VII : Comparison of data collection strategies.
Model
OCID [ 39 ]
OSD [ 40 ]
RobotPushing [ 22 ]
iTeach -HumanPlay
Fo
Fb
I0.75
C
ΔC
Fo
Fb
I0.75
C
ΔC
Fo
Fb
I0.75
C
ΔC
Fo
Fb
I0.75
C
ΔC
Stage-1 (Stage-1 fine-tuned)
MSMFormer [ 11 ]
88.2
82.9
83.2
84.8
0
82.6
60.9
77.1
73.5
0
87.0
79.5
77.3
81.3
0
28.2
21.7
30.9
26.9
0
Lu et al. [ 22 ] (OCID)
89.9
85.4
82.9
86.7
+1.9
84.6
69.0
76.2
76.7
+3.2
87.8
81.2
75.1
82.6
+1.3
24.3
21.9
28.4
24.1
-2.8
Lu et al. [ 22 ] (OSD)
89.8
84.5
83.0
86.3
+1.5
82.1
61.6
76.4
72.8
-0.7
91.0
85.6
83.8
87.4
+6.1
29.5
26.3
49.8
32.3
+5.4
iTeach -UOIS †
75.2
62.4
68.8
68.8
-16.0
74.8
54.1
68.0
65.6
-7.9
81.1
71.0
72.7
74.9
-6.4
65.3
56.9
55.9
59.4
+32.5
Appendix
TABLE VIII : Longer training (50K iterations) on aggregated HumanPlay data. The main paper uses iterative 2K-iteration fine-tuning. iTeach -UOIS † : trained from scratch on HumanPlay data. iTeach -UOIS ‡ : pretrained MSMFormer fine-tuned on HumanPlay data. Stage-2 is frozen; its gains come solely from improved Stage-1.
Component
Description
Hardware
Robot
Fetch mobile manipulator with RGB-D sensing
Onboard compute
Lenovo Legion laptop with RTX 4090 GPU (16 GB VRAM)
MR interface
Microsoft HoloLens 2 for gaze and voice based annotation and visualization
Networking
Robot–Laptop link
Wired Ethernet for RGB-D data and perception outputs
Appendix
TABLE IX : System configuration.
Group
Pipeline Stage
Command
Action
Visualization
Failure observation
Stream
Start live RGB-D streaming with predicted masks and bounding boxes overlaid on the HoloLens display
Capture
HumanPlay recording
Begin Capture
Start recording an RGB-D video sequence
Stop Capture
End the current recording
Annotation
Final-frame labeling
True Label
Register gaze point as a positive prompt
False Label
Register gaze point as a negative prompt
Erase Label
Remove all point prompts for the current object
Appendix
TABLE X : Voice commands in the iTeach MR interface, mapped to pipeline stages. The user’s hands remain free throughout.
Category
Challenge
Mitigation
Power
Laptop battery drain during GPU training
External power / periodic charging
Power
HoloLens battery limitation
External power bank during long sessions
Thermal
HoloLens overheating
Short pauses between sessions
Deployment
Unity / UWP build compatibility
Fixed dev environment or pre-built app
Networking
Wi-Fi hotspot + ROS communication
Onboard laptop as centralized hub
Compute
Concurrent inference + training
Onboard RTX 4090 GPU
Appendix
TABLE XI : Practical deployment challenges and mitigations.
Fig. 13 : Practical challenges encountered during deployment.
Fig. 14 : Four scenes annotated by P7, two per table surface.
Measure
All ( n=12 )
Expert ( n=6 )
Non-expert ( n=6 )
Box IoU (%)
94.87 (0.49)
94.95 (0.40)
94.78 (0.59)
Completion time (s)
351.19 (36.74)
354.22 (36.86)
348.16 (39.85)
NASA-TLX (/100)
21.04 (12.56)
20.42 (9.88)
21.67 (15.77)
HumanPlay time (s)
14.17 (3.15)
Prompts / object
1.03 (median 1, max 3)
Appendix
TABLE XII : User study results. Values are mean (SD). HumanPlay time and prompts per object are reported as aggregates across all participants.
Fig. 15 : Per-participant annotations, P1–P6 (continued on next page).
Fig. 16 : Per-participant annotations, P7–P12 (continued from previous page). All 48 annotated frames with bounding boxes and instance masks numbered in labeling order. Odd participants (P1, P3, P5, P7, P9, P11) follow W–W–B–B order; even participants (P2, P4, P6, P8, P10, P12) follow B–B–W–W. Corner tags indicate table color.