Towards Accurate End-Effector Localization for UMI-Style Robotic Manipulation Teaching
Authors: Junjie Zhang, Deteng Zhang, Zhisong Xu, Bo Sun, Liuyang Li, Yihong Tian, Jie Yin
Organizations: Chongqing University · Northwestern Polytechnical University · The University of Tokyo · Jiaxing University · Sichuan University · Beijing Institute of Technology · Shanghai Jiao Tong University
Robot demonstration learning requires accurate and temporally complete end-effector localization during close-range manipulation and camera occlusion. Existing SLAM benchmarks emphasize navigation motions, whereas manipulation datasets prioritize policy learning over localization evaluation. We introduce MILD, a Manipulation-Interface Localization Dataset with real-world and simulation sequences. The real-world subset provides 86 sensor sequences from Insta360 X5 and Insight9 across 15 repeated tabletop tasks, calibration assets, and a per-execution robot end-effector reference trajectory. The simulation subset, MILD-Sim, extends task coverage in Isaac Sim for controlled manipulation-replay studies. Benchmarking visual-inertial and fiducial-aided systems on instrumented real-world recordings reveals large differences in both TCP-relative trajectory error and temporal coverage, even under the same nominal task. To support marker-augmented teaching workspaces without a pre-surveyed fiducial map, we present AprilVINS, which combines fisheye visual-inertial estimation with sequence-local AprilTag geometry and separates prior admission from guarded export of the jointly optimized state. On Insta360 AprilTag4 recordings, AprilVINS(full) under a unified protocol with sequence-specific profiles reaches millimeter-level SE(3)-aligned TCP-relative APE RMSE with high time completion and lower reported error than the tested routes under their respective protocols, whereas fisheye VIO without tag factors remains at centimeter scale. Ablations separate accuracy from exportability, and a MILD-Sim replay study provides task-specific tolerance references for interpreting those error magnitudes. Together, MILD and AprilVINS provide a diagnostic benchmarking framework for UMI-style demonstration collection. Code, datasets, and evaluation manifests will be released upon acceptance.
Figures & tables
Fig. 1: Representative MILD recordings from Insta360 X5 (fisheye), Insight9 (stereo/RGB-D), and MILD-Sim (simulated fisheye) across pick-and-place, wiping, and rearrangement tasks. The montage shows the multi-sensor, cross-domain coverage of the benchmark.
Fig. 2: MILD acquisition overview. Platform: AGIbot G2 with Insta360 X5 and Insight9 under teleoperation. Scene: plain tabletop, tablecloth, and ArUco/AprilTag marker layouts. Task: 15 tabletop families with executed TCP references; orange curves show trajectories and green/blue/red markers denote start, pick/place, and end events.
Item
Analemma
Circular
Zigzag
Bookshelf
Box
Pick-and- place
Wiping
Throwing
Total
Real-world number
9
9
11
15
7
23
12
0
86
Simulation number
1
1
1
0
0
4
1
1
9
number
10
10
12
15
7
27
13
1
95
Time (min)
5.4
4.9
17.9
31.0
16.0
41.0
28.9
0.2
145.3
Size (GB)
9.6
8.7
27.9
62.6
35.8
100.2
62.7
0.33
307.83
TABLE I: MILD dataset overview by task family.
Fig. 3: MILD-Sim overview. Scene: Isaac Sim workbench and Franka manipulation setup. Task: representative pick-and-place, wiping, and throwing trajectories with executed end-effector paths.
Fig. 4: AprilVINS pipeline. Fisheye images and IMU feed a VINS-Fisheye backbone with point, plane, and line features in a fixed-lag sliding window. A parallel fiducial front end detects AprilTags, estimates sequence-local tag layout via spherical PnP, and supplies admitted tag-pose priors to the window. Guarded continuation then filters the jointly optimized body state before TCP conversion.
Route
Analemma
Circular
Zigzag
Bookshelf 01
Bookshelf 02
Box 01
Box 02
Pick-and- place 01
Pick-and- place 02
Pick-and- place 03
Pick-and- place 04
Pick-and- place 05
Pick-and- place 06
Wiping 01
Wiping 02
VINS-Fisheye
0.0818/0.98
0.0385/0.98
0.0526/0.99
0.0598/0.99
0.0803/1.00
0.0503/0.99
0.0671/0.87
0.0423/0.99
0.0314/0.99
0.0284/0.99
0.0691/0.99
0.0433/0.92
0.0349/0.99
0.0264/1.00
0.0358/0.99
DM-VIO
0.0504/1.00 †
0.0328/1.00 †
×
0.2399/0.97
591.5738/0.99
0.1425/1.00 †
×
0.1479/0.98
0.1064/0.98
0.7871/1.00 †
0.3276/1.00 †
×
0.2939/1.00 †
×
×
ORB-SLAM3
×
×
×
×
×
×
×
×
×
×
×
×
×
×
×
OpenVINS
3.0149/0.78
×
60.7054/0.88
2.9991/0.91
29.7722/0.92
1.0523/0.92
43.3075/0.96
2.6814/0.89
2.6765/0.88
3.4719/0.89
11.4182/0.94
18.3236/0.96
6.8308/0.94
17.7003/0.94
10.2853/0.95
TagSLAM(single-anchor)
0.0703/1.00
0.0404/1.00
0.3199/1.00
0.0347/1.00
0.0608/1.00
0.1104/1.00
0.0896/1.00
0.0553/1.00
0.0286/1.00
0.0404/1.00
0.2516/1.00
0.4332/1.00
0.0304/1.00
0.0260/1.00
0.0947/1.00
UcoSLAM
4.4863/0.92 †
0.4330/0.96 †
0.3402/0.43
×
0.6350/0.17
0.9210/0.98 †
0.3342/0.56
×
0.3035/0.98
0.3684/0.98
0.7185/0.99 †
0.7328/0.64
×
0.2856/0.99
0.3525/0.72 †
TABLE III: Task-level localization on MILD: APE RMSE (m) / time completion.
Fig. 5: Task-difficulty map from non-AprilVINS baselines. Points show the mean APE RMSE and mean time completion of the three lowest-RMSE evaluable routes per task.
Method
Markers
Analemma
Circular
Zigzag
Bookshelf 01
Bookshelf 02
Wiping 02
UcoSLAM
ArUco2
×
0.027/0.02
0.251/0.37
0.237/0.98
0.574/0.19
×
ArUco4
×
0.433/0.96
0.340/0.43
×
0.635/0.17
0.352/0.72
OpenVINS(ArUco)
ArUco2
×
×
×
3.050/0.90
4.826/0.93
×
ArUco4
2.725/0.78
×
×
2.885/0.90
4.603/0.92
×
TagSLAM(actual-count)
AprilTag2
0.063/1.00
×
0.072/1.00
0.035/1.00
×
0.026/1.00
AprilTag4
×
×
×
×
×
×
TABLE IV: Marker-configuration sensitivity on Insta360 X5: APE RMSE (m) / time completion.
Route
Analemma
Circular
Zigzag
Bookshelf 01
Bookshelf 02
Box 01
Pick-and- place 04
Pick-and- place 05
Pick-and- place 06
Wiping 01
Wiping 02
Body/waist motion
No
No
Yes
Yes
No
Yes
Yes
Yes
Yes
Yes
Yes
Insight9 SDK VIO
0.097/1.00
0.015/1.00
0.014/1.00
0.034/1.00
0.063/1.00
0.565/1.00
0.019/1.00
0.025/1.00
0.023/1.00
0.201/1.00
0.029/1.00
PICO Controller
0.043/1.00
0.019/1.00
0.077/1.00
0.140/1.00
0.071/1.00
0.052/1.00
0.043/1.00
0.029/1.00
0.045/1.00
0.025/1.00
0.026/1.00
AprilVINS(full)
0.0049/1.00
0.0053/1.00
0.0059/1.00
0.0031/1.00
0.0076/1.00
0.0051/1.00
0.0044/1.00
0.0036/1.00
0.0028/1.00
0.0021/1.00
0.0033/1.00
TABLE V: Teaching-device trajectories: APE RMSE (m) / time completion.
Fig. 6: AprilVINS ablation visualizations. Left: local trajectory comparisons on four MILD tasks; dashed curves are executed TCP references and solid curves are estimates. Right: output-gate continuity check in the YZ plane; curves split at timestamp gaps larger than 0.15 s.
Fig. 7: 3D practical-route trajectory comparisons for three Table V tasks. Dashed curves are route-specific executed references; solid curves are estimates.
Fig. 8: Replay success versus translational APE–RMSE on six MILD-Sim tasks. Dashed lines mark the zero-noise baseline and the first bin at or below 50% success; error bars show Wilson 95% confidence intervals.
Achieving precise positioning of the mobile manipulator's base is essential for successful manipulation actions that follow. Most of the RGB-based navigation systems only guarantee coarse, meter-level accuracy, making them less suitable for the precise positioning phase of mobile manipulation. This gap prevents manipulation policies from operating within the distribution of their training demonstrations, resulting in frequent execution failures. We address this gap by introducing an object-centric imitation learning framework for last-meter navigation, enabling a quadruped mobile manipulator robot to achieve manipulation-ready positioning using only RGB observations from its onboard cameras. Our method conditions the navigation policy on three inputs: goal images, multi-view RGB observations from the onboard cameras, and a text prompt specifying the target object. A language-driven segmentation module and a spatial score-matrix decoder then supply explicit object grounding and relative pose reasoning. Using real-world data from a single object instance within a category, the system generalizes to unseen object instances across diverse environments with challenging lighting and background conditions. To comprehensively evaluate this, we introduce two metrics: an edge-alignment metric, which uses ground truth orientation, and an object-alignment metric, which evaluates how well the robot visually faces the target. Under these metrics, our policy achieves 74.58% success in edge-alignment and 89.42% success in object-alignment when positioning relative to unseen target objects. These results show that precise last-meter navigation can be achieved at a category-level without depth, LiDAR, or map priors, enabling a scalable pathway toward unified mobile manipulation. Project page: https://rpm-lab-umn.github.io/category-level-last-meter-nav/
Real-robot evaluation is essential for understanding whether learned manipulation policies can operate reliably outside curated demonstrations. This need is particularly pressing for Universal Manipulation Interface (UMI)-style policies, whose performance depends on the coupling between wrist-view observations, action representation, data collection, and physical deployment. Existing real-world benchmarks have made important progress, but they are not designed around this UMI data-to-deployment setting. We present UMI-Bench 1.0, a local-first real-robot benchmark for standardized evaluation of UMI-style manipulation policies. To the best of our knowledge, this is the first benchmark dedicated to real-world evaluation of UMI-based manipulation models. UMI-Bench aligns data collection, scene reset, policy execution, result logging, and task-factor analysis within a unified protocol. By making the full evaluation process reproducible and auditable, UMI-Bench provides a practical testbed for measuring how UMI-trained policies generalize to real physical manipulation.
Shi Jin, Yuntian Wang, Yuhui Duan +16
Soochow University · Lumos Robotics · Fudan University +5
Robot learning from real-world demonstrations is currently constrained by data scaling. Universal Manipulation Interface (UMI) provides an efficient robot-free data collection interface, yet current UMI-style pipelines often collect redundant demonstrations and lack global scene context. To improve data efficiency, we present EgoGuide, a collection interface that records synchronized wrist and head/egocentric observations and couples them with online visual-geometric data quality guidance. We also introduce a Gated Egocentric Residual Policy for robust learning from a viewpoint-varying egocentric camera, allowing head/egocentric context to correct ambiguous local observations while preserving stable wrist-view control. Real-world experiments show that EgoGuide reduces the required number of data episodes and improves data efficiency. The residual policy further improves robustness under visual occlusion. Project Page: https://silicx.github.io/EgoGuide
Yue Xu, Mingtao Nie, Tianle Li +4
Shanghai Jiao Tong University · Beijing Institute for General Artificial Intelligence (BIGAI) · Shanghai Innovation Institute