Portable mobile-manipulation demonstrations can help alleviate data scarcity for embodied intelligence, but obtaining reliable, low-cost, and robot-free motion supervision from RGB observations remains challenging. Existing approaches often rely on teleoperation or specialized devices equipped with additional sensing hardware, while directly using estimated visual odometry (VO) trajectories can introduce inconsistencies due to accumulated drift and imperfect motion supervision. We present the Visual-Odometry-Conditioned Mobile Manipulation Interface (VOMMI), a portable demonstration collection and learning framework that connects portable RGB demonstrations to vision-language-action (VLA) post-training through offline trajectory reconstruction and online visual-motion conditioning. VOMMI synchronizes body and hand views to capture navigation context and local object interactions without requiring human-robot kinematic correspondence calibration. R2-VO refines offline demonstration trajectories using sparse geometric anchors and produces causal local-motion tokens over multiple prediction horizons for online policy conditioning. An action-group residual adapter incorporates these tokens only into the base branch. Experiments use a 500-trajectory portable for each task, with 75 trajectories held out for RGB-VO evaluation, and 200 robot demonstrations as references. Our policy, post-trained only on portable demonstrations, achieves 18.2% lower base-velocity error than a policy trained with robot-collected demonstrations, while maintaining comparable end-effector translation accuracy. Offline reconstruction reduces absolute trajectory errors for the body and hand streams by 24.6% on average relative to the best evaluated baseline for each stream. The complete system improves the mean success rate by 8.3 percentage points over OpenPI 0.5 across three real-robot tasks.
Figures & tables
Fig. 1: VOMMI. We build upon portable manipulation interfaces with a low-cost RGB-based perception system to enable accessible mobile demonstration collection and integration with VLA frameworks.
Fig. 2: Portable data collection interface. Independent Body and Hand RGB views are timestamped into a robot-compatible episode.
Method
Camera(s)
View
Price (USD)
UMI [ 5 ]
GoPro
Wrist
$350–450
iPhUMI [ 7 ]
iPhone
Wrist
$1,000–1,200
FastUMI [ 6 ]
GoPro+RealSense T265
Wrist
$650–900
Mobile UMI [ 9 ]
Dual RGB
Wrist+ego
$800–1,200
VOMMI (ours)
2 fisheye RGB
Wrist+ego
$100–200
TABLE I: Representative hardware requirements and estimated acquisition cost for portable data collection methods
Fig. 3: Overview of VOMMI. The online branch estimates multi-horizon local motion from RGB and conditions the base-action residual. The offline branch combines the causal estimator with frozen VGGT anchors, initializes the Body trajectory at its first valid pose, and applies planar Body compensation to the Hand stream before forming demonstration targets. Reference pose supervision available only during data collection.
Body SE(2)
Hand SE(3)
Local (mean/P90)
Accumulated
Local (mean/P90)
Accumulated
Method
Trans. (cm)
Yaw ( ∘ )
XY ATE (cm)
XY End (cm)
Trans. (cm)
Rot. ( ∘ )
3D ATE (cm)
3D End (cm)
XY ATE (cm)
XY End (cm)
Ours (online)
0.26/0.54
0.092 /0.227
15.65
19.29
0.60/ 1.08
0.346/0.717
71.82
60.42
71.51
60.30
Ours (offline)
0.24/0.58
0.093/ 0.224
5.97
5.33
0.60 /1.08
0.340/0.704
50.47
56.68
50.23
56.66
DPVO
1.24/2.86
0.412 0.537
28.09
27.86
14.48/21.90 †
3.887/8.591 †
220.28
219.21
220.16
219.21
DROID-SLAM
1.50/3.21
0.309/0.662
83.86
62.67
2.85/4.62 †
2.358/4.834 †
124.54
109.75
124.16
109.45
TABLE II: Local and accumulated motion errors on VOMMI trajectories. Lower is better.
Fig. 4: Camera trajectories reconstruction on held-out episode trail. The online stream and reference are projected to the ground plane. Offline Body scores in Table II use the common x – y state.
Fig. 5: Qualitative M20S rollouts on three long-horizon mobile-manipulation tasks. The upper rows show representative stages from navigation and approach to object interaction and task completion. The bottom row shows representative failure cases, marked with red crosses.
Method
Barriers–fruit
Room–bottle
Cabinet–box
OpenPI 0.5
35
50
65
GR00T N1.6
20
25
10
X-VLA
25
35
30
VOMMI (ours)
65
60
50
TABLE III: Complete-task success rate (%) over 20 trials per task.
Policy
Base MAE (m/s)
Yaw MAE (rad/s)
EE trans. mean (cm)
EE rot. mean ( ∘ )
EE trans. max (cm)
M20-trained
0.044
0.107
0.155
0.372
4.32
VOMMI-R
0.036
0.104
0.170
0.453
4.55
VOMMI-A
0.079
0.204
0.569
1.091
13.96
TABLE IV: Open-loop action prediction errors.
Variant
Residual target
Overall
Navigation
Port. manip.
Parent
None
3.86±0.06
3.82±0.06
3.99±0.19
Base-VO
Base (VO)
3.91±0.03
3.90±0.07
3.95±0.16
All-Action-VO
All actions (VO)
3.83±0.05
3.82±0.11
3.88±0.22
VOMMI
Base (VO) + manip.
3.79±0.02
3.76±0.07
3.89±0.20
TABLE V: Residual ablation: flow loss ( ×10−2 ).
Fig. 6: Open-loop comparison using global observations and first-step base actions (vx,vy,yaw) , with the residual VO turn highlighted at t=3.73 s.