General-purpose embodied manipulation hinges on a unified action representation that generalizes across embodiments and scales readily. Yet existing policies rely on embodiment-specific action spaces, making cross-embodiment demonstrations difficult to leverage at scale and limiting transfer to new embodiments and spatial variations. To this end, we introduce Universal Manipulation Representation (UMR), a unified action representation that enables zero-shot skill transfer from human demonstrations to heterogeneous robots. UMR decomposes manipulation into two functionally distinct yet geometrically linked components: embodiment-agnostic World Flow, which describes task-relevant object motion in the world frame, and Ego Trajectory, which represents end-effector motion relative to the current pose. We instantiate UMR as World--Ego Point VLA (WEPVLA), a compact 0.5B-parameter policy that learns in the unified geometric action space through a dual-stream Point Action Adapter and a unified Point Action Expert, with an SE(3) conjugation coupling the two components. To improve data efficiency, we complement UMR with a Data-Efficient Strategy (DES) that diversifies object configurations through stage-aware point-cloud editing while preserving demonstrated contact geometry. In simulation, WEPVLA achieves average success rates of 97.5% on LIBERO and 85.7% on the 10-task RLBench benchmark. In real-world experiments, a single policy trained on human demonstrations augmented by DES transfers zero-shot to diverse deployment conditions. With about 10 minutes of collected human demonstrations per task and no robot demonstrations, it achieves 91.7% average success across six evaluation settings, compared with 60.8% for HumanEgo. Code and additional materials are available at https://umr-wepvla.github.io/.
Figures & tables
Fig. 2 : Overview of the UMR pipeline. UMR reformulates the learning targets of heterogeneous manipulation data from different sources into a unified dual-stream representation consisting of World Flow and Ego Trajectory. Building on this representation, we develop WEPVLA with two dedicated branches to extract task-relevant and execution-relevant features, which are subsequently fused in a shared Action Expert for action generation. Experiments demonstrate strong generalization across robot embodiments, viewpoints, and spatial configurations.
Fig. 3 : Overview of the simulated benchmarks. (a) Ten RLBench tasks covering diverse manipulation scenarios that require object interaction and tool usage. (b) Four LIBERO suites, spanning tasks related to spatial relations, object interaction, goal specification, and long-horizon manipulation.
Close box
Close laptop
Toilet seat
Sweep to dustpan
Close fridge
Phone base
Umbrella out
Frame off
Wine rack
Water plants
Mean
ManipLLM (7B) [ 33 ]
50
80
40
20
80
35
10
25
15
20
38
OpenVLA (7B) [ 1 ]
65
40
75
60
80
20
35
15
10
10
41
π0 (2.6B) [ 34 ]
90
60
100
30
90
25
35
75
5
45
55
CogACT (7B) [ 35 ]
80
85
90
65
90
50
60
35
25
25
60
HybridVLA (7B) [ 36 ]
85
95
100
90
100
50
50
70
50
50
74
EO1 (3B) [ 37 ]
97
99
100
95
83
46
76
40
61
35
73.2
TABLE I : Success rate (%) on RLBench 10 tasks. We evaluate WEPVLA over 100 episodes per task. Results for other methods are taken from [ 10 ] .
Model size
Wrist camera
Pre- trained
Frozen VLM
Train tasks
Spatial
Object
Goal
Long
Avg.
2D VLAs
OpenVLA-OFT [ 39 ]
7B
✓
✓
×
40
97.7
98.0
96.1
95.3
96.8
SmolVLA [ 3 ]
2B
✓
✓
✓
40
93
94
91
77
88.8
GR00T-N1.6 [ 40 ]
3B
✓
✓
✓
10
97.7
98.5
97.5
94.4
97.0
π0.5 [ 2 ]
3B
✓
✓
×
10
98.8
98.2
98.0
92.4
96.9
X-VLA [ 41 ]
0.9B
✓
✓
×
40
98.2
98.6
97.8
97.6
98.1
TABLE II : Comparison of 2D VLAs and 3D-aware VLAs on LIBERO Bench. “Pretrained” denotes whether the model is fully pretrained, and “train tasks” denotes the number of LIBERO tasks the model is fine-tuned on.
Fig. 4 : Real-world setup and Eval Tasks. Left: Experimental scene and hardware, manipulated objects, and training data source. Right: Examples of evaluation tasks under diverse deployment conditions.
Table 6
Fig. 5 : Real-world comparison with HumanEgo. Both methods are trained as unified policies on 10 minutes of human demonstrations per task. Left: Success rates across six real-world evaluation (I–VI). Right: Success rates aggregated by generalization condition. Cross-Embodiment includes Tasks I–VI, Cross-Environment includes Tasks I, II, V, and VI, and Cross-Setup includes Tasks I, III, and VI. UMR consistently outperforms HumanEgo across both individual tasks and aggregated generalization settings.
Fig. 6 : Data efficiency and scaling of UMR. (a) UMR achieves higher success rates with limited human demonstrations, while DES further improves data efficiency. (b) Scaling the training data improves performance on both long-horizon tasks, increasing average success from 82% to 96%.
Robot manipulation policies are typically tied to specific robotic hand embodiments, limiting the transfer of learned behaviors across platforms with different kinematic structures. In this work, we propose the Unified Hand Action Space (UHAS), a sphere-based unified action representation for cross-embodiment dexterous manipulation. UHAS represents robotic hand actions as geometric deformations of a canonical sphere and uses a Cascade Inverse Kinematics (CIK) algorithm to map the shared representation to embodiment-specific joint configurations. Using reinforcement learning, we train dexterous manipulation policies directly in the proposed action space for in-hand cube reorientation tasks. We evaluate our method in both simulation and real-world experiments across multiple robotic hands, including the Allegro Hand, LEAP Hand, Shadow Hand, and MANO Human Hand. Experimental results demonstrate effective dexterous manipulation, zero-shot transfer to unseen hands, rapid finetuning across embodiments, and successful real-world deployment. Our experiments show that the proposed UHAS representation enables stable dexterous control and cross-embodiment policy transfer across robotic hands.
Luis Felipe Casas, Robert Teal, Keval Shah +3
Intelligent Robotics and Vision Lab, University of Texas at Dallas · Intelligent Robotics and Interactive Systems Lab, Arizona State University
Recent progress in large-scale imitation learning for robot manipulation has been driven by leveraging datasets across a wide range of robot embodiments. However, achieving significant cross-embodiment transfer is often still challenging. In this work, we study the role of using behavior-aligned representations (e.g., object bounding boxes, language motions, end-effector traces of robot motion) in vision-language-action (VLA) models to promote cross-embodiment transfer. We hypothesize that by possessing invariances across embodiments while being predictive of robot actions, these representations can help unify large-scale cross-embodiment data to enhance transfer. To assess our hypothesis, we develop a simulation-based benchmark designed to assess transfer with diverse cross-embodiment data to new embodiments. Using this benchmark, we compare different representations and ways of incorporating them. We identify that end-effector traces can be particularly beneficial for transfer, representations are generally more useful with larger prior datasets, and can be used to benefit from action-free data. We also demonstrate that they can enhance sim-to-real cross-embodiment transfer, improving task completion progress of real robot policies pre-trained on simulation data by 28%. We provide videos of our evaluations at our website: https://ajaysridhar.com/barx/.
Real-robot demonstrations are limited, motivating the use of human manipulation data collected without robots, including egocentric videos and handheld Universal Manipulation Interface (UMI) demonstrations. However, differences in viewpoint, embodiment, and available action supervision make it difficult to align representations across these sources according to manipulation motion rather than visual appearance. We introduce UMI-Bridge, which uses UMI as an intermediate domain to align representations according to action equivalence rather than pixel similarity. UMI action supervision anchors the latent representation to end-effector motion and gripper behavior, while synchronized head-wrist observations and paired ego-UMI clips support alignment across views and domains. We train a dual-view latent action model (LAM) on human manipulation data without robot demonstrations, then freeze its wrist teacher and dynamics model to regularize vision-language-action (VLA) post-training on UMI and robot data. The shared wrist interface enables this training-time supervision across both domains while preserving the policy's standard inference architecture. Across three real-robot tasks, UMI-Bridge achieves 91.7% mean success versus 73.3% for Naive Co-training with matched UMI and robot data. On two data-efficiency tasks, it surpasses a full-data Robot-only baseline using 25% of the robot demonstrations together with UMI data. It also achieves 85% and 90% success on two additional tasks learned from UMI demonstrations without task-specific robot demonstrations. These results support action-anchored latent alignment for data-efficient robot learning and UMI-to-robot task transfer.