UMR: Universal Manipulation Representation
Organizations: University of Science and Technology of China, Hefei, China. · Suzhou Artificial Intelligence Laboratory, Suzhou, China.
Abstract
General-purpose embodied manipulation hinges on a unified action representation that generalizes across embodiments and scales readily. Yet existing policies rely on embodiment-specific action spaces, making cross-embodiment demonstrations difficult to leverage at scale and limiting transfer to new embodiments and spatial variations. To this end, we introduce Universal Manipulation Representation (UMR), a unified action representation that enables zero-shot skill transfer from human demonstrations to heterogeneous robots. UMR decomposes manipulation into two functionally distinct yet geometrically linked components: embodiment-agnostic World Flow, which describes task-relevant object motion in the world frame, and Ego Trajectory, which represents end-effector motion relative to the current pose. We instantiate UMR as World--Ego Point VLA (WEPVLA), a compact 0.5B-parameter policy that learns in the unified geometric action space through a dual-stream Point Action Adapter and a unified Point Action Expert, with an conjugation coupling the two components. To improve data efficiency, we complement UMR with a Data-Efficient Strategy (DES) that diversifies object configurations through stage-aware point-cloud editing while preserving demonstrated contact geometry. In simulation, WEPVLA achieves average success rates of 97.5% on LIBERO and 85.7% on the 10-task RLBench benchmark. In real-world experiments, a single policy trained on human demonstrations augmented by DES transfers zero-shot to diverse deployment conditions. With about 10 minutes of collected human demonstrations per task and no robot demonstrations, it achieves 91.7% average success across six evaluation settings, compared with 60.8% for HumanEgo. Code and additional materials are available at https://umr-wepvla.github.io/.
Figures & tables
| Close box | Close laptop | Toilet seat | Sweep to dustpan | Close fridge | Phone base | Umbrella out | Frame off | Wine rack | Water plants | Mean | |
| ManipLLM (7B) [ 33 ] | 50 | 80 | 40 | 20 | 80 | 35 | 10 | 25 | 15 | 20 | 38 |
| OpenVLA (7B) [ 1 ] | 65 | 40 | 75 | 60 | 80 | 20 | 35 | 15 | 10 | 10 | 41 |
| (2.6B) [ 34 ] | 90 | 60 | 100 | 30 | 90 | 25 | 35 | 75 | 5 | 45 | 55 |
| CogACT (7B) [ 35 ] | 80 | 85 | 90 | 65 | 90 | 50 | 60 | 35 | 25 | 25 | 60 |
| HybridVLA (7B) [ 36 ] | 85 | 95 | 100 | 90 | 100 | 50 | 50 | 70 | 50 | 50 | 74 |
| EO1 (3B) [ 37 ] | 97 | 99 | 100 | 95 | 83 | 46 | 76 | 40 | 61 | 35 | 73.2 |
| Model size | Wrist camera | Pre- trained | Frozen VLM | Train tasks | Spatial | Object | Goal | Long | Avg. | |
| 2D VLAs | ||||||||||
| OpenVLA-OFT [ 39 ] | 7B | 40 | 97.7 | 98.0 | 96.1 | 95.3 | 96.8 | |||
| SmolVLA [ 3 ] | 2B | 40 | 93 | 94 | 91 | 77 | 88.8 | |||
| GR00T-N1.6 [ 40 ] | 3B | 10 | 97.7 | 98.5 | 97.5 | 94.4 | 97.0 | |||
| [ 2 ] | 3B | 10 | 98.8 | 98.2 | 98.0 | 92.4 | 96.9 | |||
| X-VLA [ 41 ] | 0.9B | 40 | 98.2 | 98.6 | 97.8 | 97.6 | 98.1 | |||