In embodied intelligence, the embodiment gap between robotic and human hands brings significant challenges for learning from human demonstrations. Although some studies have attempted to bridge this gap using reinforcement learning, they remain confined to merely reproducing human manipulation, resulting in limited task performance. Moreover, current methods struggle to support diverse robotic hand configurations. In this paper, we propose UniBYD, a unified framework that uses a dynamic reinforcement learning algorithm to discover manipulation policies aligned with the robot's physical characteristics. To enable consistent modeling across diverse robotic hand morphologies, UniBYD incorporates a unified morphological representation (UMR). Building on UMR, we design a dynamic PPO with an annealed reward schedule, enabling reinforcement learning to transition from offline-informed imitation of human demonstrations to online-adaptive exploration of policies better adapted to diverse robotic morphologies, thereby going beyond mere imitation of human hands. To address the severe state drift caused by the incapacity of early-stage policies, we design a hybrid Markov-based shadow engine that provides fine-grained guidance to anchor the imitation within the expert's manifold. To evaluate UniBYD, we propose UniManip, the first benchmark for cross-embodiment manipulation spanning diverse robotic morphologies. Experiments demonstrate a 44.08% average improvement in success rate over the current state-of-the-art. Our project page is https://zhanheng-creator.github.io/UniBYD.
Figures & tables
Figure 1 : Leveraging human demonstrations, UniBYD learns manipulation strategies that transcend mere imitation and are tailored to a broad spectrum of robotic hand morphologies.
Figure 2 : The framework of UniBYD. UniBYD first encodes diverse hands via UMR. It then employs a dynamic PPO with an annealed reward mechanism, which initially leverages the shadow engine for high-fidelity imitation before transitioning to autonomous exploration to discover morphology-aligned policies.
Figure 3 : Overview of action generation and object control in the shadow engine . It blends the model-predicted action and the expert-guided action to generate the final executed action Δatexec . A PD controller applies an expert object force to guide the object.
Hand Type
Metrics
Reta rget
Manip trans
DexMa china ∗
UniBYD
2 (1 hand)
SR ↑ (%)
12.27
✗
✗
78.13
PE ↓ (cm)
2.81
✗
✗
0.53
OE ↓ ( ∘ )
28.77
✗
✗
18.74
AS ↑
4.24
✗
✗
8.93
3 (1 hand)
SR ↑ (%)
4.36
✗
✗
71.81
PE ↓ (cm)
2.65
✗
✗
0.89
Table 1: Comparative results on UniManip. UniBYD consistently outperforms representative baselines across diverse hand morphologies.
Hand Type
Metrics
base
+SE
+GR
+GR +LSC
UniBYD
2 (1 hand)
SR ↑ (%)
24.94
67.06
56.19
61.50
78.13
PE ↓ (cm)
2.66
0.54
1.28
1.32
0.53
OE ↓ ( ∘ )
27.51
19.89
21.31
21.58
18.74
AS ↑
5.63
7.56
7.99
8.36
8.93
3 (1 hand)
SR ↑ (%)
19.56
51.06
65.13
67.63
71.81
PE ↓ (cm)
2.45
0.95
0.93
0.87
0.89
Table 2: Ablation study of UniBYD. The results validate the individual contributions of SE and GR, as well as the incremental gain of LSC, in enhancing performance.
Figure 4 : A visual comparison of the experimental results. UniBYD learns manipulation strategies aligned with the robot’s embodiment, thereby successfully completing the task, whereas both ManipTrans and DexMachina* fail.
Figure 5 : Training procedure of a representative task.
Figure 6 : Comparative experimental results for base and UniBYD.
Figure 9
Figure 9 : An example of task completion using a PD controller without robotic hands. The green fingertip keypoints shown in the figure represent the location of the robotic hand in the expert demonstration data.
Parameter
Value
Learning Rate
5×10−4
Mini-batch Size
1,024
Horizon Length
32
Optimization Epochs
5
Discount Factor
0.99
GAE Parameter
0.95
Table 3 : PPO Hyperparameters and Training Configuration.
Parameter
Value
Reference Object Mass ( mref )
0.03788
Baseline Prop. Gain ( Kp,cfg )
10.0
Baseline Deriv. Gain ( Kd,cfg )
3.0
Max Imitation Weight ( wmaximi )
1.0
Min Imitation Weight ( wminimi )
0.2
Max Goal Weight ( wmaxgoal )
3.0
Table 4 : Additional Dynamic Control and Curriculum Parameters.
Robotic Hand
Degrees of Freedom (DOF)
Shadow Hand
22
Allegro Hand
16
Inspire Hand
6
OHand TM
11
CasiaHand 3-Finger
10
XArm Gripper
1
Table 5 : Specifications of Supported Robotic Hands.
Figure 10 : The distribution of UniManip.
Figure 11 : The evolution of success rate and episode length over training time for a representative task: Pouring liquid from a tall blue mug into a bowl.
Figure 12 : For the same task, UniBYD can learn different manipulation policies based on the physical characteristics of different robotic hands.
Hyperparameter
Low Setting (-50%)
Default Setting
High Setting (+50%)
SE Decay Horizon ( Tdecay )
93.59%
99.83%
96.17%
Goal Reward Start Epoch
81.83%
98.17%
86.50%
Entropy Decay ( Tentropy_decay )
84.50%
98.50%
98.17%
Table 6 : Hyperparameter Sensitivity Analysis.
Method
SR ↑ (%)
OE ↓ ( ∘ )
PE ↓ (m)
AS
w/o Boundary Loss
81.57
7.50
0.17
8.97
w/o Entropy Loss
84.68
5.67
0.18
8.34
w/o Reward Annealing
75.44
5.80
0.20
7.65
UniBYD
93.00
2.92
0.10
9.12
Table 7 : Ablation study of different components on a representative task.
Figure 13 : Experimental results of the 2-fingered robotic hand in simulation.
Figure 14 : Experimental results of the 3-fingered robotic hand in simulation.
Figure 15 : Experimental results of the single 5-fingered robotic hand in simulation.
Figure 16 : Experimental results of the dual 5-fingered robotic hand in simulation.
Figure 17 : Experimental results of the 2-fingered robotic hand in the real world.
Figure 18 : Experimental results of the 3-fingered robotic hand in the real world.
Figure 19 : Experimental results of the 5-fingered robotic hand in the real world.
Robot manipulation policies are typically tied to specific robotic hand embodiments, limiting the transfer of learned behaviors across platforms with different kinematic structures. In this work, we propose the Unified Hand Action Space (UHAS), a sphere-based unified action representation for cross-embodiment dexterous manipulation. UHAS represents robotic hand actions as geometric deformations of a canonical sphere and uses a Cascade Inverse Kinematics (CIK) algorithm to map the shared representation to embodiment-specific joint configurations. Using reinforcement learning, we train dexterous manipulation policies directly in the proposed action space for in-hand cube reorientation tasks. We evaluate our method in both simulation and real-world experiments across multiple robotic hands, including the Allegro Hand, LEAP Hand, Shadow Hand, and MANO Human Hand. Experimental results demonstrate effective dexterous manipulation, zero-shot transfer to unseen hands, rapid finetuning across embodiments, and successful real-world deployment. Our experiments show that the proposed UHAS representation enables stable dexterous control and cross-embodiment policy transfer across robotic hands.
Luis Felipe Casas, Robert Teal, Keval Shah +3
Intelligent Robotics and Vision Lab, University of Texas at Dallas · Intelligent Robotics and Interactive Systems Lab, Arizona State University
General-purpose embodied manipulation hinges on a unified action representation that generalizes across embodiments and scales readily. Yet existing policies rely on embodiment-specific action spaces, making cross-embodiment demonstrations difficult to leverage at scale and limiting transfer to new embodiments and spatial variations. To this end, we introduce Universal Manipulation Representation (UMR), a unified action representation that enables zero-shot skill transfer from human demonstrations to heterogeneous robots. UMR decomposes manipulation into two functionally distinct yet geometrically linked components: embodiment-agnostic World Flow, which describes task-relevant object motion in the world frame, and Ego Trajectory, which represents end-effector motion relative to the current pose. We instantiate UMR as World--Ego Point VLA (WEPVLA), a compact 0.5B-parameter policy that learns in the unified geometric action space through a dual-stream Point Action Adapter and a unified Point Action Expert, with an SE(3) conjugation coupling the two components. To improve data efficiency, we complement UMR with a Data-Efficient Strategy (DES) that diversifies object configurations through stage-aware point-cloud editing while preserving demonstrated contact geometry. In simulation, WEPVLA achieves average success rates of 97.5% on LIBERO and 85.7% on the 10-task RLBench benchmark. In real-world experiments, a single policy trained on human demonstrations augmented by DES transfers zero-shot to diverse deployment conditions. With about 10 minutes of collected human demonstrations per task and no robot demonstrations, it achieves 91.7% average success across six evaluation settings, compared with 60.8% for HumanEgo. Code and additional materials are available at https://umr-wepvla.github.io/.
Song Liu, Linyi Li, Yanshun Zhao +11
University of Science and Technology of China, Hefei, China. · Suzhou Artificial Intelligence Laboratory, Suzhou, China.
Learning robotic manipulation from human videos is a promising solution to the data bottleneck in robotics, but the distribution shift between humans and robots remains a critical challenge. Existing approaches often produce entangled representations, where task-relevant information is coupled with human-specific kinematics, limiting their adaptability. We propose a generative framework for cross-embodiment video editing that directly addresses this by learning explicitly disentangled task and embodiment representations. Our method factorizes a demonstration video into two orthogonal latent spaces by enforcing a dual contrastive objective: it minimizes mutual information between the spaces to ensure independence while maximizing intra-space consistency to create stable representations. A parameter-efficient adapter injects these latent codes into a frozen video diffusion model, enabling the synthesis of a coherent robot execution video from a single human demonstration, without requiring paired cross-embodiment data. Experiments show our approach generates temporally consistent and morphologically accurate robot demonstrations, offering a scalable solution to leverage internet-scale human video for robot learning.
Zhiyuan Li, Wenyan Yang, Wenshuai Zhao +4
Department of Electrical Engineering and Automation, Aalto University, Finland · Department of Computer Science, Aalto University, Finland · Hong Kong University of Science and Technology, China +1