In embodied intelligence, the embodiment gap between robotic and human hands brings significant challenges for learning from human demonstrations. Although some studies have attempted to bridge this gap using reinforcement learning, they remain confined to merely reproducing human manipulation, resulting in limited task performance. Moreover, current methods struggle to support diverse robotic hand configurations. In this paper, we propose UniBYD, a unified framework that uses a dynamic reinforcement learning algorithm to discover manipulation policies aligned with the robot's physical characteristics. To enable consistent modeling across diverse robotic hand morphologies, UniBYD incorporates a unified morphological representation (UMR). Building on UMR, we design a dynamic PPO with an annealed reward schedule, enabling reinforcement learning to transition from offline-informed imitation of human demonstrations to online-adaptive exploration of policies better adapted to diverse robotic morphologies, thereby going beyond mere imitation of human hands. To address the severe state drift caused by the incapacity of early-stage policies, we design a hybrid Markov-based shadow engine that provides fine-grained guidance to anchor the imitation within the expert's manifold. To evaluate UniBYD, we propose UniManip, the first benchmark for cross-embodiment manipulation spanning diverse robotic morphologies. Experiments demonstrate a 44.08% average improvement in success rate over the current state-of-the-art. Our project page is https://zhanheng-creator.github.io/UniBYD.
Figures & tables
Figure 1 : Leveraging human demonstrations, UniBYD learns manipulation strategies that transcend mere imitation and are tailored to a broad spectrum of robotic hand morphologies.
Figure 2 : The framework of UniBYD. UniBYD first encodes diverse hands via UMR. It then employs a dynamic PPO with an annealed reward mechanism, which initially leverages the shadow engine for high-fidelity imitation before transitioning to autonomous exploration to discover morphology-aligned policies.
Figure 3 : Overview of action generation and object control in the shadow engine . It blends the model-predicted action and the expert-guided action to generate the final executed action Δatexec . A PD controller applies an expert object force to guide the object.
Hand Type
Metrics
Reta rget
Manip trans
DexMa china ∗
UniBYD
2 (1 hand)
SR ↑ (%)
12.27
✗
✗
78.13
PE ↓ (cm)
2.81
✗
✗
0.53
OE ↓ ( ∘ )
28.77
✗
✗
18.74
AS ↑
4.24
✗
✗
8.93
3 (1 hand)
SR ↑ (%)
4.36
✗
✗
71.81
PE ↓ (cm)
2.65
✗
✗
0.89
Table 1: Comparative results on UniManip. UniBYD consistently outperforms representative baselines across diverse hand morphologies.
Hand Type
Metrics
base
+SE
+GR
+GR +LSC
UniBYD
2 (1 hand)
SR ↑ (%)
24.94
67.06
56.19
61.50
78.13
PE ↓ (cm)
2.66
0.54
1.28
1.32
0.53
OE ↓ ( ∘ )
27.51
19.89
21.31
21.58
18.74
AS ↑
5.63
7.56
7.99
8.36
8.93
3 (1 hand)
SR ↑ (%)
19.56
51.06
65.13
67.63
71.81
PE ↓ (cm)
2.45
0.95
0.93
0.87
0.89
Table 2: Ablation study of UniBYD. The results validate the individual contributions of SE and GR, as well as the incremental gain of LSC, in enhancing performance.
Figure 4 : A visual comparison of the experimental results. UniBYD learns manipulation strategies aligned with the robot’s embodiment, thereby successfully completing the task, whereas both ManipTrans and DexMachina* fail.
Figure 5 : Training procedure of a representative task.
Figure 6 : Comparative experimental results for base and UniBYD.
Figure 9
Figure 9 : An example of task completion using a PD controller without robotic hands. The green fingertip keypoints shown in the figure represent the location of the robotic hand in the expert demonstration data.
Parameter
Value
Learning Rate
5×10−4
Mini-batch Size
1,024
Horizon Length
32
Optimization Epochs
5
Discount Factor
0.99
GAE Parameter
0.95
Table 3 : PPO Hyperparameters and Training Configuration.
Parameter
Value
Reference Object Mass ( mref )
0.03788
Baseline Prop. Gain ( Kp,cfg )
10.0
Baseline Deriv. Gain ( Kd,cfg )
3.0
Max Imitation Weight ( wmaximi )
1.0
Min Imitation Weight ( wminimi )
0.2
Max Goal Weight ( wmaxgoal )
3.0
Table 4 : Additional Dynamic Control and Curriculum Parameters.
Robotic Hand
Degrees of Freedom (DOF)
Shadow Hand
22
Allegro Hand
16
Inspire Hand
6
OHand TM
11
CasiaHand 3-Finger
10
XArm Gripper
1
Table 5 : Specifications of Supported Robotic Hands.
Figure 10 : The distribution of UniManip.
Figure 11 : The evolution of success rate and episode length over training time for a representative task: Pouring liquid from a tall blue mug into a bowl.
Figure 12 : For the same task, UniBYD can learn different manipulation policies based on the physical characteristics of different robotic hands.
Hyperparameter
Low Setting (-50%)
Default Setting
High Setting (+50%)
SE Decay Horizon ( Tdecay )
93.59%
99.83%
96.17%
Goal Reward Start Epoch
81.83%
98.17%
86.50%
Entropy Decay ( Tentropy_decay )
84.50%
98.50%
98.17%
Table 6 : Hyperparameter Sensitivity Analysis.
Method
SR ↑ (%)
OE ↓ ( ∘ )
PE ↓ (m)
AS
w/o Boundary Loss
81.57
7.50
0.17
8.97
w/o Entropy Loss
84.68
5.67
0.18
8.34
w/o Reward Annealing
75.44
5.80
0.20
7.65
UniBYD
93.00
2.92
0.10
9.12
Table 7 : Ablation study of different components on a representative task.
Figure 13 : Experimental results of the 2-fingered robotic hand in simulation.
Figure 14 : Experimental results of the 3-fingered robotic hand in simulation.
Figure 15 : Experimental results of the single 5-fingered robotic hand in simulation.
Figure 16 : Experimental results of the dual 5-fingered robotic hand in simulation.
Figure 17 : Experimental results of the 2-fingered robotic hand in the real world.
Figure 18 : Experimental results of the 3-fingered robotic hand in the real world.
Figure 19 : Experimental results of the 5-fingered robotic hand in the real world.
Department of Electrical Engineering and Automation, Aalto University, Finland · Department of Computer Science, Aalto University, Finland · Hong Kong University of Science and Technology, China +1