Legged robots can learn expressive whole-body skills from the motions of humans and animals. Due to the morphology gap between the source and the robot, however, the motion must be tailored to the dynamic properties of the robot. In particular, dynamic motions such as a jump require careful adjustment, since their timing and control are interdependent. We propose dense temporal motion retargeting (DTMR), which jointly optimizes timing and control within a single optimization, where dense means that the timing is adjusted for every control step. This dense formulation enables DTMR to deform only the parts of the motion that need a change in timing. The problem is solved with sampling-based model predictive control (MPC) in parallel on a GPU. We evaluate DTMR against baselines on two hours of human motion with four humanoid robots, where the results show that DTMR outperforms baseline methods, particularly on dynamic motions. We also show that allowing more temporal deformation yields more precise retargeting. We further compare DTMR with a baseline that optimizes the temporal dimension, where the result shows that DTMR retargets more precisely under the same deformation budget while being ~19x faster. Lastly, policies trained on our references transfer to a real humanoid robot.
Figures & tables
Fig. 1 : The timing of a human jump is adjusted according to the dynamic properties of the Unitree G1. The take-off is delayed (blue) and the flight is sped up (orange).
Term
Expression
Weight
root position ( ctrack )
ρ(proot−pˉroot)
28
link position ( ctrack )
∑ℓsℓρ(Δpℓ)
14
link orientation ( ctrack )
∑ℓsℓρ(ΔRℓ)
5
joint velocity ( ctrack )
21∥q˙k−2rkqˉ˙(ϕk)∥2
0.01
phase ( cϕ )
∥rk∥Qϕ2
5
torque energy ( creg )
∑i[cosh(τk,i/τirng)−1]
0.05
TABLE I : The cost terms are listed with their weights.
Setting
Value
samples ( NW )
4096
horizon ( H )
40 steps
control nodes ( Hnode )
8
annealing iterations ( N , cold / warm)
5 / 2
noise ( σ0 ), decay ( β )
0.15, 0.6
horizon envelope ( βh )
0.95
TABLE II : List of the hyperparameters used for DTMR.
Fig. 2 : (a) The dataset retargeted from two hours of human motion, by the number of clips (inner ring) and the total duration in minutes (outer ring) per motion class. (b) The mean deformation rate of each motion class on the G1 against the mean base speed of its reference clips. The more dynamic a class is, the more it is slowed down to fit the dynamics of the robot, while slow classes are sped up slightly, as the dotted trend line shows.
G1
R1
H1-2
T1
Retargeting error (mm) ↓
Q25
Q50
Q75
Q25
Q50
Q75
Q25
Q50
Q75
Q25
Q50
Q75
GMR [ 4 ]
52.7
113.4
206.5
56.7
107.5
215.1
73.4
148.4
509.7
52.0
116.2
375.2
PHUMA [ 5 ]
70.5
137.3
228.0
46.6
155.7
401.7
74.2
279.6
871.6
64.1
379.4
1627.7
OmniRetarget [ 6 ]
59.3
98.4
153.9
55.7
94.0
150.7
84.4
257.4
748.8
105.0
339.6
1306.2
SoMa-RT [ 22 ]
54.8
103.7
157.6
45.0
83.9
140.9
68.2
151.8
492.6
50.8
132.0
403.7
SoMa-RT + DTMR ( Ours )
48.5
81.4
140.3
45.3
77.6
134.5
71.4
136.3
377.2
50.3
112.8
320.2
TABLE III : Dynamic feasibility of the retargeted motions is compared across baseline methods. The retargeting error is reported in mm at the 25, 50, and 75 % quartiles over clips. Lower is better, and the best value per column is in bold.
G1
R1
H1-2
T1
DTW retargeting error (mm) ↓
Q25
Q50
Q75
Q25
Q50
Q75
Q25
Q50
Q75
Q25
Q50
Q75
GMR [ 4 ]
51.0
102.7
169.5
54.4
96.7
179.8
70.4
135.7
442.0
51.3
107.0
354.7
PHUMA [ 5 ]
66.6
119.0
184.0
45.7
141.8
318.7
73.6
258.5
686.3
63.6
373.3
1553.1
OmniRetarget [ 6 ]
55.6
90.1
139.1
53.6
86.8
131.7
83.0
230.6
594.0
104.4
336.7
1299.4
SoMa-RT [ 22 ]
52.2
93.2
141.5
43.3
75.9
123.6
65.8
141.2
392.0
50.0
117.7
359.5
SoMa-RT + DTMR ( Ours )
45.8
74.1
128.6
43.1
70.4
121.1
69.3
127.4
333.1
49.3
101.2
289.1
TABLE IV : Semantic fidelity of the retargeted motions is compared across baseline methods. To be insensitive to temporal deformation, the DTW retargeting error is reported in mm at the 25, 50, and 75 % quartiles over clips. Lower is better, and the best value per column is in bold.
G1, dynamic
G1, static
Go1
Jump-1
Jump-2
Jump-3
Flip
Dodge
Checking
Praying
Brushing
Catch-1
Catch-2
HopTurn
SideSteps
Error (mm) ↓
STMR [ 8 ]
95.8 (30.4)
85.5 (20.6)
556.0 (241.6)
1213.6 (157.8)
89.3 (10.5)
17.7 (1.5)
13.5 (2.3)
11.6 (1.2)
165.3 (130.5)
22.8 (0.7)
62.7 (8.0)
83.2 (13.4)
DTMR ( Ours )
66.5 (26.5)
40.7 (7.6)
82.2 (17.1)
111.6 (35.2)
75.0 (17.1)
14.3 (2.4)
12.0 (1.9)
12.3 (1.5)
55.3 (1.2)
20.5 (0.9)
57.9 (9.3)
76.7 (9.0)
Error rate (%) ↓
STMR [ 8 ]
6.2 (2.0)
7.3 (1.8)
39.1 (17.0)
16.4 (2.1)
4.2 (0.5)
68.5 (5.6)
34.9 (6.1)
21.0 (2.2)
4.4 (3.5)
3.4 (0.1)
4.7 (0.6)
4.4 (0.7)
TABLE V : Retargeting error of DTMR compared with the temporal retargeting baseline (STMR) [ 8 ] . We measure the DTW retargeting error on ten G1 clips and two Go1 clips. The error is given in mm, and the error rate in % of the reference root path length as mean (sd) over five seeds.
Fig. 3 : DTMR retargets motions with flight phases more successfully than the temporal retargeting baseline (STMR) [ 8 ] . In these clips, (a) STMR skips the jump and (b) fails the landing, whereas DTMR completes both.
Motion set
STMR [ 8 ]
DTMR ( Ours )
speedup
G1 (10 clips, total)
45.4 h
2.4 h
19.2×
per clip, mean
4.5 h
14 min
18.9×
per clip, range
2.0–11.8 h
7–34 min
17.1–20.7×
Go1 (2 clips, total)
2.4 h
7 min
20.8×
per clip, mean
73 min
3.5 min
20.5×
per clip, range
53–94 min
3–4 min
18.9–22.2×
TABLE VI : The wall-clock time to retarget the clips is compared between DTMR and STMR [ 8 ] on the same GPU and tracker.
Fig. 4 : DTMR can control how much the timing of a motion is deformed. A large wdϕ keeps the timing close to the source motion, and a small wdϕ deforms it more so that the robot follows the motion more precisely.
Fig. 5 : DTMR retargets precisely by deforming only the parts of the motion that need a change, on (a) a humanoid and (b) a quadruped. (a) On the G1, the retargeting error decreases as DTMR allows more temporal deformation, and DTMR is more precise than the baseline (STMR) [ 8 ] at the same absolute deformation rate ⟨∣r∣⟩ . (b) On the Go1, DTMR reaches a similar error with a smaller absolute deformation rate.
Fig. 6 : DTMR optimizes the timing densely, while the baseline [ 8 ] scales coarse segments. The source motion (middle) is retimed by the baseline (top) and by DTMR (bottom), and the colour of the curve is the deformation rate.
Fig. 7 : Policies trained on DTMR references are deployed on a Unitree G1 on (a) a contact-rich dance and (b) a dance with fast center-of-mass motion.
Retargeting human kinematic reference motion onto a robot's morphology remains a formidable challenge. Existing methods often produce physical inconsistencies, such as foot sliding, self-collisions, or dynamically infeasible motions, which hinder downstream imitation learning. We propose a bilevel optimization framework that jointly adapts reference motions to a robot's morphology while training a tracking policy using reinforcement learning. To make the optimization tractable, we derive an approximate gradient for the upper-level loss. Our framework requires only a sparse set of semantic rigid-body correspondences and eliminates the need for manual tuning by identifying optimal values for a parameterization expressive enough to preserve characteristic motion across different embodiments. Moreover, by integrating retargeting directly with physics simulation, we produce physically plausible motions that facilitate robust imitation learning. We validate our method in simulation and on hardware, demonstrating challenging motions for morphologies that differ significantly from a human, including retargeting onto a quadruped.
Imitation Learning from monocular video demonstrations provides a scalable approach for teaching complex skills to humanoid robots. However, translating human motion to humanoids requires overcoming significant morphological mismatches. Standard approaches rely on Geometric Retargeting or Indirect Dynamic Retargeting pipelines. We identify that these intermediate kinematic projections introduce a geometric bias, restricting the search space and yielding suboptimal dynamic behaviors. In this paper, we propose Direct Dynamic Retargeting (DDR), a novel single-stage framework that generates high-fidelity, dynamically feasible trajectories directly from expert videos. By formulating the problem in the task space and leveraging a sampling-based Model Predictive Control solver within a physics simulator, DDR natively optimizes over complex contact sequences while mitigating input drift. Our experiments demonstrate that bypassing the geometric bias allows DDR to outperform state-of-the-art baselines in demonstration tracking accuracy. Furthermore, we establish that providing such physically viable references to RL agents accelerates training convergence and enhances the final execution of agile and balancing behaviors. Source code will be made publicly available.
Constant Roux, Ludovic De Matteïs, Armand Jordana +4
LAAS-CNRS, Université de Toulouse, CNRS, Toulouse, France · IRT Saint-Exupéry, Toulouse, France · Artificial and Natural Intelligence Toulouse Institute (ANITI), Toulouse, France
Motion retargeting for specific robot from existing motion datasets is one critical step in transferring motion patterns from human behaviors to and across various robots. However, inconsistencies in topological structure, geometrical parameters as well as joint correspondence make it difficult to handle diverse embodiments with a unified retargeting architecture. In this work, we propose a novel unified graph-conditioned diffusion-based motion generation framework for retargeting reference motions across diverse embodiments. The intrinsic characteristics of heterogeneous embodiments are represented with graph structure that effectively captures topological and geometrical features of different robots. Such a graph-based encoding further allows for knowledge exploitation at the joint level with a customized attention mechanisms developed in this work. For lacking ground truth motions of the desired embodiment, we utilize an energy-based guidance formulated as retargeting losses to train the diffusion model. As one of the first cross-embodiment motion retargeting methods in robotics, our experiments validate that the proposed model can retarget motions across heterogeneous embodiments in a unified manner. Moreover, it demonstrates a certain degree of generalization to both diverse skeletal structures and similar motion patterns.
Zhefeng Cao, Ben Liu, Shunpeng Yang +3
Southern University of Science and Technology, Shenzhen, China · The Hong Kong University of Science and Technology, HongKong, China · LimX Dynamics, Shenzhen, China. +1