Transferring human motion to humanoid robots requires adapting the demonstrated motion to the robot morphology while preserving interactions with the environment. This is particularly challenging for loco-manipulation tasks, where contacts with the ground and manipulated objects must remain consistent despite differences in body proportions. Yet, skeletal motion alone does not fully describe these interactions, and fixing object trajectories limits the adaptation to a new embodiment. In this paper, we introduce OTR ETARGET, a unified approach to jointly retarget robot and multi-object motion from human demonstrations. Our approach represents surface interactions through signed distances, closest surface points, and relative directions, and uses entropic optimal transport to transfer these quantities across human, robot, and object geometries. We incorporate the resulting interaction targets into a constrained inverse kinematics formulation that balances contact preservation with motion style and jointly optimizes robot and object poses at each frame. This formulation accommodates robot-object and object-object interactions without rescaling the scene or the demonstration. We validate the proposed approach on OMOMO, where it achieves a robot- object interaction Jaccard score of 87% and a depth error of 8.7 mm, compared with 28% and 29.3 mm for OmniRetarget. Finally, we demonstrate transfer to a physical G1 humanoid using whole-body policies trained with reinforcement learning on the retargeted references, across motions including two-handed box pick-and-place onto a table.
Figures & tables
Fig. 2 : Pipeline overview. Each probe xc,p reads a proximity triple (d,w,n)c,p from the channel’s signed distance field. A transport plan P⋆ , computed once, transfers this target to the robot. At each frame, a constrained program balances the resulting residuals against the skeleton-style targets while jointly solving the robot pose qt and every object pose Tto . Dashed: the optional object substitution (Sec. III-E ).
Fig. 3 : Hands and knees on the floor ( crawl ). OTRetarget lays both hands flat on the floor as the demonstration does. OmniRetarget reaches the floor with the wrists bent; GMR, which represents no terrain, leaves the hands above it, and PHC drives them through it.
Terrain
Method
sc
rot ↓
self ↓
dd ↓
R ↑
P ↑
J ↑
skate ↓
clear
∘
mm
mm
%
%
%
cm/s
mm
Locomotion - AMASS (2027 sequences)
OTRetarget
1.00
6.3
0.0
1.2
99
100
99
1.8
1.3
OmniRetarget
0.76
25.5
13.7
3.3
95
100
93
2.1
5.1
GMR
0.87
7.5
0.0
23.2
47
100
47
1.3
30.9
TABLE I : Ground fidelity on locomotion (AMASS, CMU and SFU pooled) and loco-manipulation (OMOMO), each method in its released world; dataset medians, sequence counts in the block headers.
Fig. 4 : Illustration of loco-manipulation on OMOMO, OTRetarget (top) against OmniRetarget (bottom): a large convex box, then a floor lamp, a stem the mesh has no volume to hold (left); the same overhead lift in the scaled world, native with the object pinned, and native with it free (right).
Terrain
Object
Method
sc
rot ↓
R ↑
R ↑
P ↑
J ↑
drift
∘
%
%
%
%
cm
OMOMO (4421 sequences)
OTRetarget (scaled, obj. fixed)
0.75
11.1
100
97
89
82
—
native scene, obj. fixed
1.00
12.0
98
94
96
84
0.0
native scene, obj. variable
1.00
10.7
99
95
96
87
2.9
TABLE II : The object as a decision variable. The scene solved as captured (ablation: object fixed vs. variable); drift is blank on scaled rows, where it would measure the world’s scale, not the method.
Object
Method
sc
dd ↓
dw ↓
R ↑
P ↑
J ↑
mm
mm
%
%
%
OMOMO (4421 sequences)
OTRetarget
1.00
8.7
37
95
96
87
OTRetarget (scaled, obj. fixed)
0.75
8.8
38
97
89
82
OmniRetarget
0.75
29.3
102
39
62
28
TABLE III : Robot–object proximity on the OMOMO dataset (medians) and on two single captures, a large convex box and a floor lamp; the marked row solves in OmniRetarget’s scaled world with the object fixed.
Object
Variant
dd ↓
R ↑
P ↑
J ↑
drift
mm
%
%
%
cm
box × 1.0 (native)
10.0
94
89
85
9.9
ball ∅ 0.34
12.1
97
87
84
12.4
drum ∅ 0.34 × 0.36
12.2
95
88
84
11.1
capsule ∅ 0.23 × 0.36
11.8
96
88
84
13.1
TABLE IV : One capture, multiple unseen objects: the capture never touched ( largebox, OMOMO, sub3_003 ), every row is obtained with OTRetarget
Learning from demonstration (LfD) has enabled humanoid robots to acquire diverse whole-body skills, but extending this paradigm to human-object interaction (HOI) is limited by the availability of robot-compatible interaction references. We present HOI-Retarget, a contact-centric retargeting method that transfers HOI onto a humanoid robot for large-scale motion-data generation. Its windowed trajectory optimization uses every labeled contact as a target in the object frame, balancing body tracking, foot support and smoothness under the robot's kinematic limits. The method can augment a single demonstration across object sizes, absorb contacts reconstructed from monocular video, and extend to several robots manipulating one object. We publicly release the code and the retargeted motion dataset.
Jihwan Shin, Adrià López Escoriza, Junzhe He +2
Robotic Systems Lab, ETH Zürich, 8092 Zürich, Switzerland
Learning robot dexterous manipulation from human manipulation videos requires reliably retargeting human intent to executable robot actions while maintaining stable hand-object contact, which remains a key challenge in embodied intelligence. Existing retargeting methods often ignore explicit contact modeling or rely on reinforcement learning, resulting in limited accuracy and generalization. To address this, we propose ObjRetarget, a human-to-robot motion retargeting framework for learning robot dexterous manipulation from human videos, which integrates anthropomorphic arm trajectory constraints with structured hand-object geometric modeling. For arm motion, reference trajectories extracted from human videos are used for initialization, followed by anthropomorphic constraints and redundancy-aware optimization to generate natural and accurate movements. For hand manipulation, ObjRetarget represents multi-finger contacts using polytope clusters and preserves contact structure through geometric invariants to improve stability. Experiments on real robots show that ObjRetarget improves manipulation success rates and contact stability across multiple dexterous tasks, and generalizes well to different demonstrations, object poses, and task settings.
Yuanchuan Lai, Qing Gao, Ziyan Liang +2
School of Electronics and Communication Engineering, Sun Yat-sen University, Shenzhen 518107, China. · School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, Shenzhen 518172, China. · School of Computing, University of Portsmouth, Portsmouth PO1 3HE, UK.
Retargeting human kinematic reference motion onto a robot's morphology remains a formidable challenge. Existing methods often produce physical inconsistencies, such as foot sliding, self-collisions, or dynamically infeasible motions, which hinder downstream imitation learning. We propose a bilevel optimization framework that jointly adapts reference motions to a robot's morphology while training a tracking policy using reinforcement learning. To make the optimization tractable, we derive an approximate gradient for the upper-level loss. Our framework requires only a sparse set of semantic rigid-body correspondences and eliminates the need for manual tuning by identifying optimal values for a parameterization expressive enough to preserve characteristic motion across different embodiments. Moreover, by integrating retargeting directly with physics simulation, we produce physically plausible motions that facilitate robust imitation learning. We validate our method in simulation and on hardware, demonstrating challenging motions for morphologies that differ significantly from a human, including retargeting onto a quadruped.