Transferring human motion to humanoid robots requires adapting the demonstrated motion to the robot morphology while preserving interactions with the environment. This is particularly challenging for loco-manipulation tasks, where contacts with the ground and manipulated objects must remain consistent despite differences in body proportions. Yet, skeletal motion alone does not fully describe these interactions, and fixing object trajectories limits the adaptation to a new embodiment. In this paper, we introduce OTR ETARGET, a unified approach to jointly retarget robot and multi-object motion from human demonstrations. Our approach represents surface interactions through signed distances, closest surface points, and relative directions, and uses entropic optimal transport to transfer these quantities across human, robot, and object geometries. We incorporate the resulting interaction targets into a constrained inverse kinematics formulation that balances contact preservation with motion style and jointly optimizes robot and object poses at each frame. This formulation accommodates robot-object and object-object interactions without rescaling the scene or the demonstration. We validate the proposed approach on OMOMO, where it achieves a robot- object interaction Jaccard score of 87% and a depth error of 8.7 mm, compared with 28% and 29.3 mm for OmniRetarget. Finally, we demonstrate transfer to a physical G1 humanoid using whole-body policies trained with reinforcement learning on the retargeted references, across motions including two-handed box pick-and-place onto a table.
Figures & tables
Fig. 2 : Pipeline overview. Each probe xc,p reads a proximity triple (d,w,n)c,p from the channel’s signed distance field. A transport plan P⋆ , computed once, transfers this target to the robot. At each frame, a constrained program balances the resulting residuals against the skeleton-style targets while jointly solving the robot pose qt and every object pose Tto . Dashed: the optional object substitution (Sec. III-E ).
Fig. 3 : Hands and knees on the floor ( crawl ). OTRetarget lays both hands flat on the floor as the demonstration does. OmniRetarget reaches the floor with the wrists bent; GMR, which represents no terrain, leaves the hands above it, and PHC drives them through it.
Terrain
Method
sc
rot ↓
self ↓
dd ↓
R ↑
P ↑
J ↑
skate ↓
clear
∘
mm
mm
%
%
%
cm/s
mm
Locomotion - AMASS (2027 sequences)
OTRetarget
1.00
6.3
0.0
1.2
99
100
99
1.8
1.3
OmniRetarget
0.76
25.5
13.7
3.3
95
100
93
2.1
5.1
GMR
0.87
7.5
0.0
23.2
47
100
47
1.3
30.9
TABLE I : Ground fidelity on locomotion (AMASS, CMU and SFU pooled) and loco-manipulation (OMOMO), each method in its released world; dataset medians, sequence counts in the block headers.
Fig. 4 : Illustration of loco-manipulation on OMOMO, OTRetarget (top) against OmniRetarget (bottom): a large convex box, then a floor lamp, a stem the mesh has no volume to hold (left); the same overhead lift in the scaled world, native with the object pinned, and native with it free (right).
Terrain
Object
Method
sc
rot ↓
R ↑
R ↑
P ↑
J ↑
drift
∘
%
%
%
%
cm
OMOMO (4421 sequences)
OTRetarget (scaled, obj. fixed)
0.75
11.1
100
97
89
82
—
native scene, obj. fixed
1.00
12.0
98
94
96
84
0.0
native scene, obj. variable
1.00
10.7
99
95
96
87
2.9
TABLE II : The object as a decision variable. The scene solved as captured (ablation: object fixed vs. variable); drift is blank on scaled rows, where it would measure the world’s scale, not the method.
Object
Method
sc
dd ↓
dw ↓
R ↑
P ↑
J ↑
mm
mm
%
%
%
OMOMO (4421 sequences)
OTRetarget
1.00
8.7
37
95
96
87
OTRetarget (scaled, obj. fixed)
0.75
8.8
38
97
89
82
OmniRetarget
0.75
29.3
102
39
62
28
TABLE III : Robot–object proximity on the OMOMO dataset (medians) and on two single captures, a large convex box and a floor lamp; the marked row solves in OmniRetarget’s scaled world with the object fixed.
Object
Variant
dd ↓
R ↑
P ↑
J ↑
drift
mm
%
%
%
cm
box × 1.0 (native)
10.0
94
89
85
9.9
ball ∅ 0.34
12.1
97
87
84
12.4
drum ∅ 0.34 × 0.36
12.2
95
88
84
11.1
capsule ∅ 0.23 × 0.36
11.8
96
88
84
13.1
TABLE IV : One capture, multiple unseen objects: the capture never touched ( largebox, OMOMO, sub3_003 ), every row is obtained with OTRetarget
School of Electronics and Communication Engineering, Sun Yat-sen University, Shenzhen 518107, China. · School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, Shenzhen 518172, China. · School of Computing, University of Portsmouth, Portsmouth PO1 3HE, UK.