Learning from demonstration (LfD) has enabled humanoid robots to acquire diverse whole-body skills, but extending this paradigm to human-object interaction (HOI) is limited by the availability of robot-compatible interaction references. We present HOI-Retarget, a contact-centric retargeting method that transfers HOI onto a humanoid robot for large-scale motion-data generation. Its windowed trajectory optimization uses every labeled contact as a target in the object frame, balancing body tracking, foot support and smoothness under the robot's kinematic limits. The method can augment a single demonstration across object sizes, absorb contacts reconstructed from monocular video, and extend to several robots manipulating one object. We publicly release the code and the retargeted motion dataset.
Figures & tables
Fig. 2: Overview of HOI-Retarget . (a) The source is a captured or video-reconstructed HOI clip, providing human motion, an object trajectory and contact labels. (b) IK retargeting maps the human onto the robot, while the object mesh and trajectory are scaled by the robot-to-human height ratio, carrying the contact targets with them. (c) A windowed trajectory optimization with tracking, contact and smoothness costs recovers those contacts under the robot’s kinematic limits. (d) The result drives downstream policies directly, or after dynamic refinement in simulation.
Method
Source Human Motion
Object Interaction
Contact Location Preservation
Data Augmentation
Multi-Agent Source
Temporal Coupling
Dynamics Enforcement
Optimization Method
GMR [ 11 ]
✓
✗
✗
✗
✗
✗
✗
per-frame IK
OmniRetarget [ 13 ]
✓
✓
✗
✓
✗
✗
✗
per-frame SOCP
DynaRetarget [ 14 ]
✗
✓
✗
✓
✗
✓
✓
progressive SBTO
SPIDER [ 15 ]
✗
✓
✗
✓
✗
✓
✓
receding-horizon sampling
HOI-Retarget (ours)
✓
✓
✓
✓
✓
✓
✗
windowed NLP
TABLE I: Capabilities of representative retargeting methods.
Cost term
Symbol
Definition
w
Tracking
ET
Eθ+Eb+Eh+Ef+Er
Joint position
Eθ
∥θ−θ∗∥2
1
Base position
Eb
∥pbw−pbw∗∥2
1
Head position
Eh
∥pheadw−pheadw∗∥2
1
Feet position
Ef
∑i∈F∥pc,iw−pc,iw∗∥2
100
Torso orientation
Er
∥Rtorsob−Rtorsob∗∥F2
1
TABLE II: Cost terms of ( 2 ), grouped as in ( 2a ).
Fig. 3: Representative retargets across HOI sources and embodiments. Rows show the source SMPL-X motion, Unitree G1, and Unitree H2.
Metric
GMR
OmniRetarget
Ours
Body pose deviation ( ∘ ) ↓
20.0
25.0
21.0
Link vel. direction ( ∘ ) ↓
19.4
28.9
25.2
Contact-point gap (m) ↓
0.368
0.183
0.005
Rel. hand-orient. change ( ∘ ) ↓
2.3
37.6
6.1
Body jerk (m/s 3 ) ↓
66.7
193.5
49.9
Compute time (s/clip) ↓
9.1
156.2
33.7
TABLE III: Kinematic retargeting on the common OMOMO subset (3,997 clips; 3,604 for contact metrics).
Fig. 4: Qualitative comparison with OmniRetarget. HOI-Retarget recovers the labeled source contact regions on the coat rack and chair.
Fig. 5: Contact-preserving object-scale augmentation over the tested range ×0.25 – ×1.50 on the G1 and H2.
RL Tracker
SBTO
Metric
Omni.
Ours
Omni.
Ours
Body pose deviation ( ∘ ) ↓
26.5
23.5
26.0
23.1
Link vel. direction ( ∘ ) ↓
47.9
46.7
51.2
50.8∗
Object path err. (m) ↓
0.285
0.213
0.269
0.194
Contact-point gap (m) ↓
0.096
0.086
0.141
0.099
Rel. hand-orient. change ( ∘ ) ↓
31.5
25.3
38.4
28.1
TABLE IV: Dynamic-refinement quality on the common subset (68 clips; 63 for contact metrics).
Fig. 6: Example of using HOI-Retarget from video reconstruction using CARI4D. Contact refinement tool has been used to account for monocular reconstruction artifacts.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Description
(⋅)∗
Reference quantity
(⋅)h
Human source quantity
(⋅)w,(⋅)b,(⋅)o
World, base, object frame
δ(⋅)
Backward difference,
δ(⋅)t=(⋅)t−(⋅)t−1
q
Configuration (pbw,Rbw,θ)
Appendix
TABLE V: Symbols used for HOI-Retarget.
Metric
Definition
Body pose deviation ( ∘ )
Mean over frames of the mean angle between the robot’s and the human’s 16 bone directions, taken in the pelvis frame for the legs and the trunk frame for the arms.
Link vel. direction ( ∘ )
Mean angle between the robot’s and the human’s object-relative link velocity Row⊤(p˙w−p˙ow) , over the palms, ankles, head and pelvis, where both exceed 2 cm/s.
Contact-point gap (m)
Mean distance on the object between the robot’s palm and the human’s contact point ( 1 ), over frames in contact. This is the residual Ec minimizes.
Rel. hand-orient. change ( ∘ )
Mean angle between the robot’s and the human’s change in palm orientation in the object frame since the first frame of the contact segment.
Object path err. (m)
Mean distance between the robot-side and the human-side object translation, each taken relative to its own first frame.
Body jerk (m/s 3 )
Mean magnitude of the third difference of world link position, over links and frames.
Appendix
TABLE VI: Evaluation metrics reported in Tables III and IV .
Horizon
Coll.
Solved
Time (s)
RAM (MB)
Link vel. dir. ( ∘ )
Joint jerk (rad/s 3 )
Per-frame
✗
130/130
37.7
1440
40.3
588
Windowed
✗
130/130
39.4
1763
24.4
251
✓
130/130
159.9
3876
24.5
257
Full traj.
✗
130/130
45.8
2400
23.5
185
✓
87/130
210.5
8671
24.4
186
Appendix
TABLE VII: Optimization horizon, with and without the collision cost.
Learning humanoid-object interaction requires coordinating whole-body balance, locomotion, and dexterous hand contact to control both robot and object motion. Human demonstrations provide examples of coordinated interaction, but transferring these behaviors to humanoid robots requires learning how to establish and maintain effective contacts under different embodiments and dynamics. We present Weave, a unified framework for learning whole-body dexterous humanoid-object interaction from captured human demonstrations. Weave first converts captured human-object interactions into executable robot-object references through contact-aware retargeting and approach-motion completion. At its core is a contact- and geometry-aware policy that jointly commands 29 body joints and 12 actuated finger joints across multiple objects and interaction sequences. Evaluation across nine objects yields a 92.5% success rate on trained interactions and, without any additional training, 65.0% on sequences never seen during training. We additionally release ~9,000 physically executed rollouts spanning ~23 hours, providing robot-object trajectories with contact annotations for downstream interaction-policy learning and physically consistent HOI motion generation. Project website: https://xiaohu-art.github.io/Weave/
Liu Cao, Xingze Wu, Jingzhi Cui +4
1Tsinghua IIIS · 3Dalian University of Technology · 4The Chinese University of Hong Kong +1
Transferring human motion to humanoid robots requires adapting the demonstrated motion to the robot morphology while preserving interactions with the environment. This is particularly challenging for loco-manipulation tasks, where contacts with the ground and manipulated objects must remain consistent despite differences in body proportions. Yet, skeletal motion alone does not fully describe these interactions, and fixing object trajectories limits the adaptation to a new embodiment. In this paper, we introduce OTR ETARGET, a unified approach to jointly retarget robot and multi-object motion from human demonstrations. Our approach represents surface interactions through signed distances, closest surface points, and relative directions, and uses entropic optimal transport to transfer these quantities across human, robot, and object geometries. We incorporate the resulting interaction targets into a constrained inverse kinematics formulation that balances contact preservation with motion style and jointly optimizes robot and object poses at each frame. This formulation accommodates robot-object and object-object interactions without rescaling the scene or the demonstration. We validate the proposed approach on OMOMO, where it achieves a robot- object interaction Jaccard score of 87% and a depth error of 8.7 mm, compared with 28% and 29.3 mm for OmniRetarget. Finally, we demonstrate transfer to a physical G1 humanoid using whole-body policies trained with reinforcement learning on the retargeted references, across motions including two-handed box pick-and-place onto a table.
Humanoid-Object Interaction (HOI) is a fundamental capability for humanoid robots, yet it remains challenging due to the tight coupling between dynamic balance and stable interaction with diverse objects. Existing methods often require time-consuming task-specific policy training or rely on rigid trajectory replay, which limits their ability to accommodate novel interaction scenarios. In this work, we present \textit{GenHOI}, a simple yet effective framework that enables humanoid robots to perform diverse object-interaction tasks in a zero-shot manner by directly imitating a single generated video, without task-specific training or physical demonstration data. GenHOI first reconstructs the robot-object scene in simulation and renders a first-frame image, which, together with the language command, conditions the synthesis of a task-oriented interaction video. The generated video is then analyzed to identify interaction-relevant contact events and estimate hand-object contact regions, which are encoded as object-centric geometric constraints that convert visual interaction cues into physically grounded optimization priors. Guided by these priors, the reference motion recovered from the video is refined and smoothed to resolve the scale ambiguity inherent in 2D video generation, while adapting a single reference trajectory to unseen robot-object relative poses. The optimized trajectory is finally executed by a closed-loop tracking controller. We validate the proposed framework in extensive simulation and real-world experiments across diverse object-interaction tasks, including box grasping, asymmetric bimanual chair carrying, table lifting from below, and cylindrical-object enveloping.
Zhihai Bi, Qiang Zhang, Guoyang Zhao +8
The Hong Kong University of Science and Technology (Guangzhou) · Artificial General Intelligence Institute, University of Science and Technology of China · The University of Hong Kong +1