Partnered human-humanoid interaction couples locomotion with continuous physical contact. A humanoid needs to coordinate with a person's motion while responding to interaction forces and maintaining stable and natural movement. We present CoDance, a framework for learning reactive and compliant human-humanoid interaction from video. We study partnered dancing as a challenging instantiation, where a humanoid coordinates its footsteps with a moving partner and maintains continuous two-hand contact. Given a single video of two human dancers, CoDance retargets their motions into a robot reference and a moving partner. We introduce a multi-link compliance augmentation that transforms the kinematic demonstration into force-aware training data by adapting the robot reference under structured forces at both hands. Policies trained on this data follow the observed partner while preserving the demonstrated locomotion style and responding compliantly to physical interaction. In simulation, the policies adapt their footsteps to changes in the partner and reproduce approximately 80% of the wrist displacement encoded by the augmented demonstrations. On a physical humanoid, CoDance enables sustained two-hand dancing with a human partner including repeated transitions between forward and backward motions.
Figures & tables
Fig. 2 : Overview of CoDance . A paired-dancing video is converted into force-aware training data through motion retargeting and multi-link compliance augmentation. The policy learns to follow the observed partner, maintain the demonstrated locomotion style, and remain compliant at both hands. At deployment, the simulated partner is replaced by a person. Training is illustrated with the decoupled policy.
Parameter
Value
Robot stiffness K
log-uniform, 101000Nm−1
Environment stiffness
same range, independent draw
Force ceiling
140N
Displacement scale dmax
0.15m
per-event draw
U(0.03,dmax) m
Balance budget (CoM shift)
0.10m
TABLE I : Augmentation parameters.
Fig. 3 : Evaluation results on compliance interaction. Under force, both compliant policies move the wrist most of the distance encoded in the adapted reference, while the stiff tracker moves it far less. Error bars show the standard deviation over evaluation seeds.
Fig. 4 : Visualization of the free and adapted references (top) and of the decoupled policy with and without the adversarial term, at the middle of four force holds. The partner is in teal, and arrows mark the force at the left wrist (blue) and right wrist (pink), with the applied force solid and the scheduled force faded. Without the adversarial term, the pelvis turns away from the dancer’s heading, the waist and hips twist back toward the direction of travel, and the stance widens into a crouch. Top row also shows that separate retargeting does not preserve relative position in the paired motion clip.
Fig. 5 : Visualization of one real-robot dance with a human partner. The robot keeps close to the following distance for most of the dance. Top: the trajectories of the robot and the partner. Bottom: the partner’s root in the robot’s pelvis frame.
Fig. 6 : Visualization of a standing policy under force on one or both hands. The red overlay is the reference standing pose. Left: one hand loaded by a hanging weight. Middle: the two hands pulled in different directions. Right: both hands pulled outward. The arms move with the pull and the feet stay still.
TABLE V : Reward terms of the two trained policies. Tracking terms are exponential kernels exp(−e2/σ2) of the error e with the listed σ ; weights multiply the per-step term. The two policies share every term except the locomotion objective: the decoupled policy draws its locomotion style from the adversarial term, the whole-body tracker (in the style of SoftMimic) from reference tracking of the lower body of the adapted reference.
Term
Per step
History
Actor noise
Actor
Critic
Base linear velocity (IMU)
3
5
✓
Base angular velocity (IMU)
3
5
±0.2
✓
✓
Joint positions
29
5
±0.01
✓
✓
Joint velocities
29
5
±0.5
✓
✓
Previous actions
29
5
✓
✓
Projected gravity
3
5
±0.05
✓
✓
TABLE VI : Observations of the two trained policies, identical for both. Per-step sizes are stacked over the listed history. Noise is uniform in the listed range and applied to the actor only; the actor’s joint positions also carry the per-joint encoder bias of Table VII . Both networks are MLPs with hidden layers 512, 256, 128 and ELU activations. The AMP discriminator of the decoupled policy reads 2 frames of 63 features (positions and 6D orientations of 7 lower-body and torso links in the pelvis frame), which the policy itself does not see.
Setting
Value
Randomization (at episode start)
Torso centre of mass offset
uniform, x±0.025m , y±0.05m , z±0.05m
Foot friction coefficient
uniform, 0.31.2 , one draw for all foot geoms
Joint encoder bias
uniform, ±0.01rad per joint, added to the actor’s joint positions
Actor observation noise
uniform, per observation term, at the ranges the policy trained with
Reset root pose
x , y±0.15m , yaw ±0.3rad about the clip pose
TABLE VII : Randomization, force replay and episode settings, identical for both policies.
Learning humanoid-object interaction requires coordinating whole-body balance, locomotion, and dexterous hand contact to control both robot and object motion. Human demonstrations provide examples of coordinated interaction, but transferring these behaviors to humanoid robots requires learning how to establish and maintain effective contacts under different embodiments and dynamics. We present Weave, a unified framework for learning whole-body dexterous humanoid-object interaction from captured human demonstrations. Weave first converts captured human-object interactions into executable robot-object references through contact-aware retargeting and approach-motion completion. At its core is a contact- and geometry-aware policy that jointly commands 29 body joints and 12 actuated finger joints across multiple objects and interaction sequences. Evaluation across nine objects yields a 92.5% success rate on trained interactions and, without any additional training, 65.0% on sequences never seen during training. We additionally release ~9,000 physically executed rollouts spanning ~23 hours, providing robot-object trajectories with contact annotations for downstream interaction-policy learning and physically consistent HOI motion generation. Project website: https://xiaohu-art.github.io/Weave/
Liu Cao, Xingze Wu, Jingzhi Cui +4
1Tsinghua IIIS · 3Dalian University of Technology · 4The Chinese University of Hong Kong +1
Human videos offer scalable manipulation data, but the embodiment gap between human hands and robot manipulators limits their direct use. Existing video-editing methods replace hands with rendered robots, yet inaccurate interaction reconstruction and compositing can produce inconsistent grasps and implausible robot-object occlusions. We address these failures from two complementary physical aspects: interaction geometry and scene visibility. First, an interaction-aware contact reconstruction module combines hand-object segmentation with mesh-level contact prediction to recover dense 3D contacts, then converts them into temporally stabilized grasps for parallel-jaw grippers. Second, a depth-aware compositing module uses scene and robot depth to enforce physically consistent robot-object occlusions. The resulting videos preserve the interaction structure of human demonstrations in a robot-compatible form and are co-trained with robot demonstrations. Using identical human videos and robot data, we compare against robot-only training and the original Masquerade pipeline. Across four RoboTwin tasks and two Diffusion Policy visual encoders, our method achieves the highest average success rates, with especially strong gains under out-of-distribution scene variation. Real-world deployment further shows that the proposed co-training approach improves robustness to visual distractors when the task geometry is observable, while performance on depth-sensitive grasps remains limited by the single-camera setup.
Learning from demonstration (LfD) has enabled humanoid robots to acquire diverse whole-body skills, but extending this paradigm to human-object interaction (HOI) is limited by the availability of robot-compatible interaction references. We present HOI-Retarget, a contact-centric retargeting method that transfers HOI onto a humanoid robot for large-scale motion-data generation. Its windowed trajectory optimization uses every labeled contact as a target in the object frame, balancing body tracking, foot support and smoothness under the robot's kinematic limits. The method can augment a single demonstration across object sizes, absorb contacts reconstructed from monocular video, and extend to several robots manipulating one object. We publicly release the code and the retargeted motion dataset.
Jihwan Shin, Adrià López Escoriza, Junzhe He +2
Robotic Systems Lab, ETH Zürich, 8092 Zürich, Switzerland