Captured human-object interactions provide rich supervision for humanoid loco-manipulation, but they are sparse, heterogeneous, and not directly executable by robots. We introduce InterMimicGen, a self-evolving motion-imitation framework in which robot motion data and a tracking policy improve each other. First, we consolidate motion-captured human-object interaction datasets and retarget them into humanoid robot references while preserving whole-body coordination and dexterous hand-object relationships. This produces a large and diverse humanoid robot reference collection for dexterous whole-body loco-manipulation. Second, we train a physics-based generalist tracker that executes these references in simulation on a humanoid with dexterous hands, covering a scale and diversity beyond prior humanoid tracking systems for loco-manipulation. Third, we close a data flywheel: each round makes small, task-preserving changes to where an interaction takes place and how the body performs it, fine-tunes the tracker on them, and keeps only the variants whose simulated execution completes the task, which seed the next round. With more iterations, these small edits compound into broader coverage around the sparse original demonstrations while preserving task semantics and motion quality. Experiments show contact-preserving retargeting across robot configurations, broad tracking with a single generalist policy, executable motions that keep growing over augmentation rounds, and transfer to real robots. InterMimicGen provides a unified path from heterogeneous human demonstrations to a continually expanding motion resource for humanoid robot learning.
Figures & tables
Figure 1: InterMimicGen grows a finite set of human interaction captures into an expanding repertoire of executable robot motions. Left: One formulation turns the captures into executable motions across diverse objects, tasks, and robots, from humanoids to a wheeled mobile manipulator. Top right: Self-evolving motion imitation. A single motion grows round by round into many variants. Bottom right: The resulting motions run on real robots.
Figure 2: Top: Reference construction. Human-object motion capture is retargeted into whole-body dexterous references, which a physics-based tracker imitates to produce seed rollouts. Bottom: Self-evolving motion imitation. Each round augments verified references into candidates, fine-tunes the tracker on them, executes the candidates in physics, and retains only successful rollouts as seeds for the next round; Bottom right: UMAP embeddings (Sec. E.1 ) of accepted augmentations expanding.
Figure 3: Our data has broad coverage. (a-c) UMAP embeddings of human motion, object geometry, and human-object interaction of our collection, colored by source dataset. (d) How well our collection covers external motions from GRAIL ( Xie et al., 2026 ) , MOVIN ( Jang et al., 2023 ) , and ELMO ( Jang et al., 2024 ) ; a motion is covered when one of our sequences is close to it (Sec. E.1 ). (e) The same coverage as the number of our source sequences grows.
Hand contact
Penetration
Motion
Jnt. acc. (rad/s 2 ) ↓
Jnt. limit (%) ↓
Frame (%) ↑
Hand (%) ↑
Slip (m/s) ↓
Obj. (%) ↓
Depth (cm) ↓
Gnd. (%) ↓
Foot sl. (%) ↓
MPJPE (cm) ↓
EE (cm) ↓
All
Body
Fing.
Body
Fing.
Weave
96.1
97.1
0.225
67.4
2.62
31.9
62.0
9.54
12.6
12.2
6.3
19.3
0.6
33.4
Ours
98.2
98.8
0.239
37.2
2.39
0.0
12.3
5.74
5.3
2.8
3.2
2.3
0.0
0.0
Table 1: Quantitative comparison on retargeting on G1 with Inspire hands, on the 3,059 OMOMO clips. Ours shows better motion quality. Metrics are defined in Sec. E.2 .
Figure 4: Qualitative evaluation on retargeting. Captured human motion (top) and our G1 with Inspire hands reference (bottom), across interactions from different datasets.
Bimanual
Sitting
Grasping
Embodiment
SR ↑
MPJPE ↓
Obj. ↓
Rot. ↓
SR ↑
MPJPE ↓
Obj. ↓
Rot. ↓
SR ↑
MPJPE ↓
Obj. ↓
Rot. ↓
Specialists (one policy per object)
G1 + Inspire
93.8
5.19
5.79
10.4
75.4
5.94
4.76
7.4
84.9
5.30
7.17
14.1
G1 + Dex3
93.1
4.92
7.02
11.0
82.8
5.36
5.01
6.5
83.9
5.14
7.32
12.7
G1
94.4
4.52
5.88
8.5
82.8
4.93
5.26
6.2
–
–
–
–
K1
75.7
6.21
7.35
14.3
75.4
6.41
5.08
8.3
–
–
–
–
Table 2: Quantitative results on motion tracking. One generalist nearly matches the per-object specialists in body and object position error, while trailing them in success rate and object rotation. Metrics are defined in Sec. E.3 ; – marks tasks the embodiment cannot perform.
Figure 5: Physics-based tracking across objects. Executed rollouts on G1 with Inspire hands for a diverse set of objects.
Figure 6: A single capture grows into many executable variants. Each group shows the original reference followed by the verified augmentations accumulated by round 1 and by round 5.
Growth (×)
Success (%)
Quality, Orig. → Aug.
Generated
Original
Body acc.
Foot sl.
Hand jit.
Setting
R1 → R5
Orig.
Evol.
Orig.
Evol.
(rad/s 2 ) ↓
(%) ↓
(mm) ↓
Inspire, bimanual
21.8 → 150.5
52.1
98.4
100
100
11.4 → 12.4
1.18 → 1.87
1.00 → 1.17
Dex3, bimanual
21.3 → 142.0
59.3
98.7
100
100
13.9 → 16.0
1.14 → 2.21
1.12 → 1.30
Inspire, grasping
19.6 → 146.4
64.2
98.9
100
100
13.6 → 15.3
0.70 → 2.05
0.96 → 1.14
Table 3: Over five rounds, the verified references grow more than a hundredfold , and the evolved tracker executes nearly all of them without losing the original ones. Growth: verified references relative to the original references. Success: rate at which the tracker before (Orig.) and after (Evol.) self-evolution. Quality: executions of original (Orig.) and augmented (Aug.) references.
Figure 7: Real-world execution. We show the motion from our paradigm transfer to different embodiments, from Unitree G1 pulling a suitcase and G1 with Inspire hands carrying a tripod; Booster K1 lifting and placing a box and Dexmate Vega relocating a chair.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Source
Object entries
Motions
Reference h
InterAct ( Xu et al., 2025c )
117
9,030
17.82
HiPHI ( Ji et al., 2026 )
40
7,029
122.86
Consolidated
157
16,059
140.68
Appendix
Table 4: Reference collection for G1 with Inspire hands. Object entries are dataset-object pairs, so an object that appears in two source datasets counts twice. Reference hours sum the frames of the selected robot references at 50 fps and may include overlapping source segments.
Platform
Retargeting
Specialist tracking
Generalist
Augmentation loop
Sim-to- real
Excluded tasks
G1
✓
✓
–
–
✓
dexterous
G1 + Inspire
✓
✓
✓
✓
✓
none
G1 + Dex3
✓
✓
–
✓
–
none
Booster T1
✓
–
–
–
–
dexterous
Booster K1
✓
✓
–
–
✓
dexterous
Dexmate Vega
✓
–
–
–
✓
dexterous, seated
Appendix
Table 5: Studies run on each platform. Configurations with dexterous hands use all data, including the dexterous benchmark; the generalist is trained on G1 with Inspire hands.
Table 13
Setting
Value
Reference roster
16,059 motions
Parallelization
4 GPUs, 4,096 environments per GPU
Control / physics
50 Hz control, 4 physics substeps per action
PPO rollout / minibatch / epochs
32 steps / 16,384 samples / 6 epochs
Actor / critic
Separate ReLU MLPs, 1024→1024→512
Learning rate / discount / GAE
2×10−5 / 0.99 / 0.95
Appendix
Table 8: Generalist tracker settings.
Object edits
Body edits
Checked before physics
finite values
finite, object unchanged
Tracker used
fine-tuned first
current
Rollouts needed to pass
2 of 3
1 of 3
Completes the reference
✓
✓
Terminal hand contact
✓
✓ (held tasks)
No sustained fall
✓
✓
Appendix
Table 9: Acceptance rules for the two edit types. ✓: checked; –: not checked.
Figure 8: More executable variants. As in Figure 6 , for nine further captures: each group shows the original reference followed by the verified augmentations accumulated by round 1 and by round 5.
Figure 9: More retargeted interactions on G1 with Inspire hands.
Figure 10: Retargeting with dexterous hands. A bowl-holding interaction retargeted to G1 with Dex3 (left) and Inspire (right) hands.
Figure 11: Retargeting across embodiments. One suitcase interaction retargeted to G1 with Dex3, Inspire, and non-dexterous hands (left, top to bottom) and to Booster K1, Booster T1, and Dexmate Vega (right, top to bottom).
Figure 12: Retargeting comparison with OmniRetarget. OmniRetarget to G1 with rigid hands (left) and our retargeter for G1 with Inspire hands (right), at corresponding interaction phases. The dexterous hand keeps the grasp and the hand-object relationship of the demonstration.
Figure 13: Retargeting comparison with UMR. UMR to G1 (left) and our retargeter to G1 with Inspire hands (right), at corresponding interaction phases.
Metric
OmniRetarget
UMR
ULTRA
Ours
Contact preservation
Hand contact preservation (%) ↑
79.09
95.13
79.06
83.45
Penetration ( >1 cm)
Robot–object depth (cm) ↓
1.27
1.60
1.49
1.97
Foot sliding (0.15 m/s)
Sliding frames (%) ↓
0.00
7.02
N/A
7.33
Appendix
Table 10: Retargeting on G1 with non-dexterous hands , over the 537 OMOMO clips shared by all four methods. A dash denotes no qualifying event and N/A an unavailable metric; bold marks the unique best.
Figure 14: Physics-based tracking across embodiments. A box-lifting reference executed by specialist trackers on Booster K1 (top left), G1 with non-dexterous hands (top right), G1 with Inspire hands (bottom left), and G1 with Dex3 hands (bottom right).
Yield (%)
Growth (×)
Round
Frozen
Tuned
Frozen
Tuned
1
29.4
25.3
163.0
188.5
2
0.4
72.9
170.5
426.0
3
0.4
68.9
184.9
573.0
Appendix
Table 11: Frozen versus fine-tuned tracker in a separate run on a subset of the bimanual seeds with translation edits, whose counts are not directly comparable with Table 3 . Yield and growth follow Sec. E.4 . The frozen run stops after round 3.
Figure 15: Physics-based tracking with dexterous hands. A tripod grasp executed on G1 with Dex3 (left) and Inspire (right) hands.
Inspire, bimanual
Dex3, bimanual
Inspire, grasping
Round
Growth
Acc.
Foot
Hand
Growth
Acc.
Foot
Hand
Growth
Acc.
Foot
Hand
Orig.
1.0
11.4
1.18
1.00
1.0
13.9
1.14
1.12
1.0
13.6
0.70
0.96
1
21.8
11.8
1.39
1.07
21.3
15.6
1.59
1.25
19.6
15.3
0.99
1.07
2
34.8
11.3
1.39
1.09
31.6
14.8
1.77
1.23
33.3
14.4
1.42
1.09
3
80.4
12.1
1.63
1.16
73.3
14.9
1.87
1.24
77.8
14.6
1.78
1.08
4
118.1
12.9
2.14
1.22
110.8
16.4
2.41
1.33
112.8
15.7
2.42
1.16
Appendix
Table 12: Self-evolution by round , extending Table 3 . Growth (×): cumulative verified references relative to the original references. Quality of the original references (Orig.) and of the references newly accepted in each round: body joint acceleration (Acc., rad/s 2 ), foot sliding (Foot, %), and hand jitter (Hand, mm).
Figure 16: Body-motion variation around shared seeds. Each reference is embedded by the DCT coefficients of its body joint angles, taken relative to its seed and projected with PCA, so all seeds sit at the origin. Ellipses enclose 95% of a Gaussian fit.
Imitation learning is a promising approach for training humanoid robots to both walk and manipulate, but it requires a large number of demonstrations, which are time-intensive and difficult to collect via teleoperation. Existing data-generation algorithms can automatically synthesize demonstrations for manipulators, but they are ineffective on humanoids because their high-dimensional composite action spaces involve arms, legs, and torsos. We present HumanoidMimicGen, a method for generating humanoid legged loco-manipulation data. Our method adapts contact-rich whole-body skills from a handful of source demonstrations to new states, generalizing across changes in object pose. By interleaving these single- and dual-arm skills with whole-body locomotion and manipulation planning, the method generates stable, collision-free data across diverse scenes and layouts. To evaluate our approach, we introduce a new simulated loco-manipulation benchmark containing nine diverse tasks that test humanoid loco-manipulation capabilities. There, we demonstrate that HumanoidMimicGen automatically generates large datasets for imitation learning and enables a systematic study of how data generation and policy learning decisions impact model performance. We show that whole-body visuomotor policies co-trained with data generated by HumanoidMimicGen outperform those trained only on real-world data by 20%.
Learning from demonstration (LfD) has enabled humanoid robots to acquire diverse whole-body skills, but extending this paradigm to human-object interaction (HOI) is limited by the availability of robot-compatible interaction references. We present HOI-Retarget, a contact-centric retargeting method that transfers HOI onto a humanoid robot for large-scale motion-data generation. Its windowed trajectory optimization uses every labeled contact as a target in the object frame, balancing body tracking, foot support and smoothness under the robot's kinematic limits. The method can augment a single demonstration across object sizes, absorb contacts reconstructed from monocular video, and extend to several robots manipulating one object. We publicly release the code and the retargeted motion dataset.
Jihwan Shin, Adrià López Escoriza, Junzhe He +2
Robotic Systems Lab, ETH Zürich, 8092 Zürich, Switzerland
Learning humanoid-object interaction requires coordinating whole-body balance, locomotion, and dexterous hand contact to control both robot and object motion. Human demonstrations provide examples of coordinated interaction, but transferring these behaviors to humanoid robots requires learning how to establish and maintain effective contacts under different embodiments and dynamics. We present Weave, a unified framework for learning whole-body dexterous humanoid-object interaction from captured human demonstrations. Weave first converts captured human-object interactions into executable robot-object references through contact-aware retargeting and approach-motion completion. At its core is a contact- and geometry-aware policy that jointly commands 29 body joints and 12 actuated finger joints across multiple objects and interaction sequences. Evaluation across nine objects yields a 92.5% success rate on trained interactions and, without any additional training, 65.0% on sequences never seen during training. We additionally release ~9,000 physically executed rollouts spanning ~23 hours, providing robot-object trajectories with contact annotations for downstream interaction-policy learning and physically consistent HOI motion generation. Project website: https://xiaohu-art.github.io/Weave/
Liu Cao, Xingze Wu, Jingzhi Cui +4
1Tsinghua IIIS · 3Dalian University of Technology · 4The Chinese University of Hong Kong +1