Captured human-object interactions provide rich supervision for humanoid loco-manipulation, but they are sparse, heterogeneous, and not directly executable by robots. We introduce InterMimicGen, a self-evolving motion-imitation framework in which robot motion data and a tracking policy improve each other. First, we consolidate motion-captured human-object interaction datasets and retarget them into humanoid robot references while preserving whole-body coordination and dexterous hand-object relationships. This produces a large and diverse humanoid robot reference collection for dexterous whole-body loco-manipulation. Second, we train a physics-based generalist tracker that executes these references in simulation on a humanoid with dexterous hands, covering a scale and diversity beyond prior humanoid tracking systems for loco-manipulation. Third, we close a data flywheel: each round makes small, task-preserving changes to where an interaction takes place and how the body performs it, fine-tunes the tracker on them, and keeps only the variants whose simulated execution completes the task, which seed the next round. With more iterations, these small edits compound into broader coverage around the sparse original demonstrations while preserving task semantics and motion quality. Experiments show contact-preserving retargeting across robot configurations, broad tracking with a single generalist policy, executable motions that keep growing over augmentation rounds, and transfer to real robots. InterMimicGen provides a unified path from heterogeneous human demonstrations to a continually expanding motion resource for humanoid robot learning.
Figures & tables
Figure 1: InterMimicGen grows a finite set of human interaction captures into an expanding repertoire of executable robot motions. Left: One formulation turns the captures into executable motions across diverse objects, tasks, and robots, from humanoids to a wheeled mobile manipulator. Top right: Self-evolving motion imitation. A single motion grows round by round into many variants. Bottom right: The resulting motions run on real robots.
Figure 2: Top: Reference construction. Human-object motion capture is retargeted into whole-body dexterous references, which a physics-based tracker imitates to produce seed rollouts. Bottom: Self-evolving motion imitation. Each round augments verified references into candidates, fine-tunes the tracker on them, executes the candidates in physics, and retains only successful rollouts as seeds for the next round; Bottom right: UMAP embeddings (Sec. E.1 ) of accepted augmentations expanding.
Figure 3: Our data has broad coverage. (a-c) UMAP embeddings of human motion, object geometry, and human-object interaction of our collection, colored by source dataset. (d) How well our collection covers external motions from GRAIL ( Xie et al., 2026 ) , MOVIN ( Jang et al., 2023 ) , and ELMO ( Jang et al., 2024 ) ; a motion is covered when one of our sequences is close to it (Sec. E.1 ). (e) The same coverage as the number of our source sequences grows.
Hand contact
Penetration
Motion
Jnt. acc. (rad/s 2 ) ↓
Jnt. limit (%) ↓
Frame (%) ↑
Hand (%) ↑
Slip (m/s) ↓
Obj. (%) ↓
Depth (cm) ↓
Gnd. (%) ↓
Foot sl. (%) ↓
MPJPE (cm) ↓
EE (cm) ↓
All
Body
Fing.
Body
Fing.
Weave
96.1
97.1
0.225
67.4
2.62
31.9
62.0
9.54
12.6
12.2
6.3
19.3
0.6
33.4
Ours
98.2
98.8
0.239
37.2
2.39
0.0
12.3
5.74
5.3
2.8
3.2
2.3
0.0
0.0
Table 1: Quantitative comparison on retargeting on G1 with Inspire hands, on the 3,059 OMOMO clips. Ours shows better motion quality. Metrics are defined in Sec. E.2 .
Figure 4: Qualitative evaluation on retargeting. Captured human motion (top) and our G1 with Inspire hands reference (bottom), across interactions from different datasets.
Bimanual
Sitting
Grasping
Embodiment
SR ↑
MPJPE ↓
Obj. ↓
Rot. ↓
SR ↑
MPJPE ↓
Obj. ↓
Rot. ↓
SR ↑
MPJPE ↓
Obj. ↓
Rot. ↓
Specialists (one policy per object)
G1 + Inspire
93.8
5.19
5.79
10.4
75.4
5.94
4.76
7.4
84.9
5.30
7.17
14.1
G1 + Dex3
93.1
4.92
7.02
11.0
82.8
5.36
5.01
6.5
83.9
5.14
7.32
12.7
G1
94.4
4.52
5.88
8.5
82.8
4.93
5.26
6.2
–
–
–
–
K1
75.7
6.21
7.35
14.3
75.4
6.41
5.08
8.3
–
–
–
–
Table 2: Quantitative results on motion tracking. One generalist nearly matches the per-object specialists in body and object position error, while trailing them in success rate and object rotation. Metrics are defined in Sec. E.3 ; – marks tasks the embodiment cannot perform.
Figure 5: Physics-based tracking across objects. Executed rollouts on G1 with Inspire hands for a diverse set of objects.
Figure 6: A single capture grows into many executable variants. Each group shows the original reference followed by the verified augmentations accumulated by round 1 and by round 5.
Growth (×)
Success (%)
Quality, Orig. → Aug.
Generated
Original
Body acc.
Foot sl.
Hand jit.
Setting
R1 → R5
Orig.
Evol.
Orig.
Evol.
(rad/s 2 ) ↓
(%) ↓
(mm) ↓
Inspire, bimanual
21.8 → 150.5
52.1
98.4
100
100
11.4 → 12.4
1.18 → 1.87
1.00 → 1.17
Dex3, bimanual
21.3 → 142.0
59.3
98.7
100
100
13.9 → 16.0
1.14 → 2.21
1.12 → 1.30
Inspire, grasping
19.6 → 146.4
64.2
98.9
100
100
13.6 → 15.3
0.70 → 2.05
0.96 → 1.14
Table 3: Over five rounds, the verified references grow more than a hundredfold , and the evolved tracker executes nearly all of them without losing the original ones. Growth: verified references relative to the original references. Success: rate at which the tracker before (Orig.) and after (Evol.) self-evolution. Quality: executions of original (Orig.) and augmented (Aug.) references.
Figure 7: Real-world execution. We show the motion from our paradigm transfer to different embodiments, from Unitree G1 pulling a suitcase and G1 with Inspire hands carrying a tripod; Booster K1 lifting and placing a box and Dexmate Vega relocating a chair.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Source
Object entries
Motions
Reference h
InterAct ( Xu et al., 2025c )
117
9,030
17.82
HiPHI ( Ji et al., 2026 )
40
7,029
122.86
Consolidated
157
16,059
140.68
Appendix
Table 4: Reference collection for G1 with Inspire hands. Object entries are dataset-object pairs, so an object that appears in two source datasets counts twice. Reference hours sum the frames of the selected robot references at 50 fps and may include overlapping source segments.
Platform
Retargeting
Specialist tracking
Generalist
Augmentation loop
Sim-to- real
Excluded tasks
G1
✓
✓
–
–
✓
dexterous
G1 + Inspire
✓
✓
✓
✓
✓
none
G1 + Dex3
✓
✓
–
✓
–
none
Booster T1
✓
–
–
–
–
dexterous
Booster K1
✓
✓
–
–
✓
dexterous
Dexmate Vega
✓
–
–
–
✓
dexterous, seated
Appendix
Table 5: Studies run on each platform. Configurations with dexterous hands use all data, including the dexterous benchmark; the generalist is trained on G1 with Inspire hands.
Table 13
Setting
Value
Reference roster
16,059 motions
Parallelization
4 GPUs, 4,096 environments per GPU
Control / physics
50 Hz control, 4 physics substeps per action
PPO rollout / minibatch / epochs
32 steps / 16,384 samples / 6 epochs
Actor / critic
Separate ReLU MLPs, 1024→1024→512
Learning rate / discount / GAE
2×10−5 / 0.99 / 0.95
Appendix
Table 8: Generalist tracker settings.
Object edits
Body edits
Checked before physics
finite values
finite, object unchanged
Tracker used
fine-tuned first
current
Rollouts needed to pass
2 of 3
1 of 3
Completes the reference
✓
✓
Terminal hand contact
✓
✓ (held tasks)
No sustained fall
✓
✓
Appendix
Table 9: Acceptance rules for the two edit types. ✓: checked; –: not checked.
Figure 8: More executable variants. As in Figure 6 , for nine further captures: each group shows the original reference followed by the verified augmentations accumulated by round 1 and by round 5.
Figure 9: More retargeted interactions on G1 with Inspire hands.
Figure 10: Retargeting with dexterous hands. A bowl-holding interaction retargeted to G1 with Dex3 (left) and Inspire (right) hands.
Figure 11: Retargeting across embodiments. One suitcase interaction retargeted to G1 with Dex3, Inspire, and non-dexterous hands (left, top to bottom) and to Booster K1, Booster T1, and Dexmate Vega (right, top to bottom).
Figure 12: Retargeting comparison with OmniRetarget. OmniRetarget to G1 with rigid hands (left) and our retargeter for G1 with Inspire hands (right), at corresponding interaction phases. The dexterous hand keeps the grasp and the hand-object relationship of the demonstration.
Figure 13: Retargeting comparison with UMR. UMR to G1 (left) and our retargeter to G1 with Inspire hands (right), at corresponding interaction phases.
Metric
OmniRetarget
UMR
ULTRA
Ours
Contact preservation
Hand contact preservation (%) ↑
79.09
95.13
79.06
83.45
Penetration ( >1 cm)
Robot–object depth (cm) ↓
1.27
1.60
1.49
1.97
Foot sliding (0.15 m/s)
Sliding frames (%) ↓
0.00
7.02
N/A
7.33
Appendix
Table 10: Retargeting on G1 with non-dexterous hands , over the 537 OMOMO clips shared by all four methods. A dash denotes no qualifying event and N/A an unavailable metric; bold marks the unique best.
Figure 14: Physics-based tracking across embodiments. A box-lifting reference executed by specialist trackers on Booster K1 (top left), G1 with non-dexterous hands (top right), G1 with Inspire hands (bottom left), and G1 with Dex3 hands (bottom right).
Yield (%)
Growth (×)
Round
Frozen
Tuned
Frozen
Tuned
1
29.4
25.3
163.0
188.5
2
0.4
72.9
170.5
426.0
3
0.4
68.9
184.9
573.0
Appendix
Table 11: Frozen versus fine-tuned tracker in a separate run on a subset of the bimanual seeds with translation edits, whose counts are not directly comparable with Table 3 . Yield and growth follow Sec. E.4 . The frozen run stops after round 3.
Figure 15: Physics-based tracking with dexterous hands. A tripod grasp executed on G1 with Dex3 (left) and Inspire (right) hands.
Inspire, bimanual
Dex3, bimanual
Inspire, grasping
Round
Growth
Acc.
Foot
Hand
Growth
Acc.
Foot
Hand
Growth
Acc.
Foot
Hand
Orig.
1.0
11.4
1.18
1.00
1.0
13.9
1.14
1.12
1.0
13.6
0.70
0.96
1
21.8
11.8
1.39
1.07
21.3
15.6
1.59
1.25
19.6
15.3
0.99
1.07
2
34.8
11.3
1.39
1.09
31.6
14.8
1.77
1.23
33.3
14.4
1.42
1.09
3
80.4
12.1
1.63
1.16
73.3
14.9
1.87
1.24
77.8
14.6
1.78
1.08
4
118.1
12.9
2.14
1.22
110.8
16.4
2.41
1.33
112.8
15.7
2.42
1.16
Appendix
Table 12: Self-evolution by round , extending Table 3 . Growth (×): cumulative verified references relative to the original references. Quality of the original references (Orig.) and of the references newly accepted in each round: body joint acceleration (Acc., rad/s 2 ), foot sliding (Foot, %), and hand jitter (Hand, mm).
Figure 16: Body-motion variation around shared seeds. Each reference is embedded by the DCT coefficients of its body joint angles, taken relative to its seed and projected with PCA, so all seeds sit at the origin. Ellipses enclose 95% of a Gaussian fit.