Human video offers a scalable source of robot demonstrations, yet most human-to-humanoid retargeting methods assume a legged robot with human-like kinematics. This assumption does not hold for mobile-base humanoids equipped with a wheeled base, vertical lift, and two arms. Human walking must be expressed through base motion, while torso bending may require coordinated lift and arm motion. We address this mismatch with a task-conditioned framework that assigns reconstructed human motion to base, lift, and arm responsibilities before robot-specific realization. The allocator preserves the human-derived path, stabilizes heading, separates turn and translation when needed, retimes commands to satisfy base limits, and repairs lift and arm trajectories. A deployment adapter then converts the reference to 50 Hz commands using stationary-base detection, deadband and slew-rate filtering, time-consistent playback scaling, and separate linear and angular gains. We evaluate the resulting references with human-derived task-space comparisons, policy-free simulation replay, and a qualitative execution on a physical robot.
Figures & tables
Fig. 1: Embodiment-aware allocation of reconstructed human motion into arm, lift, and mobile-base intents. A single whole-body motion may require coordinated actuation across several robot subsystems.
Fig. 2: Task-conditioned embodiment-aware retargeting pipeline. Human video is reconstructed as 3D motion and decomposed into base, lift, and arm targets. These targets are optimized under feasibility, task-space, and smoothness objectives, then realized and repaired by a GMR-based adapter for simulation playback or command export.
ID
Task
Description
Primary Motion Cue
Allocation Target
01
Slow walk
Forward walking at a moderate pace
Smooth root translation
Base
02
Run
Fast forward locomotion
High root velocity
Base (speed-adaptive)
03
Side walk
Lateral walking without a dominant turn
Lateral root motion
Lateral base motion
04
Sideways-turn walk
Lateral locomotion combined with reorientation
Root translation + yaw
Base translation + rotation
05
Bend-and-reach I
Forward reach with moderate torso lowering
Torso pitch + low wrist
Lift + arms
06
Bend-and-reach II
Deeper forward reach at a lower height
Deep reach + arm extension
Lift + arms + base assist
TABLE I: TASK INDEX AND INTENDED EMBODIMENT-AWARE MOTION ALLOCATION.
Fig. 3: Example processing sequence for a bend-and-reach motion: RGB human demonstration, GVHMR 3D reconstruction, and the retargeted mobile-base humanoid reference in simulation.
Fig. 4: The same bend-and-reach reference executed in simulation and on the physical robot for Task 06. The sequence retains the intended lift and bilateral arm coordination in both environments.
Method
Locomotion
Manip./reach
All
01–04
05–09
01–09
Ours
0.355
0.295
0.322
Arm-only
1.000
0.566
0.759
Base-only
0.355
0.868
0.640
No-lift
0.355
0.440
0.402
TABLE II: Human-derived continuous task-space evaluation on scenarios 01–09. Targets are extracted from GVHMR-reconstructed SMPL motion rather than from our generated reference. Entries are normalized costs; lower is better.
Method
Loc. root
Wrist
Wrist- z
Lift short.
Primary gap
(m)
(m)
(m)
(m)
Ours
0.641
0.346
0.143
0.012
balanced
Arm-only
1.610
0.352
0.164
0.125
base/lift
Base-only
0.641
1.726
0.643
0.125
wrist/lift
No-lift
0.641
0.346
0.164
0.125
lift
TABLE III: Component-wise human-target ablation. Locomotion root is averaged over scenarios 01–04, wrist metrics over scenarios 05–09, and lift shortfall over scenarios 05, 06, and 09. Lower is better.
Group
Sc.
Completed
Body mean (m)
Anchor mean (m)
Steps
Locomotion
4
4/4
0.012
0.175
1427
Bend/reach
2
2/2
0.024
0.020
551
Base+arms
2
2/2
0.028
0.048
654
Low pick
1
1/1
0.029
0.032
315
All 01–09
9
9/9
0.020
0.097
2947
TABLE IV: Policy-free simulation replay validation. Completion means no early termination during reference playback, not learned-policy success. Lower errors are better.
Human-to-humanoid retargeting has largely been studied on legged platforms, while comparatively few wheeled-humanoid systems support coupled locomotion and manipulation from general human motion. Building on GMR's configurable general-motion retargeting and BeyondMimic's physically simulated R1 Pro learning framework, we present a reproducible pipeline that converts multi-dataset SMPLX motion into executable loco-manipulation behavior for the Galaxea R1 Pro wheeled humanoid. The robot has a planar three-wheel base, a serial torso, and two arms but no leg joints, so human lower-body motion must be redistributed across base motion and torso posture without sacrificing manipulation-relevant arm geometry. Our pipeline combines canonical body-shape preprocessing, planar-base normalization, morphology-aware differential inverse kinematics, shoulder-rooted hierarchical arm retargeting, and continuous torso substitution for bending and squatting. A reference-twist-driven planning layer then decodes planar base motion into continuous three-wheel steering and rolling commands subject to hysteresis, kinematic continuity, acceleration, and actuator-rate limits. Finally, a 21-dimensional BaseDecode policy is trained in Isaac Lab with directional joint-limit scaling, focused upper-body tracking, and a staged wheel-contact reward. The resulting system provides a complete bridge from human motion data to physically trackable wheeled-humanoid loco-manipulation rather than a visualization-only retargeter; quantitative policy comparisons remain scheduled for a later revision.
Retargeting human motion to humanoid robots is critical for teleoperation, imitation learning and human-robot interaction. However, it remains challenging because of substantial morphological discrepancies between humans and robots, including differences in skeletal topology, limb proportions and degrees of freedom, as well as the scarcity of paired motion data. This paper presents Human2Humanoid, an unsupervised motion retargeting framework that transfers human motions to humanoid robot behaviors with high fidelity. To bridge the domain gap under unpaired data, we adopt a CycleGAN-based architecture equipped with a skeleton-aware graph convolutional network to capture topology-dependent motion features. To address cross-domain scale mismatches, we introduce a morphology-invariant end-effector consistency loss that aligns normalized end-effector trajectories to preserve motion semantics across embodiments. To improve physical plausibility and reduce contact artifacts, we impose explicit physics-aware feasibility constraints to encourage reproduction of the contact patterns in the source motion. Experimental results show that the proposed method successfully retargets human motion to the Unitree G1 humanoid robot without paired data, and outperforms existing methods in both downstream controllability and physical feasibility.
Tianchen Huang, Feiyang Yuan, Junchi Gu +5
Institute of Humanoid Robots, Department of Precision Machinery and Precision Instrumentation, University of Science and Technology of China, Hefei, Anhui 230026, China
Direct transfer from human demonstration to learnable robot action is a crucial step towards scalable whole-body mobile manipulation. While human data scales better than mobile teleoperation, it requires overcoming significant embodiment gaps. Existing retargeting methods yield imprecise or inconsistent solutions, causing action multi-modality that prevents supervised policies from reliably converging. We present Whole-body-Aware Retargeting from human Pose (WARP), an offline pipeline that explicitly models embodiment differences to extract precise, unique whole-body actions. WARP leverages a closed-form Shoulder-Elbow-Wrist (SEW) geometric solver for exact end-effector tracking while preserving whole-body structural intent. Paired with lazy mobile-base control, it extracts accurate, consistent robot trajectories. Evaluations show WARP provides highly reliable data for open-loop real-world replay. To our knowledge, WARP is the first framework to achieve zero-shot whole-body mobile manipulation directly from offline human demonstrations, eliminating the need for human-in-the-loop teleoperation action data. More details on https://warp-retarget.github.io/
Zhenyang Chen, Chuizheng Kong, Chuye Zhang +4
Georgia Institute of Technology, Atlanta, Georgia 30332