Despite recent advances in humanoid locomotion, controllers optimized for command tracking and robustness tend to produce mechanical gaits, whereas controllers tied to human motion data often fail to generalize to commands outside the data distribution. This work introduces a learning framework that balances these competing objectives to synthesize real-time steerable, robust, and biomimetic locomotion policies from human data. Using an in-house curated locomotion dataset covering diverse speeds and directions, we first learn a natural locomotion prior policy through a teacher-student distillation process. Specifically, we train a full-body reference-conditioned policy with Reinforcement Learning (RL), then distill it into a lightweight prior policy conditioned solely on proprioception and a planar torso-velocity steering command. Next, we fine-tune the prior policy with multi-task RL to expand command coverage and robustness beyond the data distribution, pairing a goal-conditioned task that tracks arbitrary commands with a reference-guided task that tracks the human data as an explicit style regularizer. We validate our framework on three humanoid robots: the Boston Dynamics Atlas R1, Atlas D1, and Unitree G1. Experimental results demonstrate robust performance across real-world scenarios, including direct user-controlled locomotion in indoor and outdoor environments, and integration as the locomotion layer within hierarchical control stacks. Benchmarks against Tabula Rasa RL policies trained without human data and ablation studies confirm that our framework yields a lightweight, deployable policy that reconstructs coordinated whole-body behavior from a steering command, retaining the human gait characteristics while remaining robust and fully steerable.
Figures & tables
Figure 1 : Summary of hardware results from HuMBLE . (A) Human-like steerable locomotion on the Boston Dynamics Atlas R1 was demonstrated live to the public on stage. (B) The Atlas R1 robot following real-time user steering commands that vary continuously in speed and direction while maintaining stylistic consistency with the human reference motion dataset. (C) The Atlas D1 robot performs backward walking, circling, and sidestepping (from left to right) to demonstrate the HuMBLE policy’s coverage of the SE(2) steering command space. (D) Transition from walking to high-speed sprinting on the Unitree G1 robot. The gait transition emerges in response to an increasing in the commanded speed. (E) Integration of the HuMBLE policy as an underlying locomotion layer within hierarchical control stacks for autonomous navigation (left) and box parkour (right) tasks.
Figure 2 : Stroboscopic snapshot of Atlas R1 walking forward under different controllers. We contrast the forward walking gaits, initiated from a standing configuration, across three different locomotion controllers: (A) the HuMBLE policy, (B) an MPC controller, and (C) a Tabula Rasa RL policy. Each frame captures the transition from swing phase to ground contact to highlight the synchronized arm swinging, heel-strike, and knee-lock.
Figure 3 : Data fidelity analysis. (A) Mean absolute error (MAE) of the torso linear velocity, angular velocity, and joint positions relative to the reference motion data, comparing the teacher and the final HuMBLE deployment policies. Each bar illustrates the cumulative tracking error decomposed into linear velocity ( m/s ), angular velocity ( rad/s ), and joint positions ( rad ). The MAE is calculated by averaging over all individual state dimensions and over the trajectory length following temporal alignment via DTW. (B) Comparison of the reference motion, full-body-reference-conditioned teacher policy, and command-conditioned deployment policy for Fast Forward Walk . All plots display raw time-series data without applying DTW, where the x-axis represents time in seconds. Trajectories illustrate SE(2) torso velocities alongside selected joint angles chosen to highlight key biomimetic traits such as arm swing, hip sway, knee extension, and heel-to-toe rolling present in the human data and learned by the HuMBLE policy.
Figure 4 : Atlas R1 quantitative performance analysis. We evaluate and compare the command responsiveness, command-tracking accuracy, and robustness to external pushes for the HuMBLE and the Tabula Rasa policies. (A) Relative command-tracking error in steady-state. The SE(2) command space (longitudinal velocity, lateral velocity, and turning rate) is visualized using three 2D slices, zeroing one dimension per column, from left to right: ωturn=0 , vlat=0 and vlong=0 . Each cell in a plot compares the Tabula Rasa baseline (outer square) to the HuMBLE policy (inner circle), with lighter colors indicating better performance (lower tracking error or faster rise time). (B) Rise time in seconds. Similarly visualized as in (A) . (C) Robustness to external pushes. We evaluate stability against planar force pushes across three locomotion scenarios. The shaded regions approximate the successful recovery envelopes for the HuMBLE policy (orange) and the Tabula Rasa baseline (purple). A ◯ symbol indicates successful push recovery, while a × indicates a fall.
Boston Dynamics Atlas R1
Steady-State Error
Mean Absolute Error
Command
Sim
Real
Sim
Real
Forward Walk
(0.4, 0.0, 0.0)
(0.08, 0.02)
(0.12, 0.03)
(0.10, 0.10)
(0.11, 0.10)
(1.2, 0.0, 0.0)
(0.03, 0.02)
(0.06, 0.07)
(0.10, 0.15)
(0.10, 0.18)
(2.0, 0.0, 0.0)
(0.13, 0.11)
(0.20, 0.06)
(0.15, 0.13)
(0.18, 0.16)
Table 1 : Sim-to-real analysis. A set of representative canonical commands is employed to evaluate the sim-to-real transfer of the HuMBLE policy. The command elements vlong , vlat , ωturn , have units of m/s, m/s, and rad/s respectively. Each scenario starts with the robot standing at zero command. The command is then stepped to the target and held for a duration before stepping back to zero. We report the mean absolute error (MAE) over the entire scenario as well as the steady-state error. The linear and angular components of the tracking error are reported separately with units of m/s and rad/s , respectively. We define the linear tracking error as the max of the longitudinal and lateral tracking errors.
Figure 5 : Ablation study. We conducted ablation studies of three crucial design choices within HuMBLE . (A) Role of Teacher–Student Distillation. We compare the data fidelity of the HuMBLE prior policy obtained from the Teacher–Student distillation stage (orange) against a single-stage RL baseline (gray) that directly trains a command-conditioned policy. The DTW-aligned reference motion tracking errors are visualized as a stacked bar chart of the component-wise Mean Absolute Error (MAE). (B) Efficacy of Multi-Task RL. We evaluate data fidelity and steering command-tracking performance across varying environment allocation ratios for four canonical modalities, validating the efficacy of the multi-task RL framework in balancing these two competing objectives. (C) Impact of RL Fine-Tuning. We assess command-tracking accuracy and command coverage by measuring the steady-state relative command-tracking error of the final HuMBLE deployment policy (inner circle) following RL fine-tuning, compared directly against the prior policy (outer square) from which fine-tuning was initialized. The heatmap visualizes the relative command-tracking error in steady-state for ωturn=0 , vlat=0 , and vlong=0 from left to right. Gray squares indicate regions where the prior policy fails to robustly track the corresponding velocity command, resulting in a robot fall.
Figure 6 : Overview of the HuMBLE framework. Our approach trains the final control policy πdeploy through a two-stage process. Stage I employs a teacher–student distillation pipeline to extract a natural locomotion prior πprior . Specifically, a teacher policy πteacher trained to track full-body reference motions is distilled into the student policy πprior , which maps only proprioceptive observations and a SE(2) command to joint-level actions. In Stage II, the prior policy is fine-tuned via multi-task RL to enhance responsiveness to arbitrary user steering commands while preserving the stylistic qualities of the reference motion data, ultimately forming the final deployment policy πdeploy .
Figure 7 : Distribution of the human reference motion dataset. We illustrate the distribution of the annotated commands for our in-house curated human motion dataset, retargeted to Atlas R1. While the dataset also contains high-speed sprinting maneuvers, this visualization focuses exclusively on the standard walking command distribution to clearly illustrate the most densely populated region. (A) Scatter plots illustrating the distribution of annotated velocity commands across four locomotion regimes (fast, normal, slow walk and sidestep). For visual clarity, the number of data points is decimated by a factor of 10. (B) Histograms displaying the marginal proportions of the dataset frames binned by longitudinal velocity, lateral velocity, and turning rate. The y-axis denotes the proportion relative to the dataset size N=528,214 .
I
Inertial (world) frame
B
Torso (floating base) frame
T
IMU frame
Ek
k -th end-effector frame
nIMU
Number of IMUs
nj
Number of joints
ne
Number of end effectors
Table 9
Figure S1 : Extended Data Fidelity Analysis Results. Comprehensive trajectory comparison of the raw SE(2) torso velocities and joint angle trajectories for the simulated Unitree G1 robot across all eight canonical behaviors.
Figure S2 : Atlas D1 quantitative performance analysis. We provide the same performance evaluations for the simulated Atlas D1 robot as presented for the Atlas R1 robot in Figure 4.
Figure S3 : Unitree G1 quantitative performance analysis. We provide the same performance evaluations for the simulated Unitree G1 robot as presented for the Atlas R1 robot in Figure 4. Note that the G1 policy is evaluated under a different command range than the Atlas robots, as spatial and temporal scaling factors are applied during motion retargeting to accommodate the G1’s smaller kinematic scale.
Torso linear velocity ( m/s )
Torso angular velocity ( rad/s )
Joint positions ( rad )
Mean time warping ( ms )
Teacher
Deploy.
Teacher
Deploy.
Teacher
Deploy.
Teacher
Deploy.
Slow Forward Walk
0.0341
0.0388
0.0421
0.0852
0.0294
0.0491
11.5929
27.3009
Normal Forward Walk
0.0441
0.0472
0.0541
0.0979
0.0376
0.0547
8.6093
58.7417
Fast Forward Walk
0.0895
0.0847
0.0841
0.1438
0.0402
0.0598
4.0594
20.9901
Slow Backward Walk
0.0383
0.0550
0.0435
0.1149
0.0326
0.0549
8.0731
55.8140
Normal Backward Walk
0.0418
0.0643
0.0496
0.1478
0.0328
0.0555
8.1499
62.8103
Table S1 : Extended data fidelity analysis metrics. Comprehensive tracking performance comparison between the teacher and deployment policies across eight canonical motion trajectories for the Unitree G1 robot. Reported values indicate the Mean Absolute Error (MAE) for torso linear velocity, torso angular velocity, and joint positions, alongside the normalized mean time warping metrics converted to physical time units ( ms ). We note that the mean time warping metric based on DTW statistics can be computed as the total warp area divided by the length of the trajectory. Specifically, the warp area is defined by summing the absolute index differences between the aligned time pairs along the optimal warping path. To provide an intuitive temporal scale, we scale this discrete frame offset by the simulation control loop timestep, thereby converting the mean time warping metric into physical time units of ms .
Boston Dynamics Atlas D1
Steady-State Error
Mean Absolute Error
Command
Sim
Real
Sim
Real
Forward Walk
(0.4, 0.0, 0.0)
(0.03, 0.01)
(0.20, 0.02)
(0.08, 0.12)
(0.14, 0.08)
(1.2, 0.0, 0.0)
(0.08, 0.08)
(0.11, 0.00)
(0.10, 0.18)
(0.16, 0.21)
(2.0, 0.0, 0.0)
(0.31, 0.01)
(0.28, 0.17)
(0.23, 0.17)
(0.23, 0.29)
Table S2 : Atlas D1 and Unitree G1 Sim-to-real analysis results. We provide the same sim-to-real analysis results for the Boston Dynamics D1 and Unitree G1 robot as presented for the Atlas R1 robot in Table 1.We note that the G1 policy is evaluated under a different command than the Atlas robots as spatial and temporal scaling factors are applied during motion retargeting to accommodate the G1’s smaller kinematic scale.
Term Name
Definition
Noise
Dim.
Teacher
Deploy
Critic
Proprioceptive observations oprop
IMU linear velocity
TvIT
N(0,0.1)
3×nIMU
✓
✓ ∗
✓ †
IMU angular velocity
TωIT
N(0,0.1)
3×nIMU
✓
✓
✓ †
IMU projected gravity
TgI
N(0,0.015)
3×nIMU
✓
✓
✓ †
Joint positions
qj
N(0,0.005)
nj
✓
✓
✓ †
Joint velocities
q˙j
N(0,0.25)
nj
✓
✓
✓ †
Table S3 : Summary of observation terms. This table details the components of the observation space. No scaling or clipping is applied to these raw values. Checkmarks ( ✓ ) indicates the inclusion of a specific term within the corresponding observation vectors—teacher observations oteacher , deployment observations odeploy , and critic observation ocritic .
Term
Condition
Unitree G1
Atlas R1
Atlas D1
Body-ground penetration
Penetration
> 0.03m
> 0.03m
> 0.03m
Large contact force
Contact force
> 6000N
> 12000N
> 12000N
Torso height deviation from reference ∗
Deviation
> 0.3m
> 0.4m
> 0.4m
Torso projected gravity deviation from reference ∗
Deviation angle
> 0.6rad
> 0.6rad
> 0.6rad
Foot self collision
If collision between two feet is detected
Table S4 : Termination conditions. Criteria marked with an asterisk ( ∗ ) are applied exclusively to tasks utilizing reference motion data, whereas those marked with a dagger ( † ) are applied exclusively to the goal-conditioned task in Stage II.
Term
Value
Unitree G1
Atlas R1
Atlas D1
Feet static friction
U(0.6,1.3)
U(0.6,1.0)
U(0.6,1.0)
Feet dynamic friction
U(0.5,0.9)
U(0.5,0.9)
U(0.5,0.9)
Feet restitution
U(0.0,0.2)
U(0.0,0.2)
U(0.0,0.2)
Body mass scale factor
For all bodies
U(0.9,1.15)
–
–
Torso mass perturbation
–
U(−8.0,8.0)\text{,}\mathrm{kg}$$
U(−8.0,8.0)\text{,}\mathrm{kg}$$
Table S5: Domain randomization terms.
Term
Notation
Hyperparameter
Unitree G1
Atlas R1
Atlas D1
Reference Tracking rtrack
Ref. torso position tracking
rt,r
ct,rσr,min
–
1.0 0.4
1.0 0.4
Ref. torso orientation tracking
rt,Φ
ct,ΦσΦ,min
–
1.0 0.5
1.0 0.5
Ref. torso linear velocity tracking
rt,v
ct,vσv,min
1.0 0.2
1.0 0.2
1.0 0.2
Ref. torso angular velocity tracking
rt,ω
ct,ωσω,min
1.0 0.2
1.0 0.2
1.0 0.2
Table S6 : Summary of rewards. Formal definitions for the reward terms are available in Section S3 and S4 of the Supplementary Materials.
P gain
D gain
Nominal angle \left[\text{,}\mathrm{rad}\right]
Action scaling factors
Waist roll
150.0
3.0
0.0
0.1
Waist pitch
150.0
3.0
0.0
0.1
Waist yaw
100.0
2.0
0.0
0.1
Hip roll
40.0
3.0
0.0
0.2
Hip pitch
40.0
3.0
-0.1
0.2
Hip yaw
40.0
3.0
0.0
0.1
Table S7 : Unitree G1 joint gains and action specifications. The action vector output by the policy is element-wise multiplied by the action scaling factors (action scalers) and offset by the nominal joint configurations to yield the desired target positions for the low-level, joint-level PD controllers, utilizing the specific gain parameters detailed below.
Reinforcement learning can produce robust humanoid controllers, but each new task is typically trained as a separate policy with its own reward design and training process. Motion imitation provides an alternative source of motor competence by training policies to track retargeted human motions, yet the resulting controllers remain reference trackers and are not directly usable as task policies. We propose a three-stage pipeline that turns motion-imitation skills into a reusable hybrid motion prior (HMP) for humanoid locomotion. First, an expert policy is trained to imitate retargeted human motion-capture clips. Second, the expert is distilled into a frozen architecture composed of a proprioceptive encoder, a residual vector-quantized (RVQ) codebook, and an action decoder. Third, task-level policies are trained to solve locomotion tasks by selecting discrete codebook entries while the HMP remains frozen. We evaluate the method on velocity tracking, point-goal navigation, and fall-recovery velocity tracking in simulation, and deploy the velocity-tracking policy on a real Unitree G1 robot. The distillation process preserves the tracking behavior of the expert, while the resulting HMP can be reused without retraining as the action interface for different downstream locomotion policies. The learned HMP reveals an interpretable codebook structure in which the number of active RVQ stages modulates the available gait patterns. We further show that training the codebook with the rotation trick improves latent organization and reduces downstream falls compared with a standard straight-through estimator.
Humanoid robot motion learning requires not only task-oriented control policies but also physically feasible and natural behaviors that can be transferred to real robots. However, robot-feasible motion data are often scarce: raw human demonstrations may be incompatible with the robot morphology, open-source clips vary in quality, and simulation-collected robot trajectories still require feasibility checking. To address these challenges, we propose a data-centric training and deployment pipeline that integrates motion data curation, real-to-sim model adaptation, AMP-based reinforcement learning, and sim-to-real deployment. We validate the framework on the Booster T1 robot and further provide preliminary cross-platform validation on Booster K1.
Penghui Chen, Tinglong Zheng, Yufeng Zhang +1
Department of Automation, Tsinghua University, Beijing, China · Booster Robotics Technology Co., Ltd, Beijing, China · School of Mechanical and Electronic Control Engineering, Beijing Jiaotong University, Beijing, China
This paper presents MuGen, a data-driven framework for learning and deploying multi-skill locomotion on humanoid robots. MuGen enables a robot to perform expressive motions like humans under the guidance of example motion sequences. To achieve this, we employ vector-quantized autoencoders (VQ-VAEs) trained with model-based reinforcement learning, resulting in a generative representation of locomotion that captures key patterns of human motion from hours of heterogeneous human performance data. We employ a teacher-student learning framework and develop a new policy distillation strategy to enable a deployable student policy learning this efficient latent representation. This policy allows the robot to track and mimic unseen human motions and further enables the robot to reuse the learned latent space for other tasks. We demonstrate the effectiveness of our framework through a diverse set of motions and accurate execution.