Despite advances in reinforcement and imitation learning, achieving both precise control and dynamic whole-body behaviors in long-horizon tasks remains challenging. Existing approaches typically follow two paradigms: coupled whole-body policies for global coordination and decoupled policies for modular precision. However, effectively combining these two types of controllers to achieve agility, robustness, and precision remains challenging. In this work, we propose BAT, an online policyswitching framework that dynamically selects between coupled and decoupled whole-body RL generalist controllers according to the evolving motion context. BAT employs two complementary switching predictors: OpVQ-VAE provides motion-token-based predictions with strong generalization, while OpHRL provides closed-loop, state-aware predictions. A Token-Familiarity Router (TFR) selects between their predictions, favoring OpHRL for familiar token sequences and OpVQ-VAE for unfamiliar ones when their predictions disagree. BAT achieves 85.2% success on seen long-horizon motion combinations, while retaining 71.3% success on unseen motions. BAT also outperforms existing whole-body controllers on individual motions and demonstrates successful zero-shot deployment on the Unitree G1 humanoid.
Figures & tables
Fig. 1 : Conceptual overview of BAT . The decoupled policy provides stability and robustness, while the coupled policy enables agile whole-body motion. BAT adaptively switches between them to combine both strengths.
Fig. 2 : Complementary policy strengths across motion families. πD generally achieves higher upper-body tracking rewards, while πC achieves higher lower-body tracking rewards and greater running/jogging stability. The 5,171 motions with valid BABEL sequence-level action labels are grouped into 12 motion groups. Bars show mean πD−πC differences in fall-free rate (top, percentage points) and episode-mean tracking rewards (bottom: green, upper-body; orange, lower-body; purple, overall). Positive values favor πD ; negative values favor πC . n denotes the number of motions.
Fig. 3 : Overview of BAT. (1) Option-Aware VQ-VAE (OpVQ-VAE) learns discrete motion tokens that encode motion structure and controller preference through reconstruction, option-prediction, and auxiliary objectives. After freezing the encoder and codebook, an LSTM–MLP head learns transition-aware option prediction from individual and composed motions. (2) Option-Guided Hierarchical RL (OpHRL) uses cross-simulator switch-boundary supervision from DOp for behavior-cloning initialization and guided, proprioception-aware RL refinement. The supervision is obtained through full-chain Isaac Gym rollouts and MuJoCo validation (Sec. V-A ). A low-level policy manager executes the selected controller. (3) Token-Familiarity Router (TFR) estimates causal token-sequence familiarity using 4-gram likelihood and selects between the predictions of OpVQ-VAE and OpHRL. EMA smoothing and hysteresis are applied to the familiarity score and routing decision, yielding the final selection between the decoupled policy πD and coupled policy πC .
Fig. 4 : Token familiarity resolves predictor disagreements by selecting OpHRL for familiar sequences and OpVQ-VAE for unfamiliar ones. (a) Switching traces for a training chain (top) and a Test-U chain containing unseen motions (bottom). Highlighted regions show disagreements: BAT follows OpHRL when gt=1 (familiar) and OpVQ-VAE when gt=0 (unfamiliar). When both predictors agree, their common prediction is used. Dashed vertical lines mark motion-clip boundaries. (b) Moving-window token-NLL distributions for Train, Test-C (seen motions in unseen combinations), and Test-U (unseen motions in unseen combinations). The hysteresis thresholds τon=1.75 and τoff=2.69 correspond to the 95th and 99.5th percentiles of the training EMA-smoothed window-NLL distribution. (c) Per-chain fractions of unfamiliar windows, with medians of 0.0%, 3.1%, and 62.1%, respectively, indicating that the router responds primarily to unseen motions.
Success (%) ↑
Reward ↑
Mean Switches
Category
Method
Train
Test-C
Test-U
Train
Test-C
Test-U
Train
Test-C
Test-U
Ours
BAT
85.2
70.0
71.3
127.9
47.10
115.5
1.43
1.93
2.16
Ablations
OpVQ-VAE
77.8
65.7
70.6
124.0
58.88
115.5
1.04
1.01
1.03
OpHRL
80.8
68.6
41.3
117.0
47.17
66.1
1.54
1.83
3.61
BC
71.8
62.9
40.6
107.0
57.34
66.1
2.39
2.21
3.84
Continuous Op-AE
78.8
67.1
61.9
51.65
45.10
45.09
0.88
0.80
1.40
TABLE I: Performance comparison across three evaluation splits: Train (650 chains) uses seen motions in seen combinations, Test-C (70 chains) uses seen motions in unseen combinations, and Test-U (160 chains) uses unseen motions in unseen combinations. Mean Switches is reported as a diagnostic of switching frequency. Best and second-best distinct Success and Reward values across all methods, including offline oracles, are indicated by bold and underlined values, respectively; ties receive the same formatting. Offline oracles replay precomputed controller selections and do not guarantee successful execution.
Metric
GMT [ 3 ]
FALCON [ 1 ]
TWIST [ 6 ]
SONIC [ 4 ]
BAT (Ours)
Success ( n / total)
7149 / 8261
7735 / 8261
7659 / 8261
7519 / 8261
7901 / 8261
Success Rate (%)
86.54
93.63
92.71
91.02
95.64
Mean Reward
0.6347
0.7677
0.8002
0.7343
0.8468
Upper-body tracking
0.8695
0.9776
0.9483
0.8858
0.9643
Lower-body tracking
0.8717
0.7522
0.9081
0.9040
0.8593
Critical Force — Forward (N)
80
190
180
>200
190
TABLE II : Comparison of motion tracking performance and perturbation robustness across diverse motions. BAT consistently selects the decoupled policy πD for the standing and static motions used in the perturbation tests.
Fig. 5 : Hardware deployment of BAT on the Unitree G1 humanoid. (a)–(d) Long-horizon sequences involving online switching between the decoupled policy πD (blue) and coupled policy πC (red). (e) Single-motion tasks with adaptive policy selection, where BAT automatically selects the more suitable low-level policy based on motion characteristics.
Large-scale humanoid motion-tracking controllers are commonly improved by reallocating training effort: difficult motions are sampled more often, isolated into smaller subsets, or assigned to specialized experts. We show that this view is incomplete. In strong whole-body-control baselines, a residual set of feasible training clips remains unsolved even under targeted training, especially for high-dynamic transitions and balance-critical motions. These failures arise not only from insufficient exposure, but from a mismatch between the motion demands and the effective capability induced by the default training recipe. We propose Athena-WBC, a compact teacher-student pipeline with capability-aligned policy experts for long-tail humanoid whole-body control. Dynamic experts use a tracking-focused, constraint-aware objective that removes conservative effort and temporal-control penalties while preserving physical feasibility constraints; balance experts use a gravity curriculum to improve early-training survivability. The resulting privileged teachers are motion-routed for DAgger distillation and then compressed into a single controller with deployable observations followed by RL fine-tuning. Experiments on a full-size humanoid show improved recovery of training-set long-tail motions and better held-out tracking than a strong SONIC-recipe baseline, using only a small number of experts.
Scaling humanoid whole-body control toward general-purpose deployment requires large human motion corpora and training experience shared across robot bodies. Existing methods usually train one policy per robot, leaving motion experience isolated across embodiments. We introduce X-WBC, a cross-embodiment foundation framework that separates relatively shared human motion semantics from embodiment-specific physical execution. Human-centered command tokens align full human motion, robot reference motion, and sparse VR observations. A causal Transformer learns reusable temporal structure from mixed multi-robot rollouts, while lightweight robot-specific modules map the shared representation to each robot's proprioception and action space. Across nine simulated embodiments, external motions, and four real robots, experiments show that joint training improves tracking, the aligned representation supports consistent control across command sources, and the learned policy remains competitive beyond the training corpus. These results support heterogeneous humanoids as joint data sources and establish cross-embodiment joint training as a practical route toward whole-body control foundation models.
Juntong Zhang, Chun Gu, Li Zhang
Tongji University · Shanghai Innovation Institute · School of Data Science, Fudan University
Embodied AI systems, particularly humanoid robots deployed in real world scenarios require whole-body control policies that are both task-responsive and physically smooth. However, smoothness is not uniform across the body: lower body must remain sufficiently reactive, while the upper body must be tightly regulated to preserve stability. Existing reinforcement learning approaches typically impose smoothness through auxiliary terms in the reward function, which compete with task objectives, treating the body as uniform and provide no direct control over the physical quantities responsible for smooth behavior. We introduce DeCap (Decoupled Constraint-aware policy), a constrained reinforcement learning algorithm that decouples whole-body smoothness into separate upper- and lower-body constraint groups, each formulates smoothness as explicit constraints on physical motion limits. To improve constraint satisfaction near feasibility boundaries, DeCap incorporates a bounded barrier penalty that activates proactively as limits are approached and remains bounded at the constraint limit. On real-world humanoid whole-body control task, DeCap reduces upper-body action rate by 2.50x and acceleration by 2.18x relative to reward-based smoothness policies, while also improving lower-body smoothness and reducing transient motion. We demonstrate that a fixed set of smoothness constraints transfers across diverse terrains, alleviating the need of extensive reward tuning.
Utsav Panchal, Denis Kleyko, Unal Artan +1
AI, Robotics and Cybersecurity Center (ARC) and Department of Computer Science, Örebro University, Sweden