cs.ROApr 1, 2026

BAT: Balancing Agility and Stability via Online Policy Switching for Long-Horizon Whole-Body Humanoid Control

Authors: Donghoon Baek, Sang-Hun Kim, Sehoon Ha

Organizations: Georgia Institute of Technology, Atlanta, GA, 30308, USA · Samsung Research

Abstract

Despite advances in reinforcement and imitation learning, achieving both precise control and dynamic whole-body behaviors in long-horizon tasks remains challenging. Existing approaches typically follow two paradigms: coupled whole-body policies for global coordination and decoupled policies for modular precision. However, effectively combining these two types of controllers to achieve agility, robustness, and precision remains challenging. In this work, we propose BAT, an online policyswitching framework that dynamically selects between coupled and decoupled whole-body RL generalist controllers according to the evolving motion context. BAT employs two complementary switching predictors: OpVQ-VAE provides motion-token-based predictions with strong generalization, while OpHRL provides closed-loop, state-aware predictions. A Token-Familiarity Router (TFR) selects between their predictions, favoring OpHRL for familiar token sequences and OpVQ-VAE for unfamiliar ones when their predictions disagree. BAT achieves 85.2% success on seen long-horizon motion combinations, while retaining 71.3% success on unseen motions. BAT also outperforms existing whole-body controllers on individual motions and demonstrates successful zero-shot deployment on the Unitree G1 humanoid.

Figures & tables

Explore similar work

CardsList
  1. X-WBC: A Cross-Embodiment Foundation Model for Humanoid Whole-Body Control

    Sep 14, 2026Juntong Zhang, Chun Gu, Li ZhangWhole-Body ControlCross-Embodiment Transfer

  2. Smoothness as a Constraint for Stable Humanoid Locomotion

    Sep 21, 2026Utsav Panchal, Denis Kleyko, Unal Artan +1Full-Body Humanoid ControlWhole-Body Control