cs.ROOct 5, 2026

I-BFM: Reward-Conditioned Robust Humanoid Interaction via Unsupervised Reinforcement Learning

Authors: Ziqi Han, Yitang Li, Junhan Sun, Fanrong Dong, Yaojie Shen, Lei Ye, Zetong Jing, Yongqi Zhang, +3 more

Organizations: Tongji University · Tsinghua University · Zhejiang University · RoboParty Lab · Harbin Institute of Technology · ShanghaiTech University

Abstract

Behavioral foundation models (BFMs) have recently shown that a single humanoid policy can support diverse whole-body control, but extending such generality to physical interaction remains challenging. We introduce I-BFM, to our knowledge the first BFM for humanoid-object interaction. Rather than relying on task-specific policies or reference tracking, I-BFM learns a shared representation of the coupled dynamics among the humanoid, objects, and their contacts using forward-backward representations and unsupervised reinforcement learning. Given a downstream task reward, the same policy can be directly conditioned on a latent command to execute closed-loop interaction without task-specific policy optimization. To improve interaction control over different time scales, we further train the policy with both short-horizon interaction targets and longer-horizon goal targets. A single I-BFM policy performs carrying, pushing, and kicking, while also supporting goal reaching, motion tracking, stylistic control, and long-horizon task chaining. More importantly, it remains effective after large deviations from nominal execution: on Carry, I-BFM achieves 94.3% nominal success and retains 89.3% success after robot falls, compared with 1.3% for a planning-based baseline. Real-world experiments on a Unitree G1 further demonstrate diverse loco-manipulation behaviors, rapid recovery from interaction failures and external disturbances, and task chaining without task-specific retraining.

Figures & tables

Explore similar work

CardsList
  1. CrossBFM: Distilling a Shared Latent Behavior Space Across Humanoid Embodiments

    Sep 29, 2026Tan-Dzung Do, Tuan Dat Phuong, Nico Bohlinger +6Cross-Embodiment TransferFull-Body Humanoid Control

  2. FAME: Force-Adaptive RL for Expanding the Manipulation Envelope of a Full-Scale Humanoid

    Mar 9, 2026Niraj Pudasaini, Yutong Zhang, Jensen Lavering +3Dexterous ManipulationWrist

  3. Scaling Behavior Foundation Model for Humanoid Robots

    Jul 16, 2026Weishuai Zeng, Kangning Yin, Xiaojie Niu +15Full-Body Humanoid ControlBehavioral Foundation Models