cs.ROJul 7, 2026

ThorArena: Benchmarking Humanoid Physical Interaction with Human Motion-Force Demonstrations

Authors: Chenhao YuHongwu WangWeitao ZhangYouhao HuJiachen ZhangGangyang LiAlois KnollShaqi Luo

Organizations: Beijing Academy of Artificial Intelligence · Technical University of Munich

Abstract

Humanoid robots are increasingly expected to perform contact-rich tasks that require not only accurate whole-body motion but also robust physical interaction with surrounding objects and humans. Although recent advances in humanoid motion imitation and whole-body control have achieved remarkable tracking performance, existing datasets and benchmarks primarily focus on kinematic motion while largely overlooking synchronized interaction forces. As a result, current evaluations fail to capture how external interaction forces affect tracking accuracy, stability, and control robustness. In this paper, we present ThorArena, a benchmark for evaluating force-aware humanoid interaction based on human demonstrations with synchronized motion and force measurements. We collect a real-world interaction dataset that simultaneously captures whole-body human motion and forces exerted by both hands across six representative physical interaction tasks. Based on these demonstrations, we propose force-aware evaluation metrics that jointly assess whole-body tracking accuracy, robustness under different force levels, control effort, and episode survival through the Force-Aware Tracking Score (FATS) and complementary diagnostic metrics. We further establish a unified benchmark protocol that replays recorded interaction forces in simulation and provides a standardized evaluation interface for different humanoid control policies. Experiments on representative whole-body control policies demonstrate that force-aware evaluation reveals substantial performance differences that remain largely hidden under conventional no-force evaluation. ThorArena provides a practical and reproducible framework for studying force-aware humanoid interaction and offers a new benchmark for evaluating contact-rich humanoid behaviors.

Explore similar work

Oct 30, 2025cs.RO

Thor: Towards Human-Inspired Whole-Body Reactions for Intense Contact-Rich Environments

Maintaining whole-body stability and motion tracking under large interaction forces remains challenging for humanoids. We present Thor, a reinforcement learning framework for forceful humanoid loco-manipulation. Thor jointly trains lower-body, waist, and upper-body policies with shared whole-body observations and body-specific rewards to coordinate locomotion and force adaptation, waist posture regulation, and upper-body motion tracking. We further introduce a force-adaptive torso-tilt (FAT2) objective that derives a load-dependent horizontal center-of-mass offset reference from quasi-static moment balance. Capacity-matched simulation blations show that the three-policy architecture improves tracking under large external force disturbances, while real-world ablations demonstrate that FAT2 increases peak pulling capability. On the Unitree G1, Thor achieves mean peak dual-hand pulling forces of 167.7 N and 145.5 N during backward and forward locomotion, exceeding the best-performing baseline by 68.9% and 74.7%, respectively. Real-world demonstrations include opening a fire door with one hand using approximately 60 N of pulling force and towing a 1.7-ton passenger car.
Gangyang Li, Hongzhe Shi, Qing Shi +5
Jun 16, 2026cs.RO

HumanoidArena: Benchmarking Egocentric Hierarchical Whole-body Learning

Humanoid robots promise whole-body interaction in human-centered environments, but scalable policy learning remains difficult because task-level decision-making and whole-body dynamic execution are tightly coupled. A practical solution is hierarchical control, where a high-level policy predicts intermediate whole-body actions and low-level general motion trackers (GMTs) execute them as stable humanoid motion. However, existing benchmarks rarely evaluate the policy-tracker interface itself, leaving open whether intermediate whole-body actions are executable, robust under task distribution shifts, and transferable across different GMT backends. We introduce HumanoidArena, a simulation-first benchmark for egocentric hierarchical whole-body learning. The benchmark formulates policy learning as a hierarchical decision making problem: a high-level policy converts egocentric vision, proprioception, and instructions into a compact whole-body action, which is subsequently executed by a low-level GMT. Instead of treating the legs as planar transport tools, HumanoidArena emphasizes interactions where lower-body coordination is structurally necessary in task completion. We therefore design 7 leg-critical HOI/HSI tasks in which success requires foot placement, balance maintenance, posture adjustment, and whole-body reorientation. To further diagnose the hierarchical system, we evaluate policies from two complementary perspectives: perturbation-conditioned generalization and GMT-conditioned transfer. Experiments show that hierarchical control enables learned policies to solve diverse leg-critical interactions, but performance is strongly tracker-conditioned and cross-GMT transfer remains fragile. These results position HumanoidArena as a benchmark for studying transferable intermediate action representations and scalable egocentric whole-body policy learning.
Taowen Wang, Zikang Xie, Bin Yang +13
Aug 13, 2026cs.RO

HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark

Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with what people perceive in videos. Kinematic errors average per-frame pose differences but miss the physical artifacts that matter most, particularly unstable support and incorrect contacts such as foot skating and mistimed touch-downs. Meanwhile, widely used test suites are small and lack the diversity needed to stress contact-rich, long-horizon behaviors. We introduce HumanTracker to make humanoid tracking evaluation both perceptually aligned and scalable. The HumanTracker benchmark contains approximately 153 hours of optical motion trajectories from multiple professional performers, organized into four motion families with text labels for fine-grained diagnosis. We further propose HumanScore, a preference-aligned metric trained on 12K motion pairs containing 24K motions. Across representative state-of-the-art trackers, HumanScore better predicts human preferences and reveals contact and stability failures that kinematic metrics often miss.
Dairu Liu, Zekun Qi, Jiayu Zeng +11