cs.ROOct 1, 2026

HumanoidToolBench: Benchmarking Humanoid Tool Use from Selection to Mobile Execution

Authors: Kyochul Jang, Seohyeon Park, Ohchul Kwon, Sangjun Park, Junhyeok Choi, Seungyeop Yi, Chaeyun Kim, Sangkyu Lee, +4 more

Organizations: Seoul National University · University of Massachusetts Amherst · Google Research

Abstract

As robotic hardware and learning methods advance, humanoids need tools to perform tasks beyond their inherent physical limits. Successful tool use requires selecting a suitable tool and coordinating manipulation and, when needed, locomotion to complete the task. Existing benchmarks do not jointly evaluate these capabilities on a humanoid. We introduce HumanoidToolBench, an 18-task benchmark spanning three scenarios, three execution levels, and two tool-set modes, together with ToolBook, a dataset of 3.1k demonstrations collected in simulation and on a real Unitree G1. Evaluation of seven policies in simulation and three on the real robot reveals substantial gaps between selecting a suitable tool and completing the task. Focused GR00T N1.7 probes show reduced selection accuracy on unseen tools and continued task execution under unrelated instructions. Code and data are available at https://snu-pi.github.io/HumanoidToolBench/.

Figures & tables

Explore similar work

CardsList
  1. IMBench: A Benchmark for Intuitive Robotic Manipulation

    Jul 17, 2026Anurag Maurya, Sukhvansh Jain, Prajwal Avhad +9Robotic ManipulationRobot Policies

  2. CodeActionBench: Evaluating Agentic Code-as-Policy for Embodied Manipulation

    Sep 27, 2026Yiheng Lyu, Xueying Jiang, Wenhao Li +2

  3. ATOM-Bench: A Real-World Benchmark for Atomic Skills and Compositional Generalization in Manipulation Policies

    Jun 15, 2026Zenan Wu, Bingqing Wei, Lu Liu +8Robotic Manipulation PoliciesCompositional Generalization