Organizations: Institute for AI, Peking University. · School of Psychological and Cognitive Sciences, Peking University. · Yuanpei College, Peking University. · Beijing Key Laboratory of Brain-Computer Interface and Mental Health Modulation, Peking University. · State Key Laboratory for General AI, Peking University. · Embodied Intelligence Lab, PKU-Wuhan Institute for AI. · XS Vision.
The ability to design a tool for a task marks a level of intelligence beyond merely understanding, selecting, or using one. Existing methods for robotic tool design typically optimize a tool's continuous shape and action within a structure that is prescribed or generated beforehand, so the structure itself stays outside the physical optimization loop. We study task-driven tool design from scratch, where tool structure, shape, and action are all derived from the desired physical outcome. Here we show that the three elements can be designed jointly by HOT, a hierarchical optimization whose upper level searches over discrete tool structures with BASS, while lower-level physical optimization evaluates their task behavior and returns milestone progress as behavioral evidence for the search, ultimately providing jointly optimized shape and action. On four tool-use tasks with distinct physical functions, HOT discovers functional structures after evaluating only a small fraction of search spaces containing up to 56 million structures, and the subsequent refinement of their geometry lowers the task loss on all tasks while preserving success, through deformations that are functionally interpretable. Once 3D printed, the tools accomplish all tasks on a real robot with the actions found in simulation. Designing tools from required physical effects, rather than a catalog of known tools, is a step toward the open-ended tool making seen in humans and animals.
Figures & tables
Fig. 1 : The HOT framework. (a) A task specifies a scene and a desired physical outcome. (b) Tools are composed by a constructive grammar from a fixed handle and a library of searchable primitives. (c) Equivalent construction states are canonicalized into a dag that aggregates their behavioral evidence. In Stage 1, bass selects a promising partial state by Bayesian lookahead (d), samples a complete structure (e), evaluates it under action-only optimization (f), and propagates the achieved milestones back as behavioral feedback (g). Stage 2 co-refines the shape and action of each retained structure (h).
Fig. 2 : Tools designed by HOT and their use in simulation. Each row shows two representative successful tool designs. For each design, we list the primitives selected by bass with their counts, the Stage 1 structure they compose, and its Stage 2 refinement, where the heat map shows the absolute vertex displacement from the Stage 1 shape in millimeters; the frames on the right show the refined tool rolled out on its task. Rows (a)–(d) correspond to SweepBalls , TorqueBolt , ScoopBalls , and Hammer&ExtractNail .
Task
Design Space Size
FSI
Core-h
SweepBalls
1,296,198
10,142 (7.82‰)
916.63
TorqueBolt
1,296,198
12,263 (9.46‰)
170.59
ScoopBalls
55,784,462
11,363 (0.20‰)
1,129.84
Hammer&ExtractNail
30,389,168
10,320 (0.34‰)
1,248.33
TABLE I : Stage 1 structural search cost: Design space size (terminal states in DAG), FSI of bass , and time consumed (Core-h) for each task.
Lookahead step(s)
FSI ( ↓ ) with calibration budget
9,800
980
0
N=2 (lookahead)
12,263
11,103
> 40,000
N=1 (greedy)
12,320
> 40,000
> 40,000
TABLE II : FSI results of method variations on TorqueBolt .
Fig. 3 : Milestone achievement during bass structural search across the four tasks. Each curve shows the rolling fraction of evaluated structures that reach a given milestone level, computed over a window of 100 consecutive evaluations. The first 9,800 evaluations constitute the random calibration phase, after which bass uses milestone feedback to guide structural search. Green ticks mark evaluations that achieve all milestones.
Fig. 4 : Real-world tool use. The tools are 3D printed, mounted on a RealMan RM75-B arm, which executes the rolled-out trajectories.
While humans readily repurpose a book, a stone, or a shoe to drive a nail, robots trained on specific tools fail to transfer the same function to novel ones -- a gap we formalize as functional generalization. Functionally equivalent tools share visually recognizable functional intent, such as where contact can occur and how a contact region should move to the target. However, this perceptual similarity does not directly carry over to action space, where each tool demands a different motor pattern to realize the function. To bridge this gap, we explore intermediate representations including affordance images, human video prompts, functional videos and object masks, and 2D keypoint trajectories, finding that keypoint trajectories best balance functional expressiveness and action groundability. Building on this, we present FuncBridge, a two-stage framework that decouples functional reasoning from action execution: learning to predict generalizable keypoint trajectories from action-free data, then grounding them into robot actions with limited demonstrations. Across a benchmark spanning ten tools and three functions, including hitting, sweeping, and hooking, FuncBridge consistently outperforms state-of-the-art methods on unseen tools in both simulation and the real world.
Chuhao Zhou, Liquan Wang, Shuxin Cao +5
MARS Lab, Nanyang Technological University · Georgia Institute of Technology
We propose a causal reasoning framework for creative robot tool use where a suitable tool for a task is correctly identified for use beyond its primary objectives. The proposed framework first discovers the causal relationships between the tool and the task by conducting simulated experiments in a dynamics model. We decouple the causal discovery problem into two complementary components: VLM-based feature suggestion and counterfactual tool generation via targeted geometric and physical feature perturbations. Then, novel objects are classified based on identified causal features, and the tool use skill is transferred via keypoint matching conditioned on the identified causal features. By reconstructing the task in a dynamics model, our approach grounds tool use in the physics of the problem. We illustrate our approach in reaching a distant object with different sticks, scooping candies from a bowl using diverse items, and using different boxes or crates as stepping platforms to retrieve an object from a high shelf. Our baseline comparisons show that identifying causal features and grounding them in physical tool properties leads to more reliable tool selection and stronger skill keypoint transfer.
M. Tuluhan Akbulut, Varun Satheesh, Ahmed Jaafar +5
As robotic hardware and learning methods advance, humanoids need tools to perform tasks beyond their inherent physical limits. Successful tool use requires selecting a suitable tool and coordinating manipulation and, when needed, locomotion to complete the task. Existing benchmarks do not jointly evaluate these capabilities on a humanoid. We introduce HumanoidToolBench, an 18-task benchmark spanning three scenarios, three execution levels, and two tool-set modes, together with ToolBook, a dataset of 3.1k demonstrations collected in simulation and on a real Unitree G1. Evaluation of seven policies in simulation and three on the real robot reveals substantial gaps between selecting a suitable tool and completing the task. Focused GR00T N1.7 probes show reduced selection accuracy on unseen tools and continued task execution under unrelated instructions. Code and data are available at https://snu-pi.github.io/HumanoidToolBench/.
Kyochul Jang, Seohyeon Park, Ohchul Kwon +9
Seoul National University · University of Massachusetts Amherst · Google Research