The ability to manipulate tools significantly expands the set of tasks a robot can perform. Yet, tool manipulation represents a challenging class of dexterity, requiring grasping thin objects, in-hand object rotations, and forceful interactions. Since collecting teleoperation data for these behaviors is challenging, sim-to-real reinforcement learning (RL) is a promising alternative. However, prior approaches typically require substantial engineering effort to model objects and tune reward functions for each task. In this work, we propose SimToolReal, taking a step towards generalizing sim-to-real RL policies for tool manipulation. Instead of focusing on a single object and task, we procedurally generate a large variety of tool-like object primitives in simulation and train a single RL policy with the universal goal of manipulating each object to random goal poses. This approach enables SimToolReal to perform general dexterous tool manipulation at test-time without any object or task-specific training. We demonstrate that SimToolReal outperforms prior retargeting and fixed-grasp methods by 37% while matching the performance of specialist RL policies trained on specific target objects and tasks. Finally, we show that SimToolReal generalizes across a diverse set of everyday tools, achieving strong zero-shot performance over 120 real-world rollouts spanning 24 tasks, 12 object instances, and 6 tool categories.
Figures & tables
Fig. 2: Overview of SimToolReal . (Top) Training in Simulation: We train a goal-conditioned RL policy in simulation that manipulates a wide variety of procedurally-generated objects to randomly sampled goal poses. (Bottom) Inference in Real: We deploy this policy zero-shot on real-world tools from DexToolBench , following tool trajectories from human videos.
Fig. 3: Real-World Deployment. (Left) Human Video Processing: We collect an RGB-D human video and process it using vision foundation models. We use SAM 3D [ 16 ] to generate a metric-scale object mesh and segment a 3D grasp bounding box. Then, we use FoundationPose [ 80 ] to extract a sequence of 6D goal poses. (Right) Inference-Time Pipeline: Our LSTM policy takes in proprioception, object pose, grasp bounding box, and goal pose, and it outputs joint position targets for the 29-DoF dexterous robot (arm + hand).
Fig. 4: Generalization to Unseen DexToolBench Tools and Tasks in the Real World. We evaluate our policy in the real world on unseen tool-use tasks in DexToolBench . Our evaluations span 24 unique task trajectories across 6 different object categories and 12 object instances. Each bar corresponds to 1 task trajectory on 1 object instance. We report the average Task Progress across 5 rollouts. Despite not being trained on these objects or trajectories, our policy demonstrates strong generalization to diverse tools of varying masses and geometries.
Fig. 5: Comparison against Baselines in the Real World. We compare SimToolReal against baselines on two variations of sweeping a table with a brush: with and without requiring tool rotation based on the initial states shown on the left. Average Task Progress is indicated in parentheses. SimToolReal succeeds on both variations, performing dexterous in-hand tool rotations in the harder variation. Fixed Grasp succeeds on the simpler variation of this task without tool rotation. However, when rotation is required, enforcing a fixed grasp causes the arm to collide with the table while tracking the target trajectory. Kinematic Retargeting fails to reason about contact forces, and is unable to grasp the brush in both variations.
Fig. 6: Comparison against Specialists. We compare SimToolReal in simulation against specialist policies trained on a single object (Obj A) and trajectory (Traj A). We train one specialist policy for each of the 6 object categories in DexToolBench and report the average Task Progress across these categories. While the specialists succeed on their training setup (Obj A / Traj A), performance degrades under deviation in the trajectory (Obj A / Traj B) or the object (Obj B / Traj A). SimToolReal has high zero-shot performance across all variants, despite not being trained on these objects or trajectories.
Fig. 7: Training Objective Drives Generalization. In simulation, we evaluate (Left) the episode reward on procedurally-generated objects and (Right) the zero-shot Task Progress on unseen DexToolBench tools on different policy checkpoints throughout training. The strong correlation between the two curves validates our core hypothesis: improving random goal-pose reaching performance on diverse object primitives drives corresponding gains in generalization to unseen tool-use tasks.
Fig. 8: Ablation of RL Training Components. We compare the training reward across environment steps of SimToolReal against ablations averaged across 5 seeds. Replacing SAPG [ 71 ] with PPO [ 66 ] or not using Asymmetric Critic [ 60 ] results in a significant performance drop, highlighting their importance.
Fig. 9: Object Keypoints. (Left) Visualization of 4 pose keypoints in the local object frame. (Right) Visualization of the keypoint distances used to compute the distance d(ot,g) .
Fig. 10: Representative Examples of DexToolBench Tasks. Visual breakdown of representative tasks in DexToolBench across the 6 tool categories, highlighting the diversity of objects and manipulation tasks.
Fig. 11: Visualization of Policy Observations during Real-World Deployment. (Top) Image frames from a real-world rollout of the brush manipulation task. (Bottom) The corresponding visualization of the policy observations at each timestep, consisting of the robot state, estimated object pose, and goal pose (green). When the distance between the current pose and goal pose is sufficiently small, the goal pose is updated to the next pose in the goal sequence. Note that this visualization is a rendering of the policy’s inputs, not a physics-based simulation.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Value
Parameter
Value
Environment & Control
SAPG Hyperparameters
Simulation / Control Frequency
120 / 60 Hz
Actor Network
LSTM[1024] + MLP[1024,1024,512,512]
Num. Environments
24576
Critic Network
MLP[1024,1024,512,512]
Episode Length
600 steps
Learning Rate
1×10−4
Obj. Pos. Range ( x,y )
±10 cm
Minibatch Size
98,304
Table Height Range ( z )
±1 cm
SAPG Block Size
4096
Appendix
TABLE I: Simulation environment and SAPG training hyperparameters.
Fig. 12: Metric-Scale Mesh and Grasp Bounding Box Acquisition. We present a semi-automated pipeline to extract a metric-scale mesh and grasp bounding box from a single RGB-D video scan. (1) We first segment the target object in the initial frame and reconstruct a 3D mesh using SAM 3D [ 16 ] , injecting the captured depth map to ensure metric accuracy. (2) To establish a canonical coordinate frame, we virtually render the mesh from multiple views and use SAM 2 [ 65 ] to segment the geometry into “handle” and “head” regions based on user prompts. (3) The final grasp bounding box is derived from the handle’s geometry: it is centered on the handle, with the positive x -axis oriented toward the head, ensuring consistent alignment with policy training.
Category
Instance
Trajectory
R1
R2
R3
R4
R5
Avg
Hammer
Claw
Swing Down
100
100
100
100
100
100
Swing Side
100
100
77.5
100
100
95.5
Mallet
Swing Down
100
100
33.3
100
88.9
84.4
Swing Side
65.6
81.3
100
78.1
62.5
77.5
Marker
Sharpie
Draw Smile
100
100
100
100
0.0
80.0
Write C
100
100
100
32.0
56.0
77.6
Appendix
TABLE II: Detailed Real-World Evaluation Results. We report the Task Progress (%) for each of the 5 rollouts across all 24 object-task variations. The specific tools correspond to the instances shown in Fig. 4 .
Fig. 13: DexToolBench Objects. (Left): Real-world objects. (Right): SAM 3D [ 16 ] generated meshes of these objects.
Fig. 14: Kinematic Retargeting Pipeline. From the RGB-D human video, we use SAM 2 [ 65 ] for hand masks and HaMeR [ 58 ] for hand pose prediction. Next, we use ICP registration to align the hand pose prediction with the segmented hand point cloud to obtain accurate 3D hand poses. Lastly, we perform IK-based retargeting of the arm and hand to match the human wrist pose and fingertip positions.
Fig. 15: Fixed Grasp Baselines. We first use our SimToolReal policy to grasp and lift the object. We then attempt to follow the trajectory using a fixed grasp. Option 1 (Damped Least Squares): Tracks the target poses but causes severe collisions with the table. Option 2 (Collision-Free Trajectory Optimization): Avoids collisions but fails to reach the target poses.
Fig. 16: Detailed Comparison against Specialists. We provide a breakdown of the results in Fig. 6 . Each plot compares SimToolReal to a specialist policy trained on a single object and trajectory (Obj A / Traj A) for a single object category.
Category
Obj A
Traj A
Obj B
Traj B
Brush
Red
Sweep
Blue
Sweep
Brush
Forward
Brush
Right
Eraser
Flat
Wipe
Handle
Wipe
Eraser
Smile
Eraser
C
Hammer
Mallet
Swing
Claw
Swing
Hammer
Down
Hammer
Side
Appendix
TABLE III: Specialist Objects and Trajectories. Objects and trajectories used for specialist policy training and evaluation.
Sim-to-real transfer remains a critical bottleneck for deploying dexterous manipulation policies learned in simulation to real-world robots. Existing approaches rely on manually designed domain randomization or task-specific adaptation, limiting their generalizability across diverse manipulation scenarios. We present DexSim2Real, an integrated framework that leverages vision-language foundation models to bridge the sim-to-real gap for dexterous manipulation. Our system combines three components: (1) Foundation Model-Guided Domain Randomization (FM-DR), which uses a vision-language model as a visual realism critic to optimize simulation parameters via closed-loop CMA-ES, complementing text-based approaches like DrEureka with direct visual feedback; (2) a Tactile-Visual Cross-Attention Policy (TVCAP) that adapts cross-attention visuo-tactile fusion to zero-shot sim-to-real RL; and (3) a Progressive Skill Curriculum (PSC) that builds on LLM-based task decomposition with a difficulty scheduler tailored to contact-rich dexterous tasks. Extensive experiments on six challenging manipulation tasks with blinded evaluation demonstrate that DexSim2Real achieves a 78.2% average real-world success rate, outperforming DrEureka and DeXtreme while reducing the sim-to-real performance gap to only 8.3%.
Zijian Zeng, Fei Ding, Huiming Yang +2
Tsinghua University · Alibaba Group · Bengbu University +1
Articulated tool manipulation remains a major challenge in dexterous robotics due to the need to coordinate internal degrees of freedom and contact-rich interactions. While prior work has largely focused on rigid objects, articulated tool use remains underexplored because of its physical complexity and the difficulty of learning functional grasping and manipulation policies. We present Mana (Manipulation Animator), a general sim-to-real framework that reinterprets dexterous manipulation as an animation problem. Inspired by computer animation, Mana employs a coarse-to-fine pipeline that transforms procedurally-generated grasp keyframes into manipulation trajectories through motion planning and reinforcement learning. The data generation process is largely automatic, requiring only a few mouse clicks to specify functional affordances (<1 minute per tool). Across four articulated tools spanning different scales and joint types, Mana achieves zero-shot sim-to-real transfer for both grasping and in-hand manipulation, demonstrating a scalable approach to dexterous articulated tool use.
Human-like dexterous hands with multiple fingers offer human-level manipulation capabilities but remain difficult to train the control policies that can deploy on real hardware due to contact-rich physics and imperfect actuation. We present a sim-to-real reinforcement learning method that leverages dense tactile feedback combined with joint torque sensing to explicitly regulate physical interactions. To enable effective sim-to-real transfer, we introduce (i) a computationally fast tactile simulation that computes distances between dense virtual tactile units and the object via parallel forward kinematics, providing high-rate, high-resolution touch signals needed by RL; (ii) a current-to-torque calibration that eliminates the need for torque sensors on dexterous hands by mapping motor current to joint torque; and (iii) actuator dynamics modeling with randomization to account for non-ideal torque-speed effects and bridge the actuation gaps. Using an asymmetric actor-critic PPO pipeline, we train policies entirely in simulation and deploy them directly to a five-finger hand. The resulting policies demonstrate two essential human-hand skills: (1) command-based controllable grasp force tracking and (2) reorientation of objects in the hand, both of which are robustly executed without fine-tuning on the robot. By combining tactile and torque in the observation space with scalable sensing and actuation modeling, our system provides a practical solution to achieve reliable dexterous manipulation. To our knowledge, this is the first demonstration of controllable grasping on a multi-finger dexterous hand trained entirely in simulation and transferred zero-shot on real hardware.
Zhe Zhao, Zhibin Li, Yilin Ou +1
State Key Laboratory of Networking and Switching Technology Beijing University of Posts and Telecommunications, China · University College London, United Kingdom