Reinforcement learning (RL) in simulation can train dexterous manipulation policies without robot demonstrations, but training a single generalist policy with task-agnostic rewards faces a severe exploration problem: approaching, grasping, and reorienting diverse objects with many degrees of freedom is difficult to discover from scratch. Prior works make exploration tractable with high-quality robot demonstrations, per-task reward shaping, or by restricting policies to narrow modes of behavior. We propose X-Reset, a framework that instead resolves exploration with human hand-object demonstrations. Rather than imitating or tracking retargeted human motion, X-Reset kinematically retargets hand-object states to noisy robot states, filters out states that are unstable in simulation, and samples the remainder as resets during RL training with general-purpose object-centric rewards. The resulting policy depends only on object state and goal, with demonstrations entering training through the reset distribution. We show that X-Reset trains generalist policies on 20 objects across three embodiments---a 22-DoF hand on two different arms and a parallel-jaw gripper---and resolves the exploration challenges of RL from scratch. X-Reset scales with the number of training objects, generalizes to unseen objects, can learn from imperfect hand-pose estimates, and transfers behaviors zero-shot from sim-to-real.
Figures & tables
Figure 1 . Overview of X-Reset . We introduce X-Reset , a simple recipe to pair generalist object-centric RL in simulation with kinematically retargeted human hand-object states as resets. Learning to approach, grasp, and reorient objects from scratch results in a challenging exploration problem, while X-Reset directly exposes policies to high-value states from which exploration is easy. With a common reward formulation, we train generalist policies for diverse robot embodiments and demonstrate sim-to-real transfer.
Figure 2 . X-Reset Pipeline. (Top) First, we kinematically retarget human hand-object states to noisy robot states. Then, we spawn states into simulation and retain a bank of stable states. (Bottom) During RL training, we randomly sample diverse resets leveraging the retargeted states, and train an RL policy which consumes object keypoints at inference time and can be deployed in the real world.
Figure 3 . X-Reset Supplements RL Training. We report reorientation success at ϵ= 2 cm on 20 DexYCB objects (7 Easy, 7 Medium, 6 Hard). Evaluation uses 20 demonstrated start/goal pairs per object. Pale bars show RL baselines trained from scratch; saturated caps add the X-Reset gain ( p=0.9 ), with gain labels in percentage points. Avg. weights all objects equally. X-Reset improves both reward formulations across embodiments and object groups, with the best recipe being low-bias object-only rewards seeded with X-Reset for exploration on average. Error bars show 3-seed RL training SE.
Figure 4 . Scaling and Generalization. We report reorientation success at ϵ= 2 cm on 12 EBench objects. Scratch baseline trains on the EBench roster for its full 10k epochs. Blue connected markers show zero-shot transfer after 10k epochs of X-Reset pretraining on 1–20 DexYCB objects, outperforming training from scratch on EBench with as little as 10 DexYCB training objects. Orange diamonds (FT) show 5k epochs of 20-object X-Reset pretraining followed by 5k epochs of EBench fine-tuning, showing X-Reset is a strong pre-training foundation. Error bars show 3-seed RL training SE.
Figure 5 . Impact of Demo-Reset Probability. We measure Avg. Episode Return from ρdefault at train-time for UR7e+Sharpa with Obj-Only Reward. Higher demo-reset probability accelerates early learning, with all X-Reset policies converging in a similar band. Curves show 3-seed means and shaded bands show ±1 SE.
Figure 6 . Learning from Noisy Hand Poses. (Top) We visualize retargeted Sharpa hand-object states under varying levels of noise. (Bottom) Avg. episode return from ρdefault of X-Reset trained with noisy hand poses.
Figure 7 . Sim-to-Real Evaluation. (Top) UR7e+Sharpa filmstrips show each object manipulated to goal poses in the air. (Bottom) Grasp and reorientation successes out of 10 physical trials per object for Obj-Only Reward+ X-Reset .