Dexterous manipulation poses a challenging high-dimensional optimization problem, as useful behaviors require coordinated motion across many hand joints. In reinforcement learning (RL) and sampling-based trajectory optimization, exploration commonly relies on independent robot joint perturbations, making coordinated behaviors difficult to discover. Prior work reduces this search space for grasp learning using low-dimensional spaces of coordinated joint motions learned from human hand data, but this restricts the expressivity required for general manipulation. Some combine learned and joint-space actions to restore expressivity, but this increases dimensionality and introduces redundancy. We study these effects across diverse manipulation settings, varying action dimensionality, exploration strategy, and the source of human data. Our experiments suggest that human-motion priors are most effective when used to structure exploration rather than change the action representation. Motivated by this finding, we propose EigenDEXplore, which induces correlated exploration by adding perturbations along human-derived eigenvectors to independent joint-space noise, leaving the action space unchanged. Across multiple dexterous hands, EigenDEXplore consistently outperforms joint-space and learned action-space baselines in grasping, in-hand reorientation, and contact-rich manipulation. These gains span unstructured and reference-guided RL, trajectory optimization, and sim-to-real deployment, and are largest in settings with less reward shaping and curriculum design.
Figures & tables
Figure 1: Action spaces and exploration noise. Joint-Space perturbs joints independently, which can produce unnatural hand configurations. Eigen-Space encourages coordinated motion but constrains actions to a learned PCA subspace. Eigen-Residual restores full joint-space expressivity by combining joint and eigen actions, but it introduces redundancy in the action parameterization. EigenDEXplore combines joint and eigen noise while retaining a joint-space action representation.
Figure 2: Constructing robot principal components from human motion. We retarget EgoSuite human hand motion to each robot hand and apply PCA in its joint space. We visualize the first three principal components for Sharpa, Wuji, Allegro, and XHand.
Figure 3: SimToolReal. EigenDEXplore achieves the highest mean consecutive successes and reward in dexterous tool manipulation. (a) Fixed-budget performance at 9B frames. Error bars show 95% CIs and shading shows ± 1 SE across five seeds. (b) Extended training with a success-tolerance curriculum using one median run per method.
Figure 4: DeXtreme. EigenDEXplore achieves the highest mean consecutive successes for in-hand reorientation. Left: performance at 2B frames with DR Level fixed at −0.5 . Right: DR Level during training. Error bars show 95% CIs and shading shows ± 1 SE across five seeds.
Figure 5: DexMachina. With curriculum and auxiliary rewards, EigenDEXplore performs similarly to Joint-Space in most settings (a). However, with task reward alone, EigenDEXplore substantially improves object tracking on higher-DoF Allegro and XHand (b, left). Representative task-reward-only rollouts (b, right). Bars show mean ADD-AUC with 95% CIs across five seeds.
Figure 6: SPIDER across hands. EigenDEXplore reduces mean retargeting cost across all hands, with the largest reductions on higher-DoF hands and the smallest on lower-DoF hands. Δ cost (%) is relative to the Joint-Space baseline. Error bars show 95% CIs over ten seed-wise task averages.
Figure 7: Dataset and exploration analysis. (a) Dataset selection: EgoSuite achieves the highest mean consecutive success and training reward in SimToolReal at 9B frames. (b) Exploration design: Sharpa SPIDER ablations vary the retained PCA variance (and thus the number of PCA components k ) with and without joint-space noise. The selected 90% setting with joint-space noise achieves the lowest mean retargeting cost. Shading shows ±1 SE and error bars show 95% CIs.
Figure 8: Real-world SimToolReal. EigenDEXplore achieves higher task progress on tool-use. For each task, representative rollouts (left) and task progress, the fraction of waypoints reached (right). Boxes show quartiles/medians, diamonds show means, and dots show individual trials.
Figure 9: Real-world DeXtreme. EigenDEXplore increases mean consecutive successes from 8.45 to 11.90 over 40 trials per method. Left: hardware setup. Right: consecutive successes per trial. Boxes show quartiles/medians, diamonds show means, and dots show individual trials.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 10: Delta action spaces. Delta variants of Joint-Space and Eigen-Residual perform comparably to their absolute counterparts in SimToolReal but fall behind them in DeXtreme. (a) SimToolReal at 9B frames, with consecutive successes (left) and reward (right). (b) DeXtreme at 2B frames, with consecutive successes at DR Level fixed at −0.5 (left) and DR Level during training (right). Error bars show 95% CIs and shading shows ± 1 SE across five seeds.
Figure 11: Exploration design across datasets. Priors fitted to GRAB, ARCTIC, and EgoSuite follow the same trends, so the benefit of combining eigen and joint-space noise does not depend on the source dataset. (a) With eigen noise alone, retargeting cost decreases as more components are retained. (b) Adding joint-space noise lowers cost at nearly every threshold. Dashed blue lines mark Joint-Space. Shading shows ± 1 SE.
Figure 12: Human-derived versus random directions. The benefit of correlated exploration comes from the human-derived directions. Human-derived priors match or improve on Joint-Space, while random directions with the same eigenvalues are worse, both with eigen noise alone (a) and combined with joint-space noise (b). All eigen variants retain 90% of the variance. Error bars show 95% CIs over episodes.