Dex-One2Many: Learning Dexterous Manipulation from a Single Human Demonstration
Organizations: Seoul National University · University of Maryland, College Park · All Purpose AI · KAIST
Abstract
While learning dexterous manipulation from a single human video offers a promising alternative to costly robot demonstrations, many recent methods predominantly imitate demonstrated motions. Such strict motion matching often limits generalization to initial object poses, goal poses, and grasps not shown in the video. Alternatively, discovering a policy via reinforcement learning (RL) allows for broad generalization, but without prior guidance, it struggles with high-dimensional exploration in complex, multi-stage tasks. To address these coupled generalization and exploration challenges, we present Dex-One2Many, a real-to-sim-to-real framework that learns a generalizable dexterous manipulation policy from a single human video. Our key insight is to abstract the video into sequential scene graphs that guide RL, enabling efficient exploration while preserving broad generalizability. The graphs serve as generative constraints for sampling diverse reset states and provide dense rewards for each stage. Because the graphs constrain relations rather than exact poses, these reset states cover object poses and grasps beyond the video, while initializing each stage from them with dense rewards keeps exploration short and guided. Trained entirely in simulation, Dex-One2Many transfers zero-shot to a real multi-fingered hand. Across five tool-use and manipulation tasks, Dex-One2Many exceeds baselines by 6.5% in seen configurations, while its robust generalization widens this gap to 71% in unseen scenarios.
Figures & tables
| UR3 Wuji 1 | UR5e Sharpa | UR5e Wuji 2 | UR5e Allegro | |
| Doll | 88.4 | 98.4 | 97.2 | 97.2 |
| Can | 93.6 | 97.4 | 95.2 | 92.2 |
| Stamp | 98.6 | 96.6 | 97.0 | 98.8 |
| Hammer | 94.6 | 94.2 | 97.6 | 91.8 |
| Sweep | 98.0 | 96.6 | 97.0 | 100.0 |
| Mean | 94.6 | 96.6 | 96.8 | 96.0 |
Appendix figures & tables27 assets
Supplementary material from the paper’s appendix.
Appendix
| Step | Model | Input Output | Sec. |
| Scene-graph extraction | Gemini 3.1 Pro | video entities, parts, stage sequence | B.2 |
| Instance segmentation | SAM 3.1 | first frame, visual_phrase object mask | B.3 |
| Mesh reconstruction | SAM 3D Objects | RGB, mask, metric point map metric mesh | B.3 |
| Object pose tracking | FoundationPose++ | RGB-D stream, mesh 6D pose | D |
| Physical parameters | GPT-5 | category, size mass, friction | B.4 |
| Part localization | SAM 3 | rendered views, part name part region | B.4 |
| Predicate | Reference | Target | Distance metric | Constraint |
| pre_on_top | top support disk | bottom point | ||
| on_top | ||||
| pre_inside | interior box | root point | ||
| inside | ||||
| pre_contact | center | working point | ||
| contact |
| Task | Description | Objects (role) | Predicate chain (stage ) | Type | |
|---|---|---|---|---|---|
| Doll | Pick up a Pikachu plush toy and place it inside an orange pot. | pikachu (held), pot † (site) | 5 | pre-grasp (hand, pikachu) grasp (hand, pikachu) pre-inside (pikachu, pot) success: inside (pikachu, pot) | direct |
| Can | Pick up the coke can and place it in a bucket. | can (held), bucket † (site) | 5 | pre-grasp (hand, can) grasp (hand, can) pre-inside (can, bucket) success: inside (can, bucket) | direct |
| Stamp | Grasp the stamp and press it down onto the stack of papers. | stamp † (held), paper (site) | 5 | pre-grasp (hand, stamp) grasp (hand, stamp) pre-on_top (stamp, paper) success: on_top (stamp, paper) | direct |
| Hammer | Use a hammer to strike a peg inserted into a block. | hammer † (held), peg † (acted-on), block † (site), holder † (fixture) | 6 | fixed_to (hammer, holder) pre-grasp (hand, hammer) grasp (hand, hammer) pre-contact (hammer, peg) contact (hammer, peg) success: seated (peg, block) | tool-use |
| Sweep | Sweep the toy cat into the dustpan using the brush. | brush (held), toy (acted-on), dustpan (site) | 7 | pre-grasp (hand, brush) grasp (hand, brush) pre-contact (brush, toy) contact (brush, toy) pre-inside (toy, dustpan) success: inside (toy, dustpan) | tool-use |
| Field | Content | Used for |
| task | ||
| name , description | Task name and one-line summary | Generated first to condition the subsequent fields; not used by the pipeline |
| objects[] | ||
| id | Identifier of each task-relevant object | Nodes; key for all per-object data |
| category | Object class | Hint in the physical-parameter query |
| visual_phrase | Text description of the object’s appearance | Object segmentation |
| Task | Video [s] | Latency [s] | Cost [$] |
| Doll | 7.07 | 21.2 | 0.045 |
| Can | 5.30 | 11.5 | 0.024 |
| Stamp | 4.97 | 12.0 | 0.026 |
| Hammer | 4.00 | 18.7 | 0.039 |
| Sweep | 5.77 | 18.5 | 0.037 |
| Total | 27.11 | 81.9 | 0.171 |
| Step | Tool | Parameters |
| Convex decomposition | CoACD ( Wei et al., 2022 ) | concavity threshold 0.05; resolution 2000; MCTS 20 nodes, 150 iterations, depth 3; ; manifold preprocessing only for non-manifold input (resolution 50); hull merging on; no hull-count limit; minimum part volume ( -mv ); random seed from the system clock |
| Watertight remesh | CoACD manifold mode (OpenVDB ( Museth et al., 2013 ) ) | applied to the normalised mesh; dual-marching-cubes level set 0.1 |
| Simplification | ACVD ( Valette and Chassery, 2004 ) | 2000 vertices, gradation 1.5, manifold output |
| Step | Setting |
| Views | 34 views: elevations 8 azimuths ( steps), plus top ( ) and bottom ( ); rendered twice (textured and uniform grey); px, vertical field of view , camera distance the half bounding-box diagonal |
| Detection | SAM 3; the part string is one text concept; instance score floor 0.15, mask threshold 0.5 |
| Negative concepts | grasping only: blade, cutting edge, spike, bristles ; their masks are dilated by 12 px and removed |
| Fusion | a vertex is visible in a view if its depth is within of the bounding-box diagonal of the z-buffer; score positive votes / visible views, over views with at least one detection; vertex kept if seen in views and score |
| Transfer to mesh | each simulation-mesh face takes the maximum score of the nearest triangle of the reconstructed mesh; face kept if score ; grasping region: largest connected component |
| Outputs | per-face mask and object-frame bounding box (centre, extents) of the grasping region; for grasp synthesis on the table, the mask is restricted to faces whose lowest vertex is mm above the table |
| PPO | Network | ||
| Parallel environments | Actor recurrence | LSTM, 1 layer, width | |
| Rollout horizon | steps ( s) | Actor recurrent norm | Layer norm |
| Learning epochs | Actor MLP | ||
| Mini-batches | Critic | MLP | |
| Clip range | Activation | ELU | |
| Discount | Action distribution | Gaussian, learned | |
| Property | How it is varied |
| Physical properties | |
| Object mass | Between the minimum and maximum mass estimated for each object |
| Object static friction | Up to 10% above or below the estimated friction |
| Object dynamic friction | Up to 10% above or below the estimated friction |
| Table friction | Static: 0.3–0.6; dynamic: 0.2–0.5 |
| Table mass | 0.5–1.5 times the original mass |