Human demonstrations are a scalable data source for learning dexterous manipulation, but the embodiment gap prevents human motion from being executed directly on robots. Inverse kinematics (IK) retargets human motion to robots efficiently but ignores dynamics, often producing infeasible motions. Reinforcement learning (RL) and sampling-based model predictive control (MPC) are commonly employed to yield dynamically feasible motions, but both are sample-inefficient and sensitive to hyperparameters. RL suffers from costly and unstable training and tedious reward engineering; MPC avoids policy optimization, yet retargets each trajectory in isolation, and solving one does not make the next easier. Sampling cost grows rapidly with dataset size and task difficulty. We hypothesize that dynamically feasible trajectories concentrate near a low-dimensional manifold shared across demonstrations, so that retargeting can be reduced to sampling from that manifold, conditioned on human motion, rather than solving a fresh optimization problem for every demonstration. We propose \textbf{Generative Neural Retargeting} (GNR), which uses a flow matching model to sample feasible trajectories. GNR outperforms MPC with only 8.5% of the samples required by MPC, achieving a success rate of 56.20% compared to 27.20% for MPC. GNR can be used for scalable and efficient retargeting of large-scale, long-horizon, and millimeter precision human demonstrations: by applying GNR within a real-to-sim data engine, we produce a dexterous manipulation dataset with dense contact-force labels, spanning 223k demonstrations and 3.3k object geometries.
Figures & tables
Figure 1: We propose Generative Neural Retargeting ( GNR ), a scalable and generalizable retargeting method that retargets large-scale, long-horizon and high-precision human demonstrations for five-fingered dexterous manipulation. We apply GNR with a real-to-sim data engine that reconstructs human motion from egocentric videos or a wearable exoskeleton with motion capture to scale up both diverse and difficult high-precision tasks.
Figure 2: The human-to-robot data engine. (a) A pipeline retargets motion-capture recordings for the difficulty axis and egocentric demonstrations for the diversity axis into candidate robot trajectories with simulation-derived contact-force labels. (b) The discovery and scaling phases repeatedly apply (a) under varied conditions and filter candidates using the same success criterion Γ ( Section 4.2 ), first identifying retargetable segments and then expanding their physical and geometric coverage. The resulting pairs train GNR to accelerate subsequent retargeting.
Overall
MPC-success
MPC-failure
Method
S
core-s
SR (%)
SR (%)
Epos (mm)
Erot ( ∘ )
Efinger ( ∘ )
SR (%)
Epos (mm)
Erot ( ∘ )
Efinger ( ∘ )
IK
1
39.0
18.30
50.00
34.92
16.38
0.10
6.46
54.52
34.59
0.13
DIAL-MPC
3,000
2689.3
27.20
100.00
25.33
10.98
1.15
0.00
36.29
21.73
1.49
MPPI
3,000
1849.6
27.70
71.32
28.04
11.96
1.22
11.40
36.92
20.59
2.07
iCEM
3,000
3862.1
34.70
78.68
21.96
7.07
1.74
18.27
26.04
10.92
2.14
CMA-ES
3,000
2912.7
32.00
72.79
23.21
8.55
1.02
16.76
32.89
14.55
1.08
Table 1: Performance comparison on HOT3D , split into MPC-success and MPC-failure datasets.
Figure 3: GNR performance on HOT3D across three OOD subsets. For each subset, we report GNR ’s SRs on MPC-success, MPC-failure, and all demonstrations.
Figure 4: GNR ’s performance on high-precision and long-horizon tasks. On ShapeFilter , open-loop GNR improves with the number of samples S , and GNR warm starts reduce MPC sampling requirements. On NIST , open-loop sampling alone is insufficient, while using GNR to initialize MPC ( GNR + MPC) outperforms MPC from scratch at matched sampling budget.
Figure 5: GNR ’s efficiency in data generation. For each task, left: test-time scaling using GNR trained with different seed dataset sizes N . Right: CPU cost reduction as dataset size ∣D∣ grows. Costs include MPC seed generation and GNR generation.
Method
Samples
SR (%) ↑
core-s ↓
ShapeFilter
MPC
2,500
93.7
10,124
GNR
77
94.7
678
NIST
MPC
3,440
92.2
795,777
GNR + MPC
1,976
89.1
266,199
Table 2: Comparing efficiency and SR using MPC and GNR .
Appendix figures & tables42 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Demonstrations
Geometries
Precision (mm)
Horizon
DexTrack ( Liu et al., 2025b )
3,585
257
100
short
AdaDexTrack ( Adalibieke et al., 2026 )
2,765
50
100
short
ManipTrans ( Li et al., 2025 )
3,300
1,200
30-80
short
DexMachina ( Mandi et al., 2025 )
7
5
10 to 90
short to long
SPIDER ( Pan et al., 2025 )
2,885
103
100
short
Do-As-I-Do ( Paliwal et al., 2026 )
500
N/A
100
short
Appendix
Table 3: Recent physics-based retargeting methods on the two scaling axes: diversity (number of demonstrations and distinct geometries) and difficulty (task precision and average demonstration length).
Benchmark
Success criterion
ShapeFilter
Object centroid lies in the insertion region [−80,80]×[0,160]×[−80,80] mm.
NIST
Horizontal error <1 mm, screwing down by at least 8 mm, and axis error <0.003 rad.
HOT3D
Object position error <150 mm and orientation error <40∘ . Mean capsule penetration ≤5 mm. Median squeeze force after grasp release ≥1 N. Per-finger force 95th percentile ≤100 N.
SPIDER
Mean object-position error ≤0.10 m and mean object-orientation error ≤0.50 rad.
Appendix
Table 4: Success criteria of different benchmarks.
Initial and final noise scales 1.0 ; exploitation fraction 0.15 ; exploitation noise 0.5 .
Appendix
Table 5: DIAL-MPC settings for the main experiments. ns denotes samples per iteration and M the maximum number of MPC iterations; Samples gives nsM for each phase.
Method
Benchmark
ns
M
Samples
Sampling settings
MPPI
HOT3D
300
10
3,000
Softmax-weighted update using DIAL-MPC sampling noise settings.
CEM
ShapeFilter
32 500
6
192 3,000
Elite ratio 0.10 ; mean/variance smoothing 0.10 ; white Gaussian noise; fixed sample count; optimizer reset at the start of each optimization.
iCEM
ShapeFilter HOT3D
32 300
6 10
192 3,000
Elite ratio 0.10 ; smoothing 0.10 ; colored-noise exponent 2.0 . HOT3D : fixed sample count; reuse 30% of elite samples; reset the optimizer at the start of each optimization; return the best trajectory rather than the average of elite trajectories.
CMA-ES
ShapeFilter HOT3D
32 300
6 10
192 3,000
Full covariance over optimized control dimensions. ShapeFilter : 29-dimensional control; initialized from the DIAL-MPC noise scale. HOT3D : hand control with the arm fixed; initial noise scale 0.5 ; relative noise scales from DIAL-MPC; simulate the selected trajectory once more (3,001 rollouts total).
Appendix
Table 6: Sampling-based optimizer settings.
Edit
Implementation
Variants per object
Regeneration
Regenerate from the reference image with a new random seed
20
Deformation
Deform the mesh by different amounts along each axis
60
Occlusion completion
Mask part of the reference image and complete the shape
20
Rotation
Rotate about the vertical axis, then regenerate
20
Stretch
Scale the original mesh along a single axis
80
Total
200
Appendix
Table 7: Geometry augmentations and variants per object.
Quantity
Original geometries
Modified geometries
Total
Evaluated demonstrations
186,310
173,840
360,150
DIAL-MPC successes
47,426
47,918
95,344
DIAL-MPC failures
138,884
125,922
264,806
Additional GNR successes ( S=256 )
74,391
52,745
127,136
MPC or GNR successes
121,817
100,663
222,480
Source motions
780
728
780
Appendix
Table 8: HOT3D statistics for the 360,150 evaluated demonstrations. Columns separate demonstrations using original and modified object geometries; original geometries are augmented in pose, timing, mass, and friction. Trajectory counts and distinct geometry counts are reported separately.
Figure 6: Objects in HOT3D .
Figure 7: Examples of geometry variants. Each row uses the same original object.
Figure 8: DG-5F grasp configurations at a shared camera view and scale.
Figure 9: Test-time scaling on 360k HOT3D demonstrations: 95k MPC-success and 264k MPC-failure demonstrations. The curves show GNR ’s SR across three OOD subsets.
Figure 10: HOT3D trajectory quality. From left to right: median object-position, object-orientation, and finger-joint angular errors on the randomly sampled 1,000-demonstration test set, followed by median finger-control jerk. GNR approximately halves jerk while preserving comparable object tracking.
Figure 11: ShapeFilter spatial generalization. The training (gray) and test (blue) quadrants are shown at left. At right, we show the test-time scaling curves of GNR on the test split.
Method
nsM or S
Success (%)
DIAL-MPC
32×6
48.4
DIAL-MPC
500×6
92.9
CEM
32×6
50.0
CEM
500×6
82.8
iCEM
32×6
65.6
CMA-ES
32×6
40.6
Appendix
Table 10: ShapeFilter success rates. Optimizer budgets are samples per iteration ns× MPC iterations M .
Samples S
State
Geometry
Geometry + state
1
50.42
50.93
53.48
2
67.06
66.38
69.10
4
78.44
80.31
78.61
8
87.10
88.62
86.59
16
92.36
93.72
94.06
32
95.93
96.77
97.11
Appendix
Table 11: ShapeFilter conditioning ablation. We report success rates (%) across different numbers of GNR samples S using different conditioning.
Figure 12: Performance of GNR using either residual or absolute prediction with one, two, or four training corners. Shadows show 95% confidence intervals.
Seed set size N
Quantity
Samples S
1
2
4
8
16
32
64
128
256
50
Success (%)
13.0
17.2
33.4
47.4
66.4
83.1
93.2
96.7
99.5
Cost (k core-h)
0.235
0.263
0.313
0.390
0.501
0.627
0.740
0.832
0.890
100
Success (%)
21.0
30.5
47.5
66.7
81.3
91.1
96.4
98.7
99.3
Cost (k core-h)
0.438
0.463
0.504
0.560
0.626
0.688
0.745
0.789
0.828
200
Success (%)
33.4
44.3
57.7
71.9
80.0
89.1
94.9
96.8
99.0
Appendix
Table 12: ShapeFilter dataset-generation efficiency. Dataset success rates include the successful seeds and apply the test-set success rate to the remaining demonstrations. Costs include MPC seed generation, GNR inference, and simulation, in thousands of core-h.
Figure 13: ShapeFilter generation cost versus dataset size. MPC generates all ∣D∣ demonstrations; GNR uses MPC for N successful seeds and generates the remaining ∣D∣−N with configurations that meet or exceed the best MPC success rate. Numbers show the reduction in total cost.
Figure 14: ShapeFilter trajectory errors and jerk comparison. GNR has lower jerk; GNR +MPC has the lowest object and fingertip tracking errors.
N=100
N=200
N=294
MPC samples
Success (%)
Cost (k core-h)
MPC samples
Success (%)
Cost (k core-h)
MPC samples
Success (%)
Cost (k core-h)
1,736
64.0
35.90
1,736
79.3
57.34
1,736
91.9
77.41
1,752
70.7
36.79
1,752
83.4
57.89
1,752
93.3
77.61
1,768
74.1
37.42
1,784
89.0
58.91
1,784
95.5
77.97
1,784
79.7
38.17
1,816
90.3
59.08
1,816
97.2
78.37
1,816
80.9
39.81
1,848
93.8
60.10
1,912
97.8
78.95
Appendix
Table 13: NIST dataset-generation efficiency. Dataset success rates include the successful seeds and estimated successes on the remaining demonstrations. Each MPC budget uses the schedule with the highest success rate. GNR generation costs include MPC seed generation, GNR inference, and GNR -initialized MPC, in thousands of CPU core-h (App. F.2 ).
Figure 15: NIST generation cost versus dataset size. MPC generates all ∣D∣ demonstrations; GNR+MPC uses MPC for N successful seeds and generates the remaining ∣D∣−N using evaluated configurations with similar dataset success rates. Numbers show the reduction in total cost.
Task
Method
Sample budget
Core-h / episode
Wall h / episode
HOT3D
DIAL-MPC
3,000
1.0777
0.2427
HOT3D
iCEM
3,000
1.4213
0.3282
HOT3D
CMA-ES
3,000
1.2272
0.2877
HOT3D
GNR , S=256
256
0.1094
0.0383
ShapeFilter
MPC
3,000
3.2001
0.0472
Appendix
Table 14: Measured per-episode generation cost.
Task
Generation
Generation compute ( 103 core-h)
Coverage
MPC data
GNR
Total
HOT3D
all MPC
390.043
–
390.043
26.4%
HOT3D
all MPC data + GNR S=256
390.043
26.047
416.090
61.6%
HOT3D
28k MPC seeds + GNR S=256
116.112
32.623
148.735
59.7%
ShapeFilter
all MPC
8.541
–
8.541
93.70%
ShapeFilter
50 MPC seeds + GNR
0.203
0.562
0.766
94.82%
Appendix
Table 15: Dataset-generation cost in thousands of core-h and coverage. GNR costs are estimated from per-trajectory costs.
Figure 16: GNR test-time scaling across five different dexterous hands.
Method
Samples
Success (%, 95% CI)
GNR
128
90.0 [82.6, 94.5]
GNR
256
93.0 [86.3, 96.6]
GNR
384
95.0 [88.8, 97.8]
SPIDER
16,384–32,768
96.0 [90.2, 98.4]
Appendix
Table 16: Multi-embodiment success on 100 test trajectories (20 per hand), with 95% Wilson confidence intervals. Samples are counted per demonstration for GNR and give the maximum budget per optimization for SPIDER.
Figure 17: SR curves of behavior cloning policy training using data generated by different retargeters.
Figure 18: ShapeFilter success across four data-flywheel rounds.
Figure 19: ShapeFilter example in which MPC fails while GNR and GNR+MPC succeed.
Figure 20: ShapeFilter example in which MPC fails while GNR and GNR+MPC succeed.
Figure 21: A NIST example. GNR+MPC succeeds where IK, MPC, and open-loop GNR fail.
Figure 22: A NIST example. GNR+MPC succeeds where IK, MPC, and open-loop GNR fail.
Figure 23: HOT3D coffee pot. GNR succeeds where MPC fails.
Figure 24: HOT3D ranch bottle with a stretched OOD geometry.
Figure 25: HOT3D barbecue sauce bottle with a stretched OOD geometry.
Figure 26: HOT3D parmesan can with an OOD deformation.
Figure 27: HOT3D dumbbell manipulation.
Figure 28: HOT3D waffle pick-and-place with an OOD rotation.
Figure 29: HOT3D holder with a stretched OOD geometry.
Figure 30: HOT3D wooden spoon with an OOD deformation.
Figure 31: HOT3D whiteboard marker. MPC makes the reconstructed human motion physically feasible.
Figure 32: HOT3D whiteboard eraser. MPC makes the reconstructed human motion physically feasible.
Figure 33: HOT3D dinosaur toy. MPC makes the reconstructed human motion physically feasible.
Human hand-object demonstrations provide a scalable source of data for dexterous robot learning, but transferring them across embodiments requires physically feasible retargeting. Existing physics-based methods typically optimize each demonstration independently, leading to either limited success under finite simulation budgets or training costs that grow with dataset size. We introduce FlashDexRetarget, an RL framework for multi-reference dexterous retargeting. We formulate retargeting as multi-reference tracking, jointly learning a single policy across many demonstrations with off-policy RL and geometric supervision of the demonstrated interactions. This shared training formulation amortizes optimization across references while enabling the policy to track diverse hand-object interactions. On a 50-motion benchmark from TACO, OakInk2, and HOT3D using XHand and Sharpa Wave Hand as target embodiments, FlashDexRetarget retargets 90% of demonstrations using about 30 GPU-hours, compared with about 46% at about 3,000 GPU-hours for CHORD. This corresponds to about 100 times lower training compute and a 44-percentage-point improvement in retargeting success. Ablations examine the key design choices, while experiments with up to 1,000 motions and real-world replay further demonstrate the scalability and practical applicability of our method.
Recent work in humanoid whole-body control has found success with a simple recipe: retarget human motion to robot kinematic references, then train policies via reinforcement learning (RL) to track them. But how does this recipe transfer to dexterous manipulation? The answer is not obvious, as manipulation involves complex, contact-rich dynamics and requires delicate regulation of contact modes and forces. We present REGRIND, a minimalist retargeting-guided RL pipeline that learns dexterous manipulation policies from a single human demonstration. REGRIND retargets human hand-object motion to a robot reference that preserves hand-object spatial and contact relationships, trains a residual RL policy in simulation to track object-centric keypoints along that reference, and transfers the resulting policy zero-shot to hardware with careful system identification. The resulting policies produce fluid, human-like behavior on two different multi-fingered hands across contact-rich tool-use tasks, including operating a pair of scissors and turning a screwdriver. Through systematic hardware experiments, we identify and analyze the key factors that govern sim-to-real transfer in dexterous manipulation, offering practical guidance for retargeting-based learning in contact-rich settings. Videos and code are available at https://yunhaifeng.com/REGRIND.
Human demonstrations offer rich examples of precise dexterous manipulation and a promising source of robot training data. However, high-fidelity reproduction of demonstrated motions and hand-object interactions across robot embodiments remains challenging under physical constraints. We present DexForge, a differentiable physics-grounded framework for converting human video demonstrations into high-fidelity robot trajectories. We reconstruct spherical-Gaussian object models and hand-object motion from visual observations, then build a differentiable simulator combining efficient Gaussian collision detection with existing differentiable dynamics. Based on this simulator, DexForge combines contact-aware kinematic retargeting with force-aware dynamics retargeting: robot-adapted stable contacts guide kinematic reference construction and subsequent gradient-based control refinement for precise physical motion reproduction. Experiments on 130 DexYCB and HOT3D demonstrations across seven dexterous hands show success-rate gains of approximately 35-53 percentage points over the baseline, with object position and orientation tracking errors on successful trajectories reduced by approximately 34-67% and 71-78%, respectively. Further experiments demonstrate open-loop transfer to MuJoCo and real-robot execution. Our project page is available at https://wmz1226.github.io/DexForge/
Meizhong Wang, Kun Cao, Ruiqi Ni +2
Tongji University · Shanghai Research Institute for Intelligent Autonomous Systems · Purdue University +1