Human demonstrations are a scalable data source for learning dexterous manipulation, but the embodiment gap prevents human motion from being executed directly on robots. Inverse kinematics (IK) retargets human motion to robots efficiently but ignores dynamics, often producing infeasible motions. Reinforcement learning (RL) and sampling-based model predictive control (MPC) are commonly employed to yield dynamically feasible motions, but both are sample-inefficient and sensitive to hyperparameters. RL suffers from costly and unstable training and tedious reward engineering; MPC avoids policy optimization, yet retargets each trajectory in isolation, and solving one does not make the next easier. Sampling cost grows rapidly with dataset size and task difficulty. We hypothesize that dynamically feasible trajectories concentrate near a low-dimensional manifold shared across demonstrations, so that retargeting can be reduced to sampling from that manifold, conditioned on human motion, rather than solving a fresh optimization problem for every demonstration. We propose \textbf{Generative Neural Retargeting} (GNR), which uses a flow matching model to sample feasible trajectories. GNR outperforms MPC with only 8.5% of the samples required by MPC, achieving a success rate of 56.20% compared to 27.20% for MPC. GNR can be used for scalable and efficient retargeting of large-scale, long-horizon, and millimeter precision human demonstrations: by applying GNR within a real-to-sim data engine, we produce a dexterous manipulation dataset with dense contact-force labels, spanning 223k demonstrations and 3.3k object geometries.
Figures & tables
Figure 1: We propose Generative Neural Retargeting ( GNR ), a scalable and generalizable retargeting method that retargets large-scale, long-horizon and high-precision human demonstrations for five-fingered dexterous manipulation. We apply GNR with a real-to-sim data engine that reconstructs human motion from egocentric videos or a wearable exoskeleton with motion capture to scale up both diverse and difficult high-precision tasks.
Figure 2: The human-to-robot data engine. (a) A pipeline retargets motion-capture recordings for the difficulty axis and egocentric demonstrations for the diversity axis into candidate robot trajectories with simulation-derived contact-force labels. (b) The discovery and scaling phases repeatedly apply (a) under varied conditions and filter candidates using the same success criterion Γ ( Section 4.2 ), first identifying retargetable segments and then expanding their physical and geometric coverage. The resulting pairs train GNR to accelerate subsequent retargeting.
Overall
MPC-success
MPC-failure
Method
S
core-s
SR (%)
SR (%)
Epos (mm)
Erot ( ∘ )
Efinger ( ∘ )
SR (%)
Epos (mm)
Erot ( ∘ )
Efinger ( ∘ )
IK
1
39.0
18.30
50.00
34.92
16.38
0.10
6.46
54.52
34.59
0.13
DIAL-MPC
3,000
2689.3
27.20
100.00
25.33
10.98
1.15
0.00
36.29
21.73
1.49
MPPI
3,000
1849.6
27.70
71.32
28.04
11.96
1.22
11.40
36.92
20.59
2.07
iCEM
3,000
3862.1
34.70
78.68
21.96
7.07
1.74
18.27
26.04
10.92
2.14
CMA-ES
3,000
2912.7
32.00
72.79
23.21
8.55
1.02
16.76
32.89
14.55
1.08
Table 1: Performance comparison on HOT3D , split into MPC-success and MPC-failure datasets.
Figure 3: GNR performance on HOT3D across three OOD subsets. For each subset, we report GNR ’s SRs on MPC-success, MPC-failure, and all demonstrations.
Figure 4: GNR ’s performance on high-precision and long-horizon tasks. On ShapeFilter , open-loop GNR improves with the number of samples S , and GNR warm starts reduce MPC sampling requirements. On NIST , open-loop sampling alone is insufficient, while using GNR to initialize MPC ( GNR + MPC) outperforms MPC from scratch at matched sampling budget.
Figure 5: GNR ’s efficiency in data generation. For each task, left: test-time scaling using GNR trained with different seed dataset sizes N . Right: CPU cost reduction as dataset size ∣D∣ grows. Costs include MPC seed generation and GNR generation.
Method
Samples
SR (%) ↑
core-s ↓
ShapeFilter
MPC
2,500
93.7
10,124
GNR
77
94.7
678
NIST
MPC
3,440
92.2
795,777
GNR + MPC
1,976
89.1
266,199
Table 2: Comparing efficiency and SR using MPC and GNR .
Appendix figures & tables42 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Demonstrations
Geometries
Precision (mm)
Horizon
DexTrack ( Liu et al., 2025b )
3,585
257
100
short
AdaDexTrack ( Adalibieke et al., 2026 )
2,765
50
100
short
ManipTrans ( Li et al., 2025 )
3,300
1,200
30-80
short
DexMachina ( Mandi et al., 2025 )
7
5
10 to 90
short to long
SPIDER ( Pan et al., 2025 )
2,885
103
100
short
Do-As-I-Do ( Paliwal et al., 2026 )
500
N/A
100
short
Appendix
Table 3: Recent physics-based retargeting methods on the two scaling axes: diversity (number of demonstrations and distinct geometries) and difficulty (task precision and average demonstration length).
Benchmark
Success criterion
ShapeFilter
Object centroid lies in the insertion region [−80,80]×[0,160]×[−80,80] mm.
NIST
Horizontal error <1 mm, screwing down by at least 8 mm, and axis error <0.003 rad.
HOT3D
Object position error <150 mm and orientation error <40∘ . Mean capsule penetration ≤5 mm. Median squeeze force after grasp release ≥1 N. Per-finger force 95th percentile ≤100 N.
SPIDER
Mean object-position error ≤0.10 m and mean object-orientation error ≤0.50 rad.
Appendix
Table 4: Success criteria of different benchmarks.
Initial and final noise scales 1.0 ; exploitation fraction 0.15 ; exploitation noise 0.5 .
Appendix
Table 5: DIAL-MPC settings for the main experiments. ns denotes samples per iteration and M the maximum number of MPC iterations; Samples gives nsM for each phase.
Method
Benchmark
ns
M
Samples
Sampling settings
MPPI
HOT3D
300
10
3,000
Softmax-weighted update using DIAL-MPC sampling noise settings.
CEM
ShapeFilter
32 500
6
192 3,000
Elite ratio 0.10 ; mean/variance smoothing 0.10 ; white Gaussian noise; fixed sample count; optimizer reset at the start of each optimization.
iCEM
ShapeFilter HOT3D
32 300
6 10
192 3,000
Elite ratio 0.10 ; smoothing 0.10 ; colored-noise exponent 2.0 . HOT3D : fixed sample count; reuse 30% of elite samples; reset the optimizer at the start of each optimization; return the best trajectory rather than the average of elite trajectories.
CMA-ES
ShapeFilter HOT3D
32 300
6 10
192 3,000
Full covariance over optimized control dimensions. ShapeFilter : 29-dimensional control; initialized from the DIAL-MPC noise scale. HOT3D : hand control with the arm fixed; initial noise scale 0.5 ; relative noise scales from DIAL-MPC; simulate the selected trajectory once more (3,001 rollouts total).
Appendix
Table 6: Sampling-based optimizer settings.
Edit
Implementation
Variants per object
Regeneration
Regenerate from the reference image with a new random seed
20
Deformation
Deform the mesh by different amounts along each axis
60
Occlusion completion
Mask part of the reference image and complete the shape
20
Rotation
Rotate about the vertical axis, then regenerate
20
Stretch
Scale the original mesh along a single axis
80
Total
200
Appendix
Table 7: Geometry augmentations and variants per object.
Quantity
Original geometries
Modified geometries
Total
Evaluated demonstrations
186,310
173,840
360,150
DIAL-MPC successes
47,426
47,918
95,344
DIAL-MPC failures
138,884
125,922
264,806
Additional GNR successes ( S=256 )
74,391
52,745
127,136
MPC or GNR successes
121,817
100,663
222,480
Source motions
780
728
780
Appendix
Table 8: HOT3D statistics for the 360,150 evaluated demonstrations. Columns separate demonstrations using original and modified object geometries; original geometries are augmented in pose, timing, mass, and friction. Trajectory counts and distinct geometry counts are reported separately.
Figure 6: Objects in HOT3D .
Figure 7: Examples of geometry variants. Each row uses the same original object.
Figure 8: DG-5F grasp configurations at a shared camera view and scale.
Figure 9: Test-time scaling on 360k HOT3D demonstrations: 95k MPC-success and 264k MPC-failure demonstrations. The curves show GNR ’s SR across three OOD subsets.
Figure 10: HOT3D trajectory quality. From left to right: median object-position, object-orientation, and finger-joint angular errors on the randomly sampled 1,000-demonstration test set, followed by median finger-control jerk. GNR approximately halves jerk while preserving comparable object tracking.
Figure 11: ShapeFilter spatial generalization. The training (gray) and test (blue) quadrants are shown at left. At right, we show the test-time scaling curves of GNR on the test split.
Method
nsM or S
Success (%)
DIAL-MPC
32×6
48.4
DIAL-MPC
500×6
92.9
CEM
32×6
50.0
CEM
500×6
82.8
iCEM
32×6
65.6
CMA-ES
32×6
40.6
Appendix
Table 10: ShapeFilter success rates. Optimizer budgets are samples per iteration ns× MPC iterations M .
Samples S
State
Geometry
Geometry + state
1
50.42
50.93
53.48
2
67.06
66.38
69.10
4
78.44
80.31
78.61
8
87.10
88.62
86.59
16
92.36
93.72
94.06
32
95.93
96.77
97.11
Appendix
Table 11: ShapeFilter conditioning ablation. We report success rates (%) across different numbers of GNR samples S using different conditioning.
Figure 12: Performance of GNR using either residual or absolute prediction with one, two, or four training corners. Shadows show 95% confidence intervals.
Seed set size N
Quantity
Samples S
1
2
4
8
16
32
64
128
256
50
Success (%)
13.0
17.2
33.4
47.4
66.4
83.1
93.2
96.7
99.5
Cost (k core-h)
0.235
0.263
0.313
0.390
0.501
0.627
0.740
0.832
0.890
100
Success (%)
21.0
30.5
47.5
66.7
81.3
91.1
96.4
98.7
99.3
Cost (k core-h)
0.438
0.463
0.504
0.560
0.626
0.688
0.745
0.789
0.828
200
Success (%)
33.4
44.3
57.7
71.9
80.0
89.1
94.9
96.8
99.0
Appendix
Table 12: ShapeFilter dataset-generation efficiency. Dataset success rates include the successful seeds and apply the test-set success rate to the remaining demonstrations. Costs include MPC seed generation, GNR inference, and simulation, in thousands of core-h.
Figure 13: ShapeFilter generation cost versus dataset size. MPC generates all ∣D∣ demonstrations; GNR uses MPC for N successful seeds and generates the remaining ∣D∣−N with configurations that meet or exceed the best MPC success rate. Numbers show the reduction in total cost.
Figure 14: ShapeFilter trajectory errors and jerk comparison. GNR has lower jerk; GNR +MPC has the lowest object and fingertip tracking errors.
N=100
N=200
N=294
MPC samples
Success (%)
Cost (k core-h)
MPC samples
Success (%)
Cost (k core-h)
MPC samples
Success (%)
Cost (k core-h)
1,736
64.0
35.90
1,736
79.3
57.34
1,736
91.9
77.41
1,752
70.7
36.79
1,752
83.4
57.89
1,752
93.3
77.61
1,768
74.1
37.42
1,784
89.0
58.91
1,784
95.5
77.97
1,784
79.7
38.17
1,816
90.3
59.08
1,816
97.2
78.37
1,816
80.9
39.81
1,848
93.8
60.10
1,912
97.8
78.95
Appendix
Table 13: NIST dataset-generation efficiency. Dataset success rates include the successful seeds and estimated successes on the remaining demonstrations. Each MPC budget uses the schedule with the highest success rate. GNR generation costs include MPC seed generation, GNR inference, and GNR -initialized MPC, in thousands of CPU core-h (App. F.2 ).
Figure 15: NIST generation cost versus dataset size. MPC generates all ∣D∣ demonstrations; GNR+MPC uses MPC for N successful seeds and generates the remaining ∣D∣−N using evaluated configurations with similar dataset success rates. Numbers show the reduction in total cost.
Task
Method
Sample budget
Core-h / episode
Wall h / episode
HOT3D
DIAL-MPC
3,000
1.0777
0.2427
HOT3D
iCEM
3,000
1.4213
0.3282
HOT3D
CMA-ES
3,000
1.2272
0.2877
HOT3D
GNR , S=256
256
0.1094
0.0383
ShapeFilter
MPC
3,000
3.2001
0.0472
Appendix
Table 14: Measured per-episode generation cost.
Task
Generation
Generation compute ( 103 core-h)
Coverage
MPC data
GNR
Total
HOT3D
all MPC
390.043
–
390.043
26.4%
HOT3D
all MPC data + GNR S=256
390.043
26.047
416.090
61.6%
HOT3D
28k MPC seeds + GNR S=256
116.112
32.623
148.735
59.7%
ShapeFilter
all MPC
8.541
–
8.541
93.70%
ShapeFilter
50 MPC seeds + GNR
0.203
0.562
0.766
94.82%
Appendix
Table 15: Dataset-generation cost in thousands of core-h and coverage. GNR costs are estimated from per-trajectory costs.
Figure 16: GNR test-time scaling across five different dexterous hands.
Method
Samples
Success (%, 95% CI)
GNR
128
90.0 [82.6, 94.5]
GNR
256
93.0 [86.3, 96.6]
GNR
384
95.0 [88.8, 97.8]
SPIDER
16,384–32,768
96.0 [90.2, 98.4]
Appendix
Table 16: Multi-embodiment success on 100 test trajectories (20 per hand), with 95% Wilson confidence intervals. Samples are counted per demonstration for GNR and give the maximum budget per optimization for SPIDER.
Figure 17: SR curves of behavior cloning policy training using data generated by different retargeters.
Figure 18: ShapeFilter success across four data-flywheel rounds.
Figure 19: ShapeFilter example in which MPC fails while GNR and GNR+MPC succeed.
Figure 20: ShapeFilter example in which MPC fails while GNR and GNR+MPC succeed.
Figure 21: A NIST example. GNR+MPC succeeds where IK, MPC, and open-loop GNR fail.
Figure 22: A NIST example. GNR+MPC succeeds where IK, MPC, and open-loop GNR fail.
Figure 23: HOT3D coffee pot. GNR succeeds where MPC fails.
Figure 24: HOT3D ranch bottle with a stretched OOD geometry.
Figure 25: HOT3D barbecue sauce bottle with a stretched OOD geometry.
Figure 26: HOT3D parmesan can with an OOD deformation.
Figure 27: HOT3D dumbbell manipulation.
Figure 28: HOT3D waffle pick-and-place with an OOD rotation.
Figure 29: HOT3D holder with a stretched OOD geometry.
Figure 30: HOT3D wooden spoon with an OOD deformation.
Figure 31: HOT3D whiteboard marker. MPC makes the reconstructed human motion physically feasible.
Figure 32: HOT3D whiteboard eraser. MPC makes the reconstructed human motion physically feasible.
Figure 33: HOT3D dinosaur toy. MPC makes the reconstructed human motion physically feasible.