Learning generalizable and robust behavior cloning policies requires large volumes of high-quality robotics data. While human demonstrations (e.g., through teleoperation) serve as the standard source for expert behaviors, acquiring such data at scale in the real world is prohibitively expensive. This paper introduces ExpertGen, a framework that automates expert policy learning in simulation to enable scalable sim-to-real transfer. ExpertGen first initializes a behavior prior using a diffusion policy trained on imperfect demonstrations, which may be synthesized by large language models or provided by humans. Reinforcement learning is then used to steer this prior toward high task success by optimizing the diffusion model's initial noise while keep original policy frozen. By keeping the pretrained diffusion policy frozen, ExpertGen regularizes exploration to remain within safe, human-like behavior manifolds, while also enabling effective learning with only sparse rewards. Empirical evaluations on challenging manipulation benchmarks demonstrate that ExpertGen reliably produces high-quality expert policies with no reward engineering. On industrial assembly tasks, ExpertGen achieves a 90.5% overall success rate, while on long-horizon manipulation tasks it attains 85% overall success, outperforming all baseline methods. The resulting policies exhibit dexterous control and remain robust across diverse initial configurations and failure states. To validate sim-to-real transfer, the learned state-based expert policies are further distilled into visuomotor policies via DAgger and successfully deployed on real robotic hardware.
Figures & tables
Figure 2: ExpertGen training pipeline: generative modeling of imperfect behavior priors using a state-based diffusion policy (Phase 1); steering diffusion policy in massively parallel simulation using FastTD3 (phase 2); visual policy distillation using DAgger from expert teachers (phase 3).
Figure 3: Illustrations of the tasks in our experiments. (A) Real-world manipulation tasks. (B) Industrial assembly tasks from AutoMate. (C) Long-horizon tasks from AnyTask .
Table 1: The success rates (%) of the evaluated state-based policies on AnyTask benchmark. Bold number indicates the best number across all the approaches.
Figure 4: The success rates (%) of the evaluated approaches on assets from the AutoMate benchmark. ExpertGen outperforms all other baselines with an overall success of 91.1%. By introducing x-y noise, the diffusion policy demonstrates higher success rates compared to the no x-y noise.
Methods
Lift Banana
Open Drawer
Push Pear
Open
Force
Open
Force
Open
Force
Diffusion Policy
72.5 (7.6 ↓ )
1.9 (79.9 ↓ )
78.2 (8.7 ↓ )
58.0 (28.9 ↓ )
47.1 (1.9 ↓ )
15.8 (33.1 ↓ )
Residual RL
83.8 (1.0 ↓ )
46.7 (37.5 ↓ )
96.9 (2.6 ↓ )
99.7 (0.2 ↑ )
57.1( 0.7 ↑ )
38.4 (18.1 ↓ )
ExpertGen
99.4 (0.3 ↓ )
56.6 (43.1 ↓ )
99.9 (0.0 ↓ )
99.7 (0.2 ↓ )
85.4 (1.2 ↑ )
38.7 (45.3 ↓ )
Table 2: Success rates (%) with two perturbations: random gripper opening and random external force applied to the end-effector. Number in parentheses indicate the performance drop ( ↓ ) / increase ( ↑ ) compared to no perturbation. Bold number indicates the best success rates.
Table 3: The Wasserstein-2 distance between the evaluated policies and base diffusion policies.
Observation
Method
Number of Demos.
Lift Banana
Lift Brick
Push Pear to Center
Open Drawer
Point-cloud
ExpertGen
DAgger
75.0%
–
65.0%
85.0%
AnyTask [ 8 ]
20k
73.3%
–
16.7%
42.5%
RGB
ExpertGen
20k
80.0%
65.0%
45.0%
65.0%
AnyTask [ 8 ]
20k
80.0%
55.0%
35.0%
10.0%
Table 4: Real-world evaluation results of visuomotor policies. Policies are trained from either ExpertGen or AnyTask [ 8 ] scripted-policy rollouts. While AnyTask achieve reasonable performance on simple pick tasks, they struggle in contact-rich manipulation, highlighting the importance of ExpertGen ’s RL-refined teacher policies for robust sim-to-real transfer.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Success rate (%) distribution of imperfect prior diffusion policy (left) and ExpertGen (right) over different initial object configurations for Push Pear to Center .
Deficiencies
Base Policy
ExpertGen
Limited state coverage
39.2 (30.6 ↓ )
76.8 (23.2 ↓ )
Joint dynamics mismatch
31.4 (41.2 ↓ )
99.2 (0.1 ↓ )
Appendix
Table 5: Success rates under prior deficiencies. Arrows indicate performance drops compared to the setting without prior deficiency.
Figure 6: Success rates of base diffusion policies (left) and the ExpertGen policies (right) trained with 50 , 200 , 500 , and 1000 imperfect demonstrations on (a) Lift Banana , (b) Open Drawer , (c) Push Pear to Center , and (d) Put Object in Closed Drawer .
Figure 7: Success rates (%) of the evaluated state-based policies on the Stack Banana on Can task. Diffusion Policy w/ SkillMimicGen trains a diffusion policy on human demonstrations augmented with SkillMimicGen [ 7 ] , while ExpertGen w/ SkillMimicGen further refines the diffusion policy trained on SkillMimicGen data. Using human motions as behavior priors significantly improves the performance of both the diffusion policy and ExpertGen .
Hyperparameters
AnyTask
AutoMate
Generative Behavior Prior Modeling
Architecture
U-Net [ 19 ]
U-Net
Number of groups
8
8
Kernel size
5
5
Number of channels
[128, 256]
[128, 256]
Diffusion step embedding dim.
16
16
Appendix
Table 6: ExpertGen hyperparameters. The right two columns indicate the hyerparameters for AnyTask and AutoMate, respectively.
Hyperparameters
AnyTask
Expert Policy Acquisition
Receding horizon
8
Number of steps
64
Learning rate
5e-4
Learning rate schedule
fixed
Discount factor
0.995
Appendix
Table 7: ExpertGen -PPO hyperparameters.
Hyperparameters
AnyTask
Residual RL
Warm-start steps
10240
Action scale
0.02
Receding horizon
8
FastTD3
Critic learning rate
3e-4
Appendix
Table 8: Residual RL hyperparameters.
Hyperparameters
AnyTask
Warm-start steps
10240
Action scale
0.02
Receding horizon
8
Appendix
Table 9: SMP hyperparameters.
Figure 8: An overview of the selected eight tasks from AnyTask benchmark.
State
Dimension
End-effector pose
16
robot arm joint positions
9
Two object positions
6
Two object 6D rotations
12
Articulated object position
3
Articulated object 6D rotation
6
Appendix
Table 10: State condition and action definitions for state-based diffusion policy for the AnyTask benchmark.
Figure 9: Visualization of the selected nine assets from AutoMate benchmark covering three difficulty levels (easy, medium, and hard).
State
Dimension
End-effector pose
16
Held positions
3
Held 6D rotation
6
Task embedding
32
Action
Dimension
End-effector position
3
Appendix
Table 11: State condition and action definitions for state-based diffusion policy for the AutoMate benchmark.
Figure 10: Scripted policy for Lift Banana
Figure 11: Scripted policy for Lift Brick
Figure 12: Scripted policy for Lift Peach
Figure 13: Scripted policy for Open Drawer
Figure 14: Scripted policy for Push Pear to Center
Figure 15: Scripted policy for Stack Banana on Can
Figure 16: Scripted policy for Put Object in Closed Drawer
Figure 17: Scripted policy for Put Object in Closed Drawer
Figure 18: Scripted policy for AutoMate disassembly tasks.
Figure 19: Scripted policy for AutoMate disassembly tasks.