SMART: Zero-Shot Sim-to-Real Articulated Object Manipulation via Large-Scale Synthetic Pretraining
Authors: Jicong Ao, Shuhan Jiang, Yuling Zhong, Yanwen Liu, Yuhan Gao, Jiangyuan Zhao, Yang Zhang, Shiqiang Zhu, +2 more
Organizations: Institute of Artificial Intelligence, China Telecom · Zhejiang University · Technical University of Munich · Harbin Institute of Technology · Shanghai Jiao Tong University · Tsinghua University · Gamma Robotics (γ)
The ability to interact with articulated objects is essential for embodied intelligent systems, but collecting large-scale real-world demonstrations for these interactions remains challenging due to the precise contact and constraint-following motions involved. Although simulation provides a promising alternative, existing synthetic data efforts cover limited articulated-object categories, while general-purpose synthesis pipelines lack explicit designs for part-level semantics and articulation constraints, hindering agentic task generation and scalable synthesis of high-quality articulated-manipulation demonstrations. To bridge this gap, we introduce SMART, a scalable system leveraging large-scale Synthesized Manipulation demonstrations for ARTiculated-object manipulation. At its core, we develop SMART-Sim, a simulation platform with articulation-aware design that enables effective task generation and efficient demonstration collection. Building on SMART-Sim, we apply agentic task generation and design a scalable distributed synthesis system, using them to synthesize SMART-Data, comprising over 1M demonstrations across 44 atomic task types, 5 robot setups, and 2,507 articulated objects. The vision-language-action (VLA) model pretrained on SMART-Data shows competitive performance on simulation benchmarks and achieves zero-shot sim-to-real transfer and scalable performance in real-world articulated-object manipulation tasks. This highlights the potential of synthetic demonstrations in providing effective and scalable supervision for improving VLA model performance in contact-rich articulated-object manipulation.
Figures & tables
Figure 1 : SMART synthesizes large-scale simulation data for articulated-object manipulation by combining diverse embodiments and articulated objects, rich manipulation skills, and comprehensive domain randomization.
Figure 2 : Overview of SMART-Sim . SMART-Sim provides versatile robot embodiment and asset support, flexible annotation and skills for motion generation, efficient demonstration collection, and comprehensive domain randomization features.
Feature
SMART-Sim (Ours)
Robo- Casa365
Molmo- Space
Arti- Bench
RoboTwin 2.0
RL- Bench
Behavior- 1K
Humanoid- Gen
Mani- Skill 2
LIBERO
InternData-A1
Scenes
1,452
120
230,000+
1
1
1
50
20
–
20
227
Embodiments
7
1
2
1
5
1
12
1
1
1
4
(Rigid) Objects
9,116
2,509
130,000+
–
687
28
9318
–
2144
–
3185
Articulated Objects
2,507
20
–
–
44
–
–
4
–
–
321
Realistic Physics
✓
✓
✓
✓
✓
✓
✓
✓
✓
✓
✓
Realistic Rendering
✓
✗
✗
✗
✓
✗
✓
✓
✓
✗
✓
Table 1 : Comparison of common robot simulation frameworks. Some frameworks only list the total number of objects and do not distinguish between rigid and articulated objects.
Figure 3 : The illustration of asset annotations, with (a), (b), and (c) showing the action frames on different articulated objects, and (d) showing the annotation on a Robotiq end effector. The orange arrows show the direction of the articulation constraints between parts.
Figure 4 : The articulation-aware domain randomization effects supported by SMART-Sim .
Figure 5 : Illustration of our agentic task generation and distributed data synthesis.
Figure 6 : Statistics of SMART-Data over robot setups, task complexity levels, skills, articulated object categories, and task durations.
Dataset
Traj. (All)
Traj. (Arti.)
Skill
Task
Scene
Embodiment
Collection Method
RoboCasa365
655k
415k
8
100
120
1
Teleoperation & Augmentation
RoboTwin 2.0
100k
12k
–
50
1
5
Autonomous
Articubot
42.3k
42.3k
2
1
1
1
Autonomous
MolmoBot
1.7M
125.6k
2
1
1
1
Autonomous
InternData-A1
630k
74.4k
18
70
227
4
Autonomous
SMART-Data
1M
1M
23
44(Atomic)
1,122
5
Autonomous
Table 2 : Comparison of robotic simulation datasets with substantial articulated-object manipulation data. For SMART-Data , we only report the number of atomic tasks, as the high diversity of composite tasks makes a consistent quantitative comparison challenging.
Method
Spatial
Object
Goal
Long
Average
GR00T N1
94.4
97.6
93.0
90.6
93.9
π0 (bs=32, 30K steps)
96.8
98.8
95.8
85.2
94.2
Qwen3-VL- π (bs=32, 30K steps)
95.2
99.0
96.2
88.4
94.7
InternVLA-M1
98.0
99.0
93.8
92.6
95.9
π0.5 (bs=256, 30K steps)
98.8
98.2
98.0
92.4
96.9
SMART-VLA(Real) (bs=32, 30K steps)
97.8
99.8
98.0
95.6
97.8
Table 3 : Evaluation results (success rates) on LIBERO. Training budgets are shown in gray parentheses. The best result in each column is shown in bold .
Method
Atomic Seen
Composite Seen
Composite Unseen
Average
π0
34.6
6.1
1.1
14.8
π0.5
39.6
7.1
1.2
16.9
GR00T N1.6
51.1
9.4
1.7
21.9
GR00T N1.5
50.7
14.8
2.7
23.9
SMART-VLA(w/o S1)
43.3
9.4
4.4
20.0
SMART-VLA
52.8
23.1
7.5
28.8
Table 4 : Evaluation results (success rates) on RoboCasa365. Best results are in bold. SMART-VLA(Real) is omitted here because the official report does not provide results for the variant without CRL.
Figure 7 : Post-training performance curve of SMART-VLA and SMART-VLA(w/o S1) on LIBERO. Average success rates are reported every 5K post-training steps up to the standard 30K-step training budget.
Figure 8 : Overview of the real-world robot platforms, including a dual-arm RealMan RM75 platform (left), an AC1 dual-arm platform (center), and an R1Pro platform (right). All platforms have two wrist cameras and one center-mounted head camera.
Figure 9 : Overview of the twelve real-world articulated-object manipulation tasks.
Figure 10 : Real-world evaluation across twelve articulated-object manipulation tasks on three robot platforms.
NRMSJ ↓
NMAV ↓
Method
RM75
AC1
R1Pro
Overall
RM75
AC1
R1Pro
Overall
SMART-VLA
0.0258
0.1006
0.0448
0.0581
0.0416
0.1046
0.0520
0.0673
SMART-VLA(Real)
0.0456
0.1177
0.0157
0.0637
0.0584
0.1209
0.0253
0.0721
SMART-VLA(w/o S1)
0.0357
0.3703
0.0424
0.1592
0.0481
0.3757
0.0486
0.1674
π0.5
0.1636
0.0705
0.0225
0.0913
0.0868
0.2445
0.1539
0.1625
Table 5 : The average of Normalized Root-Mean-Square Jerk (NRMSJ, normalized using quantiles) and Normalized Mean Action Variation (NMAV, normalized using quantiles) across AC1, RM75, R1Pro, and all twelve tasks. Lower values ( ↓ ) indicate smoother and more consistent trajectories.
Figure 11 : The success rates of SMART-VLA , SMART-VLA(w/o S1) , and π0.5 on AC1 tasks at seen and unseen environment heights.
Figure 12 : Performance scaling experiment results on AC1 and R1Pro tasks.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 13 : The illustration of annotation approaches for (a) rigid objects and (b) articulated objects.
Aspect
Parameter
Range
Purpose
Spatial
Object position ( x , y , z )
±0.1 m ( x , y ), ±0.03 m ( z )
Prevent the policy from memorizing a fixed placement and generalize to arbitrary object positions on the tabletop
Object orientation (roll, pitch, yaw)
±15∘ (roll, pitch), ±5∘ (yaw)
Cover varied grasp and contact directions so that the policy does not overfit a single pose
Robot initial joint configuration
±5∘
Make the policy robust to initial configuration offsets and narrow the gap to real deployment
Camera extrinsics
±0.05 m (position), ±5∘ (orientation)
Tolerate camera mounting errors and ease visual sim-to-real transfer
Physical
Part friction
±0.15
Adapt to contact and slipping behavior under different surface conditions
Part density
±0.5
Cover the dynamics difference caused by part mass variation
Appendix
Table 6 : Domain randomization parameters of SMART-Sim .
Figure 14 : Representative samples from SMART-Data across diverse scenes, objects, and manipulation tasks.
Platform Name
RM75
R1Pro
AC1
Franka (Single)
Franka (Dual-Arm)
Flexiv Rizon 4s
Marvin M6s
Unitree H1-2
Degrees of Freedom
16
16
14
8
16
8
16
36
Camera Number
3
3
3
2
3
2
3
3
End-Effector Type
Robotiq 2F-85
Self-designed Gripper
AC1 Gripper
Franka Hand
Franka Hand
Robotiq 2F-85
DAS Gripper V3
X Hand
Center Camera Type
RealSense L515
RealSense D435
RealSense D405
RealSense D435
RealSense D435
RealSense D405
RealSense L515
RealSense L515
Wrist Camera Type
RealSense D405
RealSense D405
RealSense D405
RealSense D435
RealSense D435
RealSense D405
DAS Fisheye
RealSense D405
Appendix
Table 7 : Robot platform details. Franka is listed as single-arm and dual-arm configurations but is counted as one robot embodiment in the main text.
Hyperparameters
Pretraining
Post-training (Sim2Real)
Global Batch Size
2048
32
VLM Learning Rate
5e-5
1e-5
Action Expert Learning Rate
-
1e-4
Learning Rate Schedule
Cosine Decay
Cosine Decay
Minimum Learning Rate
5e-6
1e-6
Warmup Ratio
0.01
0.01
Appendix
Table 8 : Hyperparameters used in pretraining and post-training. For pretraining, we adopt data packing [ Zhang et al., 2026b ] to form training batches, where the global batch size is estimated according to the average number of samples contained in each data pack.
Scaling data volume and diversity is critical for generalizing embodied intelligence. While synthetic data generation offers a scalable alternative to expensive physical data acquisition, transferring robotic manipulation policies from simulation to the real world (sim-to-real) remains a formidable challenge due to the domain gap. This paper presents HyperSim, a holistic framework spanning from synthetic data generation to policy training and seamless real-world deployment. To systematically bridge the sim-to-real gap, HyperSim is realized through three core pillars: high-fidelity environment synthesis, adversarial trajectory generation, and sim-and-real co-training. Collectively, these modules address domain discrepancies by enhancing visual fidelity, expanding data coverage, and enforcing domain-invariant representations. We rigorously validate HyperSim through a large-scale empirical study involving 400 real-world task executions across two representative manipulation models. Assessed across three fine-grained metrics, our complete pipeline achieves remarkable sim-to-real success rates of 80% and 95% with ACT and π_{0}, respectively. Furthermore, policies trained on our adversarial trajectories exhibit significantly enhanced robustness against dynamic uncertainties, achieving a 35% higher completion rate under physical perturbations.
Junyi Dong, Haotian Luo, Ziwei Xu +11
CloudRobo Lab, Huawei Cloud Computing Technologies Co.,Ltd. · Shanghai Jiao Tong University · The University of Hong Kong
The development of robust and generalizable robot learning models is critically contingent upon the availability of large-scale, diverse training data and reliable evaluation benchmarks. Collecting data in the physical world poses prohibitive costs and scalability challenges, and prevailing simulation benchmarks frequently suffer from fragmentation, narrow scope, or insufficient fidelity to enable effective sim-to-real transfer. To address these challenges, we introduce Genie Sim 3.0, a unified simulation platform for robotic manipulation. We present Genie Sim Generator, a large language model (LLM)-powered tool that constructs high-fidelity scenes from natural language instructions. Its principal strength resides in rapid and multi-dimensional generalization, facilitating the synthesis of diverse environments to support scalable data collection and robust policy evaluation. We introduce the first benchmark that pioneers the application of LLM for automated evaluation. It leverages LLM to mass-generate evaluation scenarios and employs Vision-Language Model (VLM) to establish an automated assessment pipeline. We also release an open-source dataset comprising more than 10,000 hours of synthetic data across over 200 tasks. Through systematic experimentation, we validate the robust zero-shot sim-to-real transfer capability of our open-source dataset, demonstrating that synthetic data can server as an effective substitute for real-world data under controlled conditions for scalable policy training. For code and dataset details, please refer to: https://github.com/AgibotTech/genie_sim.
Bridging the sim-to-real gap is a core challenge in deploying learned manipulation policies. Sim-to-real learning is attractive because it can replace expensive real robot demonstrations with scalable synthetic data, yet world-action models have not previously been shown to transfer from simulation to real robotic manipulation. We study whether a world-action model can be trained from synthetic priors and deployed zero-shot in the real world. To this end, we build upon Cosmos Policy, a video diffusion model adapted for visuomotor control. We construct simulation environments with extensive domain randomization and generate demonstrations using the AnyTask motion planning pipeline. We evaluate our approach across object lifting, drawer opening, and pick-and-place tasks using ∼800 synthetic demonstrations per task and no real demonstrations. When deployed zero-shot on a Franka Robot, our policy attains a 35% average success rate. To our knowledge, this represents the first successful sim-to-real transfer of a world-action model for robotic manipulation.