SMART: Zero-Shot Sim-to-Real Articulated Object Manipulation via Large-Scale Synthetic Pretraining
Authors: Jicong Ao, Shuhan Jiang, Yuling Zhong, Yanwen Liu, Yuhan Gao, Jiangyuan Zhao, Yang Zhang, Shiqiang Zhu, +2 more
Organizations: Institute of Artificial Intelligence, China Telecom · Zhejiang University · Technical University of Munich · Harbin Institute of Technology · Shanghai Jiao Tong University · Tsinghua University · Gamma Robotics (γ)
The ability to interact with articulated objects is essential for embodied intelligent systems, but collecting large-scale real-world demonstrations for these interactions remains challenging due to the precise contact and constraint-following motions involved. Although simulation provides a promising alternative, existing synthetic data efforts cover limited articulated-object categories, while general-purpose synthesis pipelines lack explicit designs for part-level semantics and articulation constraints, hindering agentic task generation and scalable synthesis of high-quality articulated-manipulation demonstrations. To bridge this gap, we introduce SMART, a scalable system leveraging large-scale Synthesized Manipulation demonstrations for ARTiculated-object manipulation. At its core, we develop SMART-Sim, a simulation platform with articulation-aware design that enables effective task generation and efficient demonstration collection. Building on SMART-Sim, we apply agentic task generation and design a scalable distributed synthesis system, using them to synthesize SMART-Data, comprising over 1M demonstrations across 44 atomic task types, 5 robot setups, and 2,507 articulated objects. The vision-language-action (VLA) model pretrained on SMART-Data shows competitive performance on simulation benchmarks and achieves zero-shot sim-to-real transfer and scalable performance in real-world articulated-object manipulation tasks. This highlights the potential of synthetic demonstrations in providing effective and scalable supervision for improving VLA model performance in contact-rich articulated-object manipulation.
Figures & tables
Figure 1 : SMART synthesizes large-scale simulation data for articulated-object manipulation by combining diverse embodiments and articulated objects, rich manipulation skills, and comprehensive domain randomization.
Figure 2 : Overview of SMART-Sim . SMART-Sim provides versatile robot embodiment and asset support, flexible annotation and skills for motion generation, efficient demonstration collection, and comprehensive domain randomization features.
Feature
SMART-Sim (Ours)
Robo- Casa365
Molmo- Space
Arti- Bench
RoboTwin 2.0
RL- Bench
Behavior- 1K
Humanoid- Gen
Mani- Skill 2
LIBERO
InternData-A1
Scenes
1,452
120
230,000+
1
1
1
50
20
–
20
227
Embodiments
7
1
2
1
5
1
12
1
1
1
4
(Rigid) Objects
9,116
2,509
130,000+
–
687
28
9318
–
2144
–
3185
Articulated Objects
2,507
20
–
–
44
–
–
4
–
–
321
Realistic Physics
✓
✓
✓
✓
✓
✓
✓
✓
✓
✓
✓
Realistic Rendering
✓
✗
✗
✗
✓
✗
✓
✓
✓
✗
✓
Table 1 : Comparison of common robot simulation frameworks. Some frameworks only list the total number of objects and do not distinguish between rigid and articulated objects.
Figure 3 : The illustration of asset annotations, with (a), (b), and (c) showing the action frames on different articulated objects, and (d) showing the annotation on a Robotiq end effector. The orange arrows show the direction of the articulation constraints between parts.
Figure 4 : The articulation-aware domain randomization effects supported by SMART-Sim .
Figure 5 : Illustration of our agentic task generation and distributed data synthesis.
Figure 6 : Statistics of SMART-Data over robot setups, task complexity levels, skills, articulated object categories, and task durations.
Dataset
Traj. (All)
Traj. (Arti.)
Skill
Task
Scene
Embodiment
Collection Method
RoboCasa365
655k
415k
8
100
120
1
Teleoperation & Augmentation
RoboTwin 2.0
100k
12k
–
50
1
5
Autonomous
Articubot
42.3k
42.3k
2
1
1
1
Autonomous
MolmoBot
1.7M
125.6k
2
1
1
1
Autonomous
InternData-A1
630k
74.4k
18
70
227
4
Autonomous
SMART-Data
1M
1M
23
44(Atomic)
1,122
5
Autonomous
Table 2 : Comparison of robotic simulation datasets with substantial articulated-object manipulation data. For SMART-Data , we only report the number of atomic tasks, as the high diversity of composite tasks makes a consistent quantitative comparison challenging.
Method
Spatial
Object
Goal
Long
Average
GR00T N1
94.4
97.6
93.0
90.6
93.9
π0 (bs=32, 30K steps)
96.8
98.8
95.8
85.2
94.2
Qwen3-VL- π (bs=32, 30K steps)
95.2
99.0
96.2
88.4
94.7
InternVLA-M1
98.0
99.0
93.8
92.6
95.9
π0.5 (bs=256, 30K steps)
98.8
98.2
98.0
92.4
96.9
SMART-VLA(Real) (bs=32, 30K steps)
97.8
99.8
98.0
95.6
97.8
Table 3 : Evaluation results (success rates) on LIBERO. Training budgets are shown in gray parentheses. The best result in each column is shown in bold .
Method
Atomic Seen
Composite Seen
Composite Unseen
Average
π0
34.6
6.1
1.1
14.8
π0.5
39.6
7.1
1.2
16.9
GR00T N1.6
51.1
9.4
1.7
21.9
GR00T N1.5
50.7
14.8
2.7
23.9
SMART-VLA(w/o S1)
43.3
9.4
4.4
20.0
SMART-VLA
52.8
23.1
7.5
28.8
Table 4 : Evaluation results (success rates) on RoboCasa365. Best results are in bold. SMART-VLA(Real) is omitted here because the official report does not provide results for the variant without CRL.
Figure 7 : Post-training performance curve of SMART-VLA and SMART-VLA(w/o S1) on LIBERO. Average success rates are reported every 5K post-training steps up to the standard 30K-step training budget.
Figure 8 : Overview of the real-world robot platforms, including a dual-arm RealMan RM75 platform (left), an AC1 dual-arm platform (center), and an R1Pro platform (right). All platforms have two wrist cameras and one center-mounted head camera.
Figure 9 : Overview of the twelve real-world articulated-object manipulation tasks.
Figure 10 : Real-world evaluation across twelve articulated-object manipulation tasks on three robot platforms.
NRMSJ ↓
NMAV ↓
Method
RM75
AC1
R1Pro
Overall
RM75
AC1
R1Pro
Overall
SMART-VLA
0.0258
0.1006
0.0448
0.0581
0.0416
0.1046
0.0520
0.0673
SMART-VLA(Real)
0.0456
0.1177
0.0157
0.0637
0.0584
0.1209
0.0253
0.0721
SMART-VLA(w/o S1)
0.0357
0.3703
0.0424
0.1592
0.0481
0.3757
0.0486
0.1674
π0.5
0.1636
0.0705
0.0225
0.0913
0.0868
0.2445
0.1539
0.1625
Table 5 : The average of Normalized Root-Mean-Square Jerk (NRMSJ, normalized using quantiles) and Normalized Mean Action Variation (NMAV, normalized using quantiles) across AC1, RM75, R1Pro, and all twelve tasks. Lower values ( ↓ ) indicate smoother and more consistent trajectories.
Figure 11 : The success rates of SMART-VLA , SMART-VLA(w/o S1) , and π0.5 on AC1 tasks at seen and unseen environment heights.
Figure 12 : Performance scaling experiment results on AC1 and R1Pro tasks.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 13 : The illustration of annotation approaches for (a) rigid objects and (b) articulated objects.
Aspect
Parameter
Range
Purpose
Spatial
Object position ( x , y , z )
±0.1 m ( x , y ), ±0.03 m ( z )
Prevent the policy from memorizing a fixed placement and generalize to arbitrary object positions on the tabletop
Object orientation (roll, pitch, yaw)
±15∘ (roll, pitch), ±5∘ (yaw)
Cover varied grasp and contact directions so that the policy does not overfit a single pose
Robot initial joint configuration
±5∘
Make the policy robust to initial configuration offsets and narrow the gap to real deployment
Camera extrinsics
±0.05 m (position), ±5∘ (orientation)
Tolerate camera mounting errors and ease visual sim-to-real transfer
Physical
Part friction
±0.15
Adapt to contact and slipping behavior under different surface conditions
Part density
±0.5
Cover the dynamics difference caused by part mass variation
Appendix
Table 6 : Domain randomization parameters of SMART-Sim .
Figure 14 : Representative samples from SMART-Data across diverse scenes, objects, and manipulation tasks.
Platform Name
RM75
R1Pro
AC1
Franka (Single)
Franka (Dual-Arm)
Flexiv Rizon 4s
Marvin M6s
Unitree H1-2
Degrees of Freedom
16
16
14
8
16
8
16
36
Camera Number
3
3
3
2
3
2
3
3
End-Effector Type
Robotiq 2F-85
Self-designed Gripper
AC1 Gripper
Franka Hand
Franka Hand
Robotiq 2F-85
DAS Gripper V3
X Hand
Center Camera Type
RealSense L515
RealSense D435
RealSense D405
RealSense D435
RealSense D435
RealSense D405
RealSense L515
RealSense L515
Wrist Camera Type
RealSense D405
RealSense D405
RealSense D405
RealSense D435
RealSense D435
RealSense D405
DAS Fisheye
RealSense D405
Appendix
Table 7 : Robot platform details. Franka is listed as single-arm and dual-arm configurations but is counted as one robot embodiment in the main text.
Hyperparameters
Pretraining
Post-training (Sim2Real)
Global Batch Size
2048
32
VLM Learning Rate
5e-5
1e-5
Action Expert Learning Rate
-
1e-4
Learning Rate Schedule
Cosine Decay
Cosine Decay
Minimum Learning Rate
5e-6
1e-6
Warmup Ratio
0.01
0.01
Appendix
Table 8 : Hyperparameters used in pretraining and post-training. For pretraining, we adopt data packing [ Zhang et al., 2026b ] to form training batches, where the global batch size is estimated according to the average number of samples contained in each data pack.