cs.ROOct 6, 2026

SMART: Zero-Shot Sim-to-Real Articulated Object Manipulation via Large-Scale Synthetic Pretraining

Authors: Jicong Ao, Shuhan Jiang, Yuling Zhong, Yanwen Liu, Yuhan Gao, Jiangyuan Zhao, Yang Zhang, Shiqiang Zhu, +2 more

Organizations: Institute of Artificial Intelligence, China Telecom · Zhejiang University · Technical University of Munich · Harbin Institute of Technology · Shanghai Jiao Tong University · Tsinghua University · Gamma Robotics (γ)

Abstract

The ability to interact with articulated objects is essential for embodied intelligent systems, but collecting large-scale real-world demonstrations for these interactions remains challenging due to the precise contact and constraint-following motions involved. Although simulation provides a promising alternative, existing synthetic data efforts cover limited articulated-object categories, while general-purpose synthesis pipelines lack explicit designs for part-level semantics and articulation constraints, hindering agentic task generation and scalable synthesis of high-quality articulated-manipulation demonstrations. To bridge this gap, we introduce SMART, a scalable system leveraging large-scale Synthesized Manipulation demonstrations for ARTiculated-object manipulation. At its core, we develop SMART-Sim, a simulation platform with articulation-aware design that enables effective task generation and efficient demonstration collection. Building on SMART-Sim, we apply agentic task generation and design a scalable distributed synthesis system, using them to synthesize SMART-Data, comprising over 1M demonstrations across 44 atomic task types, 5 robot setups, and 2,507 articulated objects. The vision-language-action (VLA) model pretrained on SMART-Data shows competitive performance on simulation benchmarks and achieves zero-shot sim-to-real transfer and scalable performance in real-world articulated-object manipulation tasks. This highlights the potential of synthetic demonstrations in providing effective and scalable supervision for improving VLA model performance in contact-rich articulated-object manipulation.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. HyperSim: A Holistic Sim-To-Real Framework For Robust Robotic Manipulation

    May 26, 2026Junyi Dong, Haotian Luo, Ziwei Xu +11Sim-To-Real Reinforcement LearningRobotic Manipulation Policies

  2. Efficient Sim-to-Real Transfer of World-Action Models from Synthetic Priors

    Jun 30, 2026Zixing Wang, Kausik Sivakumar, Jinghuan Shang +5Sim-To-Real Reinforcement LearningRobotic Manipulation