cs.ROMay 27, 2026

World Models for Robotic Manipulation: A Survey

Authors: Fangyuan WangZiyuan WangGuorui PeiMengshi ZhangCanxi LiangJun HuZhongxuan LiJinsong Wu+10 more

Organizations: Department of Mechanical Engineering, The Hong Kong Polytechnic University, Kowloon, Hong Kong SAR, China · Department of Mechanical Engineering and Automation, Harbin Institute of Technology, Shenzhen, China · School of Advanced Engineering, Great Bay University, Dongguan, Guangdong, China · College of Robotics Science and Engineering, Taiyuan University of Technology, Taiyuan, China · School of Data Science, City University of Hong Kong (Dongguan), Dongguan, Guangdong, China · Department of Mechatronic Engineering, Guangdong Polytechnic Normal University, Guangdong, China · School of Computing and Data Science, The University of Hong Kong, Hong Kong SAR, China · School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore · College of Mechanical and Electrical Engineering, Northeast Forestry University, Harbin, China · Greater Bay Area National Center of Technology Innovation, Guangzhou, China · Department of Industrial and Systems Engineering, The Hong Kong Polytechnic University, Kowloon, Hong Kong SAR · Department of Production Engineering, KTH Royal Institute of Technology, Stockholm, Sweden

Abstract

Robotic manipulation depends on the ability to anticipate how actions reshape objects, contacts, and scene geometry before execution. Learned world models provide this capability by predicting task-relevant future evolution under robot intervention, yet the term now spans latent dynamics models, action-conditioned video generators, three- and four-dimensional scene predictors, physics-informed simulators, and predictive modules inside vision-language-action systems. This breadth has fragmented the literature and obscured the design choices that matter for manipulation. We survey world models for robotic manipulation through three questions: what future representation is predicted, how prediction is connected to action, and when prediction is used in the robot-learning pipeline. We operationally define a world model as an action-conditioned predictive system and distinguish it from perception modules, inverse models, policies, rewards, and value functions. We then organize existing work into five representation families, develop a functional taxonomy that separates integrated prediction-action models from explicit predictive planners, and characterize infrastructure roles including synthetic experience generation, candidate filtering, search-based evaluation, learned environments, and outcome verification. We further map these roles across pretraining, post-training, and inference adaptation, review 34 manipulation datasets, and synthesize evaluation protocols for predictive fidelity, task performance, and simulator reliability. This survey shows that world models are evolving from task-specific dynamics predictors into predictive infrastructure for robot learning, while exposing open challenges in contact modeling, hallucination control, action alignment, and benchmarking under closed-loop use.

Explore similar work

CardsList