cs.AISep 23, 2026

Training Object Permanence in World Models

Authors: Haotian Zhang, Fengyuan Yu, Dezhi Luo, Haoran Sun, Zehong Zhao, Qingying Gao, Yihan Li, Siyuan An, +23 more

Organizations: University of Southern California · Carnegie Mellon University · University of Michigan · Johns Hopkins University · University of California, San Diego · University of California, Los Angeles · Columbia University · University of Toronto · University of Bristol · University of California, Berkeley · University of Waterloo · Friedrich-Alexander-Universität Erlangen · University of Oxford · New York University · Stanford University · Harvard University

Abstract

Object permanence and solidity are hallmarks of human cognitive priors. Recent studies show that video generation models, a paradigmatic class of current world models, have begun to show emerged reasoning abilities, making them ideal candidates for building human-like physical intelligence. Do video models have emerged object permanence in them? If not, could we train them with a core-cognition inspired dataset? We introduce WROP (World Reasoning with Object Permanence), a data infrastructure of 150 hand-designed cognitive science inspired tasks, divided into six cognitive categories. We build Blender generators that randomize speed, lighting, camera angle, and other nuisance parameters while preserving each task's cognitive structure, yielding 10,000+ samples per task. We release a 1.5M-sample training corpus and a 300-question exam. On this exam we evaluate 14 video models: 3 reference-to-video, 7 edit, and 4 continuation, among which PWM-WROP, our 16B world model. In a blind pairwise Elo study, PWM-WROP ranks first among continuation models and third overall, behind only a statistical tie between two reference-to-video models. We release the data, exam, model answers, scores, weights, and PWM, our native-PyTorch training stack on AWS Trainium2.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors

    May 11, 2026Keming Wu, Yijing Cui, Wenhan Xue +11Long Video GenerationVideo Generation

  2. PhysMind: From Video to Executable Worlds for Training-Free Physical Reasoning

    Aug 5, 2026Chen Yang, Shenxiang Zeng, Haoyang Zhao +6Video World Models

  3. Astronex-World 1.0: Real-Time Interactive World Model Foundation

    Sep 17, 2026Xin Zhou, Cong MiaoVideo World ModelsAction Space