cs.CVOct 7, 2026

STRIKE: Learning Visual State Transitions for Physical World Modeling

Authors: Wenbin Teng, Tianshuo Xu, Depu Meng, Yuelei Li, Quentin Herau, Yihan Hu, Yajie Zhao, Wei Zhan

Organizations: Applied Intuition · University of Southern California · University of California, Berkeley

Abstract

Physical world modeling requires predicting how interactions change a scene, not merely generating coherent motion. We propose STRIKE, a framework that separates visual state transition learning from dense video generation. We construct event-aligned supervision by extracting observed states from training videos and pairing them with transition descriptions and temporal offsets. An image-based transition model learns to predict the next scene configuration from the current image, a local transition specification, and elapsed time. At inference, a pretrained vision-language planner predicts time transition specifications, and recursive application of the learned transition model produces a sequence of future visual states. A separately trained dynamic model then generates the complete rollout conditioned on these states and their temporal locations. Experiments on Physics-IQ Verified, PhyGenBench, Pisa-Experiments, and RoboTwin2.0 show improvements of STRIKE over the corresponding video-backbone baselines in benchmark measures of physical consistency and manipulation-video fidelity. These results support learned visual state transitions as an effective intermediate representation for physical world modeling.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Physical Object Understanding with a Physically Controllable World Model

    May 30, 2026Rahul Venkatesh, Klemen Kotar, Lilian Naing Chen +9Video World ModelsWorld Models

  2. PhiZero: A World Model Built Around Physical Language

    Jul 30, 2026Shuyao Shang, Yuqi Wang, Ruopeng Gao +4Video World ModelsWorld Models

  3. PhysMind: From Video to Executable Worlds for Training-Free Physical Reasoning

    Aug 5, 2026Chen Yang, Shenxiang Zeng, Haoyang Zhao +6Video World Models