cs.ROSep 27, 2026

VIDEAS: Distilling Explicit Action Semantics from Demonstration Videos for World Models via Prior-Guided Simulation

Authors: Jianan Wang, Haoquan Zhai, Siyang Zhang, Bin Li, Juan Chen, Jingtao Qi, Zhuo Zhang, Enze Wang, +2 more

Organizations: College of Computer Science and Technology, National University of Defense Technology · School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences · Intelligent Game and Decision Lab (IGDL) · School of Artificial Intelligence, Shanghai Jiao Tong University

Abstract

World models learn internal representations of environment dynamics to predict future states, enabling agents to optimize action plans without physical interactions. However, developing world models that genuinely internalize underlying causal physical laws to explicitly reason about action preconditions and subsequent state transitions remains an open challenge. In this paper, we propose VIDEAS, a data distillation framework that transforms continuous physical dynamics from operational videos into explicit action semantics for foundation models. Specifically, it deconstructs visual demonstrations into discrete action trajectories and utilizes advanced vision-language models (VLMs) to extract structured knowledge encapsulating action preconditions and effects. To ensure physical consistency, we introduce a prior-guided trajectory simulation mechanism grounded within a text-based environment to rigorously validate the extracted knowledge. Notably, we incorporate negative trajectories to enrich knowledge completeness and enhance data diversity to mitigate cognitive bias. Furthermore, we present VIDEAS-WM, an 8B/9B-parameter suite of language-based world models trained on 34K high-quality samples derived from AgiBot-World dataset. Extensive experiments demonstrate that VIDEAS-WM establishes state-of-the-art performance in high-level embodied action semantic reasoning, exhibiting profound physical understanding and robust generalization across unseen scenarios.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. CausalWM: Causal Chain-of-Thought Reasoning for Embodied World Model

    Sep 19, 2026Ziming Xu, Shuang Liang, Ruobing Han +10Efficient World-Action ModelFuture Video Prediction

  2. World Action Models: The Next Frontier in Embodied AI

    May 12, 2026Siyin Wang, Junhao Shi, Zhaoyang Fu +11Recent World-Action ModelsWorld Models

  3. From World Models to World Action Models: A Concise Tutorial for Robotics

    Jul 1, 2026Xiaoxiong Zhang, Xiong Zeng, Wei ZhangWorld ModelsAction Prediction