VIDEAS: Distilling Explicit Action Semantics from Demonstration Videos for World Models via Prior-Guided Simulation
Authors: Jianan Wang, Haoquan Zhai, Siyang Zhang, Bin Li, Juan Chen, Jingtao Qi, Zhuo Zhang, Enze Wang, +2 more
Organizations: College of Computer Science and Technology, National University of Defense Technology · School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences · Intelligent Game and Decision Lab (IGDL) · School of Artificial Intelligence, Shanghai Jiao Tong University
World models learn internal representations of environment dynamics to predict future states, enabling agents to optimize action plans without physical interactions. However, developing world models that genuinely internalize underlying causal physical laws to explicitly reason about action preconditions and subsequent state transitions remains an open challenge. In this paper, we propose VIDEAS, a data distillation framework that transforms continuous physical dynamics from operational videos into explicit action semantics for foundation models. Specifically, it deconstructs visual demonstrations into discrete action trajectories and utilizes advanced vision-language models (VLMs) to extract structured knowledge encapsulating action preconditions and effects. To ensure physical consistency, we introduce a prior-guided trajectory simulation mechanism grounded within a text-based environment to rigorously validate the extracted knowledge. Notably, we incorporate negative trajectories to enrich knowledge completeness and enhance data diversity to mitigate cognitive bias. Furthermore, we present VIDEAS-WM, an 8B/9B-parameter suite of language-based world models trained on 34K high-quality samples derived from AgiBot-World dataset. Extensive experiments demonstrate that VIDEAS-WM establishes state-of-the-art performance in high-level embodied action semantic reasoning, exhibiting profound physical understanding and robust generalization across unseen scenarios.
Figures & tables
Figure 1: Conventional world models vs. VIDEAS-WM. For an infeasible action, the former ignores action preconditions and hallucinates successful state transitions, while VIDEAS-WM verifies preconditions before predicting environment evolution, correctly judging the action infeasible.
Figure 2: Overview of VIDEAS. Based on in-the-wild demonstration videos, VIDEAS first curates them into discrete action trajectories and grounds them into text-based structured environments, where it leverages VLMs to iteratively simulate each action via Consider–Decide–Transfer phases under execution priors. Moreover, it constructs negative trajectories via dependency-aware mutation to enrich precondition knowledge. Finally, the distilled 34K-sample corpus serves to train VIDEAS-WM, an 8B/9B-parameter suite of language-based world models.
Figure 3: Training corpus statistics. In (b), Succ. and Neg. denote samples derived from successful and negative trajectories respectively. (c) shows representative action types.
Model
BTSIMBENCH
LogicEnvEval
Delivery
Good
CFactuals
Unreach
Avg
Delivery
Good
CFactuals
Unreach
LBranch
Avg
Closed-Source Models
Gemini-3.1-Pro
100.00
92.00
64.00
96.00
84.00
100.00
92.00
72.00
92.00
96.00
88.00
Gemini-3-Flash
100.00
88.00
64.00
88.00
80.00
100.00
76.00
76.00
88.00
96.00
84.00
Open-Source Models
DeepSeek-V4-Pro
100.00
44.00
88.00
76.00
69.33
100.00
52.00
84.00
88.00
96.00
80.00
Table 1: Performance (%) on BTSIMBENCH and LogicEnvEval benchmarks. Each policy category contains 25 samples. Gray and green rows represent closed-source baselines and VIDEAS-WM, respectively. The best performance among open-source models is in bold .
Model
WorldPrediction-WM
COIN
CrossTask
IKEA-ASM
EPIC-KITCHENS-100
Overall
VLMs
InternVL3.5-38B
47.69
55.56
40.88
73.44
54.87
InternVL3.5-14B
47.18
45.56
35.85
51.56
45.44
InternVL3.5-8B
45.13
44.44
32.08
40.63
40.41
Qwen3.5-27B
51.28
55.56
47.80
67.19
55.82
Qwen3-VL-32B
53.85
54.44
45.28
69.27
56.45
Table 2: Performance (%) on WorldPrediction-WM benchmark. Green rows represent our VIDEAS-WM. The best performance within each category (VLMs and Socratic LLMs) is in bold .
Model
Method
BTSIMBENCH
Good
CFactuals
Unreach
Avg
Llama3.1-8B
Base
0.00
76.00
8.00
28.00
Naive Distillation
72.00
32.00
88.00
64.00
W/O NTI
88.00
28.00
100.00
72.00
VIDEAS-WM
92.00
96.00
96.00
94.67
Qwen3-8B
Base
12.00
80.00
32.00
41.33
Table 3: Results of ablation study on the prior-guided trajectory simulation mechanism (Naive Distillation) and the negative trajectory integration (W/O NTI).
Figure 4: Ablation study results on training data scales for the BTSIMBENCH benchmark.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Value
Training epochs
3
Learning rate
1×10−4
Warmup ratio
0.03
Batch size
4
Gradient accumulation steps
1
LoRA rank
32
Appendix
Table 4: Training hyperparameters.
Dataset
Number of Samples
COIN
195
CrossTask
90
IKEA-ASM
159
EPIC-KITCHENS-100
193
Total
637
Appendix
Table 5: Distribution of accessible task samples in WorldPrediction-WM.
Figure 5: Pareto frontier of accuracy versus average simulation time per robot policy. VIDEAS-WM establishes the optimal balance between reasoning capability and computational efficiency.
Figure 6: An illustrative example of the text-based environment. The complete JSON structure encompasses comprehensive fine-grained properties, which are abbreviated here for brevity.
Figure 7: The prompt structure and expected output for the consider stage.
Figure 8: The prompt structure and expected output for the decide stage.
Figure 9: The prompt structure and expected output for the capture stage.
Figure 10: The prompt structure and expected output for the transfer stage.
Figure 11: An illustrative example of negative trajectories. By removing the initial Open action, the subsequent Pick action is deliberately designed to fail due to unmet physical preconditions.
Figure 12: Case study illustrating the simulation of a robot policy from the BTSIMBENCH benchmark using VIDEAS-WM-9B. For each action within the sequence, the predicted preconditions and effects are formalized in JSON format.
Embodied world models learn to predict future physical dynamics from visual observations and control signals, where physical knowledge is implicitly entangled within latent representations. We introduce CausalWM, a 16B embodied world model that performs explicit causal chain-of-thought reasoning before future video prediction. CausalWM organizes useful variables into a reasoning trajectory, allowing the model to progressively capture causal dependencies underlying physical evolution. To train CausalWM, we collect 31K hours embodied data and develop a three-stage paradigm consisting of large-scale video pre-training, causal CoT mid-training, and multi-objective RL post-training. Despite using only a limited set of supervised CoT variables, CausalWM exhibits emergent in-context learning capabilities, enabling contextual visual feature guidance and efficient few-step generation. CausalWM achieves state-of-the-art performance across language-conditioned, action-conditioned, single-view and multi-view benchmarks, including Top-1 performance on TriWorldBench leaderboard.
Ziming Xu, Shuang Liang, Ruobing Han +10
Aether AI · University of California, San Diego · Vanderbilt University
Vision-Language-Action (VLA) models have achieved strong semantic generalization for embodied policy learning, yet they learn reactive observation-to-action mappings without explicitly modeling how the physical world evolves under intervention. A growing body of work addresses this limitation by integrating world models, predictive models of environment dynamics, into the action generation pipeline. We term this emerging paradigm World Action Models (WAMs): embodied foundation models that unify predictive state modeling with action generation, targeting a joint distribution over future states and actions rather than actions alone. However, the literature remains fragmented across architectures, learning objectives, and application scenarios, lacking a unified conceptual framework. We formally define WAMs and disambiguate them from related concepts, and trace the foundations and early integration of VLA and world model research that gave rise to this paradigm. We organize existing methods into a structured taxonomy of Cascaded and Joint WAMs, with further subdivision by generation modality, conditioning mechanism, and action decoding strategy. We systematically analyze the data ecosystem fueling WAMs development, spanning robot teleoperation, portable human demonstrations, simulation, and internet-scale egocentric video, and synthesize emerging evaluation protocols organized around visual fidelity, physical commonsense, and action plausibility. Overall, this survey provides the first systematic account of the WAMs landscape, clarifies key architectural paradigms and their trade-offs, and identifies open challenges and future opportunities for this rapidly evolving field.
Siyin Wang, Junhao Shi, Zhaoyang Fu +11
1Fudan University · 2Shanghai Innovation Institute · 3National University of Singapore
World models are increasingly used in embodied intelligence and generative simulation, yet their scope remains ambiguous across communities. This tutorial presents a design-space view of world models as action-conditioned predictive models that estimate the future evolution of task-relevant observations or states. We categorize existing methods into observation-space and state-space world models, comparing their trade-offs in visual fidelity, spatial structure, physical interpretability, and control usability. We further introduce world action models, which connect predicted futures with executable robot actions, and summarize four representative paradigms: imagine-then-execute, video-feature-conditioned action prediction, joint video-action modeling, and auxiliary video prediction for policy learning. The goal of this tutorial is to clarify the conceptual scope of world (action) models and provide a structured taxonomy for embodied prediction and control.
Xiaoxiong Zhang, Xiong Zeng, Wei Zhang
School of Automation and Intelligent Manufacturing, Southern University of Science and Technology, Shenzhen, China · LimX Dynamics, Shenzhen, China