VIDEAS: Distilling Explicit Action Semantics from Demonstration Videos for World Models via Prior-Guided Simulation
Authors: Jianan Wang, Haoquan Zhai, Siyang Zhang, Bin Li, Juan Chen, Jingtao Qi, Zhuo Zhang, Enze Wang, +2 more
Organizations: College of Computer Science and Technology, National University of Defense Technology · School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences · Intelligent Game and Decision Lab (IGDL) · School of Artificial Intelligence, Shanghai Jiao Tong University
World models learn internal representations of environment dynamics to predict future states, enabling agents to optimize action plans without physical interactions. However, developing world models that genuinely internalize underlying causal physical laws to explicitly reason about action preconditions and subsequent state transitions remains an open challenge. In this paper, we propose VIDEAS, a data distillation framework that transforms continuous physical dynamics from operational videos into explicit action semantics for foundation models. Specifically, it deconstructs visual demonstrations into discrete action trajectories and utilizes advanced vision-language models (VLMs) to extract structured knowledge encapsulating action preconditions and effects. To ensure physical consistency, we introduce a prior-guided trajectory simulation mechanism grounded within a text-based environment to rigorously validate the extracted knowledge. Notably, we incorporate negative trajectories to enrich knowledge completeness and enhance data diversity to mitigate cognitive bias. Furthermore, we present VIDEAS-WM, an 8B/9B-parameter suite of language-based world models trained on 34K high-quality samples derived from AgiBot-World dataset. Extensive experiments demonstrate that VIDEAS-WM establishes state-of-the-art performance in high-level embodied action semantic reasoning, exhibiting profound physical understanding and robust generalization across unseen scenarios.
Figures & tables
Figure 1: Conventional world models vs. VIDEAS-WM. For an infeasible action, the former ignores action preconditions and hallucinates successful state transitions, while VIDEAS-WM verifies preconditions before predicting environment evolution, correctly judging the action infeasible.
Figure 2: Overview of VIDEAS. Based on in-the-wild demonstration videos, VIDEAS first curates them into discrete action trajectories and grounds them into text-based structured environments, where it leverages VLMs to iteratively simulate each action via Consider–Decide–Transfer phases under execution priors. Moreover, it constructs negative trajectories via dependency-aware mutation to enrich precondition knowledge. Finally, the distilled 34K-sample corpus serves to train VIDEAS-WM, an 8B/9B-parameter suite of language-based world models.
Figure 3: Training corpus statistics. In (b), Succ. and Neg. denote samples derived from successful and negative trajectories respectively. (c) shows representative action types.
Model
BTSIMBENCH
LogicEnvEval
Delivery
Good
CFactuals
Unreach
Avg
Delivery
Good
CFactuals
Unreach
LBranch
Avg
Closed-Source Models
Gemini-3.1-Pro
100.00
92.00
64.00
96.00
84.00
100.00
92.00
72.00
92.00
96.00
88.00
Gemini-3-Flash
100.00
88.00
64.00
88.00
80.00
100.00
76.00
76.00
88.00
96.00
84.00
Open-Source Models
DeepSeek-V4-Pro
100.00
44.00
88.00
76.00
69.33
100.00
52.00
84.00
88.00
96.00
80.00
Table 1: Performance (%) on BTSIMBENCH and LogicEnvEval benchmarks. Each policy category contains 25 samples. Gray and green rows represent closed-source baselines and VIDEAS-WM, respectively. The best performance among open-source models is in bold .
Model
WorldPrediction-WM
COIN
CrossTask
IKEA-ASM
EPIC-KITCHENS-100
Overall
VLMs
InternVL3.5-38B
47.69
55.56
40.88
73.44
54.87
InternVL3.5-14B
47.18
45.56
35.85
51.56
45.44
InternVL3.5-8B
45.13
44.44
32.08
40.63
40.41
Qwen3.5-27B
51.28
55.56
47.80
67.19
55.82
Qwen3-VL-32B
53.85
54.44
45.28
69.27
56.45
Table 2: Performance (%) on WorldPrediction-WM benchmark. Green rows represent our VIDEAS-WM. The best performance within each category (VLMs and Socratic LLMs) is in bold .
Model
Method
BTSIMBENCH
Good
CFactuals
Unreach
Avg
Llama3.1-8B
Base
0.00
76.00
8.00
28.00
Naive Distillation
72.00
32.00
88.00
64.00
W/O NTI
88.00
28.00
100.00
72.00
VIDEAS-WM
92.00
96.00
96.00
94.67
Qwen3-8B
Base
12.00
80.00
32.00
41.33
Table 3: Results of ablation study on the prior-guided trajectory simulation mechanism (Naive Distillation) and the negative trajectory integration (W/O NTI).
Figure 4: Ablation study results on training data scales for the BTSIMBENCH benchmark.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Value
Training epochs
3
Learning rate
1×10−4
Warmup ratio
0.03
Batch size
4
Gradient accumulation steps
1
LoRA rank
32
Appendix
Table 4: Training hyperparameters.
Dataset
Number of Samples
COIN
195
CrossTask
90
IKEA-ASM
159
EPIC-KITCHENS-100
193
Total
637
Appendix
Table 5: Distribution of accessible task samples in WorldPrediction-WM.
Figure 5: Pareto frontier of accuracy versus average simulation time per robot policy. VIDEAS-WM establishes the optimal balance between reasoning capability and computational efficiency.
Figure 6: An illustrative example of the text-based environment. The complete JSON structure encompasses comprehensive fine-grained properties, which are abbreviated here for brevity.
Figure 7: The prompt structure and expected output for the consider stage.
Figure 8: The prompt structure and expected output for the decide stage.
Figure 9: The prompt structure and expected output for the capture stage.
Figure 10: The prompt structure and expected output for the transfer stage.
Figure 11: An illustrative example of negative trajectories. By removing the initial Open action, the subsequent Pick action is deliberately designed to fail due to unmet physical preconditions.
Figure 12: Case study illustrating the simulation of a robot policy from the BTSIMBENCH benchmark using VIDEAS-WM-9B. For each action within the sequence, the predicted preconditions and effects are formalized in JSON format.