The central challenge of world modeling is to learn representations that capture how the world evolves. However, existing world models predominantly represent future states without explicitly capturing the latent causes underlying their evolution, limiting their ability to reason about why and how the world changes. To address this limitation, we propose Abductive World Modeling (AWM), a framework that learns structured causal representations by abductively inferring latent causes from predicted futures. Specifically, we realize AWM through the Hierarchical Abductive State Pyramid (HASP), which organizes the inferred world state into three complementary components - Entity, Dynamic, and Relation - capturing what exists, how it changes, and how entities interact, respectively. By jointly reasoning over the current observation and its predicted future, HASP abductively infers these latent factors and integrates them into a structured state representation for downstream reasoning. To the best of our knowledge, AWM is the first framework to introduce abductive state inference into latent-space world modeling for learning structured representations of world dynamics. Experiments across physical prediction, causal reasoning, and action understanding demonstrate the effectiveness of our approach. Compared with V-JEPA, a state-of-the-art latent-space world model, AWM improves physical prediction AUROC by 10.7%, causal reasoning accuracy by 16.8%, and action Top-1 accuracy by 68.0%.
Figures & tables
Figure 1 : Overview of the AWM. Given an observed context, a frozen predictive visual backbone first predicts its future latent state. The HASP reasons over the current and predicted future representations to infer Entity, Dynamic, and Relation states for downstream prediction and reasoning.
Physion++
CLEVRER
EK100 (T1/T5)
Method
AUROC
Bal. Acc.
Acc.
Option Acc.
Question Acc.
Verb
Noun
Action
V-JEPA 2
65.20
61.57
61.60
65.46
42.20
55.74/86.81
34.65/62.67
25.67/46.94
Orca
59.15
55.94
56.64
—
—
33.75/74.39
21.54/45.36
13.08/30.46
VideoMAE v2
67.61
62.60
62.86
—
—
59.54/82.57
46.37/68.79
35.96/54.57
AWM (Ours)
72.19
66.17
65.63
70.89
49.31
69.05/91.27
47.59/75.05
43.12/67.34
Table 1 : Main comparison across datasets. All entries are percentage values reported without the percent sign, and the best value in each metric column is boldfaced. “Bal. Acc.” and “Acc.” denote balanced accuracy and accuracy, respectively, while EK100 results are reported as Top-1/Top-5. Orca and VideoMAE v2 are marked as “—” on CLEVRER because their released models do not provide the predictive module required by the CLEVRER evaluation pipeline.
Figure 2 : Qualitative visualization of Entity grounding on EK100 and CLEVRER. (a)–(c) show representative EK100 scenes, where predicted Entity regions align with ground-truth hands and manipulated objects. (d) shows consecutive CLEVRER frames, where matched Entity predictions remain aligned with the corresponding ground-truth objects over time. Green boxes denote ground-truth regions and magenta boxes denote matched predicted regions.
State
Factor
Metric
Score
Entity
Presence
AUROC
0.9634
Spatial extent
R2
0.5316
Dynamic
Speed
R2
0.7198
Signed velocity
R2
0.3805
Relation
Contact
AUROC
0.9789
Pairwise distance
R2
0.5815
Table 2 : Native factor readability on Physion++.
Figure 3 : Interaction prediction with individual HASP states on CLEVRER. Full uses all three states.
Figure 4 : Entity intervention on Physion++ OCP prediction. Target masks the target Entity slot, Irrelevant masks a non-target slot, and Full retains all Entity slots. Performance is evaluated by accuracy, balanced accuracy, and AUROC ( ↑ ), showing how prediction quality changes when target-relevant or irrelevant entity information is removed.
Condition
Speed MAE ( ↓ )
Contact AUROC ( ↑ )
OCP AUROC ( ↑ )
AWM
0.0670
0.9837
0.7907
w/o Temporal
0.1120
0.9485
0.5646
Table 3 : Temporal-static latent intervention on Physion++. AWM uses the original context and future features. AWM w/o Temporal temporally averages and repeats these features at each time step, removing temporal variation while keeping the model and readouts frozen.
Condition
Pair-dist. MAE ( ↓ )
TTC MAE ( ↓ )
Contact AUROC ( ↑ )
AWM
0.0196
0.1306
0.9633
w/o Contact Pair
0.0225
0.1376
0.9488
w/o Non-contact Pair
0.0208
0.1349
0.9518
Table 4 : Relation-pair intervention on CLEVRER. AWM uses all Relation pairs; AWM w/o Contact Pair removes the target contact pair; and AWM w/o Non-contact Pair removes an unrelated pair. We report geometric and interaction prediction performance.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Diagnostic
Value
Presence F1
99.77
Slot occupancy
64.52
Identity consistency
70.29
ID switch rate
4.05
Target-object accuracy
95.79
Object-type accuracy
92.86
Appendix
Table 5 : Additional diagnostics of the Physion++ Entity slots.
Condition
TP
TN
FP
FN
Visual-only
237
292
129
142
Visual + ShallowProbe
277
297
124
102
Target mask
325
181
240
54
Irrelevant mask
262
305
116
117
Random slot
279
294
127
100
Appendix
Table 6 : Physion++ OCP confusion counts under probe interventions.
Condition
Mean change
Mean absolute change
Target mask
+0.49195
0.67680
Irrelevant mask
-0.08051
0.17355
Random slot
+0.02437
0.08895
Appendix
Table 7 : Paired OCP-logit changes under probe-level interventions.
Condition
Question Correct
Question Acc. ↑
Option Correct
Option Acc. ↑
AWM
1,785/3,557
50.18%
5,099/7,114
71.68%
AWM w/o Relevant Object
1,546/3,557
43.46%
4,752/7,114
66.80%
AWM w/o Irrelevant Object
1,684/3,557
47.34%
4,948/7,114
69.55%
Appendix
Table 8 : Relevant-object intervention on CLEVRER predictive queries. AWM uses all Entity evidence; AWM w/o Relevant Object removes the object referred to by the query; and AWM w/o Irrelevant Object removes an object unrelated to the query.
Parameter
Value
Predictive Backbone
Backbone
V-JEPA 2 ViT-H
Observed frames
16
Predicted future frames
16
Patch size
16
Tubelet size
2
Appendix
Table 9 : Main configuration and hyperparameter settings of AWM. We report the predictive backbone, evaluation configuration, and computational setup used in our experiments.
Embodied world models learn to predict future physical dynamics from visual observations and control signals, where physical knowledge is implicitly entangled within latent representations. We introduce CausalWM, a 16B embodied world model that performs explicit causal chain-of-thought reasoning before future video prediction. CausalWM organizes useful variables into a reasoning trajectory, allowing the model to progressively capture causal dependencies underlying physical evolution. To train CausalWM, we collect 31K hours embodied data and develop a three-stage paradigm consisting of large-scale video pre-training, causal CoT mid-training, and multi-objective RL post-training. Despite using only a limited set of supervised CoT variables, CausalWM exhibits emergent in-context learning capabilities, enabling contextual visual feature guidance and efficient few-step generation. CausalWM achieves state-of-the-art performance across language-conditioned, action-conditioned, single-view and multi-view benchmarks, including Top-1 performance on TriWorldBench leaderboard.
Ziming Xu, Shuang Liang, Ruobing Han +10
Aether AI · University of California, San Diego · Vanderbilt University
World Models (WM) are increasingly seen as a foundation for intelligent agents that can predict, plan, and act beyond their training distribution. In this paper, we study WMs from a causal perspective across multiple levels of abstraction, ranging from perceptual observations to building a conceptual representation of the structure governing the environment dynamics. We argue that useful WMs must go beyond generative capabilities alone: they should also capture entity properties, entity-to-entity interactions, and entity-to-environment interactions that determine and explain the dynamics of a system. We provide a formal definition of Causal WMs (CWMs) grounded in the tasks they are intended to support, connecting world modelling with existing work in causal representation learning, object-centric learning, causal discovery, structural causal models, and model-based decision-making. Finally, we relate CWMs to the literature on identifiability, clarifying when the components of a WM can be recovered from data and up to which equivalence. With this, we ground WMs in representations and structures that support causal reasoning and informed decision-making.
Avinash Kori, Fabrizio Russo
Department of Computing, Imperial College London, London, UK
A world model matters to an agent only through the state it constructs. That state must preserve some information, discard other information, and support some future function: prediction, control, planning, memory, grounding, or counterfactual reasoning. This paper treats world-model research as latent state design under sufficiency constraints. We propose a functional taxonomy that groups methods by what their latent state is for, rather than by architecture or application domain: predictive embedding, recurrent belief state, object/causal structure, latent action interface, grounded planning interface, and memory substrate. These roles expose distinctions that architecture-based groupings hide, including the gap between predictive sufficiency and control sufficiency, and the gap between passive video prediction and counterfactual action modeling. The taxonomy supports an evaluation framework that judges a model by the sufficiency constraint its latent state was built to satisfy. We compare methods along seven axes: representation, prediction, planning, controllability, causal/counterfactual support, memory, and uncertainty. We use the resulting matrix as a diagnostic for what a latent state preserves, discards, and enables. The conclusion that follows is that an actionable world model is the one whose state construction matches the task, not the one that preserves the most information.