The central challenge of world modeling is to learn representations that capture how the world evolves. However, existing world models predominantly represent future states without explicitly capturing the latent causes underlying their evolution, limiting their ability to reason about why and how the world changes. To address this limitation, we propose Abductive World Modeling (AWM), a framework that learns structured causal representations by abductively inferring latent causes from predicted futures. Specifically, we realize AWM through the Hierarchical Abductive State Pyramid (HASP), which organizes the inferred world state into three complementary components - Entity, Dynamic, and Relation - capturing what exists, how it changes, and how entities interact, respectively. By jointly reasoning over the current observation and its predicted future, HASP abductively infers these latent factors and integrates them into a structured state representation for downstream reasoning. To the best of our knowledge, AWM is the first framework to introduce abductive state inference into latent-space world modeling for learning structured representations of world dynamics. Experiments across physical prediction, causal reasoning, and action understanding demonstrate the effectiveness of our approach. Compared with V-JEPA, a state-of-the-art latent-space world model, AWM improves physical prediction AUROC by 10.7%, causal reasoning accuracy by 16.8%, and action Top-1 accuracy by 68.0%.
Figures & tables
Figure 1 : Overview of the AWM. Given an observed context, a frozen predictive visual backbone first predicts its future latent state. The HASP reasons over the current and predicted future representations to infer Entity, Dynamic, and Relation states for downstream prediction and reasoning.
Physion++
CLEVRER
EK100 (T1/T5)
Method
AUROC
Bal. Acc.
Acc.
Option Acc.
Question Acc.
Verb
Noun
Action
V-JEPA 2
65.20
61.57
61.60
65.46
42.20
55.74/86.81
34.65/62.67
25.67/46.94
Orca
59.15
55.94
56.64
—
—
33.75/74.39
21.54/45.36
13.08/30.46
VideoMAE v2
67.61
62.60
62.86
—
—
59.54/82.57
46.37/68.79
35.96/54.57
AWM (Ours)
72.19
66.17
65.63
70.89
49.31
69.05/91.27
47.59/75.05
43.12/67.34
Table 1 : Main comparison across datasets. All entries are percentage values reported without the percent sign, and the best value in each metric column is boldfaced. “Bal. Acc.” and “Acc.” denote balanced accuracy and accuracy, respectively, while EK100 results are reported as Top-1/Top-5. Orca and VideoMAE v2 are marked as “—” on CLEVRER because their released models do not provide the predictive module required by the CLEVRER evaluation pipeline.
Figure 2 : Qualitative visualization of Entity grounding on EK100 and CLEVRER. (a)–(c) show representative EK100 scenes, where predicted Entity regions align with ground-truth hands and manipulated objects. (d) shows consecutive CLEVRER frames, where matched Entity predictions remain aligned with the corresponding ground-truth objects over time. Green boxes denote ground-truth regions and magenta boxes denote matched predicted regions.
State
Factor
Metric
Score
Entity
Presence
AUROC
0.9634
Spatial extent
R2
0.5316
Dynamic
Speed
R2
0.7198
Signed velocity
R2
0.3805
Relation
Contact
AUROC
0.9789
Pairwise distance
R2
0.5815
Table 2 : Native factor readability on Physion++.
Figure 3 : Interaction prediction with individual HASP states on CLEVRER. Full uses all three states.
Figure 4 : Entity intervention on Physion++ OCP prediction. Target masks the target Entity slot, Irrelevant masks a non-target slot, and Full retains all Entity slots. Performance is evaluated by accuracy, balanced accuracy, and AUROC ( ↑ ), showing how prediction quality changes when target-relevant or irrelevant entity information is removed.
Condition
Speed MAE ( ↓ )
Contact AUROC ( ↑ )
OCP AUROC ( ↑ )
AWM
0.0670
0.9837
0.7907
w/o Temporal
0.1120
0.9485
0.5646
Table 3 : Temporal-static latent intervention on Physion++. AWM uses the original context and future features. AWM w/o Temporal temporally averages and repeats these features at each time step, removing temporal variation while keeping the model and readouts frozen.
Condition
Pair-dist. MAE ( ↓ )
TTC MAE ( ↓ )
Contact AUROC ( ↑ )
AWM
0.0196
0.1306
0.9633
w/o Contact Pair
0.0225
0.1376
0.9488
w/o Non-contact Pair
0.0208
0.1349
0.9518
Table 4 : Relation-pair intervention on CLEVRER. AWM uses all Relation pairs; AWM w/o Contact Pair removes the target contact pair; and AWM w/o Non-contact Pair removes an unrelated pair. We report geometric and interaction prediction performance.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Diagnostic
Value
Presence F1
99.77
Slot occupancy
64.52
Identity consistency
70.29
ID switch rate
4.05
Target-object accuracy
95.79
Object-type accuracy
92.86
Appendix
Table 5 : Additional diagnostics of the Physion++ Entity slots.
Condition
TP
TN
FP
FN
Visual-only
237
292
129
142
Visual + ShallowProbe
277
297
124
102
Target mask
325
181
240
54
Irrelevant mask
262
305
116
117
Random slot
279
294
127
100
Appendix
Table 6 : Physion++ OCP confusion counts under probe interventions.
Condition
Mean change
Mean absolute change
Target mask
+0.49195
0.67680
Irrelevant mask
-0.08051
0.17355
Random slot
+0.02437
0.08895
Appendix
Table 7 : Paired OCP-logit changes under probe-level interventions.
Condition
Question Correct
Question Acc. ↑
Option Correct
Option Acc. ↑
AWM
1,785/3,557
50.18%
5,099/7,114
71.68%
AWM w/o Relevant Object
1,546/3,557
43.46%
4,752/7,114
66.80%
AWM w/o Irrelevant Object
1,684/3,557
47.34%
4,948/7,114
69.55%
Appendix
Table 8 : Relevant-object intervention on CLEVRER predictive queries. AWM uses all Entity evidence; AWM w/o Relevant Object removes the object referred to by the query; and AWM w/o Irrelevant Object removes an object unrelated to the query.
Parameter
Value
Predictive Backbone
Backbone
V-JEPA 2 ViT-H
Observed frames
16
Predicted future frames
16
Patch size
16
Tubelet size
2
Appendix
Table 9 : Main configuration and hyperparameter settings of AWM. We report the predictive backbone, evaluation configuration, and computational setup used in our experiments.