Organizations: School of Computing Technologies, RMIT University, Australia · Faculty of Information Technology, Monash University, Australia · KStelrix Star Dynamics Lab, KStelrix, China · School of Computing and Information Systems, The University of Melbourne, Australia
World models aim to learn representations of real-world environments and predict their future evolution. Recent object-centric world models have made expressive progress by representing visual scenes as sets of object-level latent states, but object-object relations are often captured only implicitly, which limits explicit relational and temporal structure modeling and object-centric dynamic memory modeling. To address such challenges, we propose World-As-Graph (WAG), a graph-based object-centric world model that introduces relational inductive bias into JEPA-style predictive representation learning. The proposed WAG contains two main modules: (1) Relation-aware structure induction, which constructs time-varying latent graphs from object-centric slots and designs relation-aware object masking policies to guide relational object representation learning in latent space; (2) Object-centric memory transition, which maintains and updates object-level dynamic states by combining relational information from neighboring objects with historical memory, enabling effective autoregressive future prediction. Extensive experiments on both visual reasoning and robotic manipulation tasks could demonstrate the superior performance of our proposed WAG.
Figures & tables
Figure 1: Comparison of our proposed World-as-Graph ( WaG ) with existing object-centric world models (with no explicit graph structures (C1) and limited temporal memory (C2)).
Figure 2: Overall framework of our proposed WaG .
Model
Average
Counterfactual (%)
Explanatory (%)
Predictive (%)
Descriptive (%)
per que. (%)
per opt.
per que.
per opt.
per que.
per opt.
per que.
VideoSAUR Encoder
OC-JEPA
82.79
79.53
47.68
92.88
80.58
86.15
75.04
89.59
C-JEPA
89.40
88.67
68.81
96.62
90.74
93.03
86.93
92.84
WaG I (ours)
90.79
90.06
72.22
97.79
93.96
92.92
86.81
93.71
ΔI Improv. ↑
+1.39
+1.39
+3.41
+1.17
+3.22
-0.11
-0.12
+0.87
Table 1: VQA accuracy (%) comparison on CLEVRER using VideoSAUR and SAVi encoders. WaG I and WaG B denote instance-wise masking and batch-shared masking, respectively. ΔI and ΔB indicate the absolute improvements over C-JEPA in percentage points (pp). The best results are shown in bold, and the second-best results are underlined.
# Token ×d
Model
Success Rate (%)
196×384
DINO-WM
91.33
DINO-WM-Reg.
88.00
6×128
OC-DINO-WM (ref.)
60.67
OC-JEPA
76.00 ↑+15.33
C-JEPA
88.67 ↑+28.00
WaG (ours)
90.70 ↑+30.03
Table 2: PushT planning success rates across different world-model token budgets.
Table 5
Figure 3: Performance improvement of different latent dynamic graph construction strategies over C-JEPA.
Figure 4: Visualization of physical interaction modeling on CLEVRER.(a) Physical interaction from the last observed to future frames. (b) Latent dynamic graphs constructed by WaG with K=3 , showing evolving object relations. A/B/C correspondences are inferred. (c) Ground-truth object motion changes. (d) Latent prediction errors of C-JEPA and WaG , with lower errors achieved by WaG during interaction-driven transitions.
Figure 5: Analysis of the number of masked objects M and graph neighborhood size K .
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Property
PushT
CLEVRER
Train split
18,685 trajectories
10,000 videos
Val / test split
21 trajectories
5,000 / 5,000 videos
Object slots
4
7
Task type
2D pushing manipulation
Video QA / causal reasoning
Downstream head
CEM planner
ALOE-style
Evaluation metric
Success rate
QA accuracy
Appendix
Table A1: Dataset statistics.
Table 10
Masking
Model
Average per que. (%)
Counterfactual (%)
Explanatory (%)
Predictive (%)
per opt.
per que.
per opt.
per que.
per opt.
per que.
Random
WaG I
90.94
90.20
72.83
97.69
93.54
93.97
88.75
WaG B
91.14
90.31
72.76
97.71
93.64
94.21
89.34
Δ (B − I)
+0.20
+0.11
− 0.07
+0.02
+0.10
+0.24
+0.59
Relational Centrality
WaG I
90.79
90.06
72.22
97.79
93.96
92.92
86.81
WaG B
90.85
89.96
72.20
97.31
92.80
93.52
87.94
Appendix
Table A4: Comparison of masking strategies for WaG I vs. WaG B on the VideoSAUR encoder (all runs at λfutu=1 ) and λmask=1 .
Relation-Aware Structure Induction
Object-Centric Memory Transition
Performance (%) ↑
Variants
DyGraph
Rel. Mask
TGNN
Mem. Pred.
Avg. per que.
Counter.
Explan.
Predict.
Descrip.
C-JEPA
×
×
×
×
89.40
68.81
90.74
86.93
92.84
w/o DyGraph
×
×
×
✓
90.63
72.68
93.38
89.15
93.36
w/o Rel. Mask
✓
×
✓
✓
90.94
72.83
93.54
88.75
93.75
w/o TGNN
✓
✓
×
✓
90.97
73.42
93.52
90.33
93.59
w/o Mem. Pred.
✓
✓
✓
×
90.21
71.50
92.46
85.55
93.34
Appendix
Table A5: Overall ablation study of key components in our proposed WaG with VideoSAUR encoder. Counter. and Predict. denote per-question accuracy for counterfactual and predictive questions, respectively.
Masking Strategy
VideoSAUR Encoder
SAVi Encoder
Avg. per. que. (%)
Counter. (%)
Predict. (%)
Avg. per. que. (%)
Counter. (%)
Predict. (%)
Random [w/o DyG.]
89.40
68.81
86.93
83.88
60.19
77.25
Random [w/ DyG.]
90.94 (+1.54)
72.83 (+4.02)
88.75 (+1.82)
92.28 (+8.40)
73.75 (+13.56)
91.14 (+13.89)
Relational Centrality
91.37 (+1.97)
73.76 (+4.95)
89.54 (+2.61)
92.72 (+8.84)
75.45 (+15.26)
91.99 (+14.74)
Temporal Dynamics
90.47 (+1.07)
70.16 (+1.35)
87.71 (+0.78)
91.95 (+8.07)
73.51 (+13.32)
89.71 (+12.46)
Appendix
Table A6: Comparison of different object masking strategies under VideoSAUR and SAVi representations using per-question accuracy. Values in parentheses denote absolute improvements over Random [w/o DyG.] (i.e., C-JEPA) in percentage points, where w/o DyG. and w/ DyG. indicate without and with latent dynamic graph construction, respectively.
Model
Average
Counterfactual (%)
Explanatory (%)
Predictive (%)
Descriptive (%)
per que. (%)
per opt.
per que.
per opt.
per que.
per opt.
per que.
C-JEPA
89.40
88.67
68.81
96.62
90.74
93.03
86.93
92.84
Random Graph (IID)
90.65
89.52
71.02
97.28
92.60
94.24
89.23
93.77
Random Graph (Subset)
90.88
90.24
72.68
97.55
93.27
93.94
88.73
93.74
WaG B (ours)
91.37
90.55
73.76
97.84
94.13
94.43
89.54
94.05
Appendix
Table A7: VQA accuracy comparison of different random-graph sampling strategies and WaG B using the VideoSAUR encoder.
Figure A1: Analysis of graph neighborhood size K and MLP model size with width.
Figure A2: Analysis of the number of masked objects M and graph neighborhood size K . Average per-question and counterfactual VQA accuracy (%) are reported under VideoSAUR and SAVi representations.
Figure A3: Computational efficiency and hyperparameter sensitivity analysis. (a) Computational cost measured in GFLOPs. (b) Memory consumption. (c) Running time. (d) Performance sensitivity to the hyperparameter λfutu .
As one of the mainstream models of artificial intelligence, world models allow agents to learn the representation of the environment for efficient prediction and planning. However, classical world models based on flat tensors face several key problems, including noise sensitivity, error accumulation and weak reasoning. To address these limitations, many recent studies use graph structure to decompose the environment into entity nodes and interactive edges, and model virtual environments in a structured space. This paper systematically formalizes and unifies these emerging graph-based works under the concept of graph world models (GWMs). To the best of our knowledge, GWMs have not yet been explicitly defined and surveyed as a unified research paradigm. Furthermore, we propose a taxonomy based on relational inductive biases (RIB), categorizing GWMs by the specific structural priors they inject: (1) spatial RIB for topological abstraction; (2) physical RIB for dynamic simulation; and (3) logical RIB for causal and semantic reasoning. For each model category, we outline the key design principles, summarize representative models, and conduct comparative analyses. We further discuss open challenges and future directions, including dynamic graph adaptation, probabilistic relational dynamics, multi-granularity inductive biases, and the need for dedicated benchmarks and evaluation metrics for GWMs.
Jiawei Liu, Senqiao Yang, Mingjun Wang +2
The Chinese University of Hong Kong, Hong Kong, China · Tsinghua University, Beijing, China
World models infer latent states of an environment to capture its underlying dynamics and predict future evolution. Many real-world environments, however, are inherently relational and observed as evolving graphs, where entities, relations, and their properties change over time. Prior graph-related world models use graph structures to organize internal states or support task-specific reasoning, rather than treating an evolving graph itself as the modeled world. We instead study graph world modeling (GWM), where graph evolution itself constitutes the world dynamics. We formulate graph world modeling over observed graph evolution, latent graph states, and heterogeneous graph-transition predictions. Based on this formulation, we construct GWM-Zero, a benchmark covering node-, edge-, and graph-level transitions over eight temporal graph datasets. We propose WorldGraph, which combines a state-aware graph transformer for multi-granularity structural and transition-conditioned evolution modeling with transition-aware GRPO using dynamic grouping and structure-aware verifiable rewards. Extensive experiments on GWM-Zero show that WorldGraph consistently outperforms representative graph representation, temporal graph learning, graph pretraining, and graph world-model baselines across all three transition granularities.
Zezhong Ding, Yipeng Li, Xike Xie
School of Artificial Intelligence and Data Science, University of Science and Technology of China (USTC) · Data Darkness Lab, Suzhou Institute for Advanced Research, USTC · School of Biomedical Engineering, USTC
The central challenge of world modeling is to learn representations that capture how the world evolves. However, existing world models predominantly represent future states without explicitly capturing the latent causes underlying their evolution, limiting their ability to reason about why and how the world changes. To address this limitation, we propose Abductive World Modeling (AWM), a framework that learns structured causal representations by abductively inferring latent causes from predicted futures. Specifically, we realize AWM through the Hierarchical Abductive State Pyramid (HASP), which organizes the inferred world state into three complementary components - Entity, Dynamic, and Relation - capturing what exists, how it changes, and how entities interact, respectively. By jointly reasoning over the current observation and its predicted future, HASP abductively infers these latent factors and integrates them into a structured state representation for downstream reasoning. To the best of our knowledge, AWM is the first framework to introduce abductive state inference into latent-space world modeling for learning structured representations of world dynamics. Experiments across physical prediction, causal reasoning, and action understanding demonstrate the effectiveness of our approach. Compared with V-JEPA, a state-of-the-art latent-space world model, AWM improves physical prediction AUROC by 10.7%, causal reasoning accuracy by 16.8%, and action Top-1 accuracy by 68.0%.