Organizations: School of Artificial Intelligence and Data Science, University of Science and Technology of China (USTC) · Data Darkness Lab, Suzhou Institute for Advanced Research, USTC · School of Biomedical Engineering, USTC
World models infer latent states of an environment to capture its underlying dynamics and predict future evolution. Many real-world environments, however, are inherently relational and observed as evolving graphs, where entities, relations, and their properties change over time. Prior graph-related world models use graph structures to organize internal states or support task-specific reasoning, rather than treating an evolving graph itself as the modeled world. We instead study graph world modeling (GWM), where graph evolution itself constitutes the world dynamics. We formulate graph world modeling over observed graph evolution, latent graph states, and heterogeneous graph-transition predictions. Based on this formulation, we construct GWM-Zero, a benchmark covering node-, edge-, and graph-level transitions over eight temporal graph datasets. We propose WorldGraph, which combines a state-aware graph transformer for multi-granularity structural and transition-conditioned evolution modeling with transition-aware GRPO using dynamic grouping and structure-aware verifiable rewards. Extensive experiments on GWM-Zero show that WorldGraph consistently outperforms representative graph representation, temporal graph learning, graph pretraining, and graph world-model baselines across all three transition granularities.
Figures & tables
Figure 1: World Model ( Ha and Schmidhuber, 2018 ) vs. Our Proposed GWM
Figure 2: Overview of WorldGraph . The state-aware graph transformer constructs latent world states by integrating multi-granularity graph structure with transition-conditioned history. The transition-aware RL algorithm learns heterogeneous graph transitions through dynamic grouping and structure-aware verifiable rewards across node-, edge-, and graph-level tasks.
Method
Type
TGBN-Trade
TGBN-Genre
TGBN-Reddit
Add. F1
Rem. F1
Sem. F1
NDCG@10
Add. F1
Rem. F1
Sem. F1
NDCG@10
Add. F1
Rem. F1
Sem. F1
NDCG@10
GCN
MPNN
0.4718
0.1540
0.5741
0.6357
0.7112
0.8488
0.6375
0.2469
0.9706
0.8976
0.6284
0.2161
GAT
MPNN
0.5248
0.1385
0.6275
0.6386
0.7253
0.8598
0.6523
0.2699
0.9720
0.8972
0.6440
0.2237
GraphSAGE
MPNN
0.5864
0.0000
0.6595
0.6246
0.7256
0.8538
0.6496
0.2780
0.9705
0.8902
0.6435
0.2519
SGFormer
Graph Transformer
0.6874
0.0308
0.6732
0.6324
0.7252
0.8564
0.6497
0.2822
0.9731
0.9024
0.6445
0.2413
NodeFormer
Graph Transformer
0.7069
0.0878
0.6669
0.6336
0.7080
0.8394
0.6470
0.2344
0.8011
0.7329
0.5409
0.1275
Table 1: Overall Performance on the 3 Graph-Transition Tasks. Add., Rem., and Sem. denote addition, deletion, and property-change F1, respectively. Higher is better for F1, Macro-F1, and NDCG@10, while lower is better for MAE and RMSE. The best results are shown in bold and the second best results are underlined . We use an existing Controller design ( Dinella et al., 2020 ) to enable the baselines in the first three categories to predict node-, edge-, and graph-level changes.
Figure 3: Ablation Study on Node-Level Tasks/TGBN-Trade and Graph-Level Tasks/Enron. The left panel reports the equal-weight average of the 4 TGBN-Trade metrics, while the right panel reports Macro-F1.
Figure 4: Sensitivity Analysis on Graph-Level Tasks/Flights. Performance remains remarkably stable across varying maximum hop numbers Lmax , random-walk numbers M , and GRPO rollout budgets B .
Figure 5: Comparison with L 3 P on PointMaze and FetchPickAndPlace. Solid curves show the mean test success rate and shaded regions show the standard deviation.
Env.
Model
1 Step
5 Steps
10 Steps
H@1
MRR
H@1
MRR
H@1
MRR
Pong
C-SWM
35.00
51.45
11.33
27.24
6.67
19.03
+ WorldGraph
40.33
57.68
26.33
45.85
16.00
34.62
SI
C-SWM
63.67
75.49
41.33
56.66
29.67
46.16
+ WorldGraph
68.00
79.68
62.67
76.04
61.00
75.71
Table 2: Comparison with C-SWM on Multi-Step Latent-State Ranking Tasks. Higher H@1 and MRR are better.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Description
Graph World
t,τ,T
Current time, historical time, and sequence length.
gt=(Vt,Et,Xt)∈G
Observed graph at time t , its node, edge, and property sets, and the graph space.
Δgt
Graph change associated with the transition from gt to gt+1 .
st∈S,fs
Latent world state, world-state space, and state model.
fΔ
Transition model that predicts Δgt from the current graph, the preceding graph change, and st .
Appendix
Table 4: Core notations used in WorldGraph.
Encoder setting
∥zv,t−zv,t(IDEAL)∥2
∥sv,t−sv,t(IDEAL)∥2
SGFormer (Single-view)
11.2743±0.3400
6.2460±0.9255
WorldGraph (Multi-view SGT)
3.8210 ± 0.5362
1.5568 ± 0.3691
Appendix
Table 5: Experimental validation of Theorem 1 on Graph-level tasks/Flights. Values are mean ± standard deviation over five checkpoints; lower is better.
Method
Mean EVR
Standard GRPO
1.5550
Transition-aware RL
1.3526
Appendix
Table 6: Experimental validation of Theorem 2 on Node-level tasks/TGBN-Trade. Mean EVR is the arithmetic mean over the three transition groups; lower is better.
Dataset
Domain
Nodes
Temporal Events
Snapshots
Edge Input
Snapshot Interval
Task(s)
TGBN-Trade
Economic interaction
255
468,245
31
Weighted
Annual
Node-Level Tasks, Edge-Level Tasks
TGBN-Genre
User preference
1,505
17,858,395
1,580
Weighted
Daily
Node-Level Tasks
TGBN-Reddit
Social interaction
11,766
27,174,118
1,090
Binary
Daily
Node-Level Tasks
UN Vote
Political interaction
201
1,035,742
72
Weighted
Annual
Edge-Level Tasks
Contact
Physical proximity
692
2,426,279
112
Binary
Six hours
Edge-Level Tasks, Graph-Level Tasks
SocialEvo
Social proximity
74
2,098,119
243
Binary
Daily
Edge-Level Tasks
Appendix
Table 7: Statistics of the eight real-world datasets used in GWM-Zero. The numbers of temporal events and snapshots are computed from the processed artifacts used by all methods.
Hyperparameter
Node-level tasks
Edge-level tasks
Graph-level tasks
Latent-state dimension
64
32
64
Hidden-state dimension
64
32
64
Transition-representation dimension
32
16
32
Maximum historical time span
8
8
8
Dropout rate
0.1
0.1
0.1
Learning rate
10−3
10−3
10−3
Appendix
Table 8: Principal hyperparameters used by WorldGraph.
Model
Time (s/epoch)
GCN
47.26
GAT
51.47
GraphSAGE
43.49
SGFormer
53.45
NodeFormer
72.11
GraphGPS
64.04
Appendix
Table 9: Mean training time per epoch (seconds) on TGBN-Genre, computed over the first five training epochs.
Dataset
Model
Step 2
Step 3
Step 5
MAE
RMSE
Macro-F1
MAE
RMSE
Macro-F1
MAE
RMSE
Macro-F1
Flights
GCN
0.9208
1.5297
0.4486
1.3097
2.0548
0.3970
2.2527
3.2212
0.3584
GAT
0.9080
1.4992
0.4314
1.1381
1.7816
0.3950
1.7601
2.5047
0.3372
GraphSAGE
0.8764
1.4712
0.4667
1.1325
1.8040
0.4418
1.6862
2.4935
0.3982
SGFormer
0.8789
1.4686
0.4735
1.1885
1.8848
0.4207
1.8938
2.7597
0.3861
NodeFormer
0.9694
1.5545
0.4004
1.3759
1.9958
0.3755
2.2803
2.9962
0.3449
Appendix
Table 10: Multi-step prediction results on the Graph-level tasks. Models recursively use their predicted graph states for steps 2 , 3 , and 5 . MAE and RMSE are lower-is-better, whereas Macro-F1 is higher-is-better. The best results are shown in bold and the second best results are underlined .
Figure 6: Sensitivity analysis on Graph-Level Tasks/Flights (Macro-F1) across the rarity weight exponent γ , degree importance exponent θ , volatility importance exponent η , and historical time span tmax .
Figure 7: Transition-conditioned image-to-graph-to-image rollout. The first row shows the ground-truth sequence, and the second row shows the images decoded from the graph states predicted by WorldGraph.
School of Computing Technologies, RMIT University, Australia · Faculty of Information Technology, Monash University, Australia · KStelrix Star Dynamics Lab, KStelrix, China +1