World action models (WAMs) have recently gained increasing attention as a framework for jointly modeling scene evolution and ego actions in autonomous driving. Most existing WAMs learn scene dynamics in pixel space by combining a video-generation backbone for future-observation prediction with an action head for ego-trajectory prediction. Pixels, however, provide only an indirect representation of these dynamics: they entangle geometry and motion with appearance, texture, and illumination, forcing the model to infer three-dimensional transformations from two-dimensional observations. We argue that point-based geometry provides a more natural state space for driving. It explicitly captures spatial structure and both rigid and non-rigid scene dynamics while remaining aligned with the 3D space of driving actions. Building on this insight, we introduce GeoWAM, a visual geometry world action model for autonomous driving. Rather than predicting future images, GeoWAM is pretrained to forecast future scene geometry, yielding representations that jointly encode spatial structure and temporal evolution. A geometry-conditioned action head then leverages these learned geometric dynamics to predict future ego-trajectories. Extensive experiments show that GeoWAM outperforms image-based alternatives, achieving a combined EPDMS of 36.6 on navhard without PDMS supervision and strong zero-shot generalization to nuScenes, with a collision rate of 0.24%. Scaling geometry pretraining with unlabeled data further improves performance, increasing the navhard score by 8.2% to 39.6 and strengthening zero-shot transfer to nuScenes, where the collision rate is reduced by 50% to 0.12%. Together, these results establish geometry as an effective state representation for autonomous driving and geometry pretraining as a general, scalable strategy for downstream planning.
Figures & tables
Figure 1 : Video-based versus geometry-based world modeling. Video world models predict future pixels, from which the 3D transformations required for planning are implicitly included and difficult to recover. Geometry world models instead predict explicit future 3D structure, whose temporal evolution directly exposes object and ego motion in a planning-aligned state space.
Figure 2 : GeoWAM pipeline. (a) The geometry encoder builds a multi-level historical memory Zt from multiview images. (b) Future-geometry queries attend to Zt to produce future geometry tokens, which are then decoded into dense point maps. (c) Future-pose queries attend to the predicted geometry and Zt to produce future pose tokens, which are then decoded into trajectories. (d) Detailed structure of pose/geometry decoder.
Abs Rel ↓
δ<1.25↑
Method/Horizon
1s
2s
3s
4s
mean
1s
2s
3s
4s
mean
Epona [ 44 ] + DVGT [ 51 ]
0.229
0.263
0.292
0.310
0.274
0.732
0.677
0.620
0.589
0.655
Cosmos 3 [ 1 ] + DVGT [ 51 ]
0.300
0.376
0.405
0.422
0.376
0.588
0.513
0.464
0.447
0.503
VGGT-World [ 29 ]
0.272
0.329
0.342
0.357
0.325
0.612
0.553
0.513
0.497
0.544
GeoWAM (Ours) w/o PAI pretraining
0.171
0.190
0.227
0.245
0.208
0.849
0.813
0.777
0.716
0.789
GeoWAM (Ours) w/ PAI pretraining
0.146
0.176
0.208
0.276
0.202
0.873
0.835
0.797
0.704
0.802
Table 1 : Future geometry prediction performance on nuScenes at different horizons. PAI denotes the NVIDIA PhysicalAI-Autonomous-Vehicles dataset [ 26 ] . Bold and underlined values indicate the best and second-best results, respectively.
Method
NC ↑
DAC ↑
DDC ↑
TLC ↑
EP ↑
TTC ↑
LK ↑
HC ↑
EC ↑
EPDMS ↑
Transfuser [ 5 ]
96.9
89.9
97.8
99.7
87.1
95.4
92.7
98.3
87.2
84.0
Hydra-MDP++ [ 20 ]
97.2
97.5
99.4
99.6
83.1
96.5
94.4
98.2
70.9
81.4
DriveSuprim [ 43 ]
97.5
96.5
99.4
99.6
88.4
96.6
95.5
98.3
77.0
83.1
ARTEMIS [ 8 ]
98.3
95.1
98.6
99.8
81.5
97.4
96.5
98.3
–
83.1
DiffusionDrive [ 24 ]
98.2
96.2
99.5
99.8
87.4
97.3
96.9
98.4
87.7
88.2
WoTE [ 22 ]
98.5
96.8
98.8
99.8
86.1
97.9
95.5
98.3
82.9
87.7
Table 2 : Open-loop planning results on the NAVSIM v2 navtest split.
Method
Stage
NC ↑
DAC ↑
DDC ↑
TLC ↑
EP ↑
TTC ↑
LK ↑
HC ↑
EC ↑
EPDMS ↑
DriveVLA-W0 [ 21 ]
S1
96.8
83.3
99.0
99.6
84.6
95.3
96.4
97.6
78.2
24.4
S2
76.8
64.3
79.9
98.3
89.2
75.0
46.8
95.8
53.1
DriveLaW [ 38 ]
S1
97.3
89.1
99.2
99.6
84.3
97.1
96.2
97.8
67.6
30.6
S2
82.5
67.6
83.5
98.1
84.8
78.5
45.8
96.4
57.3
DVGT-2 [ 52 ]
S1
97.2
91.3
98.4
99.8
84.8
95.5
95.5
97.5
71.4
31.7
S2
77.8
73.8
81.3
98.3
91.5
73.2
48.0
83.9
45.1
Table 3 : navhard leaderboard. Methods trained with reinforcement learning or PDMS-score supervision are shown in gray and marked with † . Among the remaining methods, bold and underlined values indicate the best and second-best results, respectively.
Method
nuScenes Finetune
Auxiliary Supervision
L2 (m) ↓
Collision Rate (%) ↓
1s
2s
3s
Avg.
1s
2s
3s
Avg.
ST-P3 [ 13 ]
✓
Map&Box&Depth
1.33
2.11
2.90
2.11
0.23
0.62
1.27
0.71
UniAD [ 14 ]
✓
Map&Box&Motion
0.48
0.96
1.65
1.03
0.05
0.17
0.71
0.31
OccNet [ 28 ]
✓
3D-Occ&Map&Box
1.29
2.13
2.99
2.14
0.21
0.59
1.37
0.72
OccWorld [ 46 ]
✓
3D-Occ
0.52
1.27
2.41
1.40
0.12
0.40
2.08
0.87
VAD-Tiny [ 17 ]
✓
Map&Box&Motion
0.60
1.23
2.06
1.30
0.31
0.53
1.33
0.72
Table 4 : Zero-shot motion planning performance on nuScenes.
PAI Data (hours)
navhard
nuScenes Zero-shot
EPDMS ↑
L2 (m) ↓
Collision Rate (%) ↓
S1
S2
Comb.
1s
2s
3s
Avg.
1s
2s
3s
Avg.
0
80.7
43.5
36.6
0.42
1.19
2.24
1.28
0.00
0.10
0.60
0.24
100
81.3
45.5
37.8
0.31
0.81
1.64
0.92
0.02
0.12
0.44
0.19
496
81.2
45.8
38.2
0.29
0.76
1.55
0.86
0.04
0.12
0.33
0.16
832
83.1
46.7
39.6
0.29
0.79
1.58
0.89
0.02
0.12
0.23
0.12
Table 5 : Effect of geometry-pretraining data scale on downstream planning. Left: absolute planning performance at different pretraining-data scales. Right: relative improvements over the model without geometry pretraining on navhard and nuScenes zero-shot planning.
Variant
Current Geometry
Future Geometry
PAI Pretraining
navtest EPDMS ↑
navhard EPDMS ↑
w/o Future Geometry
✓
89.2
31.7
w/o PAI Pretraining
✓
✓
90.2
36.6
Full
✓
✓
✓
90.7
39.6
Table 6 : Ablation studies on downstream planning.
Figure 3 : Ego-motion recovery from predicted future states. Each row shows current observations, PWM-generated future video, GeoWAM-predicted future geometry, and trajectories recovered from different future representation, compared with the logged ground truth.