Modeling 3D scene geometry and its evolution over time is essential for autonomous driving and robotics. A common paradigm is to use world models to predict future images or latent representations of the environment and subsequently recover geometry from these predictions. However, this paradigm does not explicitly model geometric structure and typically relies on recursive rollouts to reach longer prediction horizons, leading to error accumulation and increasing computational cost. To address these limitations, we present GeoWM, a geometry world model that directly forecasts future scene geometry at specified future horizons without recursive rollout. The key idea is to leverage a geometry foundation model to transform observed RGB frames into a geometric history, which conditions a flow-matching transformer to predict the scene geometry at a specified future horizon. We further show that a lightweight camera-motion predictor can accurately estimate the future viewpoint, and that projecting the observed geometry into the predicted viewpoint provides an effective geometric prior for future geometry forecasting. Extensive experiments on four datasets spanning urban driving, aerial flight, and dynamic manipulation demonstrate that GeoWM outperforms the evaluated world models in forecasting depth, camera pose, and 3D scene geometry, while substantially reducing inference time at longer horizons.
Figures & tables
Fig. 1 : Pixels, features, or explicit geometry? Video models forecast future frames and then reconstruct geometry; their photometric objective couples only loosely to 3D structure, and dynamic objects can be mislocalised [ 7 ] . Feature-space models forecast latents and decode them, where decoding can amplify errors [ 7 ] and autoregressive rollout accumulates them [ 8 , 9 , 10 ] . GeoWM forecasts explicit geometry directly and queries each horizon independently.
Fig. 2 : Overview of GeoWM : (a) forecasting in explicit geometry, (b) the flow-matching transformer, and (c) independent horizon queries.
Method
Next
Short
Mid
Long
All
Pose (trajectory-level)
AbsRel ↓
δ1↑
CD ↓
AbsRel ↓
δ1↑
CD ↓
AbsRel ↓
δ1↑
CD ↓
AbsRel ↓
δ1↑
CD ↓
AbsRel ↓
δ1↑
CD ↓
ATE ↓
RTE ↓
RRE ↓
KITTI next 0.21 s, short 0.41–1.03 s, mid 1.24–2.07 s, long 2.27–4.14 s
Copy-last
9.1
90.1
0.754
13.9
82.6
0.900
18.9
75.7
1.152
22.5
71.8
1.406
19.2
75.8
1.209
6.82
1.33
0.83
DINO-Foresight
15.0
81.0
–
16.1
79.1
–
19.2
72.1
–
22.6
64.1
–
20.0
69.9
–
–
–
–
VGGT-World
7.8
92.4
0.737
11.2
87.8
0.796
21.4
73.9
1.263
29.1
57.6
2.123
22.5
69.5
1.574
3.43
1.39
1.50
Cosmos-3
6.8
93.8
0.676
9.3
90.5
0.735
15.1
83.5
0.957
21.2
75.5
1.281
16.6
81.4
1.060
0.34
0.22
0.67
TABLE I : Future geometry forecasting on four datasets. AbsRel ( ×100 ) and δ1 (%) score the forecast depth, CD (m) the forecast 3D geometry after similarity alignment, and ATE (m), RTE (m), RRE (deg) the forecast camera trajectory. Next is the first predicted horizon; Short , Mid , and Long average the horizons given in each block, and All averages every horizon. Depth is median-scaled per image except on DOMINO, which is evaluated at true scale. Best and second best among copy-last, VGGT-World [ 7 ] , Cosmos-3 [ 6 ] , and GeoWM ; copy-last repeats the last observed geometry.
Method
Next
Short
Mid
Long
All
AbsRel ↓
δ1↑
CD ↓
AbsRel ↓
δ1↑
CD ↓
AbsRel ↓
δ1↑
CD ↓
AbsRel ↓
δ1↑
CD ↓
AbsRel ↓
δ1↑
CD ↓
DOMINO, novel task
Copy-last
14.1
85.4
0.032
16.0
83.1
0.039
22.6
77.6
0.055
30.9
71.6
0.075
21.9
78.6
0.053
VGGT-World
16.0
84.1
0.040
18.7
81.1
0.050
27.4
73.6
0.071
32.3
65.9
0.079
24.7
75.0
0.063
Cosmos-3
15.1
85.9
0.038
16.6
84.1
0.044
22.3
78.8
0.059
27.9
73.3
0.075
21.2
79.7
0.057
GeoWM
5.1
93.7
0.017
7.4
90.3
0.026
14.7
81.1
0.048
23.4
71.8
0.071
13.7
82.9
0.044
TABLE II : Generalisation on DOMINO to unseen tasks and robots. Same metrics, horizon groups, and ranking as the DOMINO block of Table I , which reports the seen tasks and robots.
Variant
Short
Long
All
AbsRel ↓
δ1↑
AbsRel ↓
δ1↑
AbsRel ↓
δ1↑
Copy-last
13.9
82.6
22.5
71.8
19.2
75.8
Prediction space
feature space
11.7
86.6
21.4
72.8
17.6
78.1
feature space, + depth loss
11.8
86.6
18.8
73.7
16.0
78.8
Geometric conditioning
TABLE III : Ablations on KITTI. Variants of GeoWM , evaluated on the windows and with the protocol of Table I : median-scaled AbsRel ( ×100 ) and δ1 (%), averaged over the KITTI horizon groups of that table, where All averages every horizon. N is the number of noise draws averaged at read-out.
Fig. 3 : Direct forecasting and geometric conditioning. (a) Direct prediction versus one-step recursive rollout with the same checkpoint. (b) Increase in horizon-averaged AbsRel when scaling the anchor’s predicted translation relative to α=1 ; rotation is unchanged. Comparisons use matched window subsets and equal-drive averaging on KITTI.
Fig. 4 : Qualitative forecasts. Short/long horizons are h=1/10 for TartanAir and h=4/48 for DOMINO. Error maps show signed relative depth error (white is zero); boxes mark corresponding scene regions. Clouds use calibrated intrinsics, median-scaled TartanAir depths and fixed-scale DOMINO depths; colors indicate depth. All clouds share the same view without further alignment.
Method
TFLOPs
Params
Time to forecast (s)
forecaster
backbone
total
h=1
h=5
h=10
h=20
DINO-Foresight [ 5 ]
30.6
256 M
DINO
361 M
0.03
0.11
0.22
0.42
VGGT-World [ 7 ]
1116
433 M
VGGT
1.59 B
1.9
7.7
17.4
36.9
Cosmos-3 [ 6 ]
3362
17.0 B
VGGT
18.1 B
4.9
5.8
6.7
9.3
GeoWM
12.0
206 M
VGGT
1.36 B
0.11
0.11
0.11
0.11
TABLE IV : Inference cost. Compute, parameters and the time to obtain the forecasts of different horizons.
World action models (WAMs) have recently gained increasing attention as a framework for jointly modeling scene evolution and ego actions in autonomous driving. Most existing WAMs learn scene dynamics in pixel space by combining a video-generation backbone for future-observation prediction with an action head for ego-trajectory prediction. Pixels, however, provide only an indirect representation of these dynamics: they entangle geometry and motion with appearance, texture, and illumination, forcing the model to infer three-dimensional transformations from two-dimensional observations. We argue that point-based geometry provides a more natural state space for driving. It explicitly captures spatial structure and both rigid and non-rigid scene dynamics while remaining aligned with the 3D space of driving actions. Building on this insight, we introduce GeoWAM, a visual geometry world action model for autonomous driving. Rather than predicting future images, GeoWAM is pretrained to forecast future scene geometry, yielding representations that jointly encode spatial structure and temporal evolution. A geometry-conditioned action head then leverages these learned geometric dynamics to predict future ego-trajectories. Extensive experiments show that GeoWAM outperforms image-based alternatives, achieving a combined EPDMS of 36.6 on navhard without PDMS supervision and strong zero-shot generalization to nuScenes, with a collision rate of 0.24%. Scaling geometry pretraining with unlabeled data further improves performance, increasing the navhard score by 8.2% to 39.6 and strengthening zero-shot transfer to nuScenes, where the collision rate is reduced by 50% to 0.12%. Together, these results establish geometry as an effective state representation for autonomous driving and geometry pretraining as a general, scalable strategy for downstream planning.
Driving world models serve as a pivotal technology for autonomous driving by simulating environmental dynamics. However, existing approaches predominantly focus on future scene generation, often overlooking comprehensive 3D scene understanding. Conversely, while Large Language Models (LLMs) demonstrate impressive reasoning capabilities, they lack the capacity to predict future geometric evolution, creating a significant disparity between semantic interpretation and physical simulation. To bridge this gap, we propose HERMES++, a unified driving world model that integrates 3D scene understanding and future geometry prediction within a single framework. Our approach addresses the distinct requirements of these tasks through synergistic designs. First, a BEV representation consolidates multi-view spatial information into a structure compatible with LLMs. Second, we introduce LLM-enhanced world queries to facilitate knowledge transfer from the understanding branch. Third, a Current-to-Future Link is designed to bridge the temporal gap, conditioning geometric evolution on semantic context. Finally, to enforce structural integrity, we employ a Joint Geometric Optimization strategy that integrates explicit geometric constraints with implicit latent regularization to align internal representations with geometry-aware priors. Extensive evaluations on multiple benchmarks validate the effectiveness of our method. HERMES++ achieves strong performance, outperforming specialist approaches in both future point cloud prediction and 3D scene understanding tasks. The model and code will be publicly released at https://github.com/H-EmbodVis/HERMESV2.
Xin Zhou, Dingkang Liang, Xiwu Chen +4
Huazhong University of Science and Technology · Mach Drive · University of Hong Kong
Autonomous driving requires both safe and efficient planning decisions in dynamic 3D environments. Although recent Vision/Video-Action models learn policies directly from visual observations and scale well with advances in vision transformers and large-scale training data, they often lack explicit geometric grounding and future-aware spatial guidance, limiting their ability to balance collision avoidance and driving progress. In this work, we propose GeoWorldAD, a geometry world action model that grounds trajectory planning in ego-aligned 3D space and anticipates short-horizon scene evolution with latent future geometry tokens. Present geometry provides essential spatial constraints for safe planning, while future geometry reveals how surrounding agents and ego-centric free space may evolve, reducing overly conservative decisions without sacrificing safety. To efficiently exploit these geometric cues, GeoWorldAD progressively aggregates multi-scale present geometry and latent future geometry through iterative trajectory refinement. Experiments on NAVSIM v1 and v2 demonstrate state-of-the-art performance, highlighting the effectiveness of explicit 3D geometry grounding and future geometry world modeling for safe and efficient autonomous driving.
Songyan Zhang, Jinyuan Tian, Hanbing Li +9
Nanyang Technological University · Xiaomi EV · Zhejiang University