cs.CVAug 24, 2026

GeoWAM: Visual Geometry World Action Models for Autonomous Driving

Authors: Yiren Lu, Xin Ye, Jiaming Liu, Philip Jacobson, Jin Yao, Yi-chung Chen, Liam Merino, Dhruva Dixith Kurra, +5 more

Organizations: Uber AV Labs · Case Western Reserve University

Abstract

World action models (WAMs) have recently gained increasing attention as a framework for jointly modeling scene evolution and ego actions in autonomous driving. Most existing WAMs learn scene dynamics in pixel space by combining a video-generation backbone for future-observation prediction with an action head for ego-trajectory prediction. Pixels, however, provide only an indirect representation of these dynamics: they entangle geometry and motion with appearance, texture, and illumination, forcing the model to infer three-dimensional transformations from two-dimensional observations. We argue that point-based geometry provides a more natural state space for driving. It explicitly captures spatial structure and both rigid and non-rigid scene dynamics while remaining aligned with the 3D space of driving actions. Building on this insight, we introduce GeoWAM, a visual geometry world action model for autonomous driving. Rather than predicting future images, GeoWAM is pretrained to forecast future scene geometry, yielding representations that jointly encode spatial structure and temporal evolution. A geometry-conditioned action head then leverages these learned geometric dynamics to predict future ego-trajectories. Extensive experiments show that GeoWAM outperforms image-based alternatives, achieving a combined EPDMS of 36.6 on navhard without PDMS supervision and strong zero-shot generalization to nuScenes, with a collision rate of 0.24%. Scaling geometry pretraining with unlabeled data further improves performance, increasing the navhard score by 8.2% to 39.6 and strengthening zero-shot transfer to nuScenes, where the collision rate is reduced by 50% to 0.12%. Together, these results establish geometry as an effective state representation for autonomous driving and geometry pretraining as a general, scalable strategy for downstream planning.

Figures & tables

Explore similar work

CardsList
  1. GeoWorldAD: Geometry World Action Model for Autonomous Driving

    Jul 20, 2026Songyan Zhang, Jinyuan Tian, Hanbing Li +9Autonomous Driving3D Geometry

  2. PhysWAM: Physically Consistent World Action Model for Autonomous Driving

    Sep 29, 2026Dhruv Parikh, Fengcheng Yu, Quankai Gao +11Autonomous DrivingWorld

  3. SimWAM: A Simple World Action Model for End-to-End Autonomous Driving

    Aug 7, 2026Zongchuang Zhao, Xin Zhou, Tianyang Xu +6Recent World-Action ModelsAutonomous Driving