VGGT provides a strong foundation for geometry-centric world models by recovering unified 3D scene geometry from visual observations. Although recent extensions enable temporal 3D prediction, their future evolution remains weakly conditioned on driving intentions and actions, limiting their ability to model alternative action-dependent futures. We propose VGGTWorld-VLA, an intention-conditioned extension of VGGT-World for controllable 3D world evolution in autonomous driving. First, we introduce an action--semantic conditioning mechanism that injects complementary driving semantics and ego-motion representations into the future-token stream, enabling different future geometry predictions for the same observed scene under alternative ego actions. Second, we develop a geometry--language--action bridge that adapts historical geometry, VLA semantic features, and maneuver and trajectory representations for joint conditioning of future geometry prediction. We evaluate future geometry prediction on NAVSIM, while conditioning ablations further examine the contributions of semantic and action information. Compared with the baseline, our method demonstrates competitive geometry prediction performance. Ablation studies further support the effectiveness of semantic and action conditioning. These results demonstrate the potential of semantic and action conditioning for controllable VGGT-based world prediction in autonomous driving.
Figures & tables
Figure 1: Motivation for intention-conditioned 3D world prediction. Geometric reconstruction captures the observed scene, while geometry-based world models extend this capability to future prediction, primarily from historical observations. Without explicit intention and action conditioning, the forecast may depict straight driving despite an intended left turn. Our approach incorporates driving semantics and ego actions to predict future geometry consistent with the intended maneuver.
Figure 2: Overall architecture of VGGTWorld-VLA. The geometric branch encodes historical frames from a single camera stream, while the semantic branch extracts VLA features from multi-camera observations and a prompt at the conditioning timestamp. A given ego trajectory and its maneuver descriptor are encoded as action conditions. Action and semantic tokens condition the future stream of the geometry world model, whose predictions are decoded into future depth and 3D structure.
Figure 3: Action–semantic 3D world evolution. Each of the eight dual-stream blocks is followed by action and semantic cross-attention that sequentially update the future stream, with layer-specific parameters. The historical and conditioned future streams are then jointly processed by eight single-stream blocks to predict future geometry.
Figure 4: Qualitative comparison of future depth. Top: reference depth estimated by VGGT from the recorded future frames. Bottom: our predictions conditioned on historical observations, ego actions, and driving semantics.
Method
Token MSE ↓
Cosine ↑
AbsRel ↓
RMSE ↓
δ1↑
VGGT-World
0.011065
0.769799
0.385980
0.426050
0.418000
VGGT-World-N
0.010620
0.781621
0.376136
0.436463
0.348113
Ours
0.010459
0.785120
0.392457
0.367831
0.459124
Table 1: Future geometry prediction on the NAVSIM 3,000-clip motion subset. VGGT-World-N denotes visual-only continuation training on NAVSIM. All methods use stride 1 and 10 flow sampling steps, with depth evaluated against VGGT-derived pseudo-ground truth. The best and second-best results are highlighted in bold and underlined, respectively.
Encoding
Token MSE ↓
AbsRel ↓
δ1↑
RMSE ↓
Point EPE ↓
Point Rel. EPE ↓
One-hot
0.01019
0.42476
0.45279
0.37708
0.36640
0.38590
Trajectory
0.01068
0.40515
0.40943
0.39119
0.37805
0.38657
Trajectory + velocity
0.01062
0.40097
0.41798
0.38517
0.37107
0.38471
Joint
0.01050
0.39320
0.45184
0.37093
0.35807
0.36842
Table 2: Comparison of action-representation configurations on the 3,000-clip motion subset. Joint combines one-hot, position-based, and velocity-augmented representations. The best and second-best results are highlighted in bold and underlined, respectively.
Action input
Token MSE ↓
AbsRel ↓
RMSE ↓
δ1↑
Point EPE ↓
Correct
0.010459
0.392457
0.367831
0.459124
0.354330
Zero
0.010695
0.414170
0.375252
0.459097
0.356336
Shuffled
0.010738
0.409405
0.384130
0.434404
0.377294
Wrong
0.010913
0.419565
0.396635
0.412820
0.391264
Table 3: Action-input interventions on the 3,000-clip motion subset. The observed history, semantic inputs, and initial sampling noise are fixed while the action input changes. The best and second-best results are highlighted in bold and underlined, respectively.
Figure 5: Action-conditioned future depth predictions under left-turn, straight-driving, and right-turn inputs. Red boxes highlight corresponding image regions with action-dependent differences in predicted geometry.
Conditioning
Token MSE ↓
AbsRel ↓
δ1↑
RMSE ↓
Point EPE ↓
Point Rel. EPE ↓
Joint
0.010496
0.393195
0.451837
0.370929
0.358067
0.368416
Joint + VLM
0.010459
0.392457
0.459124
0.367831
0.354330
0.362280
Table 4: Comparison of Joint action conditioning and Joint + VLM.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Configuration
Geometry input / prediction
Two historical / two future frames
VGGT representation width
1,024
Flow-transformer width / attention heads
512 / 8
Dual-stream / single-stream blocks
8 / 8
Condition injection
Action CA, then semantic CA; per-layer parameters
Continuous trajectory
8 ego-centric waypoints
Appendix
Table A1: Representation and inference configuration. The action-token count below refers to the continuous trajectory component of the action interface.
History 1
History 2
Future 1
Future 2
Sample index
k−s
k
k+s
k+2s
Nominal offset, s=1
−0.5 s
0 s
+0.5 s
+1.0 s
Historical encoder input
Yes
Yes
—
—
Full-sequence teacher input
Yes
Yes
Yes
Yes
Future target selected
—
—
Yes
Yes
Appendix
Table A2: Temporal roles within a four-frame clip. Future RGB is used to construct references; the world predictor receives historical geometry and conditions from the conditioning sample.
Metric
Joint
Joint + VLM
Token cosine ↑
0.784190
0.785120
Near-field Token MSE ↓
0.009563
0.009535
Appendix
Table A3: Additional latent measurements for the Joint and Joint + VLM configurations. The near-field measurement follows the reported evaluation mask. Bold indicates the better value within each row.
Figure A1: From observed geometry to predicted future geometry. The top row shows the recorded RGB sequence. Below, the left column shows geometry from the last observed frame, the middle column shows our future predictions, and the right column shows the corresponding future reference obtained by applying VGGT to the recorded sequence. Rows display token PCA and depth for Future 1, followed by token PCA and depth for Future 2. The last observed geometry is reused as the historical reference at both horizons. This arrangement shows how the predicted representation evolves from the observed scene toward its recorded future. Original image content and color mappings are preserved.
Figure A2: Action-conditioned depth and token visualization in a curved-road scene. The first row shows the recorded four-frame sequence. Subsequent rows show first-future depth, second-future depth, and second-future token PCA. Columns are ordered as the VGGT reference, left, straight, right, and stop. The reference corresponds to the recorded left-turn sequence; the remaining columns show predictions under the labeled conditions with shared history and Euler noise. The reference is repeated conceptually across conditions as an anchor for the recorded sequence, rather than a separate observed future for each alternative action.
Figure A3: Building-facade example with a recorded left turn. The RGB strip provides temporal context; the depth rows compare the VGGT reference with left-, straight-, and right-conditioned predictions at the first and second future frames. The facade and roadside region provide spatial anchors for inspecting changes in projected structural boundaries. History and Euler noise are shared across the action-conditioned predictions.
Figure A4: Intersection example with a recorded right turn. The right-conditioned second-future prediction follows the prominent central roadside structure visible in the reference, while alternative conditions produce different spatial depth patterns. All RGB panels are recorded frames. Reference depth is obtained from the recorded sequence, and predictions share the historical input and Euler noise.
Autonomous driving requires both safe and efficient planning decisions in dynamic 3D environments. Although recent Vision/Video-Action models learn policies directly from visual observations and scale well with advances in vision transformers and large-scale training data, they often lack explicit geometric grounding and future-aware spatial guidance, limiting their ability to balance collision avoidance and driving progress. In this work, we propose GeoWorldAD, a geometry world action model that grounds trajectory planning in ego-aligned 3D space and anticipates short-horizon scene evolution with latent future geometry tokens. Present geometry provides essential spatial constraints for safe planning, while future geometry reveals how surrounding agents and ego-centric free space may evolve, reducing overly conservative decisions without sacrificing safety. To efficiently exploit these geometric cues, GeoWorldAD progressively aggregates multi-scale present geometry and latent future geometry through iterative trajectory refinement. Experiments on NAVSIM v1 and v2 demonstrate state-of-the-art performance, highlighting the effectiveness of explicit 3D geometry grounding and future geometry world modeling for safe and efficient autonomous driving.
Songyan Zhang, Jinyuan Tian, Hanbing Li +9
Nanyang Technological University · Xiaomi EV · Zhejiang University
World-action models have recently improved autonomous driving by jointly learning future scene prediction and trajectory generation. Most existing approaches model the future primarily through RGB appearance, and recent works have begun to incorporate geometric prediction to improve spatial understanding. However, dense geometry describes the spatial layout of the entire scene without indicating which parts are most relevant to the ego vehicle's action. For driving, the model must also identify and anticipate where it can safely move and which regions may pose collision risks. Jointly modeling action-relevant regions and future geometry can provide the policy with both driving-relevant cues and their corresponding spatial structure. We therefore propose AffordDrive3D, an affordance- and geometry-aware world-action model that jointly learns future action-relevant regions and spatial structure. In order to capture the scene semantics and driving context needed for driving affordance prediction, we build AffordDrive3D on a VLM backbone to forecast drivable areas and collision-critical regions that directly affect ego motion, while predicting future geometry from RGB world-model latents. On NAVSIM, AffordDrive3D achieves state-of-the-art performance with 91.3 PDMS and 89.9 EPDMS, demonstrating the effectiveness of jointly modeling future affordances and geometry for trajectory planning.
Tianhui Cai, Xinglong Sun, Chao Fang +6
University of California, Los Angeles · NVIDIA · 342dot by Hyundai +1
Vision-Language-Action (VLA) models have advanced end-to-end autonomous driving. However, existing methods either lack comprehensive world cognition or suffer from fragmented world foresight, inherently confining these models to reactive driving. To address this limitation, we propose WCog-VLA, a novel dual-level World-Cognitive VLA framework that successfully bridges semantic world forecasting with generative world evolution to achieve proactive autonomous driving. At the semantic level, WCog-VLA unifies world cognition and reasoning by incorporating 3D spatial perception and injecting agent tokens to capture the world dynamics, while concurrently enabling Game-theoretic Chain-of-Thought (Game-CoT) reasoning. At the generative level, we introduce the Aligned Decoupled Diffusion Transformer (ADDT) as a powerful generative world model that synthesizes physically-plausible joint multi-agent trajectories. Through scene representation alignment, ADDT reduces the number of denoising steps required and thus significantly accelerates inference. To facilitate strategic reasoning, we further construct a large-scale dataset featuring 85k Game-CoT annotations. Extensive experiments on the NAVSIM benchmark demonstrate that WCog-VLA achieves a State-Of-The-Art (SOTA) PDMS score of 92.9.
Xuerun Yan, Zhexi Lian, Nuoheng Zhang +5
Tongji University, China · Nanyang Technological University, Singapore