World-action models (WAMs) jointly predict how a scene will evolve and how an agent should act, however joint generation alone does not necessarily impose a shared geometric constraint on these predictions. We present PhysWAM, a unified world-action model for autonomous driving that co-denoises multiview video, metric depth, and ego motion within a single flow-matching transformer. To ground world and action generation in measured scene geometry, we introduce Coupled Point Projection (CPP) that unprojects the generated depth into 3D points, transforms them using the generated SE(3) ego motion, and minimizes their distance to LiDAR points transformed using the recorded ego motion. This geometric constraint promotes physical consistency with the measured scene by jointly supervising generated depth and motion alongside their standard flow-matching objectives. At inference, trajectory selection relies only on a simple label-free consensus rule, with no learned scorer or simulator feedback. We evaluate PhysWAM across NAVSIM v1 and v2 planning, zero-shot closed-loop transfer, and future video and metric-depth prediction. Despite PhysWAM's simple selection procedure, it achieves strong planning performance and transfers zero-shot to unseen driving environments. It also generates accurate metric depth and temporally coherent video, with CPP improving both planning and depth prediction. Together, these results demonstrate that the geometric relationship between scene depth and ego motion provides a direct way to couple world and action generation within a simple unified model.
Figures & tables
Figure 1: Physical consistency through Coupled Point Projection (CPP). CPP jointly supervises generated depth and ego motion against measured scene geometry. Generated depth is unprojected into 3D and transformed using generated motion (left). LiDAR observations transformed using recorded motion provide the reference in the same coordinate frame (right). The resulting geometric loss supervises both predictions: errors in depth, motion, or both can distort or displace the generated scene relative to the measured reference. Red and black paths show generated and recorded motion, respectively. First and third rows show close alignment with the measured scene; second and fourth rows show larger geometric discrepancies.
Figure 2: Overview of PhysWAM . The Cosmos 3 generator [ Agarwal et al., 2026 ] jointly denoises multiview video, metric depth and ego motion, conditioned on text through its frozen reasoner; RGB and depth share a frozen VAE, and calibrated ray embeddings with mRoPE encode camera geometry and position. Current RGB and motion history stay clean; the rest is noised. Flow matching supervises the velocity outputs, whose clean estimates feed CPP, which transforms generated depth with generated motion and compares it with LiDAR under recorded motion, and the two hinges.
NAVSIM v1
NAVSIM v2
Method
Input
NC
DAC
TTC
C
EP
PDMS
NC
DAC
DDC
TLC
EP
TTC
LK
HC
EC
EPDMS
Human
–
100
100
100
99.9
87.5
94.8
–
End-to-end planners
TransFuser [ Chitta et al., 2022 ]
3 × Cam+L
97.7
92.8
92.8
100.0
79.2
84.0
96.9
89.9
97.8
99.7
87.1
95.4
92.7
98.3
87.2
76.7
LTF [ Chitta et al., 2022 ]
3 × Cam
–
97.6
91.9
99.2
99.8
87.6
97.2
96.6
98.3
86.3
83.6
Hydra-MDP++ (R34) [ Li et al., 2025b ]
3 × Cam
97.6
96.0
93.1
100
80.4
86.6
97.2
97.5
99.4
99.6
83.1
96.5
94.4
98.2
70.9
81.4
Table 1: NAVSIM v1 and v2 on navtest . Numbers of other methods are as reported in their papers. R34: ResNet-34 backbone; +L: LiDAR input. Bold: best per column, oracle row excluded. PhysWAM rows use one sample per scene unless otherwise specified; the sample-to-sample sd is 0.24 PDMS and 0.30 EPDMS.
Easy
Medium
Hard
Extreme
Overall
Method
RC
HD
RC
HD
RC
HD
RC
HD
RC
HD
UniAD [ Hu et al., 2023b ]
58.6
48.7
41.2
29.5
40.4
27.3
26.0
14.3
40.6
28.9
VAD [ Jiang et al., 2023 ]
38.7
24.3
27.0
9.9
25.5
10.4
23.0
8.2
27.9
12.3
LTF [ Chitta et al., 2022 ]
68.4
52.8
40.7
24.6
36.9
19.8
25.5
8.1
41.4
24.8
BeyondDrive [ Wang et al., 2026b ]
76.8
65.6
43.0
31.4
35.5
26.3
29.6
16.2
46.2
34.8
PhysWAM
93.5
86.9
47.4
30.1
38.6
25.2
26.1
13.4
48.9
35.5
Table 2: navhard and HUGSIM. (a) NAVSIM v2 on navhard , two-stage pseudo-simulation with reactive traffic; EPDMS as reported, with the sources of reprinted values and the stage scores in Table 7 . (b) Zero-shot closed-loop driving on HUGSIM, as reported: route completion (RC) and HD-Score by difficulty; BeyondDrive’s overall is the unweighted mean over the four levels. PhysWAM is trained on NAVSIM only, one sample per planner call. Bold: best.
PDMS
EPDMS
Off-road
Collision
Heading
4k, with CPP
80.3
78.8
9.7%
3.8%
2.7 ∘
4k, without CPP
57.0
55.7
32.7%
9.8%
8.9 ∘
navtest PDMS / EPDMS
navhard
HUGSIM RC / HD
final, with CPP
91.4 / 90.3
38.1
48.9 / 35.5
final, without CPP
89.6 / 88.4
36.2
47.5 / 33.4
Table 3: Ablations and future depth. (a) The full recipe with and without CPP after 4,000 updates on navtest (off-road, collision: share of scenes with a zero drivable-area or no-collision term; heading: median error against the recorded trajectory) and at the end of training (Tables 1 – 2 ). (b, d) LoRA setting on navtest , one sample per scene: objective terms added to flow matching, and the front view alone against all three views by driving command, with the full model of Table 1 for reference. (c) Generated depth of the full model against the LiDAR of the same future frame and camera, without scale alignment, four samples per scene; side views in Table 13 ; lower block as reported on nuScenes by GeoWAM [ Lu et al., 2026 ] .
Figure 3: Qualitative examples. Left: input views at t=0 and the command. Middle: generated views and metric depth at +2 and +4 s; the single-view model generates the front view only. Right: rollouts over the drivable area (blue), recorded traffic at 4 s (green), recorded trajectory (dashed); ego boxes at the end point or at the failure step, with the hit agent outlined. Top: with CPP the model generates a parked van a few metres off in its left view and stops short of it; without CPP it generates the van pressed against the ego and hits it at 3.4 s. Bottom: an oncoming car visible only in the left camera; the multi-view model generates it passing close and keeps the turn wide, the single-view model cuts the corner and collides at 2.3 s.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Quantity
Shape
Operation
RGB or encoded depth
(T+1)×Hv×Wv×3
Input to the shared frozen VAE.
xi
J×hv×wv×cℓ
VAE grid, indexed spatially by u .
P(xiσ)
Ni×p2cℓ
Flattened patches of the noisy grid, indexed by q .
zi
Ni×dh
Projected patches with ray, modality, and noise embeddings.
OV(ziout)
Nigen×p2cℓ
Patch velocities at generated positions.
vi
J×hv×wv×cℓ
Unpatchified velocity field; conditioning frames are zero-filled and not predicted.
Appendix
Table 4: Visual representation pathway for one stream i=(v,m) . Shapes omit the minibatch dimension. Noise and flow targets are defined on the VAE grid, before patchification; each spatial patch becomes one token.
View
AbsRel, cells
AbsRel, pixels ≤25 m
Sky error
Front
0.041
0.073
0.05%
Front-left
0.051
0.113
0.22%
Front-right
0.060
0.118
0.20%
Appendix
Table 5: Depth-label quality against LiDAR on 1,000 training windows. AbsRel is the mean absolute relative depth error; the cell-level column uses the geometric-mean LiDAR depth of latent cells with at least three returns, the pixel-level column all LiDAR returns within 25 m. Sky error is the fraction of pixels labelled sky that carry a LiDAR return.
Data
Training windows
103,281 (NAVSIM navtrain )
Frames per window
1 current + 8 future, 2 Hz
Views and resolution
front 832×468 ; front-left and front-right 416×234
Model
Initialization
Cosmos 3 Nano (15.2B parameters)
Trained parameters
generation pathway, 7.0B
Appendix
Table 6: Hyperparameters of PhysWAM .
Method
Difficulty
NC
DAC
TTC
COM
RC
HD-Score
UniAD [ Hu et al., 2023b ]
Easy
77.4
88.5
70.8
82.8
58.6
48.7
Medium
72.4
86.8
60.4
72.0
41.2
29.5
Hard
66.5
86.0
55.0
67.2
40.4
27.3
Extreme
54.4
89.6
42.9
58.5
26.0
14.3
VAD [ Jiang et al., 2023 ]
Easy
66.1
73.9
58.2
100.0
38.7
24.3
Medium
45.7
79.8
29.0
100.0
27.0
9.9
Appendix
Table 8: HUGSIM by difficulty and metric ( ×100 ); baselines as reported by the benchmark paper. Overall follows the benchmark’s reduction: per-dataset episode means, averaged with equal weight over the four source datasets within each difficulty, then weighted over the difficulties by their episode counts (80, 157, 96, 103). Bold: best per metric and difficulty among UniAD, VAD, LTF, and PhysWAM .
Dataset
Episodes
Easy
Medium
Hard
Extreme
All
nuScenes
88
94.5 / 88.4
43.7 / 27.1
59.5 / 43.1
40.6 / 27.3
56.7 / 42.9
Waymo
108
94.6 / 91.6
51.5 / 34.1
57.3 / 44.8
23.9 / 8.7
55.8 / 42.8
KITTI-360
113
87.4 / 76.5
51.6 / 39.3
14.9 / 6.1
8.0 / 0.7
41.1 / 31.0
PandaSet
127
97.4 / 91.1
42.8 / 20.1
22.5 / 6.6
32.0 / 16.9
41.3 / 24.3
Appendix
Table 9: HUGSIM by source dataset , PhysWAM : RC / HD-Score per difficulty level and the plain episode mean per dataset. The Overall values of Table 8 average these per-dataset means with equal dataset weight within each difficulty before weighting by episodes, so they are not the pooled episode mean.
Adaptation
LoRA, rank 128, α=256 , on the attention projections
Trained parameters
122.7M
Updates × batch
16,000 × 44 windows
Optimizer
AdamW, β=(0.9,0.99) , weight decay 0.05, clip 1.0
Learning rate
peak 2×10−4 , 300 warm-up updates, linear decay to 0
Weight averaging
none
Appendix
Table 10: Hyperparameters of the LoRA setting. Adapters act on the generation pathway; rows not listed are as in Table 6 .
Row
NC
DAC
DDC
TLC
EP
TTC
LK
HC
PDMS
EPDMS
flow
95.9
96.4
98.6
99.7
86.0
94.9
95.2
98.4
85.7
84.3
flow + hinge
96.8
96.6
98.5
99.8
86.8
95.9
95.3
98.6
87.2
85.7
flow + CPP
96.8
97.0
98.5
99.7
87.0
95.7
95.1
98.6
87.8
86.2
flow + hinge + CPP (multi-view)
96.9
97.2
98.5
99.7
87.2
95.8
95.1
98.6
88.2
86.5
single-view
96.6
96.7
98.6
99.7
86.8
95.6
95.2
98.5
87.1
85.6
multi-view, no depth
96.9
96.8
98.5
99.7
87.0
95.9
95.2
98.6
87.5
85.9
Appendix
Table 11: LoRA setting, all sub-scores on navtest , one sample per scene.
PDMS
EPDMS
Row
Left
Straight
Right
Left
Straight
Right
flow
81.4
87.8
81.9
80.3
86.2
81.3
flow + hinge
83.4
89.0
83.8
82.4
87.2
82.8
flow + CPP
83.5
89.9
83.9
82.1
88.1
83.0
flow + hinge + CPP (multi-view)
83.9
90.3
84.3
82.4
88.4
83.3
single-view
81.2
90.0
82.0
80.1
88.2
81.4
Appendix
Table 12: Scores by driving command on navtest (2,501 left, 8,070 straight, 1,575 right scenes), one sample per scene.
+2 s
+4 s
AbsRel ↓
δ1.25↑
AbsRel ↓
δ1.25↑
View
with
without
with
without
with
without
with
without
Front
0.175
0.193
0.814
0.802
0.232
0.253
0.742
0.743
Front-left
0.292
0.315
0.725
0.706
0.374
0.404
0.654
0.649
Front-right
0.240
0.277
0.752
0.725
0.305
0.346
0.676
0.663
Front, cells
0.144
0.154
0.856
0.844
0.194
0.205
0.795
0.795
Appendix
Table 13: Generated depth with and without CPP for the final models, on navtest : AbsRel and δ1.25 against the LiDAR of the same future frame and camera, without scale alignment, as in Table 3 a; means over four samples per scene. “Cells”: the 16-pixel cells that CPP supervises.
Clips
FVD ↓
FID ↓
DrivingGPT
512
142.6
12.8
PWM †
–
86.0
–
DriveDreamer-Policy
–
53.6
–
CoWorld-VLA
–
32.7
–
Recorded (floor)
600
91.5 ± 3.2
11.0 ± 0.3
PhysWAM
600
111.3 ± 6.8
15.4 ± 0.6
Appendix
Table 14: Front-view video quality on NAVSIM. I3D FVD over one context and eight generated frames at 2 Hz, FID over the same frames. 600 clips: generated clips of half the scenes against the recorded clips of the other half, ten splits, floor from the two recorded halves; 1,200 clips: generated against recorded clips of the same scenes. Reported rows use their own clip counts and lengths (DrivingGPT [ Chen et al., 2025b ] 12 frames; PWM [ Zhao et al., 2026a ] 10, † as reported by DriveDreamer-Policy [ Zhou et al., 2026a ] , itself 9; CoWorld-VLA [ Huang et al., 2026 ] unstated).
Quantity
Floor
PhysWAM
Excess
Yaw, video vs. motion, 4 s ( ∘ )
0.29
0.80
+0.44 [+0.39, +0.48]
turns above 45∘
6.11
–
+2.31 [+1.93, +2.86]
Cross-view ∣Δlogz∣ , +2 s, L / R
0.032 / 0.030
0.045 / 0.049
1.5 ×
Temporal ∣Δlogz∣ , 0 to 4 s
0.087
0.285
3.3 ×
Appendix
Table 15: Consistency of the generated future with its own motion, across views, and over time, on navtest ; medians. Yaw: camera rotation recovered from the generated video [ Keetha et al., 2026 ] against the generated motion, and, for the floor, from the recorded clip against the recorded trajectory; the excess is the median of the per-scene paired differences, with a 95% bootstrap interval. Depth: median ∣Δlogz∣ after warping the front depth into the side views with the recorded extrinsics, or frame 0 into frame 8 with the generated motion; floors from two LiDAR sweeps 0.5 s apart and from recorded LiDAR under recorded motion.
UniPC steps
Δ PDMS
Samples
GPU-s / scene
4
−0.6
1
3.5
8
+0.7
3
4.4
15
−0.4
1
6.0
30
0
4
9.4
Appendix
Table 16: Sampler steps and cost for the model of Tables 1 – 2 : PDMS relative to the 30-step protocol row, one sample per scene, averaged over the number of samples given (sample-to-sample sd 0.2–0.3 where repeated), and the cost of one plan on one RTX PRO 6000 GPU.
Figure 4: Additional qualitative examples. On a straight road, the model with CPP generates the overtaking bus alongside in its left view. The model trained without CPP omits it and is clipped by the overtaking car at 3.9 s. At a hotel drop-off, it generates the departing minivan too close and clips the island at 1.6 s. The single-view model swings wide over a grass median that only the right camera shows and leaves the road at 3.1 s. At a left turn, it never generates the waiting truck and hits it at 3.7 s. The full model renders a stopped minivan at a constant size, plans as if it were pulling away, and rear-ends it at 3.7 s, a world-model error. At a tight left turn, both models generate indistinguishable views, but the full model clips the inner curb by 0.4 m at 3.6 s, a plan-precision error.
Representative works
Predicted world
Connection to action
LAW, SimWAM [ Li et al., 2025c , Zhao et al., 2026b ]
Future features or video
Future prediction provides training supervision; planning does not require explicit future generation.
DriveWAM, GeoWAM [ Shi et al., 2026 , Lu et al., 2026 ]
Future video or 3D geometry
Representations of the predicted future condition trajectory generation.
Epona, DriveVA [ Zhang et al., 2025b , Liu et al., 2026b ]
Future video
Video and action generation share learned representations.
4D-WAM [ Fu et al., 2026 ]
Future video
Joint video–action learning is supervised by matching geometry recovered from generated and recorded video.
WoTE, DA-WAM [ Li et al., 2025d , Zhong et al., 2026 ]
Future scene features
A learned scorer evaluates candidate trajectories using their predicted future states.
PhysWAM (ours)
Multiview video and metric depth
Video, depth, and ego motion are co-denoised. CPP jointly supervises generated depth and motion against measured scene geometry. A label-free consensus rule selects among trajectories.
Appendix
Table 17: World modeling and its role in planning. Representative works illustrate how future prediction supports policy learning, action generation, or trajectory selection. These roles can overlap within a method.
In autonomous driving, World-Action Models (WAMs) have improved end-to-end planning by transferring video dynamics priors to action prediction, but many still couple planning with future-video generation at inference, incurring substantial computational overhead. We present SimWAM, a simple yet effective WAM that leverages future-video prediction solely as a training-time supervision signal. It co-trains a pretrained video expert and a lightweight action expert with joint flow matching. An isolated attention mask keeps action prediction independent of future frames, allowing trajectory prediction without future-frame generation at inference. This design supports multiple pretrained video backbones and independent action-expert scaling within a shared attention interface, while preserving the joint learning objective. Moreover, we apply reinforcement learning to optimize a compositional driving reward beyond trajectory imitation. Experiments show that SimWAM achieves 91.9 PDMS on NAVSIM with a favorable trade-off between accuracy and latency among world-model-based planners, while transferring zero-shot to nuScenes. It also achieves competitive planning accuracy on WOD-E2E and PhysicalAI-Autonomous-Vehicles. These results position SimWAM as a plain yet solid baseline for efficient autonomous driving. The code and model weights are available at https://github.com/H-EmbodVis/SimWAM/.
Zongchuang Zhao, Xin Zhou, Tianyang Xu +6
Huazhong University of Science & Technology · Dongfeng Research & Development Institute
World action models (WAMs) have recently gained increasing attention as a framework for jointly modeling scene evolution and ego actions in autonomous driving. Most existing WAMs learn scene dynamics in pixel space by combining a video-generation backbone for future-observation prediction with an action head for ego-trajectory prediction. Pixels, however, provide only an indirect representation of these dynamics: they entangle geometry and motion with appearance, texture, and illumination, forcing the model to infer three-dimensional transformations from two-dimensional observations. We argue that point-based geometry provides a more natural state space for driving. It explicitly captures spatial structure and both rigid and non-rigid scene dynamics while remaining aligned with the 3D space of driving actions. Building on this insight, we introduce GeoWAM, a visual geometry world action model for autonomous driving. Rather than predicting future images, GeoWAM is pretrained to forecast future scene geometry, yielding representations that jointly encode spatial structure and temporal evolution. A geometry-conditioned action head then leverages these learned geometric dynamics to predict future ego-trajectories. Extensive experiments show that GeoWAM outperforms image-based alternatives, achieving a combined EPDMS of 36.6 on navhard without PDMS supervision and strong zero-shot generalization to nuScenes, with a collision rate of 0.24%. Scaling geometry pretraining with unlabeled data further improves performance, increasing the navhard score by 8.2% to 39.6 and strengthening zero-shot transfer to nuScenes, where the collision rate is reduced by 50% to 0.12%. Together, these results establish geometry as an effective state representation for autonomous driving and geometry pretraining as a general, scalable strategy for downstream planning.
Pretrained foundation models have become an important basis for end-to-end autonomous driving. In contrast to vision-language models pretrained primarily on static image-text pairs, video generative models capture temporal dynamics and motion priors that are naturally suited for driving. We present DriveWAM, a driving world-action model that adapts a pretrained video diffusion transformer into an autoregressive video-action policy. DriveWAM organizes video and action streams into a unified temporal token sequence and trains them under a joint flow-matching objective, preserving the pretrained video-generation architecture while adapting its large-scale video priors to action generation. To incorporate high-level scene understanding, we introduce scene-evolving driving guidance, where a frozen VLM produces chunk-specific semantic intent to guide video-action generation. To keep long-horizon rollout bounded, we further introduce selective KV memory, which maintains bounded modality-aware video and action memory pools through relevance-redundancy cache selection at inference time. Experiments on NAVSIM and the PhysicalAI-Autonomous-Vehicles benchmark show that DriveWAM achieves strong planning performance, and a data-scaling study from 4k to 100k driving clips further confirms the scaling potential of world-action modeling for end-to-end autonomous driving.
Chen Shi, Jinrui Xu, Shaoshuai Shi +3
The Chinese University of Hong Kong, Shenzhen · Voyager Research, Didi Chuxing