World-action models (WAMs) jointly predict how a scene will evolve and how an agent should act, however joint generation alone does not necessarily impose a shared geometric constraint on these predictions. We present PhysWAM, a unified world-action model for autonomous driving that co-denoises multiview video, metric depth, and ego motion within a single flow-matching transformer. To ground world and action generation in measured scene geometry, we introduce Coupled Point Projection (CPP) that unprojects the generated depth into 3D points, transforms them using the generated SE(3) ego motion, and minimizes their distance to LiDAR points transformed using the recorded ego motion. This geometric constraint promotes physical consistency with the measured scene by jointly supervising generated depth and motion alongside their standard flow-matching objectives. At inference, trajectory selection relies only on a simple label-free consensus rule, with no learned scorer or simulator feedback. We evaluate PhysWAM across NAVSIM v1 and v2 planning, zero-shot closed-loop transfer, and future video and metric-depth prediction. Despite PhysWAM's simple selection procedure, it achieves strong planning performance and transfers zero-shot to unseen driving environments. It also generates accurate metric depth and temporally coherent video, with CPP improving both planning and depth prediction. Together, these results demonstrate that the geometric relationship between scene depth and ego motion provides a direct way to couple world and action generation within a simple unified model.
Figures & tables
Figure 1: Physical consistency through Coupled Point Projection (CPP). CPP jointly supervises generated depth and ego motion against measured scene geometry. Generated depth is unprojected into 3D and transformed using generated motion (left). LiDAR observations transformed using recorded motion provide the reference in the same coordinate frame (right). The resulting geometric loss supervises both predictions: errors in depth, motion, or both can distort or displace the generated scene relative to the measured reference. Red and black paths show generated and recorded motion, respectively. First and third rows show close alignment with the measured scene; second and fourth rows show larger geometric discrepancies.
Figure 2: Overview of PhysWAM . The Cosmos 3 generator [ Agarwal et al., 2026 ] jointly denoises multiview video, metric depth and ego motion, conditioned on text through its frozen reasoner; RGB and depth share a frozen VAE, and calibrated ray embeddings with mRoPE encode camera geometry and position. Current RGB and motion history stay clean; the rest is noised. Flow matching supervises the velocity outputs, whose clean estimates feed CPP, which transforms generated depth with generated motion and compares it with LiDAR under recorded motion, and the two hinges.
NAVSIM v1
NAVSIM v2
Method
Input
NC
DAC
TTC
C
EP
PDMS
NC
DAC
DDC
TLC
EP
TTC
LK
HC
EC
EPDMS
Human
–
100
100
100
99.9
87.5
94.8
–
End-to-end planners
TransFuser [ Chitta et al., 2022 ]
3 × Cam+L
97.7
92.8
92.8
100.0
79.2
84.0
96.9
89.9
97.8
99.7
87.1
95.4
92.7
98.3
87.2
76.7
LTF [ Chitta et al., 2022 ]
3 × Cam
–
97.6
91.9
99.2
99.8
87.6
97.2
96.6
98.3
86.3
83.6
Hydra-MDP++ (R34) [ Li et al., 2025b ]
3 × Cam
97.6
96.0
93.1
100
80.4
86.6
97.2
97.5
99.4
99.6
83.1
96.5
94.4
98.2
70.9
81.4
Table 1: NAVSIM v1 and v2 on navtest . Numbers of other methods are as reported in their papers. R34: ResNet-34 backbone; +L: LiDAR input. Bold: best per column, oracle row excluded. PhysWAM rows use one sample per scene unless otherwise specified; the sample-to-sample sd is 0.24 PDMS and 0.30 EPDMS.
Easy
Medium
Hard
Extreme
Overall
Method
RC
HD
RC
HD
RC
HD
RC
HD
RC
HD
UniAD [ Hu et al., 2023b ]
58.6
48.7
41.2
29.5
40.4
27.3
26.0
14.3
40.6
28.9
VAD [ Jiang et al., 2023 ]
38.7
24.3
27.0
9.9
25.5
10.4
23.0
8.2
27.9
12.3
LTF [ Chitta et al., 2022 ]
68.4
52.8
40.7
24.6
36.9
19.8
25.5
8.1
41.4
24.8
BeyondDrive [ Wang et al., 2026b ]
76.8
65.6
43.0
31.4
35.5
26.3
29.6
16.2
46.2
34.8
PhysWAM
93.5
86.9
47.4
30.1
38.6
25.2
26.1
13.4
48.9
35.5
Table 2: navhard and HUGSIM. (a) NAVSIM v2 on navhard , two-stage pseudo-simulation with reactive traffic; EPDMS as reported, with the sources of reprinted values and the stage scores in Table 7 . (b) Zero-shot closed-loop driving on HUGSIM, as reported: route completion (RC) and HD-Score by difficulty; BeyondDrive’s overall is the unweighted mean over the four levels. PhysWAM is trained on NAVSIM only, one sample per planner call. Bold: best.
PDMS
EPDMS
Off-road
Collision
Heading
4k, with CPP
80.3
78.8
9.7%
3.8%
2.7 ∘
4k, without CPP
57.0
55.7
32.7%
9.8%
8.9 ∘
navtest PDMS / EPDMS
navhard
HUGSIM RC / HD
final, with CPP
91.4 / 90.3
38.1
48.9 / 35.5
final, without CPP
89.6 / 88.4
36.2
47.5 / 33.4
Table 3: Ablations and future depth. (a) The full recipe with and without CPP after 4,000 updates on navtest (off-road, collision: share of scenes with a zero drivable-area or no-collision term; heading: median error against the recorded trajectory) and at the end of training (Tables 1 – 2 ). (b, d) LoRA setting on navtest , one sample per scene: objective terms added to flow matching, and the front view alone against all three views by driving command, with the full model of Table 1 for reference. (c) Generated depth of the full model against the LiDAR of the same future frame and camera, without scale alignment, four samples per scene; side views in Table 13 ; lower block as reported on nuScenes by GeoWAM [ Lu et al., 2026 ] .
Figure 3: Qualitative examples. Left: input views at t=0 and the command. Middle: generated views and metric depth at +2 and +4 s; the single-view model generates the front view only. Right: rollouts over the drivable area (blue), recorded traffic at 4 s (green), recorded trajectory (dashed); ego boxes at the end point or at the failure step, with the hit agent outlined. Top: with CPP the model generates a parked van a few metres off in its left view and stops short of it; without CPP it generates the van pressed against the ego and hits it at 3.4 s. Bottom: an oncoming car visible only in the left camera; the multi-view model generates it passing close and keeps the turn wide, the single-view model cuts the corner and collides at 2.3 s.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Quantity
Shape
Operation
RGB or encoded depth
(T+1)×Hv×Wv×3
Input to the shared frozen VAE.
xi
J×hv×wv×cℓ
VAE grid, indexed spatially by u .
P(xiσ)
Ni×p2cℓ
Flattened patches of the noisy grid, indexed by q .
zi
Ni×dh
Projected patches with ray, modality, and noise embeddings.
OV(ziout)
Nigen×p2cℓ
Patch velocities at generated positions.
vi
J×hv×wv×cℓ
Unpatchified velocity field; conditioning frames are zero-filled and not predicted.
Appendix
Table 4: Visual representation pathway for one stream i=(v,m) . Shapes omit the minibatch dimension. Noise and flow targets are defined on the VAE grid, before patchification; each spatial patch becomes one token.
View
AbsRel, cells
AbsRel, pixels ≤25 m
Sky error
Front
0.041
0.073
0.05%
Front-left
0.051
0.113
0.22%
Front-right
0.060
0.118
0.20%
Appendix
Table 5: Depth-label quality against LiDAR on 1,000 training windows. AbsRel is the mean absolute relative depth error; the cell-level column uses the geometric-mean LiDAR depth of latent cells with at least three returns, the pixel-level column all LiDAR returns within 25 m. Sky error is the fraction of pixels labelled sky that carry a LiDAR return.
Data
Training windows
103,281 (NAVSIM navtrain )
Frames per window
1 current + 8 future, 2 Hz
Views and resolution
front 832×468 ; front-left and front-right 416×234
Model
Initialization
Cosmos 3 Nano (15.2B parameters)
Trained parameters
generation pathway, 7.0B
Appendix
Table 6: Hyperparameters of PhysWAM .
Method
Difficulty
NC
DAC
TTC
COM
RC
HD-Score
UniAD [ Hu et al., 2023b ]
Easy
77.4
88.5
70.8
82.8
58.6
48.7
Medium
72.4
86.8
60.4
72.0
41.2
29.5
Hard
66.5
86.0
55.0
67.2
40.4
27.3
Extreme
54.4
89.6
42.9
58.5
26.0
14.3
VAD [ Jiang et al., 2023 ]
Easy
66.1
73.9
58.2
100.0
38.7
24.3
Medium
45.7
79.8
29.0
100.0
27.0
9.9
Appendix
Table 8: HUGSIM by difficulty and metric ( ×100 ); baselines as reported by the benchmark paper. Overall follows the benchmark’s reduction: per-dataset episode means, averaged with equal weight over the four source datasets within each difficulty, then weighted over the difficulties by their episode counts (80, 157, 96, 103). Bold: best per metric and difficulty among UniAD, VAD, LTF, and PhysWAM .
Dataset
Episodes
Easy
Medium
Hard
Extreme
All
nuScenes
88
94.5 / 88.4
43.7 / 27.1
59.5 / 43.1
40.6 / 27.3
56.7 / 42.9
Waymo
108
94.6 / 91.6
51.5 / 34.1
57.3 / 44.8
23.9 / 8.7
55.8 / 42.8
KITTI-360
113
87.4 / 76.5
51.6 / 39.3
14.9 / 6.1
8.0 / 0.7
41.1 / 31.0
PandaSet
127
97.4 / 91.1
42.8 / 20.1
22.5 / 6.6
32.0 / 16.9
41.3 / 24.3
Appendix
Table 9: HUGSIM by source dataset , PhysWAM : RC / HD-Score per difficulty level and the plain episode mean per dataset. The Overall values of Table 8 average these per-dataset means with equal dataset weight within each difficulty before weighting by episodes, so they are not the pooled episode mean.
Adaptation
LoRA, rank 128, α=256 , on the attention projections
Trained parameters
122.7M
Updates × batch
16,000 × 44 windows
Optimizer
AdamW, β=(0.9,0.99) , weight decay 0.05, clip 1.0
Learning rate
peak 2×10−4 , 300 warm-up updates, linear decay to 0
Weight averaging
none
Appendix
Table 10: Hyperparameters of the LoRA setting. Adapters act on the generation pathway; rows not listed are as in Table 6 .
Row
NC
DAC
DDC
TLC
EP
TTC
LK
HC
PDMS
EPDMS
flow
95.9
96.4
98.6
99.7
86.0
94.9
95.2
98.4
85.7
84.3
flow + hinge
96.8
96.6
98.5
99.8
86.8
95.9
95.3
98.6
87.2
85.7
flow + CPP
96.8
97.0
98.5
99.7
87.0
95.7
95.1
98.6
87.8
86.2
flow + hinge + CPP (multi-view)
96.9
97.2
98.5
99.7
87.2
95.8
95.1
98.6
88.2
86.5
single-view
96.6
96.7
98.6
99.7
86.8
95.6
95.2
98.5
87.1
85.6
multi-view, no depth
96.9
96.8
98.5
99.7
87.0
95.9
95.2
98.6
87.5
85.9
Appendix
Table 11: LoRA setting, all sub-scores on navtest , one sample per scene.
PDMS
EPDMS
Row
Left
Straight
Right
Left
Straight
Right
flow
81.4
87.8
81.9
80.3
86.2
81.3
flow + hinge
83.4
89.0
83.8
82.4
87.2
82.8
flow + CPP
83.5
89.9
83.9
82.1
88.1
83.0
flow + hinge + CPP (multi-view)
83.9
90.3
84.3
82.4
88.4
83.3
single-view
81.2
90.0
82.0
80.1
88.2
81.4
Appendix
Table 12: Scores by driving command on navtest (2,501 left, 8,070 straight, 1,575 right scenes), one sample per scene.
+2 s
+4 s
AbsRel ↓
δ1.25↑
AbsRel ↓
δ1.25↑
View
with
without
with
without
with
without
with
without
Front
0.175
0.193
0.814
0.802
0.232
0.253
0.742
0.743
Front-left
0.292
0.315
0.725
0.706
0.374
0.404
0.654
0.649
Front-right
0.240
0.277
0.752
0.725
0.305
0.346
0.676
0.663
Front, cells
0.144
0.154
0.856
0.844
0.194
0.205
0.795
0.795
Appendix
Table 13: Generated depth with and without CPP for the final models, on navtest : AbsRel and δ1.25 against the LiDAR of the same future frame and camera, without scale alignment, as in Table 3 a; means over four samples per scene. “Cells”: the 16-pixel cells that CPP supervises.
Clips
FVD ↓
FID ↓
DrivingGPT
512
142.6
12.8
PWM †
–
86.0
–
DriveDreamer-Policy
–
53.6
–
CoWorld-VLA
–
32.7
–
Recorded (floor)
600
91.5 ± 3.2
11.0 ± 0.3
PhysWAM
600
111.3 ± 6.8
15.4 ± 0.6
Appendix
Table 14: Front-view video quality on NAVSIM. I3D FVD over one context and eight generated frames at 2 Hz, FID over the same frames. 600 clips: generated clips of half the scenes against the recorded clips of the other half, ten splits, floor from the two recorded halves; 1,200 clips: generated against recorded clips of the same scenes. Reported rows use their own clip counts and lengths (DrivingGPT [ Chen et al., 2025b ] 12 frames; PWM [ Zhao et al., 2026a ] 10, † as reported by DriveDreamer-Policy [ Zhou et al., 2026a ] , itself 9; CoWorld-VLA [ Huang et al., 2026 ] unstated).
Quantity
Floor
PhysWAM
Excess
Yaw, video vs. motion, 4 s ( ∘ )
0.29
0.80
+0.44 [+0.39, +0.48]
turns above 45∘
6.11
–
+2.31 [+1.93, +2.86]
Cross-view ∣Δlogz∣ , +2 s, L / R
0.032 / 0.030
0.045 / 0.049
1.5 ×
Temporal ∣Δlogz∣ , 0 to 4 s
0.087
0.285
3.3 ×
Appendix
Table 15: Consistency of the generated future with its own motion, across views, and over time, on navtest ; medians. Yaw: camera rotation recovered from the generated video [ Keetha et al., 2026 ] against the generated motion, and, for the floor, from the recorded clip against the recorded trajectory; the excess is the median of the per-scene paired differences, with a 95% bootstrap interval. Depth: median ∣Δlogz∣ after warping the front depth into the side views with the recorded extrinsics, or frame 0 into frame 8 with the generated motion; floors from two LiDAR sweeps 0.5 s apart and from recorded LiDAR under recorded motion.
UniPC steps
Δ PDMS
Samples
GPU-s / scene
4
−0.6
1
3.5
8
+0.7
3
4.4
15
−0.4
1
6.0
30
0
4
9.4
Appendix
Table 16: Sampler steps and cost for the model of Tables 1 – 2 : PDMS relative to the 30-step protocol row, one sample per scene, averaged over the number of samples given (sample-to-sample sd 0.2–0.3 where repeated), and the cost of one plan on one RTX PRO 6000 GPU.
Figure 4: Additional qualitative examples. On a straight road, the model with CPP generates the overtaking bus alongside in its left view. The model trained without CPP omits it and is clipped by the overtaking car at 3.9 s. At a hotel drop-off, it generates the departing minivan too close and clips the island at 1.6 s. The single-view model swings wide over a grass median that only the right camera shows and leaves the road at 3.1 s. At a left turn, it never generates the waiting truck and hits it at 3.7 s. The full model renders a stopped minivan at a constant size, plans as if it were pulling away, and rear-ends it at 3.7 s, a world-model error. At a tight left turn, both models generate indistinguishable views, but the full model clips the inner curb by 0.4 m at 3.6 s, a plan-precision error.
Representative works
Predicted world
Connection to action
LAW, SimWAM [ Li et al., 2025c , Zhao et al., 2026b ]
Future features or video
Future prediction provides training supervision; planning does not require explicit future generation.
DriveWAM, GeoWAM [ Shi et al., 2026 , Lu et al., 2026 ]
Future video or 3D geometry
Representations of the predicted future condition trajectory generation.
Epona, DriveVA [ Zhang et al., 2025b , Liu et al., 2026b ]
Future video
Video and action generation share learned representations.
4D-WAM [ Fu et al., 2026 ]
Future video
Joint video–action learning is supervised by matching geometry recovered from generated and recorded video.
WoTE, DA-WAM [ Li et al., 2025d , Zhong et al., 2026 ]
Future scene features
A learned scorer evaluates candidate trajectories using their predicted future states.
PhysWAM (ours)
Multiview video and metric depth
Video, depth, and ego motion are co-denoised. CPP jointly supervises generated depth and motion against measured scene geometry. A label-free consensus rule selects among trajectories.
Appendix
Table 17: World modeling and its role in planning. Representative works illustrate how future prediction supports policy learning, action generation, or trajectory selection. These roles can overlap within a method.