V2X-WAM: A Cooperative World Action Model for End-to-End Autonomous Driving
Authors: Junwei You, Weizhe Tang, Can Wang, Yan Zhao, Jun Hua, Haotian Shi, Wei Zhang, Lin Wang, +1 more
Organizations: ITS Center, Research Institute of Highway Ministry of Transport, Beijing, 100088, China · State Key Lab of Intelligent Transportation System, Research Institute of Highway Ministry of Transport, Beijing, 100088, China · Department of Civil and Environmental Engineering, University of Wisconsin–Madison, Madison, WI, 53706, USA · Intelligent Transportation Systems Research Center, Wuhan University of Technology, Wuhan, 430063, China · Engineering Research Center of Transportation Information and Safety, Ministry of Education, Wuhan, 430063, China · School of Transportation, Inner Mongolia University, Hohhot, 010021, China · College of Transportation, Tongji University, Shanghai, 201804, China
Vehicle-infrastructure cooperation can complement onboard sensing with broader and more informative observations of the traffic environment, providing valuable support for end-to-end autonomous driving. However, existing cooperative driving methods mainly exploit roadside information to enhance the representation of the current scene, while the future consequences of prospective driving actions are rarely modeled explicitly. This limits the ability of the planner to anticipate how its decisions may interact with the evolving traffic environment. To address this issue, we propose V2X-WAM, a cooperative world action model that tightly couples cooperative scene understanding, action generation, and future-world reasoning. V2X-WAM constructs a reliability-aware spatiotemporal representation from vehicle- and infrastructure-side observations, while compressing infrastructure information into a compact quantized message for efficient communication. Based on the resulting cooperative representation, a multimodal planner generates prospective trajectories, which explicitly condition future occupancy and dynamic-flow prediction. The predicted world consequences are then fed back to refine the planned trajectory, forming a closed interaction between action and future-world evolution. Experiments on a large-scale real-world cooperative driving dataset demonstrate that V2X-WAM consistently improves planning accuracy and safety over representative end-to-end cooperative driving methods, while achieving stronger future-world prediction and substantially lower communication overhead. Ablation studies further validate the effectiveness of the proposed design.
Figures & tables
Figure 1: Evolution of end-to-end cooperative autonomous driving paradigms.
Figure 2: Overall architecture of the proposed V2X-WAM.
Figure 3: Action-conditioned future-world modeling with joint occupancy–flow prediction and flow-guided occupancy transport.
Objective
Hyperparameter
Value
Planning
λbest
1.00
λtop
1.25
λmode
0.35
λact
0.15
λtail
0.70
λprog
0.30
Table 1: Loss hyperparameters used for training V2X-WAM.
Method
L2 Error (m) ↓
Collision Rate (%) ↓
Trans. Cost (B/s) ↓
1s
2s
3s
Avg.
1s
2s
3s
Avg.
VAD ∗ ( Jiang et al., 2023 )
1.65
2.72
3.80
2.72
0.86
1.21
1.28
1.12
–
UniAD ∗ ( Hu et al., 2023 )
1.26
2.22
3.06
2.18
0.88
1.18
1.32
1.13
–
SparseDrive ∗ ( Sun et al., 2025 )
1.02
1.69
2.37
1.69
0.46
1.23
1.28
0.99
–
Vanilla
1.36
2.29
3.32
2.32
1.03
0.88
1.32
1.08
8.19×107
V2VNet ( Wang et al., 2020 )
1.96
2.37
3.41
2.58
0.74
0.88
1.03
0.88
8.19×107
Table 2: End-to-end planning performance and transmission cost on V2X-Seq-SPD.
Method
Occupied IoU (%) ↑
Dynamic Flow EPE (m) ↓
1s
2s
3s
Avg.
1s
2s
3s
Avg.
FIERY ( Hu et al., 2021 )
21.27
17.04
15.13
17.81
1.22
1.33
1.46
1.33
PowerBEV ( Li et al., 2023 )
23.46
18.85
18.32
20.21
1.03
1.14
1.20
1.12
StreamingFlow ( Shi et al., 2024 )
24.49
19.79
17.54
20.61
0.93
1.07
1.18
1.06
V2X-WAM (Ours)
27.90
21.69
19.63
23.07
1.01
1.09
1.16
1.08
Table 3: Future-world prediction performance on V2X-Seq-SPD.
Method
Avg. L2 Error (m) ↓
Avg. Collision (%) ↓
Avg. Occupied IoU (%) ↑
Avg. Flow EPE (m) ↓
Ego Only
1.11
0.05
19.80
1.15
w/o Action Conditioning
1.15
0.02
23.02
1.09
w/o Flow Transport
1.02
0.24
22.40
1.10
w/o Reliability
1.22
0.10
22.15
1.15
w/o Temporal Modeling
1.10
0.08
20.58
1.46
w/o World–Action Coupling
1.08
0.02
22.82
1.14
Table 4: Ablation study of V2X-WAM.
Figure 4: Qualitative planning and 3-s future-occupancy results for representative left-turn, straight, and right-turn scenes.
Figure 5: Controlled infrastructure-message analysis of future occupancy along road-aligned space–time slices.