V2X-WAM: A Cooperative World Action Model for End-to-End Autonomous Driving
Authors: Junwei You, Weizhe Tang, Can Wang, Yan Zhao, Jun Hua, Haotian Shi, Wei Zhang, Lin Wang, +1 more
Organizations: ITS Center, Research Institute of Highway Ministry of Transport, Beijing, 100088, China · State Key Lab of Intelligent Transportation System, Research Institute of Highway Ministry of Transport, Beijing, 100088, China · Department of Civil and Environmental Engineering, University of Wisconsin–Madison, Madison, WI, 53706, USA · Intelligent Transportation Systems Research Center, Wuhan University of Technology, Wuhan, 430063, China · Engineering Research Center of Transportation Information and Safety, Ministry of Education, Wuhan, 430063, China · School of Transportation, Inner Mongolia University, Hohhot, 010021, China · College of Transportation, Tongji University, Shanghai, 201804, China
Vehicle-infrastructure cooperation can complement onboard sensing with broader and more informative observations of the traffic environment, providing valuable support for end-to-end autonomous driving. However, existing cooperative driving methods mainly exploit roadside information to enhance the representation of the current scene, while the future consequences of prospective driving actions are rarely modeled explicitly. This limits the ability of the planner to anticipate how its decisions may interact with the evolving traffic environment. To address this issue, we propose V2X-WAM, a cooperative world action model that tightly couples cooperative scene understanding, action generation, and future-world reasoning. V2X-WAM constructs a reliability-aware spatiotemporal representation from vehicle- and infrastructure-side observations, while compressing infrastructure information into a compact quantized message for efficient communication. Based on the resulting cooperative representation, a multimodal planner generates prospective trajectories, which explicitly condition future occupancy and dynamic-flow prediction. The predicted world consequences are then fed back to refine the planned trajectory, forming a closed interaction between action and future-world evolution. Experiments on a large-scale real-world cooperative driving dataset demonstrate that V2X-WAM consistently improves planning accuracy and safety over representative end-to-end cooperative driving methods, while achieving stronger future-world prediction and substantially lower communication overhead. Ablation studies further validate the effectiveness of the proposed design.
Figures & tables
Figure 1: Evolution of end-to-end cooperative autonomous driving paradigms.
Figure 2: Overall architecture of the proposed V2X-WAM.
Figure 3: Action-conditioned future-world modeling with joint occupancy–flow prediction and flow-guided occupancy transport.
Objective
Hyperparameter
Value
Planning
λbest
1.00
λtop
1.25
λmode
0.35
λact
0.15
λtail
0.70
λprog
0.30
Table 1: Loss hyperparameters used for training V2X-WAM.
Method
L2 Error (m) ↓
Collision Rate (%) ↓
Trans. Cost (B/s) ↓
1s
2s
3s
Avg.
1s
2s
3s
Avg.
VAD ∗ ( Jiang et al., 2023 )
1.65
2.72
3.80
2.72
0.86
1.21
1.28
1.12
–
UniAD ∗ ( Hu et al., 2023 )
1.26
2.22
3.06
2.18
0.88
1.18
1.32
1.13
–
SparseDrive ∗ ( Sun et al., 2025 )
1.02
1.69
2.37
1.69
0.46
1.23
1.28
0.99
–
Vanilla
1.36
2.29
3.32
2.32
1.03
0.88
1.32
1.08
8.19×107
V2VNet ( Wang et al., 2020 )
1.96
2.37
3.41
2.58
0.74
0.88
1.03
0.88
8.19×107
Table 2: End-to-end planning performance and transmission cost on V2X-Seq-SPD.
Method
Occupied IoU (%) ↑
Dynamic Flow EPE (m) ↓
1s
2s
3s
Avg.
1s
2s
3s
Avg.
FIERY ( Hu et al., 2021 )
21.27
17.04
15.13
17.81
1.22
1.33
1.46
1.33
PowerBEV ( Li et al., 2023 )
23.46
18.85
18.32
20.21
1.03
1.14
1.20
1.12
StreamingFlow ( Shi et al., 2024 )
24.49
19.79
17.54
20.61
0.93
1.07
1.18
1.06
V2X-WAM (Ours)
27.90
21.69
19.63
23.07
1.01
1.09
1.16
1.08
Table 3: Future-world prediction performance on V2X-Seq-SPD.
Method
Avg. L2 Error (m) ↓
Avg. Collision (%) ↓
Avg. Occupied IoU (%) ↑
Avg. Flow EPE (m) ↓
Ego Only
1.11
0.05
19.80
1.15
w/o Action Conditioning
1.15
0.02
23.02
1.09
w/o Flow Transport
1.02
0.24
22.40
1.10
w/o Reliability
1.22
0.10
22.15
1.15
w/o Temporal Modeling
1.10
0.08
20.58
1.46
w/o World–Action Coupling
1.08
0.02
22.82
1.14
Table 4: Ablation study of V2X-WAM.
Figure 4: Qualitative planning and 3-s future-occupancy results for representative left-turn, straight, and right-turn scenes.
Figure 5: Controlled infrastructure-message analysis of future occupancy along road-aligned space–time slices.
We present OmniV2X, a generative foundation model for vehicle-to-everything (V2X) cooperative driving. The model directly interprets independent context sequences comprising multi-modal and multi-agent observations. The new design mitigates the computational cost of dense 3D perception, the vulnerability to data scarcity in cooperative scenarios, and the poor compliance with standardized messaging in existing methods that fuse multi-modal inputs into a shared representation. For training, we present an end-to-end supervised pipeline using a downstream trajectory generation loss, in which a high-capacity generative sequence planner implicitly learns to steer the model and leverage multi-modal inputs via cross-attention injection. As a foundation model, we demonstrate that OmniV2X pre-trained on large-scale single-agent planning datasets can efficiently adapt to cooperative environments by integrating the conditioning context with lightweight, standard-compliant V2X tokens. Evaluated on the DAIR-V2X-Seq dataset, OmniV2X outperforms existing end-to-end cooperative driving baselines, achieving state-of-the-art performance with less than 10% of the fine-tune V2X dataset and less than 1% of the communication bandwidth. We conduct comprehensive evaluations to demonstrate its computational efficiency and robustness under real-world constraints.
Juntong Peng, Juanwu Lu, Yupeng Zhou +3
College of Engineering, Purdue University, West Lafayette, IN 47907, USA
Vehicle-to-everything-aided autonomous driving (V2X-AD) significantly enhances driving performance through information sharing. However, existing collaborative perception methods only optimize module-level perception capabilities and fail to effectively serve the ultimate planning and control tasks. We propose an end-to-end collaborative driving system that directly optimizes planning task performance. The system employs MotionNetwork to fuse historical temporal information, utilizes attention mechanisms to efficiently compress spatial features into compact tokens, and adaptively fuses multi-agent features through an autoregressive decoder. Additionally, we introduce Mixture-of-Experts (MoE) architecture to enhance the model's representation capacity for heterogeneous features. Experiments demonstrate that our method achieves a driving score of 79.72, surpassing the state-of-the-art CoDriving baseline (77.15) by 3.33% in closed-loop evaluation while maintaining communication efficiency.
Nuoran Li, Zhang Zhang, Yueran Zhao +2
Shenzhen Automotive Research Institute National Engineering Research Center of Electric Vehicles Beijing Institute of Technology Shenzhen, China
World models (WMs) have demonstrated strong potential for end-to-end autonomous driving by learning predictive representations of future scene dynamics. However, generating future videos during inference introduces substantial computational overhead, leading many recent driving WMs to adopt a single front camera as input for efficient deployment. This design restricts spatial coverage in safety-critical maneuvers such as lane changes, merges, and turns. To address this limitation, we propose SV-WAM, a surround-view world-action model (WAM) that preserves full six-camera observations while maintaining efficient inference. SV-WAM leverages future-video prediction as dense training supervision for action learning within a shared generative model, rather than as an inference-time output. At the core of this design is an action-centered causal mask that prevents action tokens from attending to future-video tokens during joint action-video denoising. Consequently, the video branch can be discarded at deployment, enabling efficient action-only planning. Furthermore, we introduce a differentiable drivable-area compliance regularizer that penalizes vehicle-footprint corners approaching or crossing drivable boundaries, improving planning safety and boundary awareness. Extensive experiments on the closed-loop NAVSIMv2 benchmark and the open-loop nuScenes benchmark demonstrate that SV-WAM achieves state-of-the-art planning performance with low inference latency and competitive zero-shot transfer capability.
Jinyang Wang, Shiwei Li, Junjian Wang +12
Institute of Automation, Chinese Academy of Sciences · 2Chongqing Changan Technology Co., Ltd. · 3Civil Aviation University of China +2