Practical autonomous driving requires models that generalize by reasoning through spatial-temporal possibilities to exclude unsafe outcomes. While state-of-the-art (SOTA) methods use parallel planning architectures, they fail to explicitly couple speed decisions with agent behavior along the driving path, leading to suboptimal coordination. To address this, we propose a cascaded framework that transforms longitudinal planning from an independent prediction task into a path-conditioned reasoning process. On the model side, we introduce an anchor-based regression design that conditions longitudinal prediction on the lateral drive path, and reformulate longitudinal planning as 1D displacement prediction along the path. This reduces geometric uncertainty and sharpens the model's focus on interaction-driven dynamics. On the data side, we introduce a planning-oriented data augmentation strategy that simulates rare safety-critical events by programmatically inserting agents and relabeling longitudinal targets to enforce collision avoidance. Evaluated on the challenging Bench2Drive benchmark, our method achieves SOTA performance with a driving score of 89.07 and a success rate of 73.18%, demonstrating significantly improved coordination and safety. Further evaluation on Fail2Drive confirms strong generalization to rare edge cases where parallel formulations typically fail. Project page:https://yanhaowu.github.io/AlignDrive/.
Figures & tables
Figure 1: (a) Drive path (black), trajectory (blue), and longitudinal displacement (red). Path waypoints are sampled spatially, trajectory waypoints temporally, and displacements represent traveled distance along the path at fixed time intervals. (b) Comparison of E2E paradigms. Parallel planning predicts the drive path and longitudinal trajectory independently, which can lead to potential coordination inconsistencies. In the example on the right, the independently predicted longitudinal trajectory is collision-free, but applying its speed along a separately predicted lateral path could cause a collision. In contrast, our cascaded paradigm first predicts the drive path and then regresses path-conditioned longitudinal displacements. With the path prior, the model identifies the potential conflict and outputs shorter displacements, yielding to avoid collision. Perception inputs are omitted for clarity.
Figure 2: Overview of the proposed AlignDrive system, which consists of three components. The Drive Path Predictor refines queries through cross-attention with image features to encode the drive path, maps, and agents. The Planning-oriented Data Augmentation enriches scenarios by inserting additional agents and relabeling longitudinal displacements. Finally, the Longitudinal Planning Module predicts forward displacements along the drive path; combined with the path, these displacements yield the final trajectory that is both collision-aware and spatially consistent. On the right side of the figure, the black numbers denote the scores of predicted drive paths, while the red numbers represent the scores of the corresponding longitudinal planning for each drive path.
Figure 3: (a) Planning-oriented augmentation. Non-threatening agents are inserted at a distance with unchanged GT displacements, while threatening agents are placed nearby and cause adaptive shortening of GT displacements. (b) Representation encoding. Inserted agents are projected into future positions, transformed to corner representations, and encoded via a Fourier encoder (top). Reference points are sampled by displacement anchors and encoded with MLPs. For clarity, although multiple drive paths are predicted in practice, only one representative path is illustrated here (bottom).
Method
Driving Score ( ↑ )
SR (%) ( ↑ )
Driving Efficiency ( ↑ )
Comfort ( ↑ )
Expert: PDM-Lite [ 1 ]
SpaceDrive [ 17 ]
78.02
55.11
-
-
SimLingo [ 25 ]
86.02
67.27
259.23
33.67
Expert: Think2Drive [ 18 ]
VAD [ 16 ]
42.35
15.00
157.94
46.01
SparseDrive [ 28 ]
44.54
16.71
170.21
48.63
Table 1: Closed-loop results of planning in Bench2Drive. Bold and underlined numbers indicate the best performance within different expert groups.
Method
Ability (%) ↑
Mean
Merging
Overtaking
Emergency Brake
Give Way
Traffic Sign
UniAD-Base [ 12 ]
15.55
14.10
17.78
21.67
10.00
14.21
VAD [ 16 ]
18.07
8.11
24.44
18.64
20.00
19.15
DriveTransformer-Large [ 15 ]
38.60
17.57
35.00
48.36
40.00
52.10
HiP-AD [ 29 ]
65.98
50.00
84.44
83.33
40.00
72.10
AlignDrive
70.06
75.00
75.56
75.00
50.00
74.74
Table 2: Multi-Ability Results in Bench2Drive.
Method
RGB
Lidar
Privilege
Language
In-Distribution
Generalization
DS ( ↑ )
SR ( ↑ )
HM ( ↑ )
DS ( ↑ )
SR ( ↑ )
HM ( ↑ )
Privileged Methods
PlanT 2.0 [ 8 ]
×
×
✓
×
87.8
85.0
86.4
73.3
58.0
64.8
PDMLite-F2D [ 1 ]
×
×
✓
×
95.6
97.0
96.3
94.0
95.3
94.6
Multimodal Methods
Orion [ 6 ]
✓
×
×
✓
53.0
52.0
52.5
51.2
46.0
48.5
Table 3: Closed-loop results on Fail2Drive benchmark. The best results for RGB-only methods are bolded , and the best for Privileged/Multimodal methods are underlined .
Method
Parameters
Latency
Driving Score
Success Rate (%)
VAD-Base [ 16 ]
-
224.3 ms
42.35
15.00
DriveTransformer-Large [ 15 ]
646 M
221.7 ms
63.46
35.01
HiP-AD [ 29 ]
97.4 M
138.9 ms
86.77
69.09
AlignDrive
117.2 M
177.5 ms
89.07
73.18
AlignDrive-Small
83.7 M
124.5 ms
87.45
71.82
Table 4: Comparison of inference efficiency and driving performance. AlignDrive-Small is a lightweight variant with fewer decoder layers. Experiments are conducted on an RTX 3090 GPU.
Variant
LP
DP
DA
Driving Score ↑
Success Rate (%) ↑
Collision Rate (%) ↓
A
83.21
63.18
22.7
B
✓
84.85
65.45
19.5
C
✓
✓
85.82
66.81
16.3
D
✓
✓
86.54
68.92
15.7
E
✓
✓
✓
89.07
73.18
11.4
Table 5: Ablation study on AlignDrive components. LP: uses lateral path prediction to condition longitudinal planning; DP: formulates longitudinal planning as displacement regression along the drive path; DA: applies planning-oriented data augmentation
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Visualization of planning-oriented data augmentation. The top row shows non-threatening agents, while the bottom row shows threatening agents. Inserted synthetic agents are indicated with dashed boxes. Red points denote the ego vehicle’s original trajectory, and blue lines represent the adjusted longitudinal displacements after augmentation.
Hyperparameter
nuScenes
Bench2Drive
Displacement threshold δ
0.5 m
0.5 m
Insertion probability α
0.05
0.1
Safe distance dsafe
3.0 m
3.0 m
Ego trajectory horizon
3 s
3 s
Appendix
Table 6: Dataset-specific hyperparameters for nuScenes and Bench2Drive.
Method
Params
L2 ( m ) ↓
Collision (%) ↓
1 s
2 s
3 s
Avg.
1 s
2 s
3 s
Avg.
VAD-Base [ 16 ]
–
0.41
0.70
1.05
0.72
0.03
0.19
0.43
0.21
GenAD [ 37 ]
–
0.28
0.49
0.78
0.52
0.08
0.14
0.34
0.19
SparseDrive-S [ 28 ]
85 M
0.29
0.58
0.96
0.61
0.01
0.05
0.18
0.08
DriveTransformer-Large [ 15 ]
646 M
0.16
0.30
0.55
0.33
0.01
0.06
0.15
0.07
HiP-AD [ 29 ]
90 M
0.28
0.53
0.87
0.56
0.01
0.05
0.15
0.07
Appendix
Table 7: Open-loop planning evaluation results on the nuScenes validation dataset.
Method
LP
RE
DA
Driving Score ↑
Success Rate (%) ↑
Collision Rate (%) ↓
Decouple
83.21
63.18
22.7
LP + Original
✓
87.47
68.18
15.4
LP + Reencode
✓
✓
85.82
66.81
16.3
Full (AlignDrive)
✓
✓
✓
89.07
73.18
11.4
Appendix
Table 8: Ablation on longitudinal planning (LP), agent query decoding–re-encoding (RE), and planning-oriented data augmentation (DA). Decouple: no LP; LP + Original: LP with original agent queries; LP + Reencode: LP with decoded–re-encoded queries; Full (AlignDrive): LP + Reencode + DA.
Run
Driving Score ↑
Success Rate (%) ↑
Driving Efficiency ↑
Comfort ↑
Run 1
89.07
73.18
212.07
16.86
Run 2
87.80
71.36
207.85
15.25
Run 3
88.05
70.00
210.08
17.10
Average
88.30
71.50
210.00
16.40
Appendix
Table 9: Multiple simulation runs of AlignDrive on Bench2Drive benchmarks. Driving Score, Success Rate, Driving Efficiency, and Comfort are reported for each run along with the average.
Figure 5: Effect of planning-oriented data augmentation on planning performance. All augmented variants ( p=0.1,0.2,0.3,0.4 ) outperform the no-augmentation baseline.
Figure 6: Red points are predicted drive paths, while blue points show longitudinal planning outputs (trajectory waypoints for the baseline, displacement sequences for ours). Relevant vehicles are highlighted in green. The baseline collides with cross-traffic, while our method avoids it.
Figure 7: Comparison of Baseline (a) and Ours (b) in a pedestrian cut-in scenario. The baseline model fails to avoid the pedestrian, resulting in a collision, whereas our method promptly reacts and avoids the accident. The pedestrian is highlighted with a black dashed circle.
Long-horizon planning is critical for safe autonomous driving in complex scenarios. Existing methods improve planning continuity with temporal memory, but such memory may become invalid and mislead decisions when the driving command changes. Thus, selectively leveraging useful history while suppressing command-inconsistent memory remains a key challenge. To address this issue, we propose MomADv2, a reliable state-space memory framework for long-horizon end-to-end autonomous driving. At its core, MomADv2 introduces a Selective State-Space Planning Memory Query Module, which filters historical planning queries based on temporal continuity and command consistency, selects planning modes relevant to the current command, and models the evolution of planning intentions through a selective state-space mechanism. To further alleviate local trajectory deviations and error accumulation in long-horizon planning, we design a Flow-Matching Trajectory Residual Refiner. It learns a continuous residual correction field from the refined planning output to the expert trajectory, enabling fine-grained trajectory refinement while preserving the stability of anchor-based planning. Extensive experiments on closed-loop NAVSIM and Bench2Drive, as well as open-loop nuScenes, demonstrate that MomADv2 improves long-horizon planning consistency and reduces the average collision rate by 15.6% over MomAD under 6-second planning.
Ziying Song, Shengkai Zhang, Lin Liu +8
School of Artificial Intelligence (School of Software), Yanshan University · Beijing Jiaotong University · Nanyang Technological University +5
End-to-end autonomous driving has made significant progress by unifying perception, prediction, and planning within a single learning framework, achieving strong performance in short-horizon decision making. However, most existing E2E-AD methods remain confined to short-horizon planning and lack the ability to model long-term temporal dependencies, which severely limits their generalization and security in complex and highly interactive driving scenarios. In this work, we propose GraphWorld, an E2E-AD framework that explicitly enhances long-horizon planning through latent world modeling. We introduce an Ego-Centric Interaction Graph, which adaptively models critical neighboring agents based on spatial proximity, and propagates relational context to planning queries via cross-node cross-attention. We present a World-State-Conditioned Planning that learns ego-centric latent world representations by modeling interactions between an ego vehicle and surrounding agents. This latent world state captures key interaction dynamics and safety-relevant semantics, and serves as a conditioning signal to guide long-horizon, safety-aware trajectory planning. Extensive experiments on Bench2Drive, NAVSIMv1/2, and nuScenes demonstrate that GraphWorld significantly reduces collision rates and improves long-horizon planning performance, validating its effectiveness in complex driving environments.
Ziying Song, Caiyan Jia, Lin Liu +8
Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence, School of Computer Science and Technology, Beijing Jiaotong University, Beijing, China
Autonomous driving systems are steadily moving toward end-to-end paradigms to mitigate the limited adaptability of rule-based pipelines in complex traffic environments. However, most existing learning-based methods still make decisions from static representations of the current scene, without explicit future rollouts or modeling of the temporal causal dynamics in traffic interactions. This limitation often results in unstable or overly conservative planning under high-uncertainty conditions, such as occlusions and unexpected events. To overcome these challenges, we introduce OWMDrive, a generative end-to-end driving framework built upon an Occupancy World Model for multi-step 3D occupancy forecasting, which serves as a conditional prior to guide diffusion-based planning. Conditioned on both current observations and predicted future states, the planner iteratively refines trajectory candidates to generate a reinforced driving trajectory. By explicitly modeling scene evolution over future horizons, OWMDrive captures key spatiotemporal causal dependencies, which leads to more foresighted and robust trajectory generation. Extensive experiments demonstrate that OWMDrive significantly improves planning reliability and safety, especially in challenging and partially observable driving scenarios.
Junjie Cheng, Ruiqi Song, Ye Wu +3
The School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 100049, China · Waytous Inc., Qingdao 266109, China · The State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China +1