MomADv2: Reliable Temporal Memory for End-to-End Autonomous Driving
Authors: Ziying Song, Shengkai Zhang, Lin Liu, Peiliang Wu, Lei Yang, Dongyang Xu, Bin Sun, Li Wang, +3 more
Organizations: School of Artificial Intelligence (School of Software), Yanshan University · Beijing Jiaotong University · Nanyang Technological University · Tsinghua University · China Automotive Technology and Research Center Co., Ltd. · School of Mechanical Engineering, Beijing Institute of Technology · University of Macau · The University of Queensland
Long-horizon planning is critical for safe autonomous driving in complex scenarios. Existing methods improve planning continuity with temporal memory, but such memory may become invalid and mislead decisions when the driving command changes. Thus, selectively leveraging useful history while suppressing command-inconsistent memory remains a key challenge. To address this issue, we propose MomADv2, a reliable state-space memory framework for long-horizon end-to-end autonomous driving. At its core, MomADv2 introduces a Selective State-Space Planning Memory Query Module, which filters historical planning queries based on temporal continuity and command consistency, selects planning modes relevant to the current command, and models the evolution of planning intentions through a selective state-space mechanism. To further alleviate local trajectory deviations and error accumulation in long-horizon planning, we design a Flow-Matching Trajectory Residual Refiner. It learns a continuous residual correction field from the refined planning output to the expert trajectory, enabling fine-grained trajectory refinement while preserving the stability of anchor-based planning. Extensive experiments on closed-loop NAVSIM and Bench2Drive, as well as open-loop nuScenes, demonstrate that MomADv2 improves long-horizon planning consistency and reduces the average collision rate by 15.6% over MomAD under 6-second planning.
Figures & tables
Figure 1: Motivation of MomADv2. (a) Long-horizon planning requires consistent intentions across planning cycles. (b) Existing methods, represented by MomAD ( Song et al. 2025 ; Jia et al. 2025 ; Zhang et al. 2025a ) , indiscriminately reuse history, introducing new interference. (c) MomADv2 selectively preserves reliable states and refines local trajectory deviations via flow matching. (d) This reliable memory inheritance enables more accurate and safer long-horizon planning.
Figure 2: Overview of MomADv2 . Multi-view images are encoded into map and detection representations to generate initial ego planning queries. A Selective State-Space Planning Memory Query module retains valid historical queries based on scene, token, command, and buffer consistency. A Flow-Matching Trajectory Residual Refiner then predicts a conditional residual velocity field toward the expert trajectory, producing smooth, accurate, and temporally consistent long-horizon planning.
Figure 3: Flow Matching Supervision learns a query-conditioned residual velocity field that transforms the SSM-enhanced planner trajectory toward the expert trajectory. During inference, Gated Residual Refinement integrates this field with a few Euler steps and uses a bounded residual gate to improve trajectory accuracy while preserving anchor-based planning stability.
Method
Venue
L2 (m)↓
Col. Rate (%)↓
1s
2s
3s
4s
5s
6s
Avg.
1s
2s
3s
4s
5s
6s
Avg.
UniAD ( Hu et al. 2023 )
CVPR’23
0.47
0.91
1.35
1.91
2.47
3.07
1.70
0.25
0.36
0.61
0.99
1.64
2.51
1.06
SparseDrive ( Sun et al. 2024 )
CVPR’23
0.43
0.87
1.23
1.75
2.32
2.95
1.59
0.19
0.31
0.56
0.87
1.54
2.33
0.97
MomAD ( Song et al. 2025 )
CVPR’25
0.41
0.85
1.13
1.67
1.98
2.45
1.42
0.17
0.30
0.54
0.83
1.43
2.13
0.90
LAW ( Li et al. 2025c )
ICLR’25
0.40
0.87
1.16
1.71
2.03
2.61
1.46
0.19
0.33
0.57
0.86
1.51
2.31
0.96
Epona ( Zhang et al. 2025b )
ICCV’25
0.39
0.91
1.17
1.73
2.02
2.75
1.50
0.14
0.18
0.45
0.74
1.48
2.23
0.87
Table 1: Planning results for 6-second long-horizon planning on the nuScenes validation set.
Method
Venue
Open-loop
Closed-loop
Avg.L2↓
DS↑
SR (%)↑
Effi↑
Comf↑
DriveAdapter* ( Jia et al. 2023 )
ICCV’23
1.01
64.22
33.08
70.22
16.01
DriveDPO ( Shang et al. 2025 )
NeurIPS’25
-
62.02
30.62
-
-
Raw2Drive ( Yang et al. 2025 )
NeurIPS’25
-
71.36
50.24
-
-
DriveTrans* ( Jia et al. 2025 )
ICLR’25
0.62
63.46
35.01
100.64
20.78
WoTE* ( Li et al. 2025e )
ICCV’25
-
61.71
31.36
-
-
Table 2: Open-Loop and Closed-Loop results on Bench2Drive (V0.0.3) under base training set. * denotes expert feature distillation. ‘DS’ denotes Driving Score, ‘SR’ denotes Success Rate, ‘Effi’ denotes Efficiency, and ‘Comf’ denotes Comfortness.
Method
Venue
NC ↑
DAC ↑
TTC ↑
Comf. ↑
EP ↑
PDMS ↑
TransFuser ( Chitta et al. 2023 )
TPAMI’22
97.7
92.8
92.8
100
79.2
84.0
VADv2 ( Jiang et al. 2026 )
ICLR’26
97.2
89.1
91.6
100
76.0
80.9
GoalFlow ( Xing et al. 2025 )
CVPR’25
98.3
93.8
94.3
100
79.8
85.7
Hydra-MDP ( Li et al. 2024 )
Arxiv’24
98.3
96.0
94.6
100
78.7
86.5
FUMP ( Liu et al. 2025a )
Arxiv’25
98.1
96.2
94.2
100
82.0
87.8
DiffusionDrive ( Liao et al. 2025 )
CVPR’25
98.2
96.2
94.7
100
82.2
88.1
Table 3: Comparison on planning-oriented NAVSIMv1 navtest split with Closed-Loop metrics.
Method
Stage
NC ↑
DAC ↑
DDC ↑
TL ↑
EP ↑
TTC ↑
LK ↑
HC ↑
EC ↑
EPDMS ↑
TransFuser
Stage1
96.2
79.5
99.1
99.5
84.1
95.1
94.2
97.5
79.1
23.1
Stage2
77.7
70.2
84.2
98.0
85.1
75.6
45.4
95.7
75.9
DiffusionDrive
Stage1
96.0
79.7
97.4
99.5
81.3
93.1
90.8
96.8
73.8
24.2
Stage2
82.1
72.2
88.5
98.7
85.1
78.8
49.2
89.3
71.2
GuideFlow
Stage1
96.6
80.5
96.3
99.3
82.3
94.9
91.5
97.7
67.8
27.1
Stage2
87.3
76.7
88.8
99.2
84.3
85.1
49.7
93.1
44.5
Table 4: Comparison with SOTA methods on the NAVSIMv2 navhard split ( Cao et al. 2025 ) .
Method
NC ↑
DAC ↑
DDC ↑
TL ↑
EP ↑
TTC ↑
LK ↑
HC ↑
EC ↑
EPDMS ↑
TransFuser ( Chitta et al. 2023 )
96.9
89.9
97.8
99.7
87.1
95.4
92.7
98.3
87.2
76.7
DiffusionDrive ( Liao et al. 2025 )
98.2
95.9
99.4
99.8
87.5
97.3
96.8
98.3
87.7
84.5
Hydra-MDP++ ( Li et al. 2025b )
97.2
97.5
99.4
99.6
83.1
96.5
94.4
98.2
70.9
81.4
DriveSuprim ( Yao et al. 2026 )
97.5
96.5
99.4
99.6
88.4
96.6
95.5
98.3
77.0
83.1
DiffusionDriveV2 ( Zou et al. 2025 )
97.7
96.6
99.2
99.8
88.9
97.2
96.0
97.8
91.0
87.5
DriveWorld-VLA ( Liu et al. 2026 )
98.6
99.1
99.6
99.8
87.4
97.9
97.0
97.8
78.6
86.8
Table 5: Comparison with SOTA methods on the NAVSIMv2 navtest split ( Cao et al. 2025 ) .
SSM-Q
FM-Ref
NC ↑
DAC ↑
TTC ↑
Comf. ↑
EP ↑
PDMS ↑
97.7
92.8
92.8
100
79.2
84.0
✓
98.2
94.5
94.0
100
82.1
86.9
✓
✓
98.7
96.8
95.4
100
84.5
89.9
Table 6: Ablation study of the two core modules on the NAVSIMv1 navtest split. SSM-Q denotes the Selective State-Space Planning Memory Query Module, and FM-Ref denotes the Flow-Matching Trajectory Residual Refiner.
Method
Memory Strategy
Col. Rate (%) ↓
1s
2s
3s
4s
5s
6s
Avg.
MomAD
Historical query reuse
0.17
0.30
0.54
0.83
1.43
2.13
0.90
MomADv2
No historical memory
0.16
0.29
0.52
0.86
1.48
2.18
0.92
MomADv2
Naive historical query fusion
0.18
0.34
0.60
0.94
1.62
2.35
1.01
MomADv2
Command-aware memory
0.12
0.25
0.47
0.82
1.46
2.19
0.89
MomADv2
+ Temporal-continuity filtering
0.08
0.21
0.41
0.78
1.39
2.12
0.83
Table 7: Ablation study of reliable temporal memory on the nuScenes validation set. All MomADv2 variants keep the FM-Ref and only differ in the historical memory strategy.
K
NC ↑
DAC ↑
TTC ↑
Comf. ↑
EP ↑
PDMS ↑
1
98.4
95.4
94.7
100
82.4
88.2
2
98.7
96.3
95.0
100
82.8
89.1
4
98.7
96.8
95.4
100
84.5
89.9
6
98.5
95.7
94.6
100
82.3
88.5
8
98.1
95.0
94.1
100
81.9
87.7
Table 8: Ablation study of historical memory length K on the NAVSIMv1 navtest split. K denotes the number of historical planning states used in the SSM-Q.
Figure 4: Visualization of MomADv2 on nuScenes dataset.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Multi-Ability(%)↑
Merging↑
Overtaking↑
EmergencyBrake↑
GiveWay↑
TrafficSign↑
Mean↑
DriveAdapter* ( Jia et al. 2023 )
28.82
26.38
48.76
50.00
56.43
42.08
DriveDPO ( Shang et al. 2025 )
-
-
-
-
-
-
Raw2Drive ( Yang et al. 2025 )
43.35
51.11
60.00
50.00
62.26
53.34
DriveTrans* ( Jia et al. 2025 )
17.57
35.00
48.36
40.00
52.10
38.60
WoTE* ( Li et al. 2025e )
-
-
-
-
-
-
Appendix
Table 9: Multi-Ability results on Bench2Drive (V0.0.3) under base training set. ‘mmt’ refers multi-mode trajectory variant of VAD and † denotes the re-implementation. * denotes expert feature distillation.
Method
Venue
TPC (m) ↓
4s
5s
6s
UniAD ( Hu et al. 2023 )
CVPR’23
1.49
1.81
2.41
VAD ( Jiang et al. 2023 )
ICCV’23
1.55
1.73
2.17
SparseDrive ( Sun et al. 2024 )
ICRA’25
1.33
1.66
1.99
MomAD ( Song et al. 2025 )
CVPR’25
1.19
1.45
1.61
MomADv2 (Ours)
-
1.04
1.32
1.61
Appendix
Table 10: Comparison of trajectory prediction consistency on the nuScenes validation set. TPC denotes Trajectory Prediction Consistency, where lower values indicate better temporal stability.
Method
L2 (m)↓
Col. Rate (%)↓
1s
2s
3s
Avg.
1s
2s
3s
Avg.
UniAD ( Hu et al. 2023 )
0.48
0.96
1.65
1.03
0.10
0.15
0.61
0.29
VAD ( Jiang et al. 2023 )
0.41
0.70
1.05
0.72
0.11
0.24
0.42
0.26
DiffusionDrive ( Liao et al. 2025 )
0.29
0.58
0.96
0.61
0.02
0.05
0.22
0.09
DIVER ( Song et al. 2026b )
–
–
–
–
0.01
0.05
0.15
0.07
FocalAD ( Sun et al. 2026 )
0.27
0.57
0.96
0.60
0.00
0.04
0.24
0.09
Appendix
Table 11: Planning results for 3-second short-horizon planning on the nuScenes validation set.
L2 (m) ↓
Col. Rate (%) ↓
SSM-Q
FM-Ref
4s
5s
6s
4s
5s
6s
1.67
1.98
2.45
0.87
1.54
2.33
✓
1.33
1.84
2.43
0.80
1.44
2.12
✓
✓
1.32
1.82
2.40
0.72
1.33
2.03
Appendix
Table 12: Ablation study of the two core modules on the nuScenes validation set under open-loop evaluation.
Temporal Modeling
Avg. L2 ↓
Avg. Col. ↓
EPDMS ↑
FPS ↑
Concat + MLP
1.52
1.26
84.2
51.0
GRU
1.41
0.77
86.4
41.9
LSTM
1.34
0.80
86.7
39.6
Transformer
1.31
0.82
87.2
33.0
Vanilla SSM
1.25
0.81
87.0
43.6
Selective SSM / Mamba
1.21
0.76
87.9
43.0
Appendix
Table 13: Ablation study of temporal memory modeling methods. Avg. L2 and Avg. Col. are evaluated on nuScenes , and EPDMS is evaluated on NAVSIMv2 navtest .
Solver
NC ↑
DAC ↑
TTC ↑
Comf. ↑
EP ↑
PDMS ↑
FPS ↑
Euler
99.0
97.1
95.5
100
83.2
89.9
43
Heun
98.5
96.7
95.2
100
82.9
89.0
38
Appendix
Table 14: Ablation on solver choice on NAVSIMv2 navtest .
Figure 5: Qualitative visualization of MomADv2 on NAVSIM across diverse driving scenarios. Compared with GuideFlow, MomADv2 generates smoother and more temporally consistent trajectories, better adheres to lane geometry and drivable-area constraints, and enables safer interactions with surrounding traffic participants.
Figure 6: Qualitative visualization of 6-second long-horizon planning by MomADv2 on nuScenes.
Long-horizon planning enables autonomous vehicles to anticipate scene evolution and potential risks, supporting safe and stable decisions in complex interactions. However, existing methods struggle to propagate motion trends from observed history into the future. Long rollouts based on a single latent state may further attenuate useful dynamics, retain stale motion patterns, and disrupt reliable near-term plans. We introduce MomWorld, a momentum-aware latent world model for long-horizon planning. MomWorld extracts scene motion trends from historical-to-current observations and propagates latent momentum into future horizons, jointly predicting future configuration and momentum states. A learnable momentum persistence mechanism preserves stable trends, scene-conditioned momentum updates adapt future dynamics, and a scene-adaptive reset gate suppresses stale momentum under abrupt changes. We further propose MoFlow, a momentum-conditioned flow-matching module that refines a base trajectory to align with the predicted future scene evolution in only a few integration steps, with a horizon-aware residual fusion that preserves near-term planning stability while permitting stronger long-range corrections. Extensive experiments on NAVSIM, nuScenes and Bench2Drive demonstrate that MomWorld improves long-horizon planning consistency and reduces the average collision rate by 12.2% relative to MomAD over a 6-second planning horizon.
Ziying Song, Shengkai Zhang, Lei Yang +6
Nanyang Technological University · Beijing Jiaotong University · North University of China +2
End-to-end autonomous driving has made significant progress by unifying perception, prediction, and planning within a single learning framework, achieving strong performance in short-horizon decision making. However, most existing E2E-AD methods remain confined to short-horizon planning and lack the ability to model long-term temporal dependencies, which severely limits their generalization and security in complex and highly interactive driving scenarios. In this work, we propose GraphWorld, an E2E-AD framework that explicitly enhances long-horizon planning through latent world modeling. We introduce an Ego-Centric Interaction Graph, which adaptively models critical neighboring agents based on spatial proximity, and propagates relational context to planning queries via cross-node cross-attention. We present a World-State-Conditioned Planning that learns ego-centric latent world representations by modeling interactions between an ego vehicle and surrounding agents. This latent world state captures key interaction dynamics and safety-relevant semantics, and serves as a conditioning signal to guide long-horizon, safety-aware trajectory planning. Extensive experiments on Bench2Drive, NAVSIMv1/2, and nuScenes demonstrate that GraphWorld significantly reduces collision rates and improves long-horizon planning performance, validating its effectiveness in complex driving environments.
Ziying Song, Caiyan Jia, Lin Liu +8
Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence, School of Computer Science and Technology, Beijing Jiaotong University, Beijing, China
Practical autonomous driving requires models that generalize by reasoning through spatial-temporal possibilities to exclude unsafe outcomes. While state-of-the-art (SOTA) methods use parallel planning architectures, they fail to explicitly couple speed decisions with agent behavior along the driving path, leading to suboptimal coordination. To address this, we propose a cascaded framework that transforms longitudinal planning from an independent prediction task into a path-conditioned reasoning process. On the model side, we introduce an anchor-based regression design that conditions longitudinal prediction on the lateral drive path, and reformulate longitudinal planning as 1D displacement prediction along the path. This reduces geometric uncertainty and sharpens the model's focus on interaction-driven dynamics. On the data side, we introduce a planning-oriented data augmentation strategy that simulates rare safety-critical events by programmatically inserting agents and relabeling longitudinal targets to enforce collision avoidance. Evaluated on the challenging Bench2Drive benchmark, our method achieves SOTA performance with a driving score of 89.07 and a success rate of 73.18%, demonstrating significantly improved coordination and safety. Further evaluation on Fail2Drive confirms strong generalization to rare edge cases where parallel formulations typically fail. Project page:https://yanhaowu.github.io/AlignDrive/.
Yanhao Wu, Haoyang Zhang, Fei He +6
School of Software Engineering, XJTU · Horizon Robotics · Shenzhen Loop Area Institute +1