MomWorld: Momentum-Aware Latent World Model for Long-Horizon Autonomous Driving
Authors: Ziying Song, Shengkai Zhang, Lei Yang, Haozhuang Chi, Yuchen Liu, Jiangtao Su, Lin Liu, Ziyang Liu, +1 more
Organizations: Nanyang Technological University · Beijing Jiaotong University · North University of China · Dalian University of Technology · Tsinghua University
Long-horizon planning enables autonomous vehicles to anticipate scene evolution and potential risks, supporting safe and stable decisions in complex interactions. However, existing methods struggle to propagate motion trends from observed history into the future. Long rollouts based on a single latent state may further attenuate useful dynamics, retain stale motion patterns, and disrupt reliable near-term plans. We introduce MomWorld, a momentum-aware latent world model for long-horizon planning. MomWorld extracts scene motion trends from historical-to-current observations and propagates latent momentum into future horizons, jointly predicting future configuration and momentum states. A learnable momentum persistence mechanism preserves stable trends, scene-conditioned momentum updates adapt future dynamics, and a scene-adaptive reset gate suppresses stale momentum under abrupt changes. We further propose MoFlow, a momentum-conditioned flow-matching module that refines a base trajectory to align with the predicted future scene evolution in only a few integration steps, with a horizon-aware residual fusion that preserves near-term planning stability while permitting stronger long-range corrections. Extensive experiments on NAVSIM, nuScenes and Bench2Drive demonstrate that MomWorld improves long-horizon planning consistency and reduces the average collision rate by 12.2% relative to MomAD over a 6-second planning horizon.
Figures & tables
Figure 1: Motivation and comparison of MomWorld. (a) Reliable future scene evolution remains a key challenge for long-horizon planning. (b) MomAD ( Song et al., 2025 ) derives ego-centric momentum from historical and current evidence to stabilize ego dynamics. (c) MomWorld elevates momentum to a world-level latent state, initializes it from past-to-present evidence, and jointly rolls out future configuration and momentum to model world dynamics. (d) On six-second nuScenes planning, MomWorld reduces L2 from 2.45 to 2.31 m and collision rate from 2.13% to 1.97%. The corresponding averages improve from 1.42 to 1.17 m and from 0.90% to 0.79%.
Figure 2: Overview of MomWorld . Historical and current multi-view images are encoded into Scene Queries, from which MoLWM initializes latent scene state and momentum and propagates them through Latent World Rollout (LWR) to construct Future World Memory for candidate scoring and Plan selection. Guided by history-conditioned momentum and predicted future evolution, MoFlow applies bounded residual refinement to produce the temporally consistent Refined trajectory τ .
Figure 3
Figure 3: Momentum-Conditioned Flow Matching (MoFlow). Conditioned on p0 and Mt , MoFlow refines the base trajectory through residual flow and bounded horizon-aware fusion.
Method
Venue
L2 (m)↓
Col. Rate (%)↓
FPS
1s
2s
3s
4s
5s
6s
Avg.
1s
2s
3s
4s
5s
6s
Avg.
UniAD ( Hu et al., 2023 )
CVPR’23
0.47
0.91
1.35
1.91
2.47
3.07
1.70
0.25
0.36
0.61
0.99
1.64
2.51
1.06
1.8
SparseDrive ( Sun et al., 2025b )
ICRA’25
0.43
0.87
1.23
1.75
2.32
2.95
1.59
0.19
0.31
0.56
0.87
1.54
2.33
0.97
9.0
MomAD ( Song et al., 2025 )
CVPR’25
0.41
0.85
1.13
1.67
1.98
2.45
1.42
0.17
0.30
0.54
0.83
1.43
2.13
0.90
7.8
LAW ( Li et al., 2025b )
ICLR’25
0.40
0.87
1.16
1.71
2.03
2.61
1.46
0.19
0.33
0.57
0.86
1.51
2.31
0.96
19.5
Epona ( Zhang et al., 2025b )
ICCV’25
0.39
0.91
1.17
1.73
2.02
2.75
1.50
0.14
0.18
0.45
0.74
1.48
2.23
0.87
–
Table 1: Six-second planning results on the nuScenes validation set. Reported FPS uses an A100 for UniAD, an RTX 3090 for LAW, and an RTX 4090 for SparseDrive, MomAD, and GuideFlow.
Method
Venue
TPC (m)↓
4s
5s
6s
Avg.
UniAD ( Hu et al., 2023 )
CVPR’23
1.49
1.81
2.41
1.90
VAD ( Jiang et al., 2023 )
ICCV’23
1.55
1.73
2.17
1.82
SparseDrive ( Sun et al., 2025b )
ICRA’25
1.33
1.66
1.99
1.66
MomAD ( Song et al., 2025 )
CVPR’25
1.19
1.45
1.61
1.42
MomWorld (Ours)
–
0.93
1.19
1.46
1.19
Table 2: Trajectory Prediction Consistency at 4–6 s on the nuScenes validation set.
Table 7
MoLWM
MoFlow
nuScenes 6s
NAVSIM v2 navhard
FWM
LWR
SMD
MGRF
HARF
L2@3(m)↓
L2@6(m)↓
TPC@6(m)↓
Col.@3(%)↓
Col.@6(%)↓
LK↑
EP↑
EC↑
EPDMS↑
1.13
2.45
1.61
0.54
2.13
54.6
69.5
49.7
41.7
✓
1.07
2.42
1.57
0.51
2.09
54.8
72.3
50.8
42.0
✓
✓
1.01
2.39
1.54
0.48
2.06
55.0
75.6
52.1
42.2
✓
✓
✓
0.96
2.36
1.51
0.46
2.03
55.1
78.4
53.0
42.4
✓
✓
✓
✓
0.91
2.33
1.48
0.44
2.00
55.2
80.2
54.0
42.6
Table 6: Roles of Different Method Components in MomWorld. Cumulative component study on the nuScenes validation set and the NAVSIM v2 navhard split. FWM, LWR, SMD, MGRF, and HARF denote Future World Memory , Latent World Rollout , Scene-Adaptive Momentum Dynamics , Momentum-Guided Residual Flow , and Horizon-Aware Residual Fusion , respectively. LK, EP, and EC are the Stage-2 metrics reported by the navhard protocol.
Variant
nuScenes 6s
NAVSIM v1 navtest
NAVSIM v2 navhard
L2@3(m)↓
L2@6(m)↓
TPC@6(m)↓
Col.@3(%)↓
Col.@6(%)↓
PDMS↑
EPDMS↑
No historical variation
0.95
2.58
1.72
0.55
2.40
88.6
39.7
No horizon embedding in SMD
0.90
2.43
1.57
0.47
2.12
89.5
41.3
Fixed retention ( ρk≡0.9 )
0.92
2.49
1.63
0.50
2.21
89.0
40.8
No momentum proposal ( uk≡0 )
0.99
2.69
1.82
0.60
2.58
87.8
38.6
No reset gate ( gk≡0 )
0.94
2.55
1.69
0.65
2.83
88.2
37.9
Table 7: Ablation study of different designs in MoLWM across nuScenes , NAVSIM v1 navtest , and NAVSIM v2 navhard .
MoFlow Design
Steps
nuScenes 6s
NAVSIM v1 navtest
NAVSIM v2 navhard
Latency (ms)↓
L2@3(m)↓
L2@6(m)↓
TPC@6(m)↓
Col.@3(%)↓
Col.@6(%)↓
PDMS↑
EPDMS↑
Flow
Total
No MoFlow (base plan τ0 )
0
0.96
2.36
1.51
0.46
2.03
89.7
42.4
0.0
130.5
Unconditioned MGRF
4
1.01
2.53
1.67
0.57
2.38
88.3
40.0
8.4
138.9
Mt only
4
0.91
2.40
1.55
0.47
2.09
89.4
41.8
8.4
138.9
p0 only
4
0.93
2.43
1.58
0.49
2.14
89.2
41.5
8.4
138.9
p0,Mt
1
0.91
2.38
1.53
0.46
2.07
89.8
42.3
2.2
132.7
Table 8: Ablation study of different designs in MoFlow across nuScenes , NAVSIM v1 navtest , and NAVSIM v2 navhard .
Figure 4: Consecutive planning visualization on nuScenes. Comparison of MomAD ( Song et al., 2025 ) and MomWorld across consecutive frames. MomWorld yields more accurate and temporally consistent future trajectories.
Configuration
Value
Model input history
2 camera frames / 4 LiDAR sweeps
Image resolution
2048×512
Image backbone
V-99-eSE VoVNet
Planner latent / FFN width
256 / 1024
Planner attention layers / heads
3 / 8
Candidate vocabulary size
16,384
Table 9: NAVSIM implementation configuration of MomWorld.
Figure 5: Qualitative 6-second planning results on nuScenes. MomWorld generates smooth trajectories when transitioning from turns to straight driving, decelerating in dense traffic, turning left at intersections, and avoiding vehicles ahead.
Method
Venue
Closed-loop Performance
Multi-Ability (%) ↑
DS ↑
SR (%) ↑
Effi. ↑
Comf. ↑
Merge
Overtake
Emergency Brake
Give Way
Traffic Sign
Mean
TCP-traj ∗ ( Wu et al., 2022 )
NeurIPS’22
59.90
30.00
76.54
18.08
12.50
22.73
52.72
40.00
46.63
34.92
UniAD ( Hu et al., 2023 )
CVPR’23
45.81
16.36
129.21
43.58
14.10
17.78
21.67
10.00
14.21
15.55
ThinkTwice ∗ ( Jia et al., 2023b )
CVPR’23
62.44
31.23
69.33
16.22
13.72
22.93
52.99
50.00
47.78
37.48
DriveAdapter ∗ ( Jia et al., 2023a )
ICCV’23
64.22
33.08
70.22
16.01
14.55
22.61
54.04
50.00
50.45
38.33
VAD ( Jiang et al., 2023 )
ICCV’23
42.35
15.00
157.94
46.01
8.11
24.44
18.64
20.00
19.15
18.07
Table 10: Closed-loop planning and multi-ability performance on Bench2Drive . ∗ denotes expert feature distillation. Higher is better for all metrics.
Figure 6: Qualitative planning results on NAVSIM. Comparison of GuideFlow ( Liu et al., 2026a ) and MomWorld across turning, straight driving, braking, stable driving, and vehicle avoidance. MomWorld produces smoother trajectories with improved roadway compliance and safer interactions.
Method
Turning-nuScenes
Adv-nuSc
Avg. L2 (m) ↓
Col. Rate (%) ↓
Col. Rate (%) ↓
1s
2s
3s
Avg.
1s
2s
3s
Avg.
UniAD
–
–
–
–
–
0.800
4.100
6.960
3.950
VAD
–
–
–
–
–
4.460
7.590
9.080
7.050
SparseDrive
0.86
0.04
0.17
0.98
0.40
0.029
0.618
2.430
1.026
DiffusionDrive ∗
–
0.03
0.14
0.85
0.34
0.068
1.299
3.646
1.671
Table 11: Open-loop robustness on Turning-nuScenes ( Song et al., 2025 ) and Adv-nuSc ( Xu et al., 2025 ) . Turning-nuScenes additionally reports average L2 where available. Lower is better.
Method
Snow
Rain
Fog
1s
2s
3s
1s
2s
3s
1s
2s
3s
SparseDrive
0.13
0.27
0.50
0.11
0.27
0.55
0.14
0.36
0.58
DiffusionDrive ∗
0.09
0.24
0.39
0.07
0.18
0.35
0.06
0.18
0.30
MomAD
0.08
0.16
0.30
0.06
0.17
0.31
0.06
0.19
0.32
DIVER
0.07
0.13
0.25
0.05
0.16
0.27
0.04
0.16
0.25
GraphWorld
0.07
0.15
0.28
0.05
0.15
0.29
0.05
0.17
0.29
Table 12: Collision rate under three adverse-weather corruptions on nuScenes-C ( Dong et al., 2023 ) . Lower is better.
Long-horizon planning is critical for safe autonomous driving in complex scenarios. Existing methods improve planning continuity with temporal memory, but such memory may become invalid and mislead decisions when the driving command changes. Thus, selectively leveraging useful history while suppressing command-inconsistent memory remains a key challenge. To address this issue, we propose MomADv2, a reliable state-space memory framework for long-horizon end-to-end autonomous driving. At its core, MomADv2 introduces a Selective State-Space Planning Memory Query Module, which filters historical planning queries based on temporal continuity and command consistency, selects planning modes relevant to the current command, and models the evolution of planning intentions through a selective state-space mechanism. To further alleviate local trajectory deviations and error accumulation in long-horizon planning, we design a Flow-Matching Trajectory Residual Refiner. It learns a continuous residual correction field from the refined planning output to the expert trajectory, enabling fine-grained trajectory refinement while preserving the stability of anchor-based planning. Extensive experiments on closed-loop NAVSIM and Bench2Drive, as well as open-loop nuScenes, demonstrate that MomADv2 improves long-horizon planning consistency and reduces the average collision rate by 15.6% over MomAD under 6-second planning.
Ziying Song, Shengkai Zhang, Lin Liu +8
School of Artificial Intelligence (School of Software), Yanshan University · Beijing Jiaotong University · Nanyang Technological University +5
End-to-end autonomous driving has made significant progress by unifying perception, prediction, and planning within a single learning framework, achieving strong performance in short-horizon decision making. However, most existing E2E-AD methods remain confined to short-horizon planning and lack the ability to model long-term temporal dependencies, which severely limits their generalization and security in complex and highly interactive driving scenarios. In this work, we propose GraphWorld, an E2E-AD framework that explicitly enhances long-horizon planning through latent world modeling. We introduce an Ego-Centric Interaction Graph, which adaptively models critical neighboring agents based on spatial proximity, and propagates relational context to planning queries via cross-node cross-attention. We present a World-State-Conditioned Planning that learns ego-centric latent world representations by modeling interactions between an ego vehicle and surrounding agents. This latent world state captures key interaction dynamics and safety-relevant semantics, and serves as a conditioning signal to guide long-horizon, safety-aware trajectory planning. Extensive experiments on Bench2Drive, NAVSIMv1/2, and nuScenes demonstrate that GraphWorld significantly reduces collision rates and improves long-horizon planning performance, validating its effectiveness in complex driving environments.
Ziying Song, Caiyan Jia, Lin Liu +8
Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence, School of Computer Science and Technology, Beijing Jiaotong University, Beijing, China
Existing latent world models for autonomous driving have opened a promising path toward future-aware driving intelligence. However, they typically treat future latent states as prediction targets or auxiliary signals, rather than directly conditioning trajectory planning. This can entangle current and future features in latent space. In this work, we propose DriveFuture, a future-aware latent world modeling framework for autonomous driving that explicitly learns planning-oriented foresight by conditioning the current latent state modeling process on future world states. Specifically, during training, the model first predicts future latent world states from the current latent state and ego action, and then refines the prediction against the ground-truth future latent state via cross-attention. The resulting future-aware latent serves as an explicit condition for a diffusion-based trajectory planner. During inference, DriveFuture conditions on the predicted future latent state instead of the ground-truth future state. DriveFuture achieves SOTA performance on the public NAVSIM benchmarks, reaching \textbf{55.5} EPDMS on NAVSIM-v2 {\textcolor{blue}{\textit{navhard}}}, \textbf{89.9} EPDMS on NAVSIM-v2 {\textcolor{blue}{\textit{navtest}}}, and \textbf{90.7} PDMS on NAVSIM-v1 {\textcolor{blue}{\textit{navtest}}}, respectively. These results suggest that the key to latent world modeling lies not merely in simulating future states, but more importantly in conditioning current decision-making on future states. Notably, as of April 2026, DriveFuture ranks \textbf{1st} on the \href{https://huggingface.co/spaces/AGC2025/e2e-driving-navhard}{NAVSIM-v2 {\textcolor{blue}{\textit{navhard}}}} leaderboard and achieves SOTA performance on \href{https://huggingface.co/spaces/AGC2024-P/e2e-driving-navtest}{NAVSIM-v1 {\textcolor{blue}{\textit{navtest}}}}.
Yufeng Hong, Xiaotian Zhou, Yingyan Li +6
Beijing Institute of Technology · Institute of Automation, Chinese Academy of Sciences · Beihang University +5