MomWorld: Momentum-Aware Latent World Model for Long-Horizon Autonomous Driving
Authors: Ziying Song, Shengkai Zhang, Lei Yang, Haozhuang Chi, Yuchen Liu, Jiangtao Su, Lin Liu, Ziyang Liu, +1 more
Organizations: Nanyang Technological University · Beijing Jiaotong University · North University of China · Dalian University of Technology · Tsinghua University
Long-horizon planning enables autonomous vehicles to anticipate scene evolution and potential risks, supporting safe and stable decisions in complex interactions. However, existing methods struggle to propagate motion trends from observed history into the future. Long rollouts based on a single latent state may further attenuate useful dynamics, retain stale motion patterns, and disrupt reliable near-term plans. We introduce MomWorld, a momentum-aware latent world model for long-horizon planning. MomWorld extracts scene motion trends from historical-to-current observations and propagates latent momentum into future horizons, jointly predicting future configuration and momentum states. A learnable momentum persistence mechanism preserves stable trends, scene-conditioned momentum updates adapt future dynamics, and a scene-adaptive reset gate suppresses stale momentum under abrupt changes. We further propose MoFlow, a momentum-conditioned flow-matching module that refines a base trajectory to align with the predicted future scene evolution in only a few integration steps, with a horizon-aware residual fusion that preserves near-term planning stability while permitting stronger long-range corrections. Extensive experiments on NAVSIM, nuScenes and Bench2Drive demonstrate that MomWorld improves long-horizon planning consistency and reduces the average collision rate by 12.2% relative to MomAD over a 6-second planning horizon.
Figures & tables
Figure 1: Motivation and comparison of MomWorld. (a) Reliable future scene evolution remains a key challenge for long-horizon planning. (b) MomAD ( Song et al., 2025 ) derives ego-centric momentum from historical and current evidence to stabilize ego dynamics. (c) MomWorld elevates momentum to a world-level latent state, initializes it from past-to-present evidence, and jointly rolls out future configuration and momentum to model world dynamics. (d) On six-second nuScenes planning, MomWorld reduces L2 from 2.45 to 2.31 m and collision rate from 2.13% to 1.97%. The corresponding averages improve from 1.42 to 1.17 m and from 0.90% to 0.79%.
Figure 2: Overview of MomWorld . Historical and current multi-view images are encoded into Scene Queries, from which MoLWM initializes latent scene state and momentum and propagates them through Latent World Rollout (LWR) to construct Future World Memory for candidate scoring and Plan selection. Guided by history-conditioned momentum and predicted future evolution, MoFlow applies bounded residual refinement to produce the temporally consistent Refined trajectory τ .
Figure 3
Figure 3: Momentum-Conditioned Flow Matching (MoFlow). Conditioned on p0 and Mt , MoFlow refines the base trajectory through residual flow and bounded horizon-aware fusion.
Method
Venue
L2 (m)↓
Col. Rate (%)↓
FPS
1s
2s
3s
4s
5s
6s
Avg.
1s
2s
3s
4s
5s
6s
Avg.
UniAD ( Hu et al., 2023 )
CVPR’23
0.47
0.91
1.35
1.91
2.47
3.07
1.70
0.25
0.36
0.61
0.99
1.64
2.51
1.06
1.8
SparseDrive ( Sun et al., 2025b )
ICRA’25
0.43
0.87
1.23
1.75
2.32
2.95
1.59
0.19
0.31
0.56
0.87
1.54
2.33
0.97
9.0
MomAD ( Song et al., 2025 )
CVPR’25
0.41
0.85
1.13
1.67
1.98
2.45
1.42
0.17
0.30
0.54
0.83
1.43
2.13
0.90
7.8
LAW ( Li et al., 2025b )
ICLR’25
0.40
0.87
1.16
1.71
2.03
2.61
1.46
0.19
0.33
0.57
0.86
1.51
2.31
0.96
19.5
Epona ( Zhang et al., 2025b )
ICCV’25
0.39
0.91
1.17
1.73
2.02
2.75
1.50
0.14
0.18
0.45
0.74
1.48
2.23
0.87
–
Table 1: Six-second planning results on the nuScenes validation set. Reported FPS uses an A100 for UniAD, an RTX 3090 for LAW, and an RTX 4090 for SparseDrive, MomAD, and GuideFlow.
Method
Venue
TPC (m)↓
4s
5s
6s
Avg.
UniAD ( Hu et al., 2023 )
CVPR’23
1.49
1.81
2.41
1.90
VAD ( Jiang et al., 2023 )
ICCV’23
1.55
1.73
2.17
1.82
SparseDrive ( Sun et al., 2025b )
ICRA’25
1.33
1.66
1.99
1.66
MomAD ( Song et al., 2025 )
CVPR’25
1.19
1.45
1.61
1.42
MomWorld (Ours)
–
0.93
1.19
1.46
1.19
Table 2: Trajectory Prediction Consistency at 4–6 s on the nuScenes validation set.
Table 7
MoLWM
MoFlow
nuScenes 6s
NAVSIM v2 navhard
FWM
LWR
SMD
MGRF
HARF
L2@3(m)↓
L2@6(m)↓
TPC@6(m)↓
Col.@3(%)↓
Col.@6(%)↓
LK↑
EP↑
EC↑
EPDMS↑
1.13
2.45
1.61
0.54
2.13
54.6
69.5
49.7
41.7
✓
1.07
2.42
1.57
0.51
2.09
54.8
72.3
50.8
42.0
✓
✓
1.01
2.39
1.54
0.48
2.06
55.0
75.6
52.1
42.2
✓
✓
✓
0.96
2.36
1.51
0.46
2.03
55.1
78.4
53.0
42.4
✓
✓
✓
✓
0.91
2.33
1.48
0.44
2.00
55.2
80.2
54.0
42.6
Table 6: Roles of Different Method Components in MomWorld. Cumulative component study on the nuScenes validation set and the NAVSIM v2 navhard split. FWM, LWR, SMD, MGRF, and HARF denote Future World Memory , Latent World Rollout , Scene-Adaptive Momentum Dynamics , Momentum-Guided Residual Flow , and Horizon-Aware Residual Fusion , respectively. LK, EP, and EC are the Stage-2 metrics reported by the navhard protocol.
Variant
nuScenes 6s
NAVSIM v1 navtest
NAVSIM v2 navhard
L2@3(m)↓
L2@6(m)↓
TPC@6(m)↓
Col.@3(%)↓
Col.@6(%)↓
PDMS↑
EPDMS↑
No historical variation
0.95
2.58
1.72
0.55
2.40
88.6
39.7
No horizon embedding in SMD
0.90
2.43
1.57
0.47
2.12
89.5
41.3
Fixed retention ( ρk≡0.9 )
0.92
2.49
1.63
0.50
2.21
89.0
40.8
No momentum proposal ( uk≡0 )
0.99
2.69
1.82
0.60
2.58
87.8
38.6
No reset gate ( gk≡0 )
0.94
2.55
1.69
0.65
2.83
88.2
37.9
Table 7: Ablation study of different designs in MoLWM across nuScenes , NAVSIM v1 navtest , and NAVSIM v2 navhard .
MoFlow Design
Steps
nuScenes 6s
NAVSIM v1 navtest
NAVSIM v2 navhard
Latency (ms)↓
L2@3(m)↓
L2@6(m)↓
TPC@6(m)↓
Col.@3(%)↓
Col.@6(%)↓
PDMS↑
EPDMS↑
Flow
Total
No MoFlow (base plan τ0 )
0
0.96
2.36
1.51
0.46
2.03
89.7
42.4
0.0
130.5
Unconditioned MGRF
4
1.01
2.53
1.67
0.57
2.38
88.3
40.0
8.4
138.9
Mt only
4
0.91
2.40
1.55
0.47
2.09
89.4
41.8
8.4
138.9
p0 only
4
0.93
2.43
1.58
0.49
2.14
89.2
41.5
8.4
138.9
p0,Mt
1
0.91
2.38
1.53
0.46
2.07
89.8
42.3
2.2
132.7
Table 8: Ablation study of different designs in MoFlow across nuScenes , NAVSIM v1 navtest , and NAVSIM v2 navhard .
Figure 4: Consecutive planning visualization on nuScenes. Comparison of MomAD ( Song et al., 2025 ) and MomWorld across consecutive frames. MomWorld yields more accurate and temporally consistent future trajectories.
Configuration
Value
Model input history
2 camera frames / 4 LiDAR sweeps
Image resolution
2048×512
Image backbone
V-99-eSE VoVNet
Planner latent / FFN width
256 / 1024
Planner attention layers / heads
3 / 8
Candidate vocabulary size
16,384
Table 9: NAVSIM implementation configuration of MomWorld.
Figure 5: Qualitative 6-second planning results on nuScenes. MomWorld generates smooth trajectories when transitioning from turns to straight driving, decelerating in dense traffic, turning left at intersections, and avoiding vehicles ahead.
Method
Venue
Closed-loop Performance
Multi-Ability (%) ↑
DS ↑
SR (%) ↑
Effi. ↑
Comf. ↑
Merge
Overtake
Emergency Brake
Give Way
Traffic Sign
Mean
TCP-traj ∗ ( Wu et al., 2022 )
NeurIPS’22
59.90
30.00
76.54
18.08
12.50
22.73
52.72
40.00
46.63
34.92
UniAD ( Hu et al., 2023 )
CVPR’23
45.81
16.36
129.21
43.58
14.10
17.78
21.67
10.00
14.21
15.55
ThinkTwice ∗ ( Jia et al., 2023b )
CVPR’23
62.44
31.23
69.33
16.22
13.72
22.93
52.99
50.00
47.78
37.48
DriveAdapter ∗ ( Jia et al., 2023a )
ICCV’23
64.22
33.08
70.22
16.01
14.55
22.61
54.04
50.00
50.45
38.33
VAD ( Jiang et al., 2023 )
ICCV’23
42.35
15.00
157.94
46.01
8.11
24.44
18.64
20.00
19.15
18.07
Table 10: Closed-loop planning and multi-ability performance on Bench2Drive . ∗ denotes expert feature distillation. Higher is better for all metrics.
Figure 6: Qualitative planning results on NAVSIM. Comparison of GuideFlow ( Liu et al., 2026a ) and MomWorld across turning, straight driving, braking, stable driving, and vehicle avoidance. MomWorld produces smoother trajectories with improved roadway compliance and safer interactions.
Method
Turning-nuScenes
Adv-nuSc
Avg. L2 (m) ↓
Col. Rate (%) ↓
Col. Rate (%) ↓
1s
2s
3s
Avg.
1s
2s
3s
Avg.
UniAD
–
–
–
–
–
0.800
4.100
6.960
3.950
VAD
–
–
–
–
–
4.460
7.590
9.080
7.050
SparseDrive
0.86
0.04
0.17
0.98
0.40
0.029
0.618
2.430
1.026
DiffusionDrive ∗
–
0.03
0.14
0.85
0.34
0.068
1.299
3.646
1.671
Table 11: Open-loop robustness on Turning-nuScenes ( Song et al., 2025 ) and Adv-nuSc ( Xu et al., 2025 ) . Turning-nuScenes additionally reports average L2 where available. Lower is better.
Method
Snow
Rain
Fog
1s
2s
3s
1s
2s
3s
1s
2s
3s
SparseDrive
0.13
0.27
0.50
0.11
0.27
0.55
0.14
0.36
0.58
DiffusionDrive ∗
0.09
0.24
0.39
0.07
0.18
0.35
0.06
0.18
0.30
MomAD
0.08
0.16
0.30
0.06
0.17
0.31
0.06
0.19
0.32
DIVER
0.07
0.13
0.25
0.05
0.16
0.27
0.04
0.16
0.25
GraphWorld
0.07
0.15
0.28
0.05
0.15
0.29
0.05
0.17
0.29
Table 12: Collision rate under three adverse-weather corruptions on nuScenes-C ( Dong et al., 2023 ) . Lower is better.
Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence, School of Computer Science and Technology, Beijing Jiaotong University, Beijing, China