Closed-loop evaluation of end-to-end driving requires continuous rollouts that reveal how earlier decisions affect subsequent driving. However, existing benchmarks evaluate only short segments and fail to capture later consequences. Ambiguous directional commands also obscure the intended navigation objective. We introduce Odyssey, a closed-loop benchmark for long-horizon driving comprising 100 scenarios, each reconstructed from a 100-second nuPlan driving log to preserve the context of navigation maneuvers and traffic interactions. To provide a consistent navigation objective, Odyssey replaces directional commands with explicit standard-definition (SD) map routes that specify which roads to follow, while sensor-based planning determines local driving actions. Throughout these rollouts, diffusion-based refinement of 3DGS-rendered images reduces rendering artifacts along the ego trajectory. To assess how effectively planners follow these routes and prepare for upcoming maneuvers, we introduce SD Route Compliance and Pre-Lane Change Score. These assessments are complemented by RouteDS, which extends the Driving Score with penalties for SD-route deviations and failed lane preparation. We adapt state-of-the-art planners, including vision-language-action (VLA) models, and evaluate their navigation performance using these metrics. Odyssey highlights open questions in route representation and integration for E2E driving. Benchmark code and adapted baselines will be released publicly.
Figures & tables
Figure 1: Odyssey framework. (a) Previous closed-loop benchmarks suffer from limited scenario horizons, ambiguous command-based routing, and artifacts from 3DGS-only rendering. (b) Odyssey evaluates E2E planners over continuous closed-loop rollouts based on 100-second driving sequences with explicit SD-route guidance and diffusion-refined 3DGS observations, assessing sustained safe navigation toward the intended destination.
Figure 2: Limitations of existing evaluation. (a) Short rollouts cannot assess whether a prior lane change enables a later right turn. (b) Inconsistent command flips obscure navigation intent along a continuous trajectory.
Benchmark
Real-log scenarios
Sensor closed-loop
Realistic images
Navigation input
Recorded sequence duration
Rendering type
Traffic-light state
Pedestrian motion
nuScenes
✓
×
✓
Commands
-
×
-
-
NAVSIM
✓
×
✓
Commands
-
×
-
-
NAVSIMv2
✓
×
✓
Commands
-
Offline 3DGS
-
-
WOD-E2E
✓
×
✓
Commands
-
×
-
-
nuPlan
✓
×
×
HD route
15s
×
-
-
Waymax
✓
×
×
HD route
9s
×
-
-
Table 1: Comparison of driving benchmarks and simulation frameworks. ✓ , △ , × , and - denote support, partial support, no documented support, and not applicable, respectively.
Figure 3: Scenario composition and example routes in Odyssey.
Figure 4: Traffic triggering and pre-lane change evaluation. (a) When the simulated ego is slower than the logged ego, fixed-time playback (top) places traffic interactions ahead of it. By activating traffic as the ego approaches a section and keeping subsequent sections inactive, section-based triggering (bottom) aligns traffic interactions with ego progress. (b) Pre-Lane Change requires both driving in route-compatible lanes throughout the final 10m and crossing the stop line in a compatible lane. The four examples illustrate cases that satisfy both conditions, only one, or neither.
Table 6
Figure 5: Effect of diffusion refinement at a novel viewpoint. The first image is the recorded image, and the other three are rendered from a viewpoint shifted 2m to the left of the recorded camera position. GS-only rendering exhibits distorted road textures at this novel viewpoint. Pretrained Fixer improves visual quality but introduces spurious curb-like structures on the road. Fine-tuning on Odyssey suppresses these artifacts and yields smoother roads and clearer lane markings.
Figure 6: Traffic-light modeling. Our deformation field reproduces the red-to-green transition, which OmniRe fails to recover.
Method
SD-Route Guidance
Odyssey Non-reactive
Odyssey Reactive
NAVSIM Open Loop
RouteDS
SDC
PLCA
PLCS
Eff.
Comf.
RouteDS
SDC
PLCA
PLCS
Eff.
Comf.
PDMS
Vision-based planners
LTF
×
16.2
51
66.0
50.3
63.9
100
16.3
54
63.5
49.1
76.1
99
83.8
✓
20.1
95
57.1
56.1
68.2
96
23.0
93
53.6
52.6
78.0
93
85.6
DiffusionDrive
×
25.6
59
78.8
63.7
61.4
100
24.2
59
77.0
63.1
75.7
100
85.8
✓
35.9
100
82.0
80.5
63.4
99
41.0
100
81.2
80.2
71.5
99
86.9
Table 4: Comparison of baseline planners with and without SD-route guidance under non-reactive and reactive traffic. Eff. and Comf. denote Efficiency and Comfortness, respectively. Bold and underlined values indicate the best and second-best results in each metric, respectively.
Figure 7: Driving commands vs. SD-route guidance. With driving commands, SafeDrive deviates from the intended route when the command flips. With SD-route guidance, it follows the intended route.
Configuration
Odyssey Non-reactive
Odyssey Reactive
NAVSIM
RouteDS
SDC
PLCA
PLCS
Eff.
Comf.
RouteDS
SDC
PLCA
PLCS
Eff.
Comf.
PDMS
SafeDrive
Safety Scoring
44.4
82
75.0
69.8
68.3
44
48.9
79
73.8
69.5
81.3
49
90.6
IL Scoring
53.9
95
77.1
74.4
66.5
100
55.3
94
78.0
74.4
75.7
100
89.5
IL Scoring + Safety Filtering
54.4
96
76.7
72.7
65.3
100
56.1
96
77.7
74.1
78.3
100
89.5
ReCogDrive
Table 5: Effects of safety scoring and reinforcement learning on driving performance under non-reactive and reactive traffic and on NAVSIM. All models use SD-map route guidance.
Figure 8: Effects of evaluation horizon and route guidance under non-reactive traffic. (a) RouteDS decreases with longer evaluation horizons, and model rankings change at the highlighted crossings. Solid and dashed curves show the original and IL-based settings, respectively. (b) SD-route guidance improves SDC but reduces PLCA for most planners. (c) RouteDS improves with SD-route guidance, but collision scores decrease for most planners. Higher collision scores indicate fewer collision penalties.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Background
Rigid actors
Densification
Start / end iteration
2,500 / 93,500
2,500 / 93,500
Interval (iterations)
500
500
Views per scoring step
120
120
Normalized error threshold
0.1
0.1
Required score
>5
>5
Appendix
Table 6: Gaussian densification, pruning, and optimization settings.
Loss
Objective
Weight ( λ )
Appearance supervision
Lrgb
RGB L1 after color correction
0.8
Lssim
SSIM before color correction
0.2
Lcolor
Out-of-range color penalty
0.8
Lres
Frame-dependent residual magnitude
0.01
Ltemp
Temporal variation of color residuals
0.1
Appendix
Table 7: Reconstruction losses and weights.
Method
Time (ms/image) ↓
Speedup ↑
Baseline
25.0
1.0×
Optimized
6.7
3.7×
Appendix
Table 8: Fixer inference acceleration.
Figure 9: Effect of section-based triggering on nearby agent density. (a) GT Log shows nearby agent density during recorded driving. (b) Playback follows the original traffic timeline, with density decreasing as the ego slows and simulation continues. (c) Section-based triggering aligns traffic activation with ego progress, maintaining density closer to GT Log across the evaluated speeds and time intervals.
Table 17
Method
SD Route guidance
Non-reactive
Reactive
RouteDS
RC SD
P SD
P col
P off
P TL
P PLC
RouteDS
RC SD
P SD
P col
P off
P TL
P PLC
Vision-based planners
LTF
×
16.24
0.72
0.51
0.54
0.67
0.98
0.88
16.27
0.75
0.54
0.59
0.67
0.98
0.86
✓
20.13
0.99
0.95
0.43
0.63
0.99
0.81
23.01
0.99
0.93
0.48
0.63
0.98
0.80
DiffusionDrive
×
25.64
0.75
0.59
0.61
0.71
0.99
0.92
24.17
0.73
0.59
0.61
0.71
0.98
0.91
✓
35.93
0.96
1.00
0.51
0.76
0.99
0.91
40.97
0.98
1.00
0.57
0.76
0.99
0.91
Appendix
Table 11: Detailed RouteDS results under non-reactive and reactive traffic.
Figure 10: Qualitative Comparison of Fixer Refinement. We show challenging novel views with limited coverage from training observations, where 3DGS-only renderings exhibit pronounced artifacts. Fine-tuned Fixer reduces these artifacts and improves visual detail.
Figure 11: Example trajectories and SD-Routes from Odyssey. The benchmark includes lane changes in preparation for upcoming turns, routes with multiple turns, and long straight routes. These scenarios span varying route complexity and include anticipatory maneuvers that reflect the driver’s navigation intent.
Figure 12: Additional Qualitative Results of Traffic Light Modeling. Our model reproduces red-to-green transitions in sync with the traffic-light schedule in each scene.