End-to-end autonomous driving policies are commonly trained via imitation learning on logged demonstrations without observing the consequences of their own actions, leading to causal confusion in closed-loop real-world deployment. To address this issue, reinforcement learning (RL) post-training offers a promising alternative by leveraging world models as interactive training environments to enable future scene generation for policy improvement. Nevertheless, existing approaches either rely on reconstruction-based simulators, offering limited counterfactual interaction, or adopt synthetic simulators to enable long-horizon closed-loop interaction at the cost of a substantial sim-to-real gap. Recently, video world models have exhibited the ability to generate realistic multi-step future rollouts but may not faithfully reflect action conditions, resulting in action-vision mismatch. In this paper, we introduce RoXDrive, a plug-and-play closed-loop RL framework that enables reliable policy optimization by identifying action-faithful world-model rollouts, consisting of two stages: 1) Model pre-training: In addition to imitation-based policy pre-training, we devise an Action-Vision Faithfulness Evaluator for inverse dynamics estimation with our geometry-aware auxiliary trajectory supervision, enabling long-horizon assessment of whether visual dynamics faithfully reflect the conditioning ego actions. 2) Action-faithful RL post-training: Agents iteratively interact with world models to form long-horizon scene rollouts, retaining only action-faithful ones for dense safety-aware scoring and scene-level closed-loop RL post-training. Extensive experiments on nuScenes and an in-house dataset with over 130K training scenarios demonstrate consistent gains across planners, reducing safety violations by 27.6% with DiffusionDrive on nuScenes and 33.7% with Qwen3-VL on the internal data.
Figures & tables
Figure 1: (a) Conventional synthetic and reconstruction-based simulators trade off interactive freedom and real-scene fidelity. (b) Action-conditioned video world models enable realistic long-horizon interaction, but action–vision mismatch may produce unreliable consequences. (c) RoXDrive explicitly measures action-vision faithfulness, retaining reliable episodes to optimize agents.
Figure 2: Naive fine-tuning of powerful Cosmos 3 ( i.e., Cosmos 3 FT) remains insufficient for long-horizon inverse dynamics, e.g., jitter errors (top) and long-horizon trajectory drifts (bottom). Hence, RoXDrive devises geometry-aware auxiliary supervision to handle these issues.
Figure 3: Illustration of RoXDrive. In stage one (top), we train the action-vision faithfulness evaluator with our geometry-aware auxiliary trajectory supervision beyond policy pre-training. In stage two (bottom), agents iteratively interact with a frozen world model to form long-horizon scene rollouts, retaining only action-faithful rollouts for dense safety-aware scoring and closed-loop RL.
Table 4
Method
RL
Open-loop Col. (%) ↓
Closed-loop Evaluation (4,675 4s clips, 2 action steps, based on X-World)
1s
2s
Avg.
Safety
Driving Quality
Driving Score ↑
Obj. Col. ↓
Lane Viol. ↓
Progress ↑
Comfort ↑
Clearance ↑
ST-P3
✗
0.23
0.62
0.43
268
557
0.397
0.425
0.387
0.341
VAD
✗
0.07
0.17
0.12
820
824
0.255
0.681
0.418
0.271
UniAD
✗
0.62
0.58
0.60
348
870
0.408
0.805
0.415
0.430
CLEAR ‡
✓
0.11
0.23
0.17
–
–
–
–
–
–
Table 3: Comparison of open-loop (2s, collision) and closed-loop (4s) model performance on nuScenes. RL indicates reinforcement learning and ‡ denotes the official weights are unavailable.
Figure 4: Qualitative results of RoXDrive avoiding risky behaviors (see also Appendix H ).
Planner (10 actions)
Training
Safety
Driving Quality
DS ↑
Obj. Col. ↓
Lane Viol. ↓
Progress ↑
Comfort ↑
Centering ↑
Gemma-3-4B
Imitation-only
328
171
0.938
0.952
0.521
0.371
+ RoXDrive
198
127
0.865
0.940
0.539
0.540
Qwen2.5-VL-3B
Imitation-only
331
141
0.945
0.911
0.538
0.394
+ RoXDrive
206
107
0.883
0.936
0.547
0.559
Qwen3-VL-2B
Imitation-only
242
162
0.884
0.938
0.527
0.450
Table 4: Closed-loop performance on 1K in-house test scenes (80 frames, 12 Hz) based on VLAs.
Setting
Obj. Col. ↓
Lane Viol. ↓
Progress ↑
Comfort ↑
Clearance ↑
DS ↑
DiffusionDrive (official)
266
772
0.694
0.866
0.392
0.526
∙ RL w. All Rollouts
237
672
0.701
0.881
0.385
0.562
∙ RL w. Cosmos3-FT
224
667
0.706
0.882
0.388
0.568
∙ RL w. AVFE (Ours)
189
562
0.717
0.883
0.380
0.588
Group Size
Obj. Col. ↓
Lane Viol. ↓
Progress ↑
Comfort ↑
Clearance ↑
DS ↑
DiffusionDrive (official)
266
772
0.694
0.866
0.392
0.526
Table 5: Ablations of our AVFE and group size on nuScenes. More ablations are put in Appendix G .
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Qualitative inverse dynamics evaluation on representative nuScenes scenes.
Setting
Obj. Col. ↓
Lane Viol. ↓
Progress ↑
Comfort ↑
Clearance ↑
DS ↑
DiffusionDrive
266
772
0.694
0.866
0.392
0.526
η=0.25
209
624
0.691
0.869
0.382
0.565
η=0.50
205
630
0.707
0.883
0.384
0.575
η=0.75 (default)
189
562
0.717
0.883
0.380
0.588
η=1.25
184
587
0.717
0.882
0.380
0.585
η=1.50
211
627
0.709
0.884
0.381
0.576
Appendix
Table 6: Ablation studies of the filtering threshold η on nuScenes.
Figure 6: Quantitative trends on test scenes during RL based on TransFuser.
Model
ADE Median ↓
ADE Mean ↓
FDE Median ↓
FDE Mean ↓
TransFuser ( Chitta et al., 2022 )
0.060
0.075
0.116
0.146
GTRS ( Li et al., 2025d )
0.052
0.049
0.105
0.104
Appendix
Table 7: Open-loop model performance on 1K test scenes in terms of ADE and FDE.
Model
Collision Scene ↓
Lane Violation ↓
Ego Progress ↑
Comfort ↑
Centering ↑
Driving Score ↑
TransFuser ( Chitta et al., 2022 )
257
178
0.881
0.892
0.540
0.498
+ RoXDrive (Ours)
215
145
0.900
0.932
0.539
0.571
GTRS ( Li et al., 2025d )
264
164
0.930
0.917
0.532
0.522
+ RoXDrive (Ours)
214
126
0.890
0.967
0.551
0.607
Appendix
Table 8: Closed-loop model performance on 1K test scenes with a default frame horizon of 80.
Model
MF
Object Collision
Lane Violation
Progress
Comfort
Centering
Driving Score
GTRS ( Li et al., 2025d )
20
31
19
0.968
0.918
0.564
0.859
+ RoXDrive
31
17
0.944
0.966
0.565
0.866
GTRS
40
73
50
0.952
0.918
0.553
0.791
+ RoXDrive
60
36
0.922
0.966
0.556
0.814
GTRS
60
202
123
0.938
0.917
0.539
0.615
+ RoXDrive
160
97
0.901
0.967
0.555
0.682
Appendix
Table 9: Closed-loop performance of GTRS under different maximum frame horizons (MF).
(a) Reward components Rcol , Rlane , and Rdq denote the collision, lane violation, and driving-quality terms.
TransFuser
Rcol
Rlane
Rdq
Obj. Col. ↓
Lane Viol. ↓
Progress ↑
Comfort ↑
Centering ↑
DS ↑
✓
×
×
×
257
178
0.881
0.892
0.540
0.498
✓
×
×
✓
372
185
0.943
0.860
0.529
0.396
✓
✓
×
✓
241
164
0.911
0.924
0.537
0.545
✓
×
✓
✓
229
156
0.889
0.899
0.539
0.539
✓
✓
✓
✓
215
145
0.900
0.932
0.539
0.571
Appendix
Table 10: Ablation studies on reward components and the horizon lengths on the 1K test scenes.
Figure 7: Qualitative results of RoXDrive avoiding the collision on the in-house dataset.
Figure 8: Qualitative results of RoXDrive avoiding the collision on the in-house dataset.
Figure 9: Qualitative results of RoXDrive avoiding lane violations on the in-house dataset.
Figure 10: Qualitative results of RoXDrive avoiding lane violations on the in-house dataset.
Figure 11: Additional qualitative results of RoXDrive mitigating the causal confusion on nuScenes.
Figure 12: Additional qualitative results of RoXDrive mitigating the causal confusion on nuScenes.
Figure 13: Additional qualitative results of RoXDrive mitigating the causal confusion on nuScenes.