End-to-end autonomous driving policies are commonly trained via imitation learning on logged demonstrations without observing the consequences of their own actions, leading to causal confusion in closed-loop real-world deployment. To address this issue, reinforcement learning (RL) post-training offers a promising alternative by leveraging world models as interactive training environments to enable future scene generation for policy improvement. Nevertheless, existing approaches either rely on reconstruction-based simulators, offering limited counterfactual interaction, or adopt synthetic simulators to enable long-horizon closed-loop interaction at the cost of a substantial sim-to-real gap. Recently, video world models have exhibited the ability to generate realistic multi-step future rollouts but may not faithfully reflect action conditions, resulting in action-vision mismatch. In this paper, we introduce RoXDrive, a plug-and-play closed-loop RL framework that enables reliable policy optimization by identifying action-faithful world-model rollouts, consisting of two stages: 1) Model pre-training: In addition to imitation-based policy pre-training, we devise an Action-Vision Faithfulness Evaluator for inverse dynamics estimation with our geometry-aware auxiliary trajectory supervision, enabling long-horizon assessment of whether visual dynamics faithfully reflect the conditioning ego actions. 2) Action-faithful RL post-training: Agents iteratively interact with world models to form long-horizon scene rollouts, retaining only action-faithful ones for dense safety-aware scoring and scene-level closed-loop RL post-training. Extensive experiments on nuScenes and an in-house dataset with over 130K training scenarios demonstrate consistent gains across planners, reducing safety violations by 27.6% with DiffusionDrive on nuScenes and 33.7% with Qwen3-VL on the internal data.
Figures & tables
Figure 1: (a) Conventional synthetic and reconstruction-based simulators trade off interactive freedom and real-scene fidelity. (b) Action-conditioned video world models enable realistic long-horizon interaction, but action–vision mismatch may produce unreliable consequences. (c) RoXDrive explicitly measures action-vision faithfulness, retaining reliable episodes to optimize agents.
Figure 2: Naive fine-tuning of powerful Cosmos 3 ( i.e., Cosmos 3 FT) remains insufficient for long-horizon inverse dynamics, e.g., jitter errors (top) and long-horizon trajectory drifts (bottom). Hence, RoXDrive devises geometry-aware auxiliary supervision to handle these issues.
Figure 3: Illustration of RoXDrive. In stage one (top), we train the action-vision faithfulness evaluator with our geometry-aware auxiliary trajectory supervision beyond policy pre-training. In stage two (bottom), agents iteratively interact with a frozen world model to form long-horizon scene rollouts, retaining only action-faithful rollouts for dense safety-aware scoring and closed-loop RL.
Table 4
Method
RL
Open-loop Col. (%) ↓
Closed-loop Evaluation (4,675 4s clips, 2 action steps, based on X-World)
1s
2s
Avg.
Safety
Driving Quality
Driving Score ↑
Obj. Col. ↓
Lane Viol. ↓
Progress ↑
Comfort ↑
Clearance ↑
ST-P3
✗
0.23
0.62
0.43
268
557
0.397
0.425
0.387
0.341
VAD
✗
0.07
0.17
0.12
820
824
0.255
0.681
0.418
0.271
UniAD
✗
0.62
0.58
0.60
348
870
0.408
0.805
0.415
0.430
CLEAR ‡
✓
0.11
0.23
0.17
–
–
–
–
–
–
Table 3: Comparison of open-loop (2s, collision) and closed-loop (4s) model performance on nuScenes. RL indicates reinforcement learning and ‡ denotes the official weights are unavailable.
Figure 4: Qualitative results of RoXDrive avoiding risky behaviors (see also Appendix H ).
Planner (10 actions)
Training
Safety
Driving Quality
DS ↑
Obj. Col. ↓
Lane Viol. ↓
Progress ↑
Comfort ↑
Centering ↑
Gemma-3-4B
Imitation-only
328
171
0.938
0.952
0.521
0.371
+ RoXDrive
198
127
0.865
0.940
0.539
0.540
Qwen2.5-VL-3B
Imitation-only
331
141
0.945
0.911
0.538
0.394
+ RoXDrive
206
107
0.883
0.936
0.547
0.559
Qwen3-VL-2B
Imitation-only
242
162
0.884
0.938
0.527
0.450
Table 4: Closed-loop performance on 1K in-house test scenes (80 frames, 12 Hz) based on VLAs.
Setting
Obj. Col. ↓
Lane Viol. ↓
Progress ↑
Comfort ↑
Clearance ↑
DS ↑
DiffusionDrive (official)
266
772
0.694
0.866
0.392
0.526
∙ RL w. All Rollouts
237
672
0.701
0.881
0.385
0.562
∙ RL w. Cosmos3-FT
224
667
0.706
0.882
0.388
0.568
∙ RL w. AVFE (Ours)
189
562
0.717
0.883
0.380
0.588
Group Size
Obj. Col. ↓
Lane Viol. ↓
Progress ↑
Comfort ↑
Clearance ↑
DS ↑
DiffusionDrive (official)
266
772
0.694
0.866
0.392
0.526
Table 5: Ablations of our AVFE and group size on nuScenes. More ablations are put in Appendix G .
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Qualitative inverse dynamics evaluation on representative nuScenes scenes.
Setting
Obj. Col. ↓
Lane Viol. ↓
Progress ↑
Comfort ↑
Clearance ↑
DS ↑
DiffusionDrive
266
772
0.694
0.866
0.392
0.526
η=0.25
209
624
0.691
0.869
0.382
0.565
η=0.50
205
630
0.707
0.883
0.384
0.575
η=0.75 (default)
189
562
0.717
0.883
0.380
0.588
η=1.25
184
587
0.717
0.882
0.380
0.585
η=1.50
211
627
0.709
0.884
0.381
0.576
Appendix
Table 6: Ablation studies of the filtering threshold η on nuScenes.
Figure 6: Quantitative trends on test scenes during RL based on TransFuser.
Model
ADE Median ↓
ADE Mean ↓
FDE Median ↓
FDE Mean ↓
TransFuser ( Chitta et al., 2022 )
0.060
0.075
0.116
0.146
GTRS ( Li et al., 2025d )
0.052
0.049
0.105
0.104
Appendix
Table 7: Open-loop model performance on 1K test scenes in terms of ADE and FDE.
Model
Collision Scene ↓
Lane Violation ↓
Ego Progress ↑
Comfort ↑
Centering ↑
Driving Score ↑
TransFuser ( Chitta et al., 2022 )
257
178
0.881
0.892
0.540
0.498
+ RoXDrive (Ours)
215
145
0.900
0.932
0.539
0.571
GTRS ( Li et al., 2025d )
264
164
0.930
0.917
0.532
0.522
+ RoXDrive (Ours)
214
126
0.890
0.967
0.551
0.607
Appendix
Table 8: Closed-loop model performance on 1K test scenes with a default frame horizon of 80.
Model
MF
Object Collision
Lane Violation
Progress
Comfort
Centering
Driving Score
GTRS ( Li et al., 2025d )
20
31
19
0.968
0.918
0.564
0.859
+ RoXDrive
31
17
0.944
0.966
0.565
0.866
GTRS
40
73
50
0.952
0.918
0.553
0.791
+ RoXDrive
60
36
0.922
0.966
0.556
0.814
GTRS
60
202
123
0.938
0.917
0.539
0.615
+ RoXDrive
160
97
0.901
0.967
0.555
0.682
Appendix
Table 9: Closed-loop performance of GTRS under different maximum frame horizons (MF).
(a) Reward components Rcol , Rlane , and Rdq denote the collision, lane violation, and driving-quality terms.
TransFuser
Rcol
Rlane
Rdq
Obj. Col. ↓
Lane Viol. ↓
Progress ↑
Comfort ↑
Centering ↑
DS ↑
✓
×
×
×
257
178
0.881
0.892
0.540
0.498
✓
×
×
✓
372
185
0.943
0.860
0.529
0.396
✓
✓
×
✓
241
164
0.911
0.924
0.537
0.545
✓
×
✓
✓
229
156
0.889
0.899
0.539
0.539
✓
✓
✓
✓
215
145
0.900
0.932
0.539
0.571
Appendix
Table 10: Ablation studies on reward components and the horizon lengths on the 1K test scenes.
Figure 7: Qualitative results of RoXDrive avoiding the collision on the in-house dataset.
Figure 8: Qualitative results of RoXDrive avoiding the collision on the in-house dataset.
Figure 9: Qualitative results of RoXDrive avoiding lane violations on the in-house dataset.
Figure 10: Qualitative results of RoXDrive avoiding lane violations on the in-house dataset.
Figure 11: Additional qualitative results of RoXDrive mitigating the causal confusion on nuScenes.
Figure 12: Additional qualitative results of RoXDrive mitigating the causal confusion on nuScenes.
Figure 13: Additional qualitative results of RoXDrive mitigating the causal confusion on nuScenes.
Open-loop imitation learning has advanced modern autonomous driving policy architectures, but closed-loop deployment remains vulnerable to policy-induced distribution shift. Existing post-training paradigms exhibit fundamental trade-offs: closed-loop RL fine-tuning provides grounded feedback from executed actions but is constrained by the sparsity of informative events, whereas counterfactual fine-tuning provides dense supervision over candidate futures but inherits bias from imperfect future estimates. We introduce Counterfactual-to-Interactive Reinforcement Fine-Tuning (CRAFT), an on-policy framework that formulates closed-loop post-training as proxy-residual optimization. CRAFT uses group-normalized counterfactual advantages as a dense proxy for real closed-loop advantages and aligns this proxy with the closed-loop world through grounded residual correction from interaction-critical events. To stabilize adaptation, CRAFT regularizes the online policy toward an EMA teacher via asymmetric KL self-distillation. Theoretically, CRAFT decomposes the real closed-loop policy gradient into proxy and residual terms under the same visited-state distribution, reducing residual variance with an aligned proxy while mitigating proxy bias through grounded residual approximation. Empirically, CRAFT achieves the strongest closed-loop gains on Bench2Drive across hierarchical planning, vision-language-action, and vocabulary-scoring architectures. Ablations, scaling behavior, stability analyses, and transfer results further validate the complementary roles of dense counterfactual proxy and grounded residual correction. Project page: https://currychen77.github.io/CRAFT.
Keyu Chen, Nanfei Ye, Yida Wang +4
School of Vehicle and Mobility, Tsinghua University · Li Auto Inc
End-to-end autonomous driving (E2E-AD) aims to directly map raw sensor information to driving actions. Recently, with the rapid advancement of multi-modal large language models (MLLMs), researchers have proposed the paradigm of Vision-Language-Action (VLA) models for E2E-AD, where it seeks to integrate visual perception, language understanding and action prediction within a single policy. However, existing VLA-based policies largely adopts imitation learning, where it only learns to drive by optimizing distance-based metrics w.r.t. logged expert trajectories. Such distribution shift between open-loop training and closed-loop inference leads to suboptimal performance in closed-loop planning. To close this gap, we present CLEAR, a system that enables closed-loop training using Reinforcement Learning (RL) at scale for E2E-AD. We propose to learn a novel residual waypoint policy around the waypoint prior from pretrained VLA policies, effectively harnessing the knowledge within. On another front, one of the key challenges to scale up RL for vision-based policies is the number of parallel simulation environments since RL is data hungry. To that end, we design a heterogeneous pipeline that places the simulator and the VLA learner on distinct compute groups, which allows us to dramatically increase the number of simulation environments running in parallel while avoiding resource contention and maintaining training stability. We show that with a simple reward, CLEAR significantly outperforms previous methods and sets new state-of-the-art performance on the challenging benchmarks of CARLA longest6 v2 and Bench2Drive.
As autonomous vehicle capabilities advance, the safe evaluation of driving policies in long-tail scenarios remains a critical bottleneck. In closed-loop simulation, the driving policy model actively interacts with the environment, where its actions dynamically update the simulator state and directly influence the next set of generated sensor observations. While recent reconstruction-based neural simulators offer photorealism, they are fundamentally constrained by their initial captured data and struggle to generalize to highly dynamic or novel scenes. To overcome these limitations, we introduce OmniDreams, a foundation generative world model mid- and post-trained from the Cosmos diffusion model to autoregressively generate action-conditioned videos in real time. By leveraging the rich visual priors of Cosmos and mid- and post-training on 21k hours of driving scenarios, OmniDreams synthesizes complex, unobserved phenomena that are hard for traditional simulators to capture, such as extreme weather and unpredictable dynamic agent behaviors. Crucially, it autoregressively conditions its photorealistic sensor generation on past frames, the current simulator state, and immediate driving actions. Deployed in a closed-loop system with the Alpamayo 1 policy model and AlpaSim orchestrator, OmniDreams acts as a highly responsive, reactive environment, providing a scalable and comprehensive solution for training and evaluating next-generation autonomous driving policies. We additionally show preliminary results indicating that a world-action model (WAM) post-trained from OmniDreams achieves strong performance on the Physical AI Autonomous Vehicles NuRec dataset, surpassing the VLA-based Alpamayo 1.5 research policy model while using only 1/5 the total parameters. These results highlight the potential for a real-time world model like OmniDreams to also serve as a backbone for policy architectures.