Offline reinforcement learning enables reward-driven policy improvement from fixed datasets without requiring online exploration, making it particularly attractive in safety-critical domains. A central challenge, however, is distribution shift: policy optimization may favor actions that are weakly supported by the offline data, rendering value estimates unreliable. Existing approaches primarily control this shift in the policy's own action space. In interactive environments such as autonomous driving, this can be insufficient: a candidate ego trajectory may remain well supported under the marginal behavior distribution while being poorly supported jointly with the surrounding-agent behavior observed in the logged interaction. We refer to this degradation in interaction support as \emph{interaction distribution shift} (IDS), and introduce \emph{Interaction-Constrained Drive Policy} (ICDP), an offline reinforcement learning framework that explicitly controls interaction-level distribution shift. Starting from the joint data distribution over ego and surrounding-agent futures, we show that joint-support degradation decomposes exactly into an ego-support component and a residual interaction-support component. We recover the latter through contrastive density-ratio estimation, isolating interaction compatibility without explicit joint-density modeling, surrounding-agent prediction, or rollouts in reactive simulators or learned world models during policy optimization. Closed-loop evaluations on nuPlan, Interplan and real-world truck experiments show that ICDP suppresses high-value yet interaction-unsupported trajectory selections and improves performance in interaction-critical driving scenarios. Project webpage: https://mahmoud-selim.github.io/ICDP/
Figures & tables
Figure 1: Overview of ICDP . The logged scene is encoded into structured scene features used by the diffusion actor and critic. The actor generates candidate ego trajectories, which are evaluated by a pessimistic chunk-level critic and by the interaction-support module. The latter combines a joint ego–agent classifier with an ego-only classifier; their residual score isolates interaction-specific support and yields the IDS estimate used to constrain policy improvement.
Val14
Test14-Hard
Test14-Random
Planner
NR
R
NR
R
NR
R
PDM-Open ∗
53.53
54.24
33.51
35.83
52.81
57.23
GameFormer w/o refine.
13.32
8.69
7.08
6.69
11.36
9.31
PlanTF
84.27
76.95
69.70
61.61
85.62
79.58
PLUTO w/o refine. ∗
88.89
78.11
70.03
59.74
89.90
78.62
Diffusion Planner
89.87
82.80
75.99
69.22
89.19
82.93
Table 1: Closed-loop performance on nuPlan Val14, Test14-Hard, and Test14-Random under non-reactive (NR) and reactive (R) evaluation. The best result in each column is highlighted in blue.
Val14
Test14-Hard
Test14-Random
Method
NR
R
NR
R
NR
R
BC
89.8±1.2
82.4±1.09
78.42±0.69
68.21±0.95
91.98±0.62
82.71±0.53
Offline RL
89.4±1.74
83.4±2.11
78.75±0.75
70.07±1.08
90.74±0.70
83.02±1.08
ICDP
91.3±1.91
84.3±2.04
78.34±1.03
72.87±2.01
93.33±0.85
86.34±1.18
Table 2: Controlled comparison of BC, unconstrained Offline RL, and ICDP . Results are averaged over ten evaluation checkpoints spaced by 2k optimization steps and reported as mean ± std.
Figure 2: Sensitivity of ICDP to (a) ego-support regularization, (b) interaction-support regularization, and (c) the number of surrounding agents. Scores are averaged across Test14-Hard and Test14-Random under non-reactive and reactive evaluation.
Figure 3: Real-world deployment of ICDP on a full-scale autonomous truck. (a–b) Representative multi-agent interaction trial involving an overtaking vehicle, shown from external and onboard-perception viewpoints. (c–d) Standalone driving on curved road segments under real sensing and vehicle dynamics.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Value
Action horizon H
32 steps ( 3.2 s)
Sampling interval Δt
0.1 s
Discount factor γ
0.99
Collision / severe off-road return
−5
Soft / hard drivable-area threshold
0.3 / 0.5 m
TTC additive-penalty threshold
0.95 s
Appendix
Table 3: Principal reward parameters used for training.
Parameter
Value
Planning horizon
32 steps ( 3.2s )
History horizon
16 frames ( 1.5s )
Actor/critic agent capacity
32
Interaction-support agent context
48
Lane / route capacity
70 / 25
Map-query radius
100m
Appendix
Table 4: Principal implementation settings used for ICDP .
Figure 4: Qualitative trajectory samples from the learned diffusion policy across diverse nuPlan scenes. Each panel shows ten independently generated ego trajectories together with the logged ego trajectory, preferred route, and surrounding traffic. Samples are shown directly without critic-based selection or post-processing.
Figure 5: Distribution of scenario types in the processed offline dataset. The 62 recorded categories are divided into more- and less-frequent groups to make the long-tailed distribution legible.
Planner
Overall
Nudge Around
High Traffic
Jaywalk
PlanTF
47.70
49.40
58.85
33.94
PLUTO w/o refine.
58.47
71.56
67.25
25.48
Diffusion Planner
52.90
60.48
49.71
26.20
Flow Planner
61.82
72.96
67.21
43.57
ICDP
62.7
51.8
70.2
51.1
Appendix
Table 5: Closed-loop performance on InterPlan and selected interaction-critical scenario categories. The best result in each column is highlighted in blue.
Open-loop imitation learning has advanced modern autonomous driving policy architectures, but closed-loop deployment remains vulnerable to policy-induced distribution shift. Existing post-training paradigms exhibit fundamental trade-offs: closed-loop RL fine-tuning provides grounded feedback from executed actions but is constrained by the sparsity of informative events, whereas counterfactual fine-tuning provides dense supervision over candidate futures but inherits bias from imperfect future estimates. We introduce Counterfactual-to-Interactive Reinforcement Fine-Tuning (CRAFT), an on-policy framework that formulates closed-loop post-training as proxy-residual optimization. CRAFT uses group-normalized counterfactual advantages as a dense proxy for real closed-loop advantages and aligns this proxy with the closed-loop world through grounded residual correction from interaction-critical events. To stabilize adaptation, CRAFT regularizes the online policy toward an EMA teacher via asymmetric KL self-distillation. Theoretically, CRAFT decomposes the real closed-loop policy gradient into proxy and residual terms under the same visited-state distribution, reducing residual variance with an aligned proxy while mitigating proxy bias through grounded residual approximation. Empirically, CRAFT achieves the strongest closed-loop gains on Bench2Drive across hierarchical planning, vision-language-action, and vocabulary-scoring architectures. Ablations, scaling behavior, stability analyses, and transfer results further validate the complementary roles of dense counterfactual proxy and grounded residual correction. Project page: https://currychen77.github.io/CRAFT.
Keyu Chen, Nanfei Ye, Yida Wang +4
School of Vehicle and Mobility, Tsinghua University · Li Auto Inc
Self-play reinforcement learning has recently emerged as a way to train driving policies without any human data. It uses cheap, large-scale simulations to substitute expensive, large-scale human driving demonstrations. A key limitation of this approach is that policies trained through pure self-play can learn effective but alien driving conventions incompatible with people. Previous works attempt to mitigate such behavioral misalignments through extensive reward engineering and domain randomization, which are brittle and labor-intensive. Instead of completely discarding human demonstrations, our method treats them as a regularization objective on top of a minimal safe goal-reaching reward. Like the spice in a good stew, we find that a little human data goes a long way: our method uses only 30 minutes of human demonstrations, 2500x fewer than comparable imitation learning approaches. Resulting policies coordinate with held-out human trajectories and complete training in 15 hours on a single consumer-grade GPU. Videos and full source code are available at https://spiced-self-play.com/.
Daphne Cornelisse, Julian Hunt, Zixu Zhang +4
1NYU Tandon School of Engineering · 2NYU Courant · 3Princeton University +2
Reinforcement learning (RL) has shown promising performance in autonomous driving, yet ensuring the safety of online RL policies remains challenging due to insufficient exposure to safety-critical driving scenes. The long-tailed nature of real-world traffic situations makes dangerous and rare interactions difficult to encounter through conventional sampling, limiting the ability of RL policies to learn robust safety behaviors. Existing methods improve training diversity by synthesizing challenging scenes or adversarial situations. However, these approaches typically optimize scene generation objectives separately from the evolving policy, without explicitly modeling how generated perturbations relate to the current policy's weaknesses and learning needs. In this paper, we propose Threat-guided Policy-aware Scene Perturbation (TPSP) for safe autonomous driving with online RL. TPSP introduces a policy-aware scene encoder to capture the interaction between policy behaviors and surrounding environments, enabling scene perturbation aligned with the current policy. Based on this representation, TPSP selectively perturbs critical objects rather than applying uniform modifications across the scene. Furthermore, we develop a threat-guided optimization strategy that evaluates perturbed scenes through threat-level differences between policy rollouts on original and perturbed scenes, guiding the generation of safety-critical scenes with higher training value. Comprehensive experiments demonstrate that TPSP improves safety learning efficiency, achieving strong safety performance on NAVSIM v2 with approximately 4 million kilometers of simulated driving data. Ablation studies verify that policy-aware targeted perturbations provide more informative safety-critical experiences than random or policy-unaware strategies, enabling safer driving under limited interaction budgets.
Xincong Hu, Lei Ou, Maosen Li +3
Nanjing University · 1Nanjing University, Nanjing, Jiangsu, China · 2Yinwang Intelligent Technology Co., Ltd., China +1