cs.LGOct 7, 2026

Beyond Policy Support: Interaction Constrained Offline Reinforcement Learning for Autonomous Driving

Authors: Mahmoud Selim, Cristina Cipriani, Karl Henrik Johansson

Organizations: KTH Royal Institute of Technology

Abstract

Offline reinforcement learning enables reward-driven policy improvement from fixed datasets without requiring online exploration, making it particularly attractive in safety-critical domains. A central challenge, however, is distribution shift: policy optimization may favor actions that are weakly supported by the offline data, rendering value estimates unreliable. Existing approaches primarily control this shift in the policy's own action space. In interactive environments such as autonomous driving, this can be insufficient: a candidate ego trajectory may remain well supported under the marginal behavior distribution while being poorly supported jointly with the surrounding-agent behavior observed in the logged interaction. We refer to this degradation in interaction support as \emph{interaction distribution shift} (IDS), and introduce \emph{Interaction-Constrained Drive Policy} (ICDP), an offline reinforcement learning framework that explicitly controls interaction-level distribution shift. Starting from the joint data distribution over ego and surrounding-agent futures, we show that joint-support degradation decomposes exactly into an ego-support component and a residual interaction-support component. We recover the latter through contrastive density-ratio estimation, isolating interaction compatibility without explicit joint-density modeling, surrounding-agent prediction, or rollouts in reactive simulators or learned world models during policy optimization. Closed-loop evaluations on nuPlan, Interplan and real-world truck experiments show that ICDP suppresses high-value yet interaction-unsupported trajectory selections and improves performance in interaction-critical driving scenarios. Project webpage: https://mahmoud-selim.github.io/ICDP/

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. CRAFT: Counterfactual-to-Interactive Reinforcement Fine-Tuning for Driving Policies

    May 6, 2026Keyu Chen, Nanfei Ye, Yida Wang +4Counterfactual LearningReinforcement Fine-Tuning

  2. Human-like autonomy emerges from self-play and a pinch of human data

    Jun 11, 2026Daphne Cornelisse, Julian Hunt, Zixu Zhang +4Self-PlayRandom

  3. Threat-guided Policy-aware Scene Perturbation for Safe Autonomous Driving with Online Reinforcement Learning

    Aug 11, 2026Xincong Hu, Lei Ou, Maosen Li +3Autonomous DrivingAdversarial Robustness