cs.LGSep 29, 2026

In-Distribution Imagination for Model-Based Offline Reinforcement Learning

Authors: Mintae Kim, Koushil Sreenath

Organizations: Hybrid Robotics, BAIR, UC Berkeley

Abstract

Model-based offline reinforcement learning (MBORL) improves sample efficiency through model-generated trajectories. However, accumulative model error can drive imagined trajectories outside the offline data distribution, leading to unrealistic synthetic data and unstable policy optimization. Many existing methods primarily control rollouts using transition-level uncertainty. We propose \emph{in-distribution imagination} (IDI), a rollout control framework that estimates trajectory support in a learned representation space and adaptively truncates rollouts that leave the offline trajectory manifold. Combined with trajectory-regularized RL, an extension of entropy-regularized RL, IDI consistently improves performance in limited-data settings. Experiments show that trajectory support predicts rollout failure substantially better than transition-level uncertainty, highlighting the importance of trajectory-level rollout control in MBORL.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. On Training in Imagination

    May 7, 2026Nadav Timor, Ravid Shwartz-Ziv, Micah Goldblum +2Model-Based Reinforcement LearningRollout

  2. IADD-TR: Intervention-Aware Dynamics Decoupling with Targeted Regularization for Model-Based Reinforcement Learning

    Aug 11, 2026Zefeng Liang, Jie Qiao, Ruichu Cai +2Model-Based Reinforcement LearningRegularization

  3. Counterfactual Transport Flows for Offline Conservative Trajectory Refinement

    Jun 8, 2026Lena Krieger, Xuan Zhao, Zhuo Cao +3Model-Based Reinforcement LearningLong Refinement Trajectories