cs.LGMay 28, 2025

Fully Offline Reinforcement Learning

Authors: Mattie FellowsClarisse WibaultUljad BerdicaJohannes ForkelMaike OsborneJakob N. Foerster

Organizations: Foerster Lab for AI Research (FLAIR), Department of Engineering Science, University of Oxford · Machine Learning Research Group, Department of Engineering Science, University of Oxford

Abstract

Offline RL (ORL) promises safe and sample-efficient deployment but existing methods rely on undocumented online interactions for hyperparameter tuning and lack reliable fully offline estimates of initial online performance. We introduce SOReL, a fully offline Bayesian model-based RL method that learns a posterior over dynamics, estimates policy value via predictive uncertainty, and enables complete offline hyperparameter selection. We further propose TOReL, which extends this tuning framework to arbitrary model-free and model-based ORL algorithms. We provide a regret analysis showing that Bayesian offline RL achieves the minimax-optimal parametric rate under standard regularity conditions. Together, our methods establish a practical and theoretically grounded framework for fully offline RL.

Explore similar work

CardsList
  1. Active Offline-to-Online Reinforcement Learning

    Jul 13, 2026Alper Kamil Bozkurt, Shangtong Zhang, Yuichi MotaiOffline Reinforcement LearningLimited Data