Organizations: Southern University of Science and Technology · Harbin Institute of Technology, Shenzhen Shenzhen, Guangdong, China · Southern University of Science and Technology Shenzhen, Guangdong, China · City University of Hong Kong Hong Kong SAR, China · Fudan University · Shanghai Academy of AI for Science2026 Shanghai, China · Southern University of Science and Technology Shenzhen 518055, Guangdong, China28
This paper investigates a hybrid reinforcement learning setting in tabular Markov Decision Processes (MDPs), where an agent aims to learn an optimal policy by combining online interactions with a target environment and offline data from a source environment. A central challenge is that offline data may be collected from outdated environments with shifted transition dynamics, making naive integration of historical data ineffective. To address this, we propose a unified algorithmic framework featuring two algorithms: MIN-UCB-VI for regret minimization and MAX-LCB-VI for best policy identification. Both algorithms leverage fine-grained bias information to more effectively exploit offline data under general transition shifts. We provide theoretical guarantees for our framework, including both instance-dependent and independent upper bounds on regret and sub-optimality gap. Furthermore, we establish matching lower bounds to demonstrate the optimality of our approach and validate our theoretical findings through extensive experiments.