cs.LGApr 15, 2026

Provably Efficient Offline-to-Online Value Adaptation with General Function Approximation

Authors: Shangzhe Li, Weitong Zhang

Organizations: UNC Chapel Hill

Abstract

We study value adaptation in offline-to-online reinforcement learning under general function approximation. Starting from an imperfect offline pretrained QQ-function, the learner aims to adapt it to the target environment using only a limited amount of online interaction. We first characterize the difficulty of this setting by establishing a minimax lower bound, showing that even when the pretrained QQ-function is close to optimal Q⋆Q^\star, online adaptation can be no more efficient than pure online RL on certain hard instances. On the positive side, under a novel structural condition on the offline-pretrained value functions, we propose O2O-LSVI, an adaptation algorithm with problem-dependent sample complexity that provably improves over pure online RL. Finally, we complement our theory with neural-network experiments that demonstrate the practical effectiveness of the proposed method.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning?

    Jul 29, 2026Perry Dong, Ron Polonsky, Dorsa Sadigh +1Continuous ControlRL Fine-Tuning

  2. Adaptive Policy Selection and Fine-Tuning under Interaction Budgets for Offline-to-Online Reinforcement Learning

    May 6, 2026Alper Kamil Bozkurt, Xiaoan Xu, Shangtong Zhang +2Sequential Decision MakingRL Fine-Tuning