We study value adaptation in offline-to-online reinforcement learning under general function approximation. Starting from an imperfect offline pretrained Q-function, the learner aims to adapt it to the target environment using only a limited amount of online interaction. We first characterize the difficulty of this setting by establishing a minimax lower bound, showing that even when the pretrained Q-function is close to optimal Q⋆, online adaptation can be no more efficient than pure online RL on certain hard instances. On the positive side, under a novel structural condition on the offline-pretrained value functions, we propose O2O-LSVI, an adaptation algorithm with problem-dependent sample complexity that provably improves over pure online RL. Finally, we complement our theory with neural-network experiments that demonstrate the practical effectiveness of the proposed method.
Figures & tables
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 1: Results on Linear MDP. O2O-LSVI and LSVI-UCB with a fixed true action gap of 0.10 , reference tolerance βref=0.40 , and reference error 0.44 on 3 of 60 entries. Both learners use d=5 features. (a) Mean cumulative pseudo-regret across five seeds, with a linear episode axis. (b) Mean paired regret difference; negative values favor O2O-LSVI. Shaded regions in (a) and (b) are pointwise 95% Student- t intervals with four degrees of freedom, not simultaneous confidence bands. (c) Mean fraction of entries within each reference class satisfying the containment test at a checkpoint; these are table-entry fractions, not visitation-weighted or cumulative rates. Panels (b) and (c) use logarithmic episode axes. Lines connect saved checkpoints; panel (a) also includes the known origin R0=0 .
Settings
Offline RL
Offline-to-Online Adaptation
CQL
IQL
Cal-QL
Ours
Umaze
94.0 ± 1.6
77.0 ± 0.7
76.8 ± 7.5 → 99.8 ± 0.4
85.8 ± 3.3 → 99.8 ± 0.4
Medium-Play
59.0 ± 11.2
71.8 ± 3.0
71.8 ± 3.3 → 98.8 ± 1.6
70.3 ± 2.0 → 99.3 ± 1.3
Large-Play
28.8 ± 7.8
38.5 ± 8.7
31.8 ± 8.9 → 97.3 ± 1.8
35.3 ± 4.0 → 98.5 ± 1.6
Appendix
Table 1: Empirical Results on AntMaze. We compare our method with existing offline-to-online adaptation and offline RL baselines on AntMaze. Our method achieves performance comparable to or better than existing baselines. Results are reported as D4RL scores averaged over four random seeds.
Hyperparameter
Value
Ensemble Size
5
β
50
CQL Conservative Coefficient α (Offline & Online)
5.0
Discount factor
0.99
Q Network Learning Rate
3e-4
Policy Network Learning Rate
1e-4
Appendix
Table 2: Hyperparameters. We demonstrate the hyperparameters used in the empirical experiment with the practical implementation of O2O-LSVI.