Provably Efficient Offline-to-Online Value Adaptation with General Function Approximation
Organizations: UNC Chapel Hill
Abstract
We study value adaptation in offline-to-online reinforcement learning under general function approximation. Starting from an imperfect offline pretrained -function, the learner aims to adapt it to the target environment using only a limited amount of online interaction. We first characterize the difficulty of this setting by establishing a minimax lower bound, showing that even when the pretrained -function is close to optimal , online adaptation can be no more efficient than pure online RL on certain hard instances. On the positive side, under a novel structural condition on the offline-pretrained value functions, we propose O2O-LSVI, an adaptation algorithm with problem-dependent sample complexity that provably improves over pure online RL. Finally, we complement our theory with neural-network experiments that demonstrate the practical effectiveness of the proposed method.
Figures & tables
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Settings | Offline RL | Offline-to-Online Adaptation | ||
|---|---|---|---|---|
| CQL | IQL | Cal-QL | Ours | |
| Umaze | 94.0 1.6 | 77.0 0.7 | 76.8 7.5 99.8 0.4 | 85.8 3.3 99.8 0.4 |
| Medium-Play | 59.0 11.2 | 71.8 3.0 | 71.8 3.3 98.8 1.6 | 70.3 2.0 99.3 1.3 |
| Large-Play | 28.8 7.8 | 38.5 8.7 | 31.8 8.9 97.3 1.8 | 35.3 4.0 98.5 1.6 |
| Hyperparameter | Value |
| Ensemble Size | 5 |
| 50 | |
| CQL Conservative Coefficient (Offline & Online) | 5.0 |
| Discount factor | 0.99 |
| Q Network Learning Rate | 3e-4 |
| Policy Network Learning Rate | 1e-4 |