Deep Epistemic Value Functions for Optimistic Exploration
Organizations: ETH Zürich, Switzerland · Max Planck Institute for Intelligent Systems, Germany
Abstract
Principled exploration in reinforcement learning requires an agent to quantify its epistemic uncertainty and act to resolve it. Uncertainty over the value function provides a natural signal for exploration, yet existing deep approximations remain brittle and perform inconsistently. The central challenge is therefore to scale these ideas robustly. We conduct a systematic empirical study of how epistemic uncertainty is represented, propagated, and optimized in deep epistemic value functions, and uncover distinct failure modes along each of these axes. These findings motivate DEVOTE, a model-free reinforcement learning algorithm that controls how uncertainty generalizes beyond observed data, stabilizes its temporal propagation, and preserves adaptation to the resulting non-stationary exploration objective. Across reward-free exploration and challenging continuous-control tasks, DEVOTE reaches novel states more effectively and achieves higher task return than strong model-free and model-based exploration baselines. These results provide evidence that deep epistemic value functions are a promising path toward scalable, principled exploration.
Figures & tables
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Environment | Length scale |
|---|---|
| DeepSea | |
| PointMaze | |
| Cheetah | |
| AntMaze | |
| PickAndPlace |
| Cartpole anchor | Cart position | Pole angle | Cart velocity |
|---|---|---|---|
| Upright, balanced | 0 | 0 | 0 |
| Near top, offset | 0 | 0.6 | 0 |
| Mid-swing, climbing | 0.4 | 1.5 | 0.5 |
| Rail-pinned, hanging | |||
| Racing upright | 0 | 0 | 6.0 |
| Rail-charging, hanging | 1.5 | 4.0 |