cs.LGSep 28, 2026

Deep Epistemic Value Functions for Optimistic Exploration

Authors: Leander Diaz-Bone, Marco Bagatella, Jonas Hübotter, Andreas Krause

Organizations: ETH Zürich, Switzerland · Max Planck Institute for Intelligent Systems, Germany

Abstract

Principled exploration in reinforcement learning requires an agent to quantify its epistemic uncertainty and act to resolve it. Uncertainty over the value function provides a natural signal for exploration, yet existing deep approximations remain brittle and perform inconsistently. The central challenge is therefore to scale these ideas robustly. We conduct a systematic empirical study of how epistemic uncertainty is represented, propagated, and optimized in deep epistemic value functions, and uncover distinct failure modes along each of these axes. These findings motivate DEVOTE, a model-free reinforcement learning algorithm that controls how uncertainty generalizes beyond observed data, stabilizes its temporal propagation, and preserves adaptation to the resulting non-stationary exploration objective. Across reward-free exploration and challenging continuous-control tasks, DEVOTE reaches novel states more effectively and achieves higher task return than strong model-free and model-based exploration baselines. These results provide evidence that deep epistemic value functions are a promising path toward scalable, principled exploration.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Information-Based Exploration via Random Features for Reinforcement Learning

    Jul 20, 2026Waris Radji, Odalric-Ambrym MaillardStochastic ExplorationExploration

  2. SHAPO: Sharpness-Aware Policy Optimization for Safe Exploration

    Jun 8, 2026Kaustubh Mani, Yann Pequignot, Vincent Mai +1Frictive Policy OptimizationExploration

  3. ELEMENT: Episodic and Lifelong Exploration via Maximum Entropy

    Dec 5, 2024Hongming Li, Zhao Yang, Xiaoxuan Liang +2Stochastic ExplorationEntropy Regularized Reinforcement Learning