Principled exploration in reinforcement learning requires an agent to quantify its epistemic uncertainty and act to resolve it. Uncertainty over the value function provides a natural signal for exploration, yet existing deep approximations remain brittle and perform inconsistently. The central challenge is therefore to scale these ideas robustly. We conduct a systematic empirical study of how epistemic uncertainty is represented, propagated, and optimized in deep epistemic value functions, and uncover distinct failure modes along each of these axes. These findings motivate DEVOTE, a model-free reinforcement learning algorithm that controls how uncertainty generalizes beyond observed data, stabilizes its temporal propagation, and preserves adaptation to the resulting non-stationary exploration objective. Across reward-free exploration and challenging continuous-control tasks, DEVOTE reaches novel states more effectively and achieves higher task return than strong model-free and model-based exploration baselines. These results provide evidence that deep epistemic value functions are a promising path toward scalable, principled exploration.
Figures & tables
Figure 1: DEVOTE provides stable and scalable estimates of a deep epistemic value function, guides the exploratory policy toward novel and informative states, and outperforms strong RL exploration methods across a range of environments. We plot the average coverage and performance across environments outlined in Section 5 .
Figure 2: Comparison of uncertainty estimates from an Ensemble, an ENN with MLP prior (ENN MLP), and an ENN with RFN prior (ENN RFN) on a synthetic regression task. Dashed curves show the true function, dots indicate training observations, solid curves show predictive means, and shaded bands represent ±βσf . ENNs with RFN function-space priors yield better calibrated uncertainty.
Figure 3: Epistemic value estimation on DMC Cartpole Swingup across prior complexities ℓ . Left: OOD-detection AUROC of the predicted uncertainty σQ . Middle: Pearson correlation between predicted uncertainty and value-estimation error. Right: Predicted initial-state value over training, compared with the Monte Carlo reference. TUD improves both uncertainty metrics across length scales and tracks the Monte Carlo value more closely.
Figure 4: State coverage and the fraction of dormant neurons ( Sokar et al., 2023 ) in the U-shaped PointMaze environment across soft-reset strengths. Increasing the strength of soft resets increases exploration performance and avoids loss of plasticity.
Figure 5: State coverage over training for DEVOTE and the baselines in the pure-exploration environments. DEVOTE consistently drives exploration toward novel states, matching or outperforming all baselines across tasks.
Figure 6: State coverage over training when ablating the DEVOTE components introduced in Section 3 . All components contribute in the continuous exploration tasks, with their relative importance depending on the structure of the environment.
Figure 7: Normalized return over training in the high-dimensional control environments. DEVOTE achieves the strongest task performance across both environments, indicating that its exploration signal can be combined effectively with extrinsic reward.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Environment
Length scale ℓ
DeepSea
1
PointMaze
1
Cheetah
7.5
AntMaze
7.5
PickAndPlace
1
Appendix
Table 2: Environment-specific RFN length scales. For each environment, we select one value from the fixed candidate set ℓ∈{1,2.5,5,7.5} and share it across all DEVOTE variants.
Cartpole anchor
Cart position x
Pole angle θ
Cart velocity vx
Upright, balanced
0
0
0
Near top, offset
0
0.6
0
Mid-swing, climbing
0.4
1.5
0.5
Rail-pinned, hanging
−1.85
π
−0.5
Racing upright
0
0
6.0
Rail-charging, hanging
1.5
π
4.0
Appendix
Table 3: Cartpole anchor configurations. Angles are in radians; anchor pole angular velocity and held action are zero. Pole angular velocity is swept over [−12,12] .
Figure 8: Environments used for evaluation. Top: Pure-exploration environments, with task rewards removed: Behaviour Suite DeepSea ( N=50 ) ( Osband et al., 2020 ) , Spiral PointMaze, and DMC Cheetah ( Tassa et al., 2018 ) . Together, these environments span discrete exploration, navigation, and high-dimensional continuous control. Bottom: Complex-control environments: AntMaze and PickAndPlace ( Zakka et al., 2025 ) . AntMaze requires navigating around obstacles to reach the goal, while PickAndPlace requires discovering coordinated manipulation behavior.
Figure 9: Samples from a GP prior, a randomly initialized ReLU MLP prior, and an RFN prior on the synthetic one-dimensional domain. The sampled MLP functions are approximately linear over this interval, whereas the RFN distributes variability throughout the domain. The GP and RFN use an RBF length scale of one.
Figure 10: Fixed-policy cartpole evaluation with RFN length scale ℓ=1 . Columns show upright, near-top, and mid-swing anchors while pole angular velocity is varied and the remaining state coordinates and anchor action are held fixed. Top: Monte Carlo returns and TD/TUD predictions, with shaded bands showing one predicted standard deviation σQ . Bottom: estimated log visitation density. TUD reduces uncertainty in frequently visited regions while retaining uncertainty outside them; TD produces comparatively uniform uncertainty across the slices.
Figure 11: Fixed-policy cheetah evaluation across RFN prior length scales. Left: OOD-detection AUROC. Middle: error–uncertainty correlation. Right: predicted initial-state value over training compared with Monte Carlo returns. TUD improves both calibration metrics for ℓ≥2.5 , while standard TD learning exhibits pronounced overestimation at small length scales. Results use three seeds.
Figure 12: U-shaped PointMaze diagnostics without resets ( λ=0 ) and with resets ( λ=0.5 ). Maps show visitation density ρ(s) , mean absolute residual ∣r∣(s) , and mean residual bootstrap rb(s) . Each snapshot depicts a single run. Color scales are specific to each map, so the comparison concerns spatial structure rather than absolute color intensity across maps.
Figure 13: Pure-exploration performance of DEVOTE under varying RFN length scales ℓ∈{1,2.5,5,7.5} , measured by state coverage. Prior complexity has little effect in the discrete DeepSea environment, while PointMaze favors the smaller length scale ℓ=1 and Cheetah favors the larger length scale ℓ=7.5 .
Figure 14: State coverage in AntMaze and PickAndPlace, complementing the normalized returns in the main text. Coverage measures the fraction of visited bins in ant x – y position and grasped-cube x – y – z position, respectively. DEVOTE expands coverage fastest in both environments. RND also reaches full coverage in AntMaze, but achieves lower coverage than SAC in PickAndPlace.
Representation learning has enabled classical exploration strategies to be extended to deep Reinforcement Learning (RL), but often makes algorithms more complex and theoretical guarantees harder to establish. We introduce Random Feature Information Gain (RFIG), grounded in Bayesian kernel methods theory, which uses random Fourier features to approximate information gain and compute exploration bonuses in non-countable spaces. We provide error bounds on information gain approximation and avoid the black-box aspects of neural network-based uncertainty estimation, for optimism-based exploration. We present practical details that make RFIG scalable to deep RL scenarios, enabling smooth integration into standard deep RL algorithms. Experimental evaluation across diverse control and navigation tasks demonstrates that RFIG achieves competitive performance with well-established deep exploration methods while offering superior theoretical interpretation.
Waris Radji, Odalric-Ambrym Maillard
Univ. Lille, Inria, CNRS, Centrale Lille, UMR 9189-CRIStAL, F-59000 Lille, France
Safe exploration is a prerequisite for deploying reinforcement learning (RL) agents in safety-critical domains. In this paper, we approach safe exploration through the lens of epistemic uncertainty, where the actor's sensitivity to parameter perturbations serves as a practical proxy for regions of high uncertainty. We propose Sharpness-Aware Policy Optimization (SHAPO), a sharpness-aware policy update rule that evaluates gradients at perturbed parameters, making policy updates pessimistic with respect to the actor's epistemic uncertainty. Analytically we show that this adjustment implicitly reweighs policy gradients, amplifying the influence of rare unsafe actions while tempering contributions from already safe ones, thereby biasing learning toward conservative behavior in under-explored regions. Across several continuous-control tasks, our method consistently improves both safety and task performance over existing baselines, significantly expanding their Pareto frontiers.
Kaustubh Mani, Yann Pequignot, Vincent Mai +1
Universit´e de Montr´eal · Mila - Qu´ebec AI Institute · Universit´e Laval +2
Reinforcement learning agents depend on reward signals whose density is rarely under the designer's control, and when such signals are absent, an agent must generate its own drive to explore. State entropy maximization offers a principled objective for this, but existing methods break down at scale in two ways: the intrinsic reward vanishes once a state has been visited, discouraging revisits to the very gateways that lead onward, and estimating entropy over millions of accumulated observations becomes computationally prohibitive. We address both with Episodic and Lifelong Exploration via Maximum Entropy (ELEMENT), a multiscale intrinsically motivated framework for reward-free exploration that transfers to downstream tasks. ELEMENT couples lifelong entropy maximization with a complementary episodic term acting on a faster timescale. For the episodic term, we derive average episodic state entropy, an intrinsic reward that is the exact minimizer of a tractable upper bound on the reward-decomposition objective; for the lifelong term, we propose a kNN graph-based estimator that keeps entropy tractable without forgetting. ELEMENT consistently outperforms state-of-the-art intrinsic reward baselines on state coverage and unsupervised pre-training. Videos, code, and supplementary material: https://sites.google.com/view/element-rl.
Hongming Li, Zhao Yang, Xiaoxuan Liang +2
Zhejiang Lab, Huangzhou, China · Vrije Universiteit Amsterdam, Amsterdam, the Netherlands · Computational NeuroEngineering Laboratory at the University of Florida, Gainesville, USA