cs.LGOct 1, 2026

When Do Intrinsic Rewards Lead to Exploration?

Authors: Scott W. Viteri, Laura Gomezjurado Gonzalez, Clark Barrett

Organizations: Stanford University

Abstract

Intrinsic rewards are designed to guide exploration in reinforcement learning by assigning value to an agent's experience, for example through prediction error or learning progress. However, maximizing these rewards need not produce the most informative experience available. We propose a formal criterion for exploration that compares policies by the counterfactual information they acquire: how well their histories can substitute for experience under alternative policies. We construct a single, simple environment in which specified count-based, prediction-error, empowerment, and information-gain objectives have maximizing policies that are Pareto-suboptimal at acquiring counterfactual information. We explain these failures and establish conditions under which existing intrinsic rewards successfully encourage optimal exploration. We also construct an objective that assigns a higher value whenever exploration strictly improves under our criterion.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Jul 18, 2026cs.LG

Principled Direction-Free Intrinsic Motivation through Model-Free Epistemic Free-Energy Estimators

Across environments with mixed sources of uncertainty, unsupervised reinforcement learning requires intrinsic motivation that does not precommit to a particular direction of surprise. Surprise minimization is scoped by design to ``unstable'' environments. Prediction-error curiosity rewards total expected surprise, including irreducible noise. Bandit or mixture switching between surprise-minimizing and surprise-maximizing rewards reintroduces non-stationarity by construction. We propose a single intrinsic reward, stationary within each window, derived from the novelty contribution of a preference-free Expected Free Energy objective, expressed in reward-maximization form. Our claim is that parameter information gain, the expected surprise of the next state minus its irreducible part, is the appropriate intrinsic signal in both high-entropy and low-entropy components of the state space. Maximizing it seeks exactly the surprise the model can explain away. In regions of unresolved dynamics, this epistemic term drives exploration. As dynamics become resolved, the epistemic term vanishes, while an aleatoric penalty favors lower-variance transitions, all without fitting an explicit next-state predictor. A pseudocount supplies epistemic value, a probe-based penalty captures aleatoric variance, and a short-horizon gate protects informative successors. A window-based freeze of all reward-defining objects yields a stationary Bellman operator, explicit bounds on learning targets, and a conditional uniform-concentration result for the nonparametric estimators under mixing, smoothness, bandwidth, and capacity assumptions. In active-inference terms, the agent is preference-free where novelty is retained, standard likelihood ambiguity vanishes under full observability, a nonstandard transition-entropy penalty is added, and surprise minimization emerges in resolved regions of the state space.
Dec 5, 2024cs.LG

ELEMENT: Episodic and Lifelong Exploration via Maximum Entropy

Reinforcement learning agents depend on reward signals whose density is rarely under the designer's control, and when such signals are absent, an agent must generate its own drive to explore. State entropy maximization offers a principled objective for this, but existing methods break down at scale in two ways: the intrinsic reward vanishes once a state has been visited, discouraging revisits to the very gateways that lead onward, and estimating entropy over millions of accumulated observations becomes computationally prohibitive. We address both with Episodic and Lifelong Exploration via Maximum Entropy (ELEMENT), a multiscale intrinsically motivated framework for reward-free exploration that transfers to downstream tasks. ELEMENT couples lifelong entropy maximization with a complementary episodic term acting on a faster timescale. For the episodic term, we derive average episodic state entropy, an intrinsic reward that is the exact minimizer of a tractable upper bound on the reward-decomposition objective; for the lifelong term, we propose a kkNN graph-based estimator that keeps entropy tractable without forgetting. ELEMENT consistently outperforms state-of-the-art intrinsic reward baselines on state coverage and unsupervised pre-training. Videos, code, and supplementary material: https://sites.google.com/view/element-rl.
Sep 7, 2026cs.LG

Efficient Exploration Is Enough

This work introduces an alternative view of efficient exploration and studies its theoretical and empirical implications in the absence of extrinsic rewards. Specifically, we define efficient explorers as agents that prioritize generating generalizable experience, i.e., data that supports learning models capable of predicting and adapting across the environment. This allows us to analyze efficient exploration through the lens of prediction and generalization. Theoretically, we demonstrate that optimally efficient explorers naturally schedule their trajectories to visit the most informative and learnable regions first. Empirically, we show that optimizing for these agents gives rise to an automatic curriculum of progressively more complex behaviors, even in relatively simple environments. These results indicate that pursuing this purely intrinsic objective alone is enough to drive the emergence of highly sophisticated behaviors. We believe that this new framework provides a principled mechanism by which agent-environment systems may sustain an open-ended process of increasingly complex behavior without external rewards, tasks, or objectives.