cs.LGOct 1, 2026

When Do Intrinsic Rewards Lead to Exploration?

Authors: Scott W. Viteri, Laura Gomezjurado Gonzalez, Clark Barrett

Organizations: Stanford University

Abstract

Intrinsic rewards are designed to guide exploration in reinforcement learning by assigning value to an agent's experience, for example through prediction error or learning progress. However, maximizing these rewards need not produce the most informative experience available. We propose a formal criterion for exploration that compares policies by the counterfactual information they acquire: how well their histories can substitute for experience under alternative policies. We construct a single, simple environment in which specified count-based, prediction-error, empowerment, and information-gain objectives have maximizing policies that are Pareto-suboptimal at acquiring counterfactual information. We explain these failures and establish conditions under which existing intrinsic rewards successfully encourage optimal exploration. We also construct an objective that assigns a higher value whenever exploration strictly improves under our criterion.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Principled Direction-Free Intrinsic Motivation through Model-Free Epistemic Free-Energy Estimators

    Jul 18, 2026Alireza Furutanpey, Schahram DustdarEntropy Regularized Reinforcement LearningIntrinsic Motivation

  2. ELEMENT: Episodic and Lifelong Exploration via Maximum Entropy

    Dec 5, 2024Hongming Li, Zhao Yang, Xiaoxuan Liang +2Stochastic ExplorationEntropy Regularized Reinforcement Learning

  3. Efficient Exploration Is Enough

    Sep 7, 2026Mikel Malagón, Jon Vadillo, Josu Ceberio +2Efficient ExplorationAgentic Discovery