cs.AIFeb 13, 2026

Calculating Mutual Information between a Reward Maximizer and its Environment

Authors: Alfred HarwoodJose FaustinoAlex Altair

Organizations: 1Dovetail Research · University of Sao Paulo

Abstract

An important question in the field of AI is the extent to which successful behaviour requires an internal representation of the world. In this work, we quantify the amount of information an optimal policy provides about the underlying environment. We consider a Controlled Markov Process (CMP) with nn states and mm actions, assuming a uniform prior over the space of possible transition dynamics. We prove that observing a deterministic policy that is optimal for any non-constant reward function then conveys exactly nlogmn \log m bits of information about the environment. Specifically, we show that the mutual information between the environment and the optimal policy is nlogmn \log m bits. This bound holds across a broad class of objectives, including finite-horizon, infinite-horizon discounted, and time-averaged reward maximization. These findings provide a precise information-theoretic lower bound on the ``implicit world model'' necessary for optimality.

Explore similar work

CardsList
  1. Imperfect World Models are Exploitable

    May 15, 2026Logan Mondal Bhamidipaty, Esmeralda S. Whitammer, David Abel +2HackingWorld Models