Empowerment captures the capacity for an agent to actively control its environment. While conceptually appealing as an information-theoretic quantity, the connection between empowerment and structurally central states that provide broad access to future outcomes has remained an open question. In this work, we link empowerment maximization and skill-learning methods to provide new geometries for interpreting and analyzing empowerment. Our analyses answer longstanding open questions on the connections between empowerment and structural centrality. Our analyses also reveal distinctions between information and reward geometries, highlighting important theoretical implications to build scalable empowerment-maximization methods. Website and code can be found at https://empowerment-geometry.github.io/.
Figures & tables
Figure 1 : Two trained empowerment-maximizing policies characterize central states in an MDP: Consider an agent (robot) on a hill that is easier to descend than to ascend. In Fig. 1(a), the potential policy ascends the empowerment landscape to the top of the hill, using \lx@glossaries@gls@linkmaingemathrmpot\hyperlinkglo:gemathrmpotEpot(s) as an intrinsic reward. In Fig. 1(b), The effective policy π eff maximizes Iπ(Z;S+∣S0) , associating each skill z with a downhill route.
Figure 2 : Skills measure the centrality of a state. We visualize the potential empowerment field E(s) (a) for each state in a tabular GridWorld. The central state has high potential empowerment, because (b) the central state gives an agent access to a maximal set of distinguishable skills. A potential empowerment policy ascends the potential empowerment landscape to the center state in (c), from which it can execute different skills. See for the formal statement. See Appendix I for tabular experiment details.
Figure 3 : Potentially empowered states are central states in the polytope. Tabular intuition for . (a) We can use a probability simplex Δ(S) to visualize the discounted state occupancy measures of each skill (circles). Orange circles indicate the best set of skills starting from state s1 , and teal circles indicate the best skills starting from state s2 . from starting states s1 and s2 in arbitrary MDP. The skills starting from s1 define a smaller polytope ( orange points ) in the simplex than s2 ( teal points ) as a result of their transition dynamics, so the potential empowerment is higher at s2 than s1 . (b) We use a GridWorld to illustrate the connection between centrality and tool use. An agent in the GridWorld begins at the bottom of the hallway. The agent can set down and pick up a key at the top of the hallway. The key provides the agent access to shaded regions of the MDP . (b, i) A potentially empowered agent picks up the key then navigates to the center of the room, correctly identifying the middle state with the key as the most central state. (b, ii) From the center, the agent can execute the maximal number of distinguishable skills. See §. I.3 for experimental details.
Figure 4 : Bottleneck states are high potential empowerment states. Consider GridWorlds with a door, or a “bottleneck”. (a) For short horizon behaviors where γ=0.5 , the high empowerment regions are the central door (a, left) and the center of the two sides of the room (a, right). (b) Lifting γ=0.95 increases the horizon of possible behaviors, such that the empowerment tracks the location of the door. In §. 4.2 we show that empowerment measures the reduction in temporal distance to future states, which quantifies a bottleneck effect. See §. I.2 for experimental details.
Figure 5 : Discrepancy between information geometry and reward adaptation geometry in discrete settings. (a) By going left/right from s , an agent can reach node s1 or s2 with high likelihood, so there are two easily separable skills from s . (d) Targeting nodes s1 , s2 , or s3 can make the agent inadvertently land in other nodes, so there are three somewhat separable skills from s . (b,e) The MDP s in (a) and (d) admit polytopes of feasible state occupancies in the simplex, visualized by the segment and triangle. (b) Axis u1 trades off occupancy of s1 vs. s2 . (e) Axis u1 trades off occupying s2 vs. s3 ; axis u2 trades off occupying s1 vs. {\color[rgb]{0.8672,0.5195,0.3203}s_{2}}/{\color[rgb]{0.5781,0.4688,0.375}s_{3}} . Return is linear in the state occupancy, so we can also represent reward functions as vectors in state-occupancy space corresponding to the reward weights. The color indicates the best adapted skill for that reward coordinate. The polytope in (e) has a larger Gaussian Width than that in (b), supporting better adaptation. (c,f) However, increased skill separability in (a) gives larger information radius and MI , where the black contour indicates the maximal MI . Thus, even though empowerment is higher for the 2 state system (a) versus the 3 state system (d), the 2 state system has worse adapted returns. See §. I.4 for details.
Figure 6 : The information ball can collapse in the continuous limit ( Proposition G.2 ) . Consider an arbitrary smooth reward function (level curves in gray) and some MDP that admits Gaussian skill state occupancy measures (colored distributions). While the potential empowerment of the starting state s∈S increases from subfigure (a) → (c), the skills have increasingly identical returns and, thus, provide progressively smaller adaptation benefit. Therefore, large channel capacities do not imply adaptation for continuous MDP s and reward functions. See Proposition G.2 for the theoretical result.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7 : Empowerment over contiguous rollouts concentrates at bottlenecks ( Proposition C.1 ) . The figures show the MI of the contiguous process at each start state on the 8×8 grids from Fig. 4 , with (a) a center opening and (b) a corner opening. The bottleneck openings have high potential contiguous empowerment.
Figure 8 : State refinement leaves return distribution unchanged ( Example F.1 ). Bars display the occupancy p^{\lx@glossaries@gls@link{main}{pi_{m}athrmeff}{{{}}{\hyperlink{glo:pi_mathrmeff}{\pi_{\mathrm{eff}}}}}}({\color[rgb]{0.5,0.5,0.5}s_{+}}\mid s_{0},z)=(p_{1},p_{2}) of a skill over the coarse states (a) and the occupancy over the split states (b). Each panel lists the reference mass and per-state reward prior r(s+)∼N(0,1/b(s+)) . Splitting states halves the reference mass and doubles the reward variance. Thus, the return obeys the same distribution N(0,2p12+2p22) in both MDP s.
Figure 9 : MI bound as a function of skill commitment ( ) . Consider the MDP consisting of 5 states with a central starting state, where any left/right/up/down action sends probability mass to absorbing left/right/up/down states. Actions left/right/up/down have a level of commitment t , which measures the additional mass above uniform assigned to the corresponding state. When commitment is small, probability is almost spread equally over all four states regardless of action so there is little control. In this regime, the adaptation J dominates the lower bound. As the commitment increases, actions/skills commit to states and the bound tightens.
Figure 10 : MI bound as a function of absorbing states ( ) . Consider an MDP with one starting state and K absorbing target states. Each of the K actions transitions deterministically to a distinct target. We normalize occupancies over the targets (neglecting the occupancy of the starting state). Thus, K counts the maximum possible number of total skills. The scaled MI (orange) correctly lower bounds the adaptation J (purple). The bound weakens as the number of absorbing target states in the MDP grows.
Parameters
Value
Environment size (n × m)
5 × 5
Environment actions
{Left, Right, Up, Down, Stay}
Periodic boundary conditions
No
P(random action)
0.0
Gamma for DSOM
0.95
# uniformly sampled vertex directions per s0
500
Appendix
Table 1 : Hyperparameters for the 5×5 gridworld experiments ( Fig. 2 )
Parameters
Value
Environment size (n × m)
5 × 5
Environment actions
{Left, Right, Up, Down, Stay, Pick up key at (0,2) , Drop key at (0,2) }
Periodic boundary conditions
No
P(random action)
0.0
Gamma for skill-learning
0.99
# uniformly sampled vertex directions per s0
1000
Appendix
Table 2 : Hyperparameters for the 5×5 gridworld with walls and a key
Figure 11 : High empowerment finds central states, where skills are diverse from the central states ( ) . Potential policy navigates to central states. Learned skills spread out from the central state and measure the degree of centrality of the state (in information geometry). The dotted line denotes the potential empowerment policy’s trajectory for 1−γ1 steps. Diamond denotes the handoff state s0 , where the agent randomly samples z∼p(z∣s0) and rolls out skill policy π(a∣s,z;s0) .
Table 3 : Skill occupancy distributions of the example MDP s. (a) Example X with mass on the two reachable targets. (b) Example Y with mass on the three reachable targets.
In many practical reinforcement learning environments, observations are far higher-dimensional than the variables that matter for control. In this work, we ask: can we learn representations that capture only control-relevant features of the environment? We study this question through the empowerment objective, which maximizes an agent's influence over the environment and is widely used for unsupervised skill learning. We show that empowerment agents induce two distinct representations -- forward and backward -- that capture complementary aspects of the state, and both of which are invariant to control-irrelevant features. Thus, empowerment maximization leads agents to learn an implicit, control-centric model of the world. Our analysis highlights the importance of learning representations through interaction rather than from passive datasets: interaction aimed at maximizing control is essential for learning useful invariance properties, a perspective that aligns closely with the causal learning literature.
Mahsa Bastankhah, Sophie Broderick, Benjamin Eysenbach
This paper proposes a novel method that incorporates empowerment when reasoning actions in reinforcement learning (RL), thereby achieving the flexibility of exploration-exploitation dilemma (EED). In previous methods, empowerment for promoting exploration has been provided as a bonus term to the task-specific reward function as an intrinsically-motivated RL. However, this approach introduces a delay until the policy that accounts for empowerment is learned, making it difficult to adjust the emphasis on exploration as needed. On the other hand, a trick devised for fine-tuning recent foundation models at reasoning, so-called best-of-N (BoN) sampling, allows for the implicit acquisition of modified policies without explicitly learning them. It is expected that applying this trick to exploration-promoting terms, such as empowerment, will enable more flexible adjustment of EED. Therefore, this paper investigates BoN sampling for empowerment. Furthermore, to adjust the degree of policy modification in a generalizable manner while maintaining computational cost, this paper proposes a novel BoN sampling method extended by Tsalis statistics. Through toy problems, the proposed method's cability to balance EED is verified. In addition, it is demonstrated that the proposed method improves RL performance to solve complex locomotion tasks.
Taisuke Kobayashi
National Institute of Informatics and The Graduate University for Advanced Studies, SOKENDAI Japan
We present an alternative characterization of the occupancy measure of reinforcement learning, obtained by embedding the planning criterion into the dynamics through a resetting planning process. Its stationary measure, which we term visitation measure, is the object on which the information geometry of decision making is most naturally expressed. The achievable visitation measures form a dually flat statistical manifold whose two affine charts are the visitation probabilities and the log-policies, dual under the conditional entropy. This structure makes planning-as-inference generalize from linear rewards to nonlinear functionals of the visitation, each iterate solved by one natural-gradient step, and gives the temporal-difference error the interpretation of a marginal-utility estimate. We develop the geometry and its consequences for reinforcement learning and theoretical neuroscience.
Nikola Milosevic, Asaki Kataoka, Nicolas Hinrichs +2
Neural Data Science and Statistical Computing Group Max Planck Institute for Human Cognitive and Brain Sciences Leipzig, 04103 Germany · Neural Computation Unit Okinawa Institute of Science and Technology Okinawa, 904-0495 Japan