Intrinsic motivation plays a central role in adaptive and goal-directed behavior by conferring agents reward-independent objectives and biases useful to act in noisy and uncertain environments. Active Inference addresses the problem of acting in a partially observable environment through a principled framework for belief updating and action selection. A key component of Active Inference is the specification of prior preferences, which shapes behavior by encoding desirable future outcomes. An intrinsic motivation approach called the Maximum Occupancy Principle (MOP) proposes that agents act so as to maximize occupancy over future paths of states and actions, with no preferences or epistemic targets. Despite its simple formulation, MOP gives rise to rich and adaptive behaviors that combine exploratory variability with goal-directed dynamics. In this work, we extend MOP to partially observable environments and introduce a Bellman reformulation of the Expected Free Energy for Active Inference, both incorporating belief-based inference over hidden states as part of the agent state. The Bellman formulation enables tractable offline computation via value iteration over the full belief-state space. We compare the resulting behaviors in a set of minimal experimental settings with uncertain food sources. We find that MOP agents switch between goal-directed (food seeking) behavior and exploration between different food sources, depending on their energy available and their belief state. In contrast, Active Inference agents mostly inhabit regions around a single food source, a strategy having both high pragmatic and epistemic value. We finally compare with Empowerment, which is shown to be qualitatively similar to Active Inference.
Figure 2: Space habitation heatmaps (average over 100 episodes) and a sample trajectory for MOP, EFE and Empowerment (MPOW), under independent (top) and depletion (bottom) food dynamics. Tˉ denotes average survival duration in steps.
Figure 3: Normalized policy entropy, survival and exploration for MOP and Active Inference, in the location-independent dynamics; depletion-at-visit shows similar behavior. (a) MOP exhibits a transition when the agent’s internal energy is around E=15 , switching between broad exploration at high energy and goal-directed behavior at low energy (insets are corresponding state visitation heatmaps). (b) Active Inference policy entropy for six values of the inverse temperature d . Every curve increases with energy, but d modulates the behavior: for small d the policy entropy remains high and the policy stays stochastic, whereas for large d it falls toward zero at every energy level and the policy becomes nearly deterministic (see heatmaps in Appendix A.9 ). (c) Average survival duration of the Active Inference agent as a function of d , compared with MOP. (d) Entropy of the empirical state-visitation distribution over the grid, over the same range of d , compared with MOP. Panels (c) and (d) are approximately mirror images of each other: the values of d at which its occupancy of the environment collapses (e) The same trade-off as a single curve, survival duration against the entropy of the empirical state-visitation distribution over the grid, one point per value of d , with MOP as a single point lying off the trace, indicated best survival for matched fixed state entropy, or best state entropy for matched fixed survival.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Variable
Value
Emax
30
Egain
15
Grid size
5×5
Food source locations
(1,1) and (5,5)
Initial h1
1
Initial h2
1
Appendix
Table 1: Parameters used in the experiments.
Figure 4: Value function V∗(s,b) as a function of belief b1 (with b2=0.5 fixed) for three representative states, computed with a fine grid ( Δ=0.01 , black curve) and a coarse grid ( Δ=0.1 , blue curve). The coarse grid points closely follow the fine curve in all cases, confirming that the Δ=0.1 discretization introduces negligible approximation error.
Figure 5: Behavior of an active inference agent under stochastic policy, for various values of the inverse temperature, in the case of location-independent dynamics ( G+=100 , 100 episodes per panel). One complete episode is drawn on each panel (start ∙ , end × ); the line retraces itself, so a committed agent appears as a few repeated edges rather than as a short path. In panels (b)–(f) the episodes are conditioned on the food source reached first, so that each heatmap shows a single commitment rather than a mixture of two.
Figure 6: (a, b) Per-episode behavior of the active inference agent across precision, 300 episodes per point; MOP (black) is drawn in both panels as a reference. (a) Number of trips between the two sources per 103 steps, over the range of precision d . (b) Share of time spent standing on a food source. (c) Occupancy resolved by Manhattan distance to the nearest food source, for MOP and for the agent at d=0.8 , the precision whose lifespan comes closest to MOP’s: with lifespans matched, the active inference agent still spends 41% of its time standing on a source against MOP’s 26% , while MOP retains the occupancy at distances 3 and 4 that the active inference agent has largely given up.
Figure 7: Normalized policy entropy of the EFE agent against energy level, for three terminal boundary conditions G+ , with the MOP agent (black) as a fixed reference in each panel. MOP value function has no EFE terminal, so the same curve is the appropriate comparison throughout. Curves are labeled by the inverse temperature d .