Intrinsic motivation plays a central role in adaptive and goal-directed behavior by conferring agents reward-independent objectives and biases useful to act in noisy and uncertain environments. Active Inference addresses the problem of acting in a partially observable environment through a principled framework for belief updating and action selection. A key component of Active Inference is the specification of prior preferences, which shapes behavior by encoding desirable future outcomes. An intrinsic motivation approach called the Maximum Occupancy Principle (MOP) proposes that agents act so as to maximize occupancy over future paths of states and actions, with no preferences or epistemic targets. Despite its simple formulation, MOP gives rise to rich and adaptive behaviors that combine exploratory variability with goal-directed dynamics. In this work, we extend MOP to partially observable environments and introduce a Bellman reformulation of the Expected Free Energy for Active Inference, both incorporating belief-based inference over hidden states as part of the agent state. The Bellman formulation enables tractable offline computation via value iteration over the full belief-state space. We compare the resulting behaviors in a set of minimal experimental settings with uncertain food sources. We find that MOP agents switch between goal-directed (food seeking) behavior and exploration between different food sources, depending on their energy available and their belief state. In contrast, Active Inference agents mostly inhabit regions around a single food source, a strategy having both high pragmatic and epistemic value. We finally compare with Empowerment, which is shown to be qualitatively similar to Active Inference.
Figure 2: Space habitation heatmaps (average over 100 episodes) and a sample trajectory for MOP, EFE and Empowerment (MPOW), under independent (top) and depletion (bottom) food dynamics. Tˉ denotes average survival duration in steps.
Figure 3: Normalized policy entropy, survival and exploration for MOP and Active Inference, in the location-independent dynamics; depletion-at-visit shows similar behavior. (a) MOP exhibits a transition when the agent’s internal energy is around E=15 , switching between broad exploration at high energy and goal-directed behavior at low energy (insets are corresponding state visitation heatmaps). (b) Active Inference policy entropy for six values of the inverse temperature d . Every curve increases with energy, but d modulates the behavior: for small d the policy entropy remains high and the policy stays stochastic, whereas for large d it falls toward zero at every energy level and the policy becomes nearly deterministic (see heatmaps in Appendix A.9 ). (c) Average survival duration of the Active Inference agent as a function of d , compared with MOP. (d) Entropy of the empirical state-visitation distribution over the grid, over the same range of d , compared with MOP. Panels (c) and (d) are approximately mirror images of each other: the values of d at which its occupancy of the environment collapses (e) The same trade-off as a single curve, survival duration against the entropy of the empirical state-visitation distribution over the grid, one point per value of d , with MOP as a single point lying off the trace, indicated best survival for matched fixed state entropy, or best state entropy for matched fixed survival.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Variable
Value
Emax
30
Egain
15
Grid size
5×5
Food source locations
(1,1) and (5,5)
Initial h1
1
Initial h2
1
Appendix
Table 1: Parameters used in the experiments.
Figure 4: Value function V∗(s,b) as a function of belief b1 (with b2=0.5 fixed) for three representative states, computed with a fine grid ( Δ=0.01 , black curve) and a coarse grid ( Δ=0.1 , blue curve). The coarse grid points closely follow the fine curve in all cases, confirming that the Δ=0.1 discretization introduces negligible approximation error.
Figure 5: Behavior of an active inference agent under stochastic policy, for various values of the inverse temperature, in the case of location-independent dynamics ( G+=100 , 100 episodes per panel). One complete episode is drawn on each panel (start ∙ , end × ); the line retraces itself, so a committed agent appears as a few repeated edges rather than as a short path. In panels (b)–(f) the episodes are conditioned on the food source reached first, so that each heatmap shows a single commitment rather than a mixture of two.
Figure 6: (a, b) Per-episode behavior of the active inference agent across precision, 300 episodes per point; MOP (black) is drawn in both panels as a reference. (a) Number of trips between the two sources per 103 steps, over the range of precision d . (b) Share of time spent standing on a food source. (c) Occupancy resolved by Manhattan distance to the nearest food source, for MOP and for the agent at d=0.8 , the precision whose lifespan comes closest to MOP’s: with lifespans matched, the active inference agent still spends 41% of its time standing on a source against MOP’s 26% , while MOP retains the occupancy at distances 3 and 4 that the active inference agent has largely given up.
Figure 7: Normalized policy entropy of the EFE agent against energy level, for three terminal boundary conditions G+ , with the MOP agent (black) as a fixed reference in each panel. MOP value function has no EFE terminal, so the same curve is the appropriate comparison throughout. Curves are labeled by the inverse temperature d .
Sophisticated Inference is a variant of active inference often associated with recursive belief modeling and tree search. We argue that its central computational role is simpler: within a planning horizon, it makes active inference closed-loop by allowing future actions to depend on future states and observations. This closed-loop structure can be represented in the epistemic-prior variational free energy framework. Epistemic priors supply the active-inference objective, while a joint posterior over future states and actions supplies the state-contingent control structure. We evaluate this decomposition in the Reactivity Maze, a stochastic benchmark designed to separate epistemic incentive from inner-horizon closed-loop control. The comparison includes three variational objectives with the same state-action posterior family, an action-state factorized active inference objective, Sophisticated Inference, and standard Expected Free Energy planning. The results show that neither ingredient is sufficient on its own. Methods without an epistemic component do not seek information, while methods that prevent future actions from depending on future states cannot turn information into reliable goal-reaching. By contrast, both Sophisticated Inference and full-joint epistemic-prior active inference solve the environment by combining epistemic drive with closed-loop inference. These results show that the advantage associated with Sophisticated Inference need not be specific to tree search itself. It arises from the closed-loop form of active inference, and this form can be represented in epistemic-prior variational inference when the posterior keeps future actions dependent on future states.
Wouter W. L. Nuijten, Bert de Vries
Eindhoven University of Technology, 5612 AP Eindhoven, the Netherlands · Lazy Dynamics, Utrecht, the Netherlands
Planning under uncertainty requires agents to balance goal achievement with information gathering. Active inference addresses this through the Expected Free Energy (EFE), a cost function that unifies instrumental and epistemic objectives. However, existing EFE-based methods typically employ specialized optimization procedures that are difficult to extend or analyze. In this paper, we show that EFE-based planning can be formulated as Variational Free Energy minimization on a generative model augmented with epistemic priors. Our main result demonstrates that minimizing a Variational Free Energy functional with appropriately chosen priors yields a decomposition into expected plan costs (the EFE) plus a complexity term. This formulation reinforces theoretical consistency with the Free Energy Principle by casting planning as the same inferential process that governs perception and learning. We validate our approach on three environments of increasing complexity: a deterministic T-maze, a stochastic Reactivity Maze, and a partially observable MiniGrid DoorKey-8x8 environment. The experiments demonstrate that the epistemic priors induce information-seeking behavior, that the variational formulation yields policy-based inference outperforming plan-based methods under stochastic transitions, and that temporal factorization enables scalability to environments where existing tabular active inference methods cannot operate.
Wouter W. L. Nuijten, Thijs van de Laar, Bert de Vries
Eindhoven University of Technology, Eindhoven, the Netherlands · Lazy Dynamics B.V., the Netherlands
Active inference casts decision-making as inference, with the Expected Free Energy (EFE) unifying goal-directed and information-seeking behavior. Recent work showed that EFE minimization can be written as Variational Free Energy (VFE) minimization on a generative model augmented with epistemic priors. We prove that the VFE of the augmented model can be rewritten as the VFE of the predictive model plus explicit entropy-correction terms, making the EFE contribution transparent. We then show that proper EFE-based planning requires combining these epistemic corrections with a planning correction that turns marginal inference into policy optimization, yielding a full variational characterization of EFE-based planning. This clarifies which corrections are needed for cross-entropy planning and for full EFE-based planning. The same entropy-corrected formulation leads to a detailed message-passing scheme for EFE-based planning together with simpler ablations. Experiments on three grid-world environments show that full EFE-based planning outperforms ablations that omit either the planning correction or the epistemic corrections.
Wouter W. L. Nuijten, Mykola Lukashchuk, Thijs van de Laar +1
Department of Electrical Engineering, Eindhoven University of Technology, Eindhoven, the Netherlands · Lazy Dynamics, Utrecht, the Netherlands