cs.LGJul 22, 2026

Active Inference as a Convex Markov Decision Process

Authors: Nikola MilosevicNicolás HinrichsNico Scherf

Organizations: Max Planck Institute for Human Cognitive and Brain Sciences Leipzig, Germany

Abstract

Active Inference (AIF) frames adaptive behavior as the minimization of expected free energy (EFE), combining epistemic and pragmatic objectives within a single variational principle. We frame AIF as policy optimization and show that, for closed-loop control policies, EFE minimization can be formulated as a convex Markov decision process (MDP). This perspective reveals that policy-dependent reward prediction errors transmit natural gradients of the expected free energy backwards in time rather than up a hierarchy. Finally, we show that coupling world-model learning with policy optimization gives active inference the structure of performative reinforcement learning. Together this places EFE minimization within modern reinforcement learning and optimization theory and opens a route toward principled algorithms for active inference.

Explore similar work

CardsList
  1. What Type of Inference is Active Inference?

    Jun 3, 2026Wouter W. L. Nuijten, Mykola Lukashchuk, Thijs van de Laar +1Classical PlanningFree Energy Principle

  2. Expected Free Energy-based Planning as Variational Inference

    Jun 9, 2026Wouter W. L. Nuijten, Thijs van de Laar, Bert de VriesVariational InferenceClassical Planning