Organizations: Turner Institute for Brain and Mental Health, School of Psychological Sciences, Monash University, Clayton 3800, Australia · VERSES, Los Angeles, California, USA · Cortical Labs Pty Ltd, Melbourne 3056, Australia · CIFAR Azrieli Global Scholars Program, CIFAR, Toronto, Canada
Recent and rapid advances in artificial intelligence (AI) make it increasingly important to understand the foundations of adaptive behaviour in autonomous agents, especially for building safe and efficient systems. While artificial neural networks have dominated the development of AI, recent work has begun to explore living biological neuronal networks as an alternative substrate for computation. These systems promise remarkable data and sample efficiency and rich dynamics, and may also inspire explainable and biologically plausible models. Here, we develop an experiment-informed active inference framework to model decision-making in closed-loop agents that mirror experimental setups using biological neurons. Using a generative model whose dimensions are matched to an experiment protocol, we systematically compare three decision-making schemes within this common generative model. Under matched episode counts (i.e. total data available for learning) to the in-vitro experiment, our simulations show that agents with short memory horizons reach a level of performance close to that of mouse and human cortical cultures (DishBrain platform), whereas longer memory horizons depart from it substantially. Increasing the planning horizon, by contrast, confers no comparable benefit. Because all model parameters are explicit, we can also track the quantities in our generative model that accompany this improvement, such as the risk term and the entropy of the transition and state-action mappings. Together, these results illustrate how active inference offers a formal language for comparing decision-making schemes in similar closed-loop control environments.
Figures & tables
Models and methods
Experimental groups
AI
Artificial intelligence
MCC
Mouse cortical cells
ANN
Artificial neural network
HCC
Human cortical cells
BNN
Biological neuronal network
CTL
Media-only control
SBI
Synthetic biological intelligence
RST
Rest session (spontaneous activity), no feedback
RL
Reinforcement learning
IS
In-silico control, random paddle
POMDP
Partially observable Markov decision process
Table 1: Abbreviations used in this paper. Experimental group labels follow the original DishBrain study Kagan et al. (2022) .
Figure 1: Memory horizon influences performance in CFL agents and comparison with MCC and HCC groups. Relative improvement in key performance metrics for in-silico CFL agents and in-vitro MCC and HCC cultures. Performance is evaluated by comparing the first 5 minutes (open boxes) and last 15 minutes (filled boxes) of each trial. Points on the box plots show individual sessions (in vitro) or trials (in silico); in c, g and h, where the boxes summarise per-minute (c) or per-episode (g, h) values, the points are session or trial means. Metrics include average rally length, percentage of aces (Aces are episodes with score-zero (i.e. immediate miss), lower the better), and percentage of long rallies ( ≥3 hits). a–c: CFL- T performance across discrete trial windows. d–f: Linear-regression analysis of continuous-time performance. g: Evolution of the NTE of the CL vector. h: Evolution of the average risk term Γt . i–j: Continuous-time evolution of NTE and risk. k: Relative improvement of CFL agents across memory horizons. l: Relative improvement of the experimental baseline groups. Group abbreviations: MCC, mouse cortical cells; HCC, human cortical cells; CTL, media-only control; RST, rest session, in which an active culture drove the paddle but received no sensory input; IS, in-silico control, in which the paddle was driven by random noise Kagan et al. (2022) . CFL- T denotes a counterfactual learning agent with a finite state-action update window of length T ; NTE denotes normalised total entropy. All abbreviations are collected in Table 1 .
Agent
Culture
Δ rally
95% CI
Hedges’ g
padj
Outcome
CFL-1
MCC
+0.169
[0.039, 0.336]
+0.33
0.013
agent higher
CFL-1
HCC
+0.139
[0.000, 0.312]
+0.28
0.060
not distinguishable
CFL-2
MCC
+0.161
[0.057, 0.268]
+0.45
0.005
agent higher
CFL-2
HCC
+0.131
[0.017, 0.247]
+0.37
0.032
agent higher
CFL-3
MCC
+0.588
[0.401, 0.790]
+0.86
< 0.001
agent higher
CFL-3
HCC
+0.558
[0.366, 0.765]
+0.86
< 0.001
agent higher
Table 2: Average rally length reached in the second window (6-20 min) by each in-silico agent, compared with the mouse (MCC) and human (HCC) cortical cultures. Differences and intervals come from a bootstrap that resamples whole replicates; p -values are Benjamini-Hochberg corrected across all 18 contrasts.
Figure 2: Performance of AIF-1 compared to MCC and HCC. Relative improvement in average rally length, percentage of aces, and percentage of long rallies, evaluated over the first 5 minutes (open boxes) and last 15 minutes (filled boxes). Points on the box plots show individual sessions (in vitro) or trials (in silico); in c, g and h, where the boxes summarise per-minute (c) or per-episode (g, h) values, the points are session or trial means. a–c: Discrete-window performance. d–f: Continuous-time regression. g–h: Evolution of NTE for transition dynamics and prior preferences. i–j: Continuous-time NTE evolution. k: Relative improvement of AIF-1. l: Experimental baselines. AIF-1 denotes the active inference agent with one-step planning; NTE denotes normalised total entropy. Group abbreviations: MCC, mouse cortical cells; HCC, human cortical cells; CTL, media-only control; RST, rest session, in which an active culture drove the paddle but received no sensory input; IS, in-silico control, in which the paddle was driven by random noise Kagan et al. (2022) . See Table 1 .
Figure 3: Performance of DPEFE agents compared to MCC and HCC. Structure and metric layout follow Fig. 2 . Points on the box plots show individual sessions (in vitro) or trials (in silico); in c, g and h, where the boxes summarise per-minute (c) or per-episode (g, h) values, the points are session or trial means. a–c: DP- T performance. d–f: Continuous-time regression. g–h: NTE evolution for transition and preference distributions. i–j: Continuous-time NTE evolution. k: Relative improvement of DP- T agents. l: Experimental baselines. DP- T denotes a dynamic-programming expected-free-energy agent with planning horizon T ; NTE denotes normalised total entropy. Group abbreviations: MCC, mouse cortical cells; HCC, human cortical cells; CTL, media-only control; RST, rest session, in which an active culture drove the paddle but received no sensory input; IS, in-silico control, in which the paddle was driven by random noise Kagan et al. (2022) . See Table 1 .
Figure 4: Experiment-informed generative model shared by all agents. a: The Pong court to scale ( 600×600 px, paddle 250 px). Each state factor discretises one game variable: ‘ball-x’ into 38 bins of 16 px, communicated to DishBrain as a rate code ( 4 – 41 Hz); ‘ball-y’ into 8 bins of 86 px, place-coded over 8 stimulation electrodes; and ‘paddle-y’, the paddle centre, into the same 8 bins. Filled cells mark the bins of the example ball and paddle; hatched cells lie beyond the court edge. b: The generative model over one time step. Each factor has its own transition model B and an identity likelihood A (fully observed states); the action ut∈{Stay,Up,Down} enters the ‘ball-x’ transition. c: Parameters of interest, their array shapes, and the schemes that learn and use them. All abbreviations are collected in Table 1 .
Source
Group
Replicates
Sessions
Episodes
In vitro
MCC
7 chips
110
7,317
In vitro
HCC
12 chips
138
9,528
In silico
CFL- T (each T )
100 seeds
100
7,000
In silico
AIF-1
99 seeds
99
6,930
In silico
DP- T (each T )
99 seeds
99
6,930
Table 3: Units of analysis. Episode counts are matched by design between in-vitro and in-silico groups, but the biological sessions are nested within a small number of cultures.
Group
Δ rally length
Δ % long rallies
Δ % aces
MCC
+0.147 [0.090, 0.204]
< 0.001
+3.421 [2.102, 4.740]
< 0.001
-1.580 [-3.948, 0.788]
0.274
HCC
+0.243 [0.190, 0.296]
< 0.001
+4.323 [3.426, 5.220]
< 0.001
-2.482 [-4.382, -0.583]
0.026
CFL-1
+0.147 [-0.013, 0.307]
0.139
+0.577 [-1.287, 2.440]
0.619
-2.124 [-5.276, 1.029]
0.274
CFL-2
+0.176 [0.063, 0.289]
0.007
+1.211 [-0.456, 2.877]
0.255
-0.648 [-3.859, 2.563]
0.714
CFL-3
+0.450 [0.252, 0.648]
< 0.001
+2.788 [0.872, 4.704]
0.012
-2.326 [-5.400, 0.747]
0.253
CFL-4
+1.158 [0.486, 1.830]
0.002
+5.674 [3.297, 8.051]
< 0.001
-5.314 [-8.342, -2.286]
0.002
Table S1: Change from the first (0–5 min) to the second (6–20 min) window, from a mixed-effects model with a random intercept on replicate. Each cell gives the estimate with its 95% confidence interval and the corrected p -value.
Group
Window 1
Window 2
Absolute gain
Relative improvement (%)
CFL-4
1.325
2.482
+1.157
+106.2
CFL-3
0.988
1.439
+0.450
+75.6
DP-5
0.885
1.189
+0.304
+70.9
DP-10
0.825
1.179
+0.354
+66.9
CFL-1
0.873
1.019
+0.147
+47.2
CFL-2
0.835
1.011
+0.176
+40.7
Table S2: Average rally length reached in each window, the absolute gain between them, and the relative improvement. Relative improvement normalises by each session’s own first-window mean, so it rewards a low starting point.
Group
Rally length
% long rallies
% aces
MCC
0.85 [0.81, 0.91]
9.23 [8.69, 10.49]
54.26 [52.35, 55.55]
HCC
0.88 [0.82, 0.95]
8.06 [7.17, 8.95]
53.94 [52.54, 55.07]
CFL-1
1.02 [0.90, 1.18]
11.19 [10.13, 12.28]
56.75 [54.82, 58.65]
CFL-2
1.01 [0.92, 1.11]
11.10 [9.95, 12.29]
57.09 [54.66, 59.51]
CFL-3
1.44 [1.26, 1.64]
14.10 [12.64, 15.65]
54.19 [51.80, 56.56]
CFL-4
2.47 [1.80, 3.59]
19.05 [17.08, 21.36]
47.03 [44.75, 49.25]
Table S3: Absolute performance in the second window (6–20 min), with 95% intervals from a replicate-level bootstrap.
Agent
Culture
Δ RI (%)
95% CI
Hedges’ g
padj
Outcome
CFL-1
MCC
+4.0
[-31.6, 35.6]
+0.08
0.875
inconclusive
CFL-1
HCC
-10.1
[-42.1, 21.2]
-0.11
0.725
inconclusive
CFL-2
MCC
-2.5
[-35.5, 24.9]
+0.02
0.952
inconclusive
CFL-2
HCC
-16.6
[-46.2, 10.4]
-0.21
0.481
inconclusive
CFL-3
MCC
+32.4
[-4.0, 65.2]
+0.36
0.298
inconclusive
CFL-3
HCC
+18.3
[-15.1, 50.7]
+0.19
0.482
inconclusive
Table S4: Relative improvement of each in-silico agent against the two cultures, reported for continuity with the original DishBrain analysis. Outcomes are three-way: “different” where the contrast is significant after correction, “equivalent” where the interval falls entirely within a margin of ±18 percentage points (approximately the MCC–HCC gap), and “inconclusive” otherwise.
Class
Group
Rally length (6–20 min)
Change across trial
padj
In vitro
MCC
0.85 [0.81, 0.91]
+0.15 [0.08, 0.29]
< 0.001
In vitro
HCC
0.88 [0.82, 0.95]
+0.24 [0.12, 0.36]
< 0.001
Active inference
AIF-1
0.80 [0.77, 0.84]
-0.02 [-0.09, 0.06]
0.682
Active inference
CFL-1
1.02 [0.90, 1.18]
+0.15 [0.00, 0.32]
0.093
Active inference
CFL-2
1.01 [0.92, 1.11]
+0.18 [0.07, 0.29]
0.005
Active inference
CFL-3
1.44 [1.26, 1.64]
+0.45 [0.26, 0.65]
< 0.001
Table S5: Non-active-inference baselines beside the published agents and the cultures. Absolute rally length is the second-window mean; the change across the trial is the paired within-session difference. Intervals are replicate-level bootstraps.
A
B
Δ rally
95% CI
Hedges’ g
padj
Outcome
DynaQ
CFL-1
+0.15
[-0.03, 0.29]
+0.26
0.108
comparable
DynaQ
CFL-2
+0.16
[0.04, 0.27]
+0.36
0.018
A higher
DynaQ
CFL-3
-0.27
[-0.48, -0.08]
-0.37
0.009
B higher
DynaQ
CFL-4
-1.30
[-2.43, -0.62]
-0.37
< 0.001
B higher
DynaQ
DP-5
-0.02
[-0.18, 0.12]
-0.04
0.861
comparable
DynaQ
AIF-1
+0.37
[0.29, 0.44]
+1.33
< 0.001
A higher
Table S6: Baseline contrasts on second-window rally length. Benjamini-Hochberg correction is applied within this family of twelve comparisons, which is separate from the agent-versus-culture family of the main analysis.
Agent
Parameter
TEmax
NTE 0–5
NTE 6–20
Δ
Fall (% of max)
CFL-1
CL
2671.83
0.9987
0.9956
-0.0031
1.2
CFL-2
CL
2671.83
0.9976
0.9922
-0.0054
2.0
CFL-3
CL
2671.83
0.9960
0.9884
-0.0076
2.9
CFL-4
CL
2671.83
0.9936
0.9814
-0.0122
4.3
AIF-1
B (ball-x)
414.68
0.4148
0.3616
-0.0532
83.8
DP-2
B (ball-x)
414.68
0.2758
0.2177
-0.0580
71.3
Table S7: Normalised total entropy against the analytic maximum, by time window. NTE is one when the parameter is uniform and approaches zero as it becomes fully determined, so values are comparable across agents and parameters.
Agent
Seeds
Pearson r
padj
Spearman ρ
padj
CFL-1
100
-0.205
0.238
-0.122
0.941
CFL-2
100
-0.168
0.284
-0.008
0.941
CFL-3
100
-0.054
0.671
+0.024
0.941
CFL-4
100
-0.119
0.534
+0.040
0.941
AIF-1
99
+0.062
0.671
+0.028
0.941
DP-2
99
+0.070
0.671
+0.015
0.941
Table S8: Across-seed association between the change in parameter entropy over a trial and the change in average rally length. A negative coefficient would indicate that seeds whose parameters became more informed improved more.
Setting
CFL- T
DP- T
AIF-1
Status
Episodes per trial
70
70
70
shared
Action precision
1024
1024
1024
shared
Transition learning rate
1016
1016
1016
shared
Numerical floor
10−16
10−16
10−16
shared
Prior preferences C
uniform
uniform
random categorical
differs
Initial state prior D
uniform
uniform
random categorical
differs
Table S9: Hyperparameters of the three decision-making schemes, read directly from the simulation code. All settings that the schemes share are identical; the quantities that differ are the ones that define the schemes, together with the prior specification of AIF-1.
Depth h
Normalised entropy
SD
1
0.324
0.054
2
0.513
0.066
3
0.605
0.071
4
0.666
0.075
5
0.707
0.078
10
0.810
0.088
Table S10: Predictive information of the learned transition model as a function of rollout depth. Entries are the entropy of the h -step-ahead prediction, normalised by the entropy of the uniform distribution over the same support. The value of one means the rollout carries no information about the future state. Here, normalised entropy is computed from the transition parameters of the DP-5 agent on 3 seeds, averaged over the three actions and over all starting states. The environment itself is deterministic; what decays with depth is the information carried by the agent’s estimate of it.
Agent
Condition
Seeds
0–5 min
6–20 min
Median
Δ
p
CFL-4
control
100
1.08
2.39
1.83
+1.31
< 0.001
CFL-4
obs. ε=0.05
100
2.40
6.29
2.81
+3.89
< 0.001
CFL-4
obs. ε=0.15
100
1.94
4.75
3.27
+2.81
< 0.001
CFL-4
act. η=0.10
100
1.01
1.46
1.24
+0.44
< 0.001
DP-5
control
6
0.91
1.30
1.21
+0.38
0.273
DP-5
obs. ε=0.05
6
0.85
1.12
1.10
+0.27
0.117
Table S11: Performance under observation noise and unreliable action execution. The first row of each block is the control condition, in which both manipulations are switched off and the published run is recovered.
Active inference, a neurally-inspired model for inferring actions based on the free energy principle (FEP), has been proposed as a unifying framework for understanding perception, action, and learning in the brain. Active inference has previously been used to model ecologically important tasks such as navigation and planning, but scaling it to solve complex large-scale problems in real-world environments has remained a challenge. Inspired by the existence of multi-scale hierarchical representations in the brain, we propose a model for planning of actions based on hierarchical active inference. Our approach combines a hierarchical model of the environment with successor representations for efficient planning. We present results demonstrating (1) how lower-level successor representations can be used to learn higher-level abstract states, (2) how planning based on active inference at the lower-level can be used to bootstrap and learn higher-level abstract actions, and (3) how these learned higher-level abstract states and actions can facilitate efficient planning. We illustrate the performance of the approach on several planning and reinforcement learning (RL) problems including a variant of the well-known four rooms task, a key-based navigation task, a partially observable planning problem, the Mountain Car problem, and PointMaze, a family of navigation tasks with continuous state and action spaces. Our results represent, to our knowledge, the first application of learned hierarchical state and action abstractions to active inference in FEP-based theories of brain function.
Prashant Rangarajan, Rajesh P. N. Rao
Paul G. Allen Center for Computer Science and Engineering, University of Washington, Seattle, USA
Sophisticated Inference is a variant of active inference often associated with recursive belief modeling and tree search. We argue that its central computational role is simpler: within a planning horizon, it makes active inference closed-loop by allowing future actions to depend on future states and observations. This closed-loop structure can be represented in the epistemic-prior variational free energy framework. Epistemic priors supply the active-inference objective, while a joint posterior over future states and actions supplies the state-contingent control structure. We evaluate this decomposition in the Reactivity Maze, a stochastic benchmark designed to separate epistemic incentive from inner-horizon closed-loop control. The comparison includes three variational objectives with the same state-action posterior family, an action-state factorized active inference objective, Sophisticated Inference, and standard Expected Free Energy planning. The results show that neither ingredient is sufficient on its own. Methods without an epistemic component do not seek information, while methods that prevent future actions from depending on future states cannot turn information into reliable goal-reaching. By contrast, both Sophisticated Inference and full-joint epistemic-prior active inference solve the environment by combining epistemic drive with closed-loop inference. These results show that the advantage associated with Sophisticated Inference need not be specific to tree search itself. It arises from the closed-loop form of active inference, and this form can be represented in epistemic-prior variational inference when the posterior keeps future actions dependent on future states.
Wouter W. L. Nuijten, Bert de Vries
Eindhoven University of Technology, 5612 AP Eindhoven, the Netherlands · Lazy Dynamics, Utrecht, the Netherlands
Active inference casts decision-making as inference, with the Expected Free Energy (EFE) unifying goal-directed and information-seeking behavior. Recent work showed that EFE minimization can be written as Variational Free Energy (VFE) minimization on a generative model augmented with epistemic priors. We prove that the VFE of the augmented model can be rewritten as the VFE of the predictive model plus explicit entropy-correction terms, making the EFE contribution transparent. We then show that proper EFE-based planning requires combining these epistemic corrections with a planning correction that turns marginal inference into policy optimization, yielding a full variational characterization of EFE-based planning. This clarifies which corrections are needed for cross-entropy planning and for full EFE-based planning. The same entropy-corrected formulation leads to a detailed message-passing scheme for EFE-based planning together with simpler ablations. Experiments on three grid-world environments show that full EFE-based planning outperforms ablations that omit either the planning correction or the epistemic corrections.
Wouter W. L. Nuijten, Mykola Lukashchuk, Thijs van de Laar +1
Department of Electrical Engineering, Eindhoven University of Technology, Eindhoven, the Netherlands · Lazy Dynamics, Utrecht, the Netherlands