Learning from trajectory demonstrations offers a route to active-inference control of complex systems whose dynamics are difficult to model explicitly. We introduce generative active-inference control (GenAIF), in which one generative trajectory model learns from demonstrations and measured action interventions to supply a goal-conditioned policy distribution and a state-to-observation likelihood mapping. From this control design, we derive three model requirements: (i) useful action proposals, (ii) accurate prediction under imposed actions, and (iii) probabilistic observation evidence for belief updating and expected information gain. We benchmark diffusion, autoregressive Transformers, conditional variational autoencoders (CVAEs), and flow matching in a MuJoCo manipulation task with multiple physical conditions. Diffusion delivers the strongest control across the tested dynamics, while CVAE combines comparable short-horizon prediction with much faster inference. Correct conditioning is decisive, and trajectory reuse offers further computational savings. With the same frozen models, a hidden-dynamics experiment demonstrates prompt belief adaptation after an unannounced tilt change; subsequent instability identifies sustained inference as a remaining challenge. These findings support the use of shared generative trajectory models to connect action proposal, controlled prediction, and observation evidence within GenAIF.
Figures & tables
Figure 1: Ball-pushing task at illustrative lateral tilts. The pusher moves the ball around the obstacle toward the green goal; lines illustrate routes.
Figure 2: E1 control and prediction. (a) Control success with correct or incorrect tilt inputs (zero-lateral-tilt subset: 30 episodes; all tilts: 90 per group). (b) Mean Euclidean ball-position error of the predictive mean over six steps. Whiskers and shading show 95% bootstrap intervals over 10 sampling seeds (a) and 18 source trajectories (b); bands are pointwise. AR: autoregressive Transformer.
Success /90
H=6 position
Time/action (ms)
Model
Direct
Reuse
RMSE (mm)
Direct
Reuse
Diffusion
81
68
14.295
57.123
9.794
Autoregressive
2
2
65.809
48.735
12.982
CVAE
49
61
15.264
5.643
1.236
Flow matching
43
34
24.776
199.571
37.037
Table 1: E1 control, prediction, and computation. Position RMSE uses 841 valid endpoints. Times cover model evaluation and checking per executed action, excluding simulator execution.
Model
Accuracy (%)
Log loss (nats)
Calibration error
Diffusion
98.96
0.540
0.062
Autoregressive
79.28
0.959
0.127
CVAE
99.07
0.434
0.057
Flow matching
99.07
0.474
0.069
Table 2: E1 hidden-tilt identification over 864 action–observation pairs. Lower log loss and calibration error indicate better probability estimates.
Figure 3: E2 belief tracking after a 0∘→+15∘ tilt switch. Lines show mean p(+15∘) across 12 trials; observation 11 first reflects the switch. Shading shows pointwise 95% bootstrap intervals over source episodes. The horizontal line marks the uniform prior.
Model
Identification delay median [min,max]
Identified by six
Correct at observation 30
Diffusion
3 [2,4]
12/12
5/12
Autoregressive
5 [5,8]
10/12
7/12
CVAE
4 [3,10]
10/12
0/12
Flow matching
4 [3,4]
12/12
0/12
Table 3: E2 adaptation across 12 switched trials. Delay counts observations after the switch until p(+15∘) first reaches 0.9; final correctness uses the most likely hypothesis at observation 30.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Family
Success
Feasible (%)
Proposal (ms)
Plan (ms)
BC-MDN
11/18
82.4
0.75
346
CVAE
17/18
92.1
0.39
352
Diffusion
18/18
100.0
50.05
411
Flow
14/18
83.9
3.86
354
Autoregressive
14/18
82.1
14.03
360
Appendix
Table A1: Proposal-family comparison at K=128 . Feasibility averages the fraction of feasible candidates per rollout; times are per decision.
Figure A1: Proposal-only control with MuJoCo transitions. Success versus (A) candidate budget, (B) measured planning time, and (C,D) requested CPU targets for nominal and new start–goal scenes. Diffusion exceeds the two smallest CPU targets, so pooled comparisons there are not at equal measured cost.
Success /18
Plan time (ms)
Method
K
Nominal
Transfer
Nominal
Transfer
CEM
8
0
0
44.5
44.4
CVAE
16
16
18
43.8
43.5
Appendix
Table A2: CEM and CVAE at a 50 ms CPU target. Transfer uses new start–goal scenes; planning time includes proposal generation, rollout, and scoring.
University of Tübingen, Tübingen, Germany · Division of the Social Sciences, University of Chicago, Chicago, IL, USA · Heidelberg University, Heidelberg, Germany +3