Organizations: Department of Computer Science, Virginia Tech, Alexandria, VA 22305, USA · Department of Electrical and Computer Engineering, Virginia Tech, Alexandria, VA 22305, USA
Epidemic policy planning often requires coordination between geographical regions, taking into account mobility-driven spillovers and how to make use of limited resources. Existing methods either lack action-conditioned models of coupled dynamics or cannot guarantee per-period feasibility. We present EpiMind, a graph world model framework for constrained epidemic policy planning across regions. A graph-factored recurrent state-space model generates joint policy-conditioned rollouts from regional latent beliefs, while graph-temporal ADMM optimizes regional interventions, enforces shared-resource feasibility through projection, and evaluates temporal specifications under the learned model. EpiMind reduces admission RMSE by 29% relative to graph-free dynamics modeling, plans within 1-5% of the best feasible constant policy with guaranteed shared-budget feasibility, and outperforms all deployable baselines across three resource budgets in real-context evaluation. These results demonstrate that graph-structured policy imagination with explicit constrained coordination supports effective epidemic interventions from learned dynamics.
Figures & tables
Figure 1: EpiMind framework. A parameter-shared GF-RSSM updates regional beliefs and generates graph-coupled policy rollouts. GT-ADMM coordinates and projects regional actions, rerolls the feasible allocation for model-relative STL evaluation, and executes its first action.
Model
Adm. MAE ↓
Adm. RMSE ↓
Cum. err. @5 ↓
Params
Statistical baselines
Persistence
1.563 ± 0.053
1.927 ± 0.056
0.089 ± 0.003
0
Climatology
1.129 ± 0.021
1.283 ± 0.027
0.528 ± 0.011
0
Ridge + action
0.341 ± 0.008
0.437 ± 0.003
0.068 ± 0.006
84
VARX(1) + action
0.329 ± 0.008
0.424 ± 0.001
0.066 ± 0.007
1,960
Learned dynamics models
Table 1: World-model fidelity in the synthetic environment. Values are mean ± s.d. Admission errors are measured per 100K; cumulative error is over a five-step rollout. Lower is better.
Figure 2: Policy-conditioned admission response. (a) Peak weekly admissions under alternative NPI intensities applied from a common synthetic state. (b) Predicted cumulative admissions under the same sweep at four held-out U.S. decision origins.
Objective J↓
Regret at λ=10↓
Method
λ=3
λ=10
λ=30
mean ± s.d.
Best feasible constant †
54.7
58.9
70.9
0.00 ± 0.00
PPO
60.4
64.6
76.6
5.69 ± 3.71
MPC-SEIR
108.6
112.8
124.6
53.9 ± 30.9
Independent MPC
366.3
431.1
394.1
372 ± 307
Greedy
7,186
7,189
7,197
7,130 ± 1,221
Table 2: Matched-objective planning performance. Values are mean ± s.d. Lower is better.
Setting
Feasible (%)
Max excess
Projection displacement
STL satisfaction (%)
STL robustness
Synthetic
100.0
0
0.476 ± 0.201
100.00
0.0021 ± 0.0004
Real context
100.0
0
0.0039 ± 0.0400
100.00
0.0016 ± 0.0003
Table 3: Constraint handling and model-relative verification.
Method
Adm./100K ↓
NPI burden
J†↓
ΔJ (%)
EpiMind wins ( p )
Standard ADMM
30.24
0.423
31.51
+1.14
23/36 (0.132)
Centralized MPC
30.65
0.424
31.92
+2.45
32/36 ( <0.001 )
Independent MPC
30.94
0.424
32.21
+3.40
26/36 (0.011)
Uniform allocation
31.23
0.479
32.67
+4.87
30/36 ( <0.001 )
EpidRLearn
35.64
0.500
37.14
+19.23
36/36 ( <0.001 )
PPO
35.69
0.499
37.19
+19.37
36/36 ( <0.001 )
Table 4: Realized policy performance in the real-context semi-synthetic evaluation.
Figure 3: Real-context policy comparisons. Each point reports the mean paired admission difference ΔAdm=Admcomparator−AdmEpiMind ; positive values (blue circles) favor EpiMind; horizontal bars denote 95% confidence intervals. (a) Retrospective Track A reports model-relative projections at held-out U.S. state-level decision origins. (b) Semi-synthetic Track B reports realized outcomes under known simulation dynamics. Diamonds denote the track-specific reference policy.
Figure 4: Texas case study. (a) Admission forecasts under the observed policy; the inset enlarges weeks 50–60. (b) Model-relative counterfactual trajectories under matched resource constraints; the observed trajectory provides context but is not an outcome of the unexecuted policies. (c) EpiMind’s first-step allocation compared with historical actions in Texas and neighboring states. The dashed vertical line marks the training cutoff. MPC and RL baselines appear only in panel (b).
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A1: Multi-region policy graph under shared resource constraints. Regions are represented as nodes vi in a dynamic policy graph Gt=(Vt,Et) , with edges encoding interregional mobility coupling. Each region receives allocations of vaccine, hospitalization-capacity, and fiscal resources. The joint regional allocations must satisfy graph-level resource budgets, ∑i=1Nat,ki≤Bt,k , for k∈{v,h,f} .
Challenge
Epidemiological property
Component
Computational response
(A)
Latent disease burden
GF-RSSM
Latent belief inference
Policy-dependent surveillance
GF-RSSM
Action-conditioned observation model
Delayed intervention effects
GF-RSSM
Multi-horizon policy-conditioned rollout
Cross-region spillovers
GF-RSSM
Graph attention over neighbors
(B)
Shared resource limits
GT-ADMM
Capped-simplex projection
Cross-border coordination
GT-ADMM
Neighbor-consensus signal
Appendix
Table A1: Epidemiological challenges and their computational treatment in EpiMind.
Figure A2: Extended world-model evaluation on the synthetic benchmark. (a) Held-out admission MAE for learned architectures and statistical baselines under the same evaluation protocol. Error bars for learned models show mean ± s.d. over three training seeds. (b) Relative cumulative admission error over open-loop rollout horizons H∈{5,10,20} , shown on a logarithmic scale.
Figure A3: Sensitivity to operational and model perturbations. Changes in cumulative hospitalizations relative to each seed-matched baseline, reported as mean ± s.d. The shaded band shows the largest within-condition standard deviation among the operational perturbations. Graph noise, masked regional observations, reporting delays, a mid-horizon budget cut, and regional non-compliance remain within this descriptive variability band. Epidemiological model mismatch produces the only substantially larger mean degradation.
λ
H
Regret ↓
Model effect ↓
p†
Learned
Oracle
3
4
32.5 ± 32.8
27.8 ± 3.2
4.7 ± 34.6
1.000
3
8
38.2 ± 50.0
17.6 ± 0.7
20.5 ± 50.3
1.000
3
12
32.5 ± 34.2
20.6 ± 0.2
11.9 ± 34.2
1.000
10
4
217.0 ± 229.3
103.2 ± 3.8
113.9 ± 231.3
1.000
10
8
184.9 ± 201.1
79.3 ± 0.4
105.6 ± 200.9
1.000
Appendix
Table A3: Learned- versus oracle-dynamics planning regret. Values are mean ± s.d.
Regret ↓
vs. decoder
vs. uncalibrated
λ
Calibrated
Shared dec.
Uncalib.
cells
p
cells
p
3
56.5 ± 50.5
72.8 ± 85.4
68.8 ± 42.2
5/9
1.000
7/9
0.180
10
252.8 ± 242.3
705.3 ± 586.5
364.9 ± 309.3
8/9
0.039
9/9
0.004
30
641.4 ± 474.1
2401.9 ± 906.1
882.1 ± 634.7
9/9
0.004
7/9
0.180
Appendix
Table A4: Calibration ablation. Values are mean ± s.d. over n=9 cells.
Variant
γ
η
μ
Cum. adm./100K ↓
Δ (%)
Graph-free
–
–
–
662.57
+0.88
No edge ( γ off)
–
✓
✓
659.87
+0.47
No spillover ( η off)
✓
–
✓
660.35
+0.54
No global ( μ off)
✓
✓
–
662.99
+0.94
EpiMind
✓
✓
✓
656.81
—
Appendix
Table A5: Matched-burden coordination ablation. Simulator-realized outcomes on the synthetic benchmark. All variants follow the reference configuration’s stepwise NPI-burden trajectory. Here, γ denotes neighbor consensus, η spillover sensitivity, and μ global resource consensus. Δ is the percentage change relative to the full model; lower is better.
Figure A4: Real interstate mobility graph. Row-normalized Advan device-mobility flows among ten selected high-flow U.S. states in January 2021. Rows denote origins, columns denote destinations, and color indicates each destination’s share of an origin’s outgoing travel among the displayed states. The asymmetric, nonuniform matrix provides mobility-edge weights to the graph world model.
Figure A5: Texas real-context evaluation. (a) Observed admissions, the held-out forecast under historical actions, and the model-relative rollout under EpiMind’s proposed policy. The inset enlarges weeks 50–62; shading denotes the predicted difference between the historical- and proposed-policy rollouts. (b) Forecast RMSE versus model-relative admissions projected under each model’s optimized policy. Dashed lines mark the observed-policy admission mean and the selected forecast-error reference. Lower values are preferable on both axes.
Figure A6: Multi-state forecasting and planner projections. Observed weekly admissions (gray), forecasts under the recorded policy (blue), and model-relative rollouts under EpiMind’s proposed policy (red) for ten mobility-connected U.S. states. The vertical dashed line marks the training cutoff. Planner projections represent unexecuted counterfactuals, not observed outcomes.