Artificial Intelligence agents based on large language models provide convenient natural language interfaces to scientific software, but reliability is not automatic. Here we introduce the Epydemix Agent Framework, an additive layer over Epydemix, an open-source Python library for stochastic compartmental epidemic modeling. The framework extends the library with four capabilities to facilitate interaction with an AI agent: discovery of available models and parameters, preventive validation of a declarative scenario specification, execution through tested library code, and inspectability of results. These capabilities let an agent handle the entire modeling process, from the natural-language description of the scenario to quantitative results, figures, and interpretation of findings without writing custom code. Each step reads input files and saves results in a separate output bundle, making the process auditable and reproducible. First, we show the end-to-end workflow with a case study comparing vaccination strategies for a novel respiratory virus. Second, we assessed the framework across 50 agent sessions and five modeling tasks by comparing the agent use of the framework against the direct use of the Python interface. The framework reduced turns, output tokens, and cost on most tasks, unless it trades resources for per-point reproducibility.
Figures & tables
Figure 1: Canonical four-phase workflow of the Epydemix Agent Framework. Starting from a natural-language scenario description, an AI agent progresses through four sequential phases: (1) Discover : enumerate available models and retrieve literature-sourced parameter defaults; (2) Declare : encode the scenario as a YAML configuration file and validate it against the parameter registry before any computation; (3) Run : execute simulation, calibration, or projection, writing results to a self-contained .epx output bundle; (4) Inspect : query the bundle for compact, structured summaries and compare outcomes across scenarios. The framework is stateless: each phase operates exclusively on files, with no Python objects persisting across calls.
Figure 2: The components of the Epydemix Agent Framework. The agent contract ( AGENT.md , top) is the document an agent reads first. The command-line interface (middle) is the single control surface for both an AI agent and a human researcher, and the inspection engine is part of it. The three data components (bottom) are the filesystem-backed artifacts that the framework reads and writes. The agent contract and tooling create a layer over the existing Epydemix Python API without modifying it.
Figure 3: Hospital occupancy over the 16-month simulation horizon under the four vaccination strategies, for the representative framework run reported in Section 3.1 , with the 55,000 -bed capacity threshold marked (dashed red line). Solid lines show the median across 100 stochastic simulations; shaded bands show the 90% interval.
Figure 4: Compute-budget distribution across the eight case-study runs. Top (A): total cost per run in USD, stacked by phase (solid: planning; hatched: execution), colored by condition (green: framework, run1–run4; orange: direct API, run5–run8). Bottom (B): wall-clock running time per run in seconds, stacked in the same way. The rightmost bar in each condition group shows the across-run average, with error bars giving ±1 standard deviation for each phase and the percentage of the total contributed by planning and execution annotated alongside.
Metric (CV = std/mean)
Phase
Framework
Direct API
Total cost
Planning
0.10
0.19
Execution
0.15
0.35
Total duration
Planning
0.15
0.30
Execution
0.17
0.40
Table 1: Coefficient of variation across the four case-study runs of each condition, for cost (USD) and running time (seconds), broken down by planning-phase and execution-phase values. Lower values indicate greater run-to-run consistency. The higher, less consistent value in each row is shown in bold.
Task
Wall clock time (s)
Time in API (s)
# Turns
# Output tokens
Cost ($)
SEIRHD scenarios
1.05 [0.85, 4.78]
0.98 [0.73, 5.14]
0.77 [0.60, 0.96]
0.79 [0.68, 0.91]
0.91 [0.81, 1.04]
SIR calibration
0.73 [0.59, 0.95]
0.57 [0.53, 0.70]
0.62 [0.55, 0.72]
0.58 [0.49, 0.67]
0.71 [0.65, 0.95]
Calibrate → Project
0.53 [0.45, 0.57]
0.50 [0.43, 0.55]
0.63 [0.53, 0.70]
0.45 [0.37, 0.45]
0.52 [0.49, 0.73]
School-closure sweep
2.04 [1.24, 3.27]
1.41 [1.25, 2.24]
1.38 [1.00, 2.21]
1.32 [1.18, 2.23]
1.81 [1.34, 3.34]
Measles coverage sweep
0.81 [0.42, 1.12]
0.83 [0.62, 0.93]
0.80 [0.68, 0.89]
0.71 [0.62, 0.93]
0.88 [0.67, 1.05]
Table 2: Framework performance relative to direct use of the Epydemix Python API, by task. Each cell reports the ratio of medians (framework / direct API) over 5 runs per condition, together with an exact 96.8% distribution-free confidence interval for the ratio (Wilcoxon rank-sum). Values below 1 indicate lower resource use in the framework condition. Values in bold indicate ratios whose confidence interval excludes 1 .
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Phase
Metric
With framework
Without framework
Planning
Duration (s)
240.2±33.0
504.0±168.0
Cost ($)
1.43±0.17
2.73±0.60
Execution
Duration (s)
342.2±59.0
331.0±99.0
Cost ($)
1.55±0.27
2.16±0.86
Total
Duration (s)
582.5±82.6
835.0±128.8
Cost ($)
2.98±0.38
4.89±0.78
Appendix
Table 1: Planning, execution, and total cost and duration, with-framework vs. without-framework ( n=4 per condition).
Metric
With framework
Without framework
Peak occupancy (mean ± std)
81,277±4,279
80,587±6,843
Peak occupancy (range)
78,223 – 87,318
73,630 – 87,849
Attack rate % (mean ± std)
72.5±2.3
71.6±4.1
Days over capacity
28–31 (all 4 runs)
26–31 (all 4 runs)
Appendix
Table 2: Slow rollout outcomes, with-framework vs. without-framework.
Metric
With framework
Without framework
Peak occupancy (mean ± std)
12,634±21,707
20,674±10,046
Peak occupancy (range)
1,090 – 45,161
11,187 – 34,872
Attack rate % (mean ± std)
14.5±22.8
25.9±10.6
Days over capacity
0 (all 4 runs)
0 (all 4 runs)
Appendix
Table 3: Rapid rollout outcomes, with-framework vs. without-framework.
Metric
With framework
Without framework
Peak occupancy (mean ± std)
19,944±7,428
33,697±6,470
Attack rate % (mean ± std)
33.5±8.5
42.2±6.4
Days over capacity
0 (all 4 runs)
0 (all 4 runs)
Appendix
Table 4: Combined strategy outcomes, with-framework vs. without-framework.
Parameter
Condition
Slow / school
Rapid / work
Combined / window
Vaccination rate (%/day)
With framework
0.115±0.015
1.025±0.327
0.425±0.075
Vaccination rate (%/day)
Without framework
0.181±0.064
1.028±0.357
0.464±0.097
NPI reduction (%)
With framework
62.5±13.0
47.5±10.9
82.0±15.0 days
NPI reduction (%)
Without framework
55.0±5.0
38.8±7.4
85.0±8.7 days
Appendix
Table 5: Assumed vaccination rates (slow/rapid/combined) and NPI intervention magnitudes (school/work reduction, window length), mean ± std across runs.
Setting
Value
Agent runtime
Claude Code 2.1.224, non-interactive
Model
claude-sonnet-5
Reasoning effort
medium
Tools available
Bash , Read , Write , Edit ,
Glob , Grep , BashOutput , KillShell
Tools withheld
sub-agents, task planning, web search and fetch
Appendix
Table 6: Agent harness configuration.
Task
Metric
Baseline
Framework
SEIRHD scenarios
Wall clock time (s)
160 [146, 177]
169 [142, 767]
Time in API calls (s)
146 [126, 153]
143 [107, 748]
# Turns
30 [24, 37]
23 [18, 26]
# Output tokens
12 374 [11 015, 14 188]
9 717 [9 094, 10 909]
Cost ($)
0.785 [0.698, 0.865]
0.715 [0.634, 0.774]
SIR calibration
Wall clock time (s)
214 [163, 237]
156 [126, 164]
Appendix
Table 7: Per-condition medians ( 5 runs per cell) with the observed minimum–maximum range in square brackets. Baseline is the bare Epydemix library, and Framework is the agent framework.
Human behaviour during epidemics affects infectious disease dynamics, but quantifying this remains deeply challenging. Here we introduce the Epi-LLM framework: a novel integration of agent-based modelling, real-life epigames, and large language models (LLMs) in which a synthetic society of agents reasons and adapts dynamically over an outbreak contact network. Comparing synthetic agent behaviour against a no-intervention SEIR baseline and human participant data from the AUIB epigame study, we find that LLM agents across four different architectures reduced peak active infections, with quarantine compliance peaking at 58-65% on day six of the 15-day simulation. A binomial generalised linear model showed that perceived health severity was the strongest predictor of quarantine behaviour (β=0.33,p=0.002), yielding a pseudo-R2 of 0.055, comparable to the 0.072 observed in the human trial. LLM architecture is a key determinant of epidemic dynamics: low-variance architectures offer greater internal validity for testing behavioural rules, while high-variance models may better represent real-world decision-making. Geographic labels alone do not induce culturally differentiated behaviour; explicit attitudinal parameterisation is required. This proof-of-principle work lays the groundwork for deploying the Epi-LLM framework as a scalable, risk-free simulation environment for pandemic preparedness research.
Petra Ferencz, Ava Keeling, Tobias O'Keefe +4
Big Data Institute, Li Ka Shing Center for Health Information and Discovery, University of Oxford, Oxford, United Kingdom · Leverhulme Centre for Demographic Science, Nuffield Department of Population Health, University of Oxford, Oxford, United Kingdom · Pandemic Sciences Institute, Nuffield Department of Medicine, University of Oxford, Oxford, United Kingdom +3
Agent-based modeling (ABM) has the capability to model millions of individuals and their interactions, which is useful for policy making. However, ABMs have traditionally relied on static prior, which prevents the models from adapting to real-time changes. Our research provides a novel approach to addressing this information gap. Large language models (LLMs) offer new opportunities to predict human decision-making. Here, we introduce a scalable Hybrid Agent-based and Language-driven Epidemic (HALE) modeling framework that leverages LLMs to predict human decision-making in an ABM simulation. As a proof-of-concept, we use HALE to simulate COVID-19 and its effects in Salt Lake County, UT.
Sifat Afroj Moon, Dakotah Maguire, Adam Spannaus +5
Oak Ridge National Laboratory · Computational Science and Engineering Division, Oak Ridge National Laboratory, Oak Ridge, TN 37830, USA · Geospatial Science and Human Security Division, Oak Ridge National Laboratory, Oak Ridge, TN 37830, USA +1
Agent-based model (ABM) are a kind of computer model that makes it possible to simulate a set of autonomous interacting programs called agents in a shared virtual environment. Among other application field, it has been commonly used to simulate social phenomena such as urban segregation, opinion dynamic or epidemiological crisis [1]. Recently, a research emphasis has been put on ABM to study in silico the impact of non-pharmaceutical interventions to mitigate the SARS-CoV-2 outbreak of 2020, with few of them that had a great impact on global political responses [2]. Among the model used COMOKIT [3] has been design to simulate the every-day-life of inhabitant of various cities in Vietnam and test policy interventions for various COVID-19 spread scenarios. Such endeavor required huge computational power to handle a huge number of simulation replication over a large set of parameters. In this proposal we present a python package that enables to easily generate, explore and build reports for any COMOKIT experiment to be launched over High-Performance Computing (HPC) infrastructure.
Arthur Brugière, Kévin Chapuis
UMMISCO UMI 209, SU/IRD Bondy, France · Thuyloi University 175 Tay Son, Dong Da, Hanoi, Vietnam · ESPACE-DEV, Univ Montpellier, IRD, Univ Antilles, Univ Guyane, Univ R´eunion Montpellier, France