Driving Epidemic Models with AI Agents: the Epydemix Agent Framework
Organizations: ISI Foundation, Turin, Italy · Network Science Institute, Northeastern University, Boston, MA, USA
Abstract
Artificial Intelligence agents based on large language models provide convenient natural language interfaces to scientific software, but reliability is not automatic. Here we introduce the Epydemix Agent Framework, an additive layer over Epydemix, an open-source Python library for stochastic compartmental epidemic modeling. The framework extends the library with four capabilities to facilitate interaction with an AI agent: discovery of available models and parameters, preventive validation of a declarative scenario specification, execution through tested library code, and inspectability of results. These capabilities let an agent handle the entire modeling process, from the natural-language description of the scenario to quantitative results, figures, and interpretation of findings without writing custom code. Each step reads input files and saves results in a separate output bundle, making the process auditable and reproducible. First, we show the end-to-end workflow with a case study comparing vaccination strategies for a novel respiratory virus. Second, we assessed the framework across 50 agent sessions and five modeling tasks by comparing the agent use of the framework against the direct use of the Python interface. The framework reduced turns, output tokens, and cost on most tasks, unless it trades resources for per-point reproducibility.
Figures & tables
| Metric (CV = std/mean) | Phase | Framework | Direct API |
| Total cost | Planning | 0.10 | 0.19 |
| Execution | 0.15 | 0.35 | |
| Total duration | Planning | 0.15 | 0.30 |
| Execution | 0.17 | 0.40 |
| Task | Wall clock time (s) | Time in API (s) | # Turns | # Output tokens | Cost ($) |
| SEIRHD scenarios | 1.05 [0.85, 4.78] | 0.98 [0.73, 5.14] | 0.77 [0.60, 0.96] | 0.79 [0.68, 0.91] | 0.91 [0.81, 1.04] |
| SIR calibration | 0.73 [0.59, 0.95] | 0.57 [0.53, 0.70] | 0.62 [0.55, 0.72] | 0.58 [0.49, 0.67] | 0.71 [0.65, 0.95] |
| Calibrate Project | 0.53 [0.45, 0.57] | 0.50 [0.43, 0.55] | 0.63 [0.53, 0.70] | 0.45 [0.37, 0.45] | 0.52 [0.49, 0.73] |
| School-closure sweep | 2.04 [1.24, 3.27] | 1.41 [1.25, 2.24] | 1.38 [1.00, 2.21] | 1.32 [1.18, 2.23] | 1.81 [1.34, 3.34] |
| Measles coverage sweep | 0.81 [0.42, 1.12] | 0.83 [0.62, 0.93] | 0.80 [0.68, 0.89] | 0.71 [0.62, 0.93] | 0.88 [0.67, 1.05] |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Phase | Metric | With framework | Without framework |
| Planning | Duration (s) | ||
| Cost ($) | |||
| Execution | Duration (s) | ||
| Cost ($) | |||
| Total | Duration (s) | ||
| Cost ($) |
| Metric | With framework | Without framework |
| Peak occupancy (mean std) | ||
| Peak occupancy (range) | – | – |
| Attack rate % (mean std) | ||
| Days over capacity | 28–31 (all 4 runs) | 26–31 (all 4 runs) |
| Metric | With framework | Without framework |
| Peak occupancy (mean std) | ||
| Peak occupancy (range) | – | – |
| Attack rate % (mean std) | ||
| Days over capacity | 0 (all 4 runs) | 0 (all 4 runs) |
| Metric | With framework | Without framework |
| Peak occupancy (mean std) | ||
| Attack rate % (mean std) | ||
| Days over capacity | 0 (all 4 runs) | 0 (all 4 runs) |
| Parameter | Condition | Slow / school | Rapid / work | Combined / window |
| Vaccination rate (%/day) | With framework | |||
| Vaccination rate (%/day) | Without framework | |||
| NPI reduction (%) | With framework | days | ||
| NPI reduction (%) | Without framework | days |
| Setting | Value |
| Agent runtime | Claude Code 2.1.224, non-interactive |
| Model | claude-sonnet-5 |
| Reasoning effort | medium |
| Tools available | Bash , Read , Write , Edit , |
| Glob , Grep , BashOutput , KillShell | |
| Tools withheld | sub-agents, task planning, web search and fetch |
| Task | Metric | Baseline | Framework |
| SEIRHD scenarios | Wall clock time (s) | 160 [146, 177] | 169 [142, 767] |
| Time in API calls (s) | 146 [126, 153] | 143 [107, 748] | |
| # Turns | 30 [24, 37] | 23 [18, 26] | |
| # Output tokens | 12 374 [11 015, 14 188] | 9 717 [9 094, 10 909] | |
| Cost ($) | 0.785 [0.698, 0.865] | 0.715 [0.634, 0.774] | |
| SIR calibration | Wall clock time (s) | 214 [163, 237] | 156 [126, 164] |