EpiNarrate: Agentic Generation of Grounded Narratives from Epidemiological Scenario Projections
Authors: Rituparna Datta, Srini Venkatramanan, Bryan L. Lewis, Yiqi Su, Harry Hochheiser, Lucie Contamin, Parantapa Bhattacharya, Naren Ramakrishnan, +1 more
Abstract
Generation of clear and accessible public health narratives is critical for communicating complex epidemiological projections to policymakers and the general public at large. Such narratives require more than simply reporting numbers: projections must be contextualized and quantitatively grounded across multiple dimensions. Further, projections are often derived from large ensemble datasets which combine intervention assumptions, geographic and demographic strata, outcomes, time horizons, and uncertainty quantiles. However, directly using large language models (LLMs) to summarize and contextualize such data often leads to inconsistencies, omissions, and fragile behavior. We introduce an agentic framework (EpiNarrate) for public health report generation that separates structured numerical reasoning from natural-language generation. The framework first extracts scenario axes and organizes them into a partial-order schema, enabling systematic traversal of the underlying multidimensional space. It then constructs an augmented dataset and derives valid quantitative statements through a comparison grammar that enforces semantic and arithmetic consistency. To balance coverage and non-redundancy, we introduce an interestingness-driven selection mechanism based on maximum-entropy principles. Experiments on the COVID-19 Scenario Modeling Hub demonstrate that our model produces narratives with improved factual grounding and broader coverage of salient epidemiological patterns, while preserving the style of expert-written reports.
Epidemiological models often rely on survey data to represent how individuals make health-related decisions, such as whether to vaccinate or adopt protective behaviors. However, repeated large-scale surveys are costly, time-consuming, and limited in the range of scenarios they can capture. In this work, we investigate whether large language models (LLMs) can generate synthetic survey responses that reproduce patterns observed in real populations. Using longitudinal data from the FluPaths surveys, we first identify groups associated with broadly positive or negative attitudes toward vaccination through clustering analysis. We then evaluate several LLMs using a cluster-informed prompting approach to generate synthetic survey responses across multiple epidemic waves. Across models, the synthetic data generally reproduce the distributions of demographic characteristics, vaccination-related beliefs, risk perceptions, and health behaviors observed in the survey data. However, they are less successful at capturing how these factors vary together within respondents. Some models reproduce group-level vaccination trends more reliably than others, although performance varies across waves. We also trained a classifier to distinguish real from synthetic records and found that the generated responses remained identifiable as synthetic. Overall, our findings suggest that LLM-generated survey data may provide a useful tool for exploratory data augmentation and we hope that it could support agent-based epidemic modeling approaches. However, the generated data should not be treated as a substitute for human survey data without further methodological improvements and validation.
Reliable epidemiological reasoning requires synthesizing study evidence to infer disease burden, transmission dynamics, and intervention effects at the population level. Existing medical question answering benchmarks primarily emphasize clinical knowledge or patient-level reasoning, yet few systematically evaluate evidence-grounded epidemiological inference. We present EpiQAL, to our knowledge the first diagnostic benchmark for epidemiological question answering over research literature, comprising three subsets built from open-access articles across diverse diseases. The three subsets progressively test factual recall, multi-step inference, and conclusion reconstruction under incomplete information, and are constructed through a quality-controlled pipeline combining taxonomy guidance, multi-model verification, and difficulty screening. Experiments on fifteen models spanning open-source and proprietary systems reveal that current LLMs show limited performance on epidemiological reasoning, with multi-step inference posing the greatest challenge. Model rankings shift across subsets, and scale alone does not predict success. Chain-of-Thought prompting benefits multi-step inference but yields mixed results elsewhere. EpiQAL provides fine-grained diagnostic signals for evidence-grounding, inferential reasoning, and conclusion reconstruction.
Effective HFMD surveillance requires forecasts capturing both time-series patterns and contextual drivers such as school calendars, weather, and policy or surveillance reports. In clinical settings, forecasts must be trusted and actionable; thus, beyond point accuracy, decision-makers require concise, auditable explanations of why risk is expected to rise or fall. Classical models (e.g., ARIMA and Prophet) and foundation models (e.g., Chronos, Moirai, and TimesFM) treat external covariates as numerical inputs, lacking semantic reasoning to reflect epidemiological mechanisms or resolve conflicting signals. We propose a two-agent neuro-symbolic framework that decouples contextual interpretation from probabilistic forecasting. An LLM-based Event Interpreter ingests heterogeneous signals -- school schedules, weather summaries, government reports, and clinical guidelines -- and outputs a scalar transmission-impact signal. A Forecast Generator combines this signal with historical case counts to produce point forecasts that are mapped to probabilistic predictions through Poisson/negative-binomial moment matching. We focus on one-week-ahead rolling forecasts, aligning with weekly hospital-capacity planning and the rapid, context-driven inflections typical of HFMD. We evaluate on two datasets: Hong Kong surveillance (90 target weeks in 2023--2024) and Lishui hospital visits (33 target weeks in 2024). Against traditional and foundation-model baselines, our approach achieves competitive point accuracy while providing robust 90% intervals (coverage approximately 0.85--1.00) and concise rationales. This demonstrates that integrating domain knowledge through LLM-based agents can match strong numerical forecasters while yielding interpretable, context-aware forecasts aligned with public-health decision-making.