APOD: reasoning-guided agentic population ordinary differential equation discovery for pharmacological digital twins
Authors: Romain Ferrara, Martin Soucail, Victor Gertner, Adil Moussali, Joris Cocquebert, Sandrine Oziel-Taieb, Julien Nicolas, Florence Gattacceca, +2 more
Organizations: COMPutational pharmacology and clinical Oncology (COMPO), Inria Center at University Côte d’Azur, Cancer Research Center of Marseille, Marseille, France · Department of Nuclear Medicine, Institut Paoli-Calmettes, ENETS Center of Excellence, IPC NET Center, Marseille, France · Centre for Computational Biology (CBIO), Mines Paris, PSL University, Paris, France · Department of Medical Oncology, Institut Paoli-Calmettes, ENETS Center of Excellence, IPC NET Center, Marseille, France · Université Paris-Saclay, CNRS, Institut Galien Paris-Saclay, Orsay, France · Department of Applied Mathematics and Theoretical Physics, University of Cambridge, Cambridge, United Kingdom
Establishing ordinary differential equations (ODEs) describing population data is a fundamental part of mathematical modeling in pharmacology, crucial to developing digital twins. However, doing so from sparse, noisy data is a slow, expert-driven task. Existing automated methods either search a restricted model space or ignore population inter-individual variability. Here we introduce APOD (Agentic Population ODE Discovery), a language-model agent that iteratively reasons over biological knowledge and fit diagnostics in an open-ended search space to discover a population digital twin (PDT), i.e., a shared ODE system with between-subject variability. On synthetic pharmacokinetic and tumor-dynamics benchmarks, APOD recovered ground-truth structures in 94-100% of runs, 12-fold faster in median than an established library-based search. On real cohorts it converged to valid structures, and proposed a PDT of radioligand-therapy-induced platelet dynamics that predicts thrombocytopenia from first-cycle data and simulates alternative dosing schedules that lower the predicted risk of toxicity.
Figures & tables
Table 1: Automated model-discovery approaches against the five properties a population digital twin requires. Interpretable parameters: the parameters are pharmacological quantities, such as a clearance or a carrying capacity. Population-level: subjects share a structural model and a between-subject variability model. Small cohorts: usable on a few dozen sparsely sampled subjects. Fast method: the search completes in minutes rather than hours. Open structural space: candidates are not confined to a predefined library. ✓ yes, ✗ no, – not applicable. A property that holds only in a restricted sense is scored yes. RF: random forest. XGBoost: extreme gradient boosting. RNN: recurrent neural network. ODE: ordinary differential equation. ML: machine learning. LLM: large language model. AMD: automatic model development.
Figure 1: The APOD (Agentic Population ODE Discovery) model-discovery workflow. From a one-sentence task and a residual-error model (here "Build a model of tumor growth evolution" , proportional error), the reflexion agent proposes candidate models, each made of a structural ordinary differential equation (ODE) model and a between-subject variability model. The agents see only aggregate metrics of the dataset. The model writer encodes each candidate in Mlxtran, the Monolix modelling language, and Monolix estimates it on the full dataset with the SAEM algorithm. The evaluation agent assesses goodness of fit, parameter identifiability and plausibility, then starts a new iteration (central loop) or stops and returns the selected model with its diagnostics and reasoning trace. Box colours: user actions (purple), language-model actions (orange), Monolix estimation (blue), data (green). LLM: large language model. SAEM: stochastic approximation expectation–maximization. DV: dependent variable.
Table 2: APOD recovers ground-truth structures on synthetic pharmacokinetic and tumor-dynamics benchmarks faster than a library-bounded search. Ten pharmacokinetic (PK) datasets, combining one to three compartments (cmt), intravenous (IV) or oral absorption with or without a lag, and linear or Michaelis–Menten (MM) elimination, and five tumor-dynamics (TD) growth laws (Supplementary Tables 1–4). APOD was run 10 times per dataset and the library search mlxModelFinder (MLX) with 5 estimation seeds. The ablation (5 runs per dataset) selects candidates on the corrected Bayesian information criterion (BICc) alone, without the reasoning layer (Methods). Recovery: percentage of runs recovering the generating structure, then the exact set of random effects as well. Tested models: median number of distinct structures per run or search. Time: median wall-clock time per run. ΔBICc: best APOD run minus best MLX search, a negative value favouring APOD. Cost: median application programming interface cost per run, from the recorded token usage. All datasets: recovery pooled over all runs, other columns as medians. IIV: inter-individual variability. USD: US dollars.
Table 3: Structural recovery across seven language-model backends on the synthetic benchmark. Each cell gives the percentage of runs recovering the generating structure, then the percentage also recovering the exact set of random effects. A run that did not complete counts as a miss. Each dataset was run 10 times with the reference, Claude Opus 4.5 (bold), and with the three open-weight models run locally through Ollama, and 5 times with the other models, served through an application programming interface (API). All datasets: percentages pooled over all runs. Run time and cost: medians per run, local models incurring no API cost. 20B, 8B, 24B: model size in billions of parameters. Candidates tested, failure rates and run times per dataset are in Supplementary Table 5. PK: pharmacokinetic. TD: tumor dynamics. IIV: inter-individual variability. USD: US dollars.
Figure 2: APOD models of real clinical and experimental cohorts, and their predictions for held-out subjects. Each row is one cohort, modelled from a one-sentence prompt naming neither the drug nor a model structure. (A–D) Docetaxel plasma pharmacokinetics 30 , three-compartment intravenous model, 24 patients for training and 10 held out. (E–H) Untreated tumor growth in breast-cancer xenografts 33 , Gompertz-type law, 46 mice for training and 20 held out. (I–L) Tumor growth under paclitaxel in MCF-7 breast-cancer xenografts, given as intravenous free drug or as a subcutaneously injectable polymer prodrug 55 , logistic growth with a threshold drug effect acting through an effect compartment, 74 mice for training and 32 held out. (A, E, I) Individual fits of the first 10 (A) or 20 (E, I) subjects of the training set: observations (points) and individual predictions (lines). (B, F, J) and (C, G, K) Visual predictive checks (VPC) in the training and held-out sets: observed median (red line) and 10th and 90th percentiles (blue lines) against the 90% prediction intervals (PI) of the same percentiles simulated from the model ( bands). (D, H, L) Observed versus individual-predicted values in the held-out set, with the population parameters fixed at their training values, the identity line and a Loess trend. The first three panels of a row share their axes.
Figure 3: The APOD platelet model as a predictive digital twin of thrombocytopenia under radioligand therapy. APOD discovered de novo, from a blinded one-sentence prompt, a Friberg-type myelosuppression structure. (A, B) Visual predictive checks (VPC) in the training cohort, on which the model was discovered (n = 44), and in an independent test cohort (n = 19): observed median and 10th and 90th percentiles (lines) against the 90% prediction intervals (PI) of the same percentiles simulated from the model (bands). (C) First-cycle forecasts for three test patients, with forecast errors of 4%, 8% and 22% (patients 1 to 3, mean absolute relative error): cycle-1 data used for individualisation (grey window), posterior median (line) and 90% predictive interval (band), later observations (points). (D) Platelet nadir of the 19 test patients, sorted by observed value: observed (grey), median and 90% PI of the first-cycle forecast (blue), and the 150 G/L threshold of grade-1 thrombocytopenia (dashed). (E) Patient 1, calibrated on cycle 1 and simulated under the schedule given (standard, 7,400 MBq), with one cycle omitted (omit) and with 3,700 MBq from cycle 2 (half). Points: observed counts. Upper rail: dose times, marker size proportional to the administered activity. Band: 90% PI under the standard schedule.
Automatic scientific discovery has long been a goal of computational scholars - a machine that can discover nature's secrets on its own, moving computational systems beyond data-fitting tools toward the generation and refinement of mechanistic models of the universe. Recent advances in symbolic regression (SR) and large-language-model (LLM)-based agents suggest that such systems can recover equations from data, incorporate domain priors, and automate parts of the research workflow. However, most existing approaches either focus on narrow equation-discovery benchmarks or broad end-to-end automation pipelines, while biological systems remain comparatively underexplored. Here, we introduce the MEDA system, an LLM- and SR-powered agentic framework for discovering ordinary-differential-equation (ODE) models of biological and biologically inspired dynamical systems. MEDA retrieves background knowledge, defines admissible variables, generates mechanistic constraints, proposes candidate ODEs, and fits and evaluates them. We evaluate it across canonical model retrieval, reasoning-based extrapolation to unseen variants, and open-ended discovery, with and without experimental data. Across these settings, MEDA recovered the correct state variables, achieved strong structural recovery in retrieval and extrapolation tasks, and produced biologically plausible discovery-oriented models. Ablation and robustness analyses show that knowledge-guided formalization and mechanistic constraints are load-bearing components, whereas numerical fitting alone can preserve trajectory-compatible but biologically incorrect equations.
David Krongauz, Arad Zulti, Eran Segal +1
Department of Computer Science and Applied Mathematics, Weizmann Institute of Science, Rehovot, Israel · Department of Molecular Cell Biology, Weizmann Institute of Science, Rehovot, Israel · Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, United Arab Emirates +2
Mechanistic modeling via ordinary differential equations (ODEs) provides interpretable descriptions of complex dynamics and enables inference of underlying mechanisms, which is particularly valuable in clinical settings. However, in rare diseases, both the structure and parameters of the model are typically unknown, while individual-level data is scarce, noisy, heterogeneous, and subject to privacy constraints. In such settings, population-level summary statistics provide a practical privacy-preserving data representation, while capturing heterogeneity further requires modeling parameters as distributions rather than fixed values. Yet no existing method jointly discovers ODE structure and refines parameter distributions solely from summary statistics. We present AgentODE, an end-to-end framework that addresses this gap. An LLM proposes candidate ODE structures, while a tool-augmented inference agent iteratively refines parameter distributions through a diagnosis--update loop, operating on population-level summary statistics alone. We evaluate AgentODE on three benchmark problems across different fields and two clinical datasets, including the rare disease recessive dystrophic epidermolysis bullosa (RDEB), with only 231 observations across 46 patients. AgentODE recovers functionally consistent ODE structures across all settings, and experiments on RDEB demonstrates that in sparse and noisy data settings reasoning from summary statistics promotes mechanistically principled structure discovery, whereas baselines with individual-level data access recover implausible structures despite better predictive performance. AgentODE opens new possibilities for mechanistic modeling of rare diseases directly from population-level summary statistics, where data scarcity and privacy constraints have traditionally limited such analyses.
Hanning Yang, Meropi Karakioulaki, Lennart Purucker +3
Institute of Medical Biometry and Statistics, Faculty of Medicine and Medical Center, University of Freiburg, Germany · Department of Dermatology, Medical Faculty and Medical Center, University of Freiburg, Germany · Prior Labs, University of Freiburg, Germany
We introduce Prior-Fitted Functional Flows, a generative foundation model for pharmacokinetics that enables zero-shot population synthesis and individual forecasting without manual parameter tuning. We learn functional vector fields, explicitly conditioned on the sparse, irregular data of an entire study population. This enables the generation of coherent virtual cohorts as well as forecasting of partially observed patient trajectories with calibrated uncertainty. We construct a new open-access literature corpus to inform our priors, and demonstrate state-of-the-art predictive accuracy on extensive real-world datasets.
César Ojeda, Niklas Hartung, Wilhelm Huisinga +6
University of Potsdam, Germany · Technical University Berlin, Germany · Freie Universität Berlin, Germany +2