cs.AIOct 7, 2026

DrugTargetWorld: A Synthetic Biobank for Training and Benchmarking AI Scientists

Authors: Samuel Margolis, Paul Schmiedmayer, Alan Huang, Ethan Chen, Ishan Bhattacharjee, Atman Shah, Ben Viggiano, Fang Cao, +7 more

Organizations: Department of Biomedical Data Science, Stanford University, Stanford, CA 94305, USA · Department of Medicine, Stanford University, Stanford, CA 94305, USA · Division of Computational Medicine, Department of Medicine, Stanford University, Stanford, CA 94305, USA · Brown University, Providence, RI 02912, USA

Abstract

Drug target discovery requires distinguishing molecules that causally drive disease from those that are merely associated with it. Training and evaluating AI agents to perform this workflow end-to-end is difficult because real world biobanks lack known causal ground truth and participant-level data is access controlled. We introduce DrugTargetWorld, a framework that procedurally generates simulated biobanks, or "worlds," with known but concealed causal structure. Each world contains genotypes, proteins, health records, outcomes, and synthetic magnetic resonance imaging (MRI) for 54,000 participants. Agents must construct a disease phenotype, identify causal driver proteins, infer the beneficial direction of modulation, and optionally conduct virtual 'wet lab' experiments. We evaluated nine agents in 540 episodes across 20 cardiovascular worlds and three experimental budgets. Opus 5 and GPT-5.6 Sol achieved the highest mean composite scores, 39.98 and 35.38 of 100, respectively, and both recovered 64% of causal drivers on average. However, no agent reliably distinguished misleading non-causal proteins, and performance remained limited by the integrative judgments required to connect phenotype construction, causal evidence, and intervention decisions. By making each world's causal structure known to the evaluator but hidden from the agent, DrugTargetWorld turns end-to-end drug target discovery into a scalable training and evaluation problem with verifiable reward.

Explore similar work

CardsList
  1. TxBench-PP: Analyzing AI Agent Performance on Small-Molecule Preclinical Pharmacology

    Jun 17, 2026Hannah Le, Ramesh Ramasamy, Alex Urrutia +3Drug Design

  2. CausaLab: A Scalable Environment for Interactive Causal Discovery Toward AI Scientists

    May 25, 2026Junlin Yang, Dylan Zhang, Xiangchen Song +7Causal Discovery MethodsCausal

  3. CausalDS: Benchmarking Causal Reasoning in Data-Science Agents

    Jul 9, 2026Andrej Leban, Yuekai SunData Science AgentsCausal Reasoning