Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming
Organizations: Cherise · School of Computation, Information and Technology, Technical University of Munich (TUM), Munich, Germany · Department of Engineering Science, University of Oxford, Oxford, United Kingdom · School of Medicine, TUM University Hospital, Munich, Germany · Munich Center for Machine Learning, Munich, Germany · Department of Diagnostic Radiology, National University of Singapore, Singapore · Department of Computing, Imperial College London, London, United Kingdom · Department of Computer Science, Stony Brook University, Stony Brook, USA · School of Computer Science, University of Sheffield, Sheffield, United Kingdom · School of Medicine and Health, Technical University of Munich, Munich, Germany · Department of Neurosurgery/Neuro-Oncology, State Key Laboratory of Oncology in South China, Guangdong Provincial Clinical Research Center for Cancer, Sun Yat-sen University Cancer Center, Guangzhou, China · Department of Diagnostic and Interventional Neuroradiology, TUM University Hospital, Munich, Germany
Abstract
Large language models (LLMs) are increasingly used to answer health-related questions and support healthcare workflows, yet evidence for their safety still relies heavily on static benchmarks that can rapidly become obsolete or be optimized against. Here we introduce a Dynamic, Automatic, and Systematic (DAS) red-teaming audit framework that continuously stress-tests LLMs for health across four safety-critical axes: robustness, privacy, bias/fairness, and hallucination/factual inaccuracies. Validated against board-certified clinicians with high concordance, a suite of adversarial agents autonomously mutates health-related test cases to uncover vulnerabilities in real time. Applying DAS to 15 proprietary and open-source LLMs revealed a profound gap between high static benchmark performance and low dynamic reliability--the "Benchmarking Gap". Despite median MedQA accuracy exceeding 80%, 94% of previously correct answers failed under dynamic robustness testing. This brittleness generalized to the realistic, open-ended HealthBench dataset, where top-tier models exhibited failure rates exceeding 70% and sharp shifts in model rankings across evaluations, suggesting that high scores on established static benchmarks may reflect superficial memorization. We observed similarly high failure rates across other domains: privacy leaks were elicited in 86% of scenarios, cognitive-bias priming altered recommendations in 81% of fairness tests, and hallucination rates exceeded 74% in widely used models. By converting LLM safety evaluation for health from a static checklist into a living adversarial audit, DAS provides a scalable framework for surfacing latent risks before such systems are deployed in consumer-facing health assistants, clinician-facing tools, and broader healthcare workflows. Code is available at https://github.com/JZPeterPan/DAS-Medical-Red-Teaming-Agents.