cs.CLSep 15, 2026

Challenges of Auditing: Variability in Outputs of Large Language Models for Health

Authors: Yuan PuYewon ChangFurong JiaXunjian YinJessica MaAyman AliMonica Agrawal

Organizations: Department of Computer Science, Duke University, 308 Research Drive, Durham, 27708, NC. · Department of Medicine, Duke University, 40 Duke Medicine Circle, Durham, 27710, NC. · Geriatrics and Extended Care, Durham VA Health System, 508 Fulton Street, Durham, 27705, NC. · Department of Surgery, Duke University, 2301 Erwin Road, Durham, 27710, NC. · Department of Biostatistics and Bioinformatics, Duke University, 2424 Erwin Road, Durham, 27710, NC.

Abstract

People increasingly use frontier AI models for health advice, but via different access modes (e.g., ChatGPT, ChatGPT Health, APIs) with varying settings. Here, we find systematic differences across access modes. Because evaluations typically rely on APIs while consumers interact through chatbot interfaces, these discrepancies limit evaluation validity. Our findings underscore an urgent need for model providers to enable faithful replication of consumer experiences and settings for rigorous audits.

Explore similar work

CardsList