cs.CLMay 14, 2026

Quantifying and Mitigating Premature Closure in Frontier LLMs

Authors: Rebecca HandlerSuhana BediNigam Shah

Organizations: Department of Medicine, Stanford University · Department of Biomedical Data Science, Stanford University

Abstract

Premature closure, or committing to a conclusion before sufficient information is available, is a recognized contributor to diagnostic error but remains underexamined in large language models (LLMs). We define LLM premature closure as inappropriate commitment under uncertainty: providing an answer, recommendation, or clinical guidance when the safer response would be clarification, abstention, escalation, or refusal. We evaluated five frontier LLMs across structured and open-ended medical tasks. In MedQA (n = 500) and AfriMed-QA (n = 490) questions where the correct choice had been removed, models still selected an answer at high rates, with baseline false-action rates of 55-81% and 53-82%, respectively. In open-ended evaluation, models gave inappropriate answers on an average of 30% of 861 HealthBench questions and 78% of 191 physician-authored adversarial queries. Safety-oriented prompting reduced premature closure across models, but residual failure persisted, highlighting the need to evaluate whether medical LLMs know when not to answer.

Explore similar work

CardsList
  1. CARE-Bench: Benchmarking Patient-Facing LLM Triage

    Aug 4, 2026Yining Hua, Hongbin Na, Cyrus AyubchaTriageCare