cs.AIOct 3, 2026

Decide, Ask, or Defer: Clinical LLMs under Incomplete Evidence

Authors: Mingzhan Yang, Weili Wu

Organizations: Department of Computer Science University of Texas at Dallas Richardson, TX 75080

Abstract

Clinical LLMs must decide not only what diagnosis to produce, but also whether the available evidence is sufficient for autonomous decision making. Binary DECIDE/ABSTAIN formulations merge distinct non decision states and do not explicitly evaluate information acquisition. We introduce a DECIDE/ASK/DEFER formulation together with a blinded protocol that prevents models from using evidence completeness metadata. We evaluate Qwen, Gemini, and GPT on 200 matched clinical evidence states constructed from DDXPlus. The models show substantial differences in action selection under identical evidence, with disagreement in 137 of 200 states. For Qwen, a matched targeted versus random analysis shows that selected information changes the likelihood of a subsequent autonomous decision more clearly than diagnostic correctness. Its matched DECIDE/ABSTAIN baseline further reveals a safety autonomy tradeoff: the three action policy rescues some erroneous autonomous decisions but also removes some correct autonomous deci sions. These results show that separating information acquisition from clinician deferral exposes behavior that binary abstention hides, without yielding a uniformly improved decision policy.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry

    May 6, 2026Oriana Presacan, Andreea Grama, Larisa Irimină +6Mental HealthDiagnosis

  2. Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support

    Jul 29, 2026Shayndhan Sivanathan, Shravan Nageswaran, Mehdi Zadem +10TriageClinical Decision Support

  3. Quantifying and Mitigating Premature Closure in Frontier LLMs

    May 14, 2026Rebecca Handler, Suhana Bedi, Nigam ShahLarge Language Models FailClosure