cs.AIOct 1, 2026

OpenMTB-Audit: Exposing Over-Refusal and Clinical Expert Perspectives in LLM-Based Molecular Tumor Board Safety Evaluation

Authors: Negin Ashrafi, Jia Luo, Stacey M. Frumm, Roxana Daneshjou

Organizations: Department of Biomedical Data Science, Stanford University, Stanford, CA, USA · Lowe Center for Thoracic Oncology, Dana-Farber Cancer Institute, and Harvard Medical School, Boston, MA, USA

Abstract

Molecular tumor boards integrate genomic findings, clinical context, and therapeutic evidence to support precision oncology. As AI enters this workflow, a key safety challenge is distinguishing truly unsupported recommendations from evidence-supported options that still require oncologist review because of incomplete information, poor ECOG performance status, or other clinical caveats. We introduce OpenMTB-Audit, an open-source benchmark of 500 synthetic non-small cell lung cancer cases spanning five adversarial error categories and four safety labels: Supported, Partially Supported, Unsupported, and Insufficient Information. Across eight large language model configurations, we identify pervasive over-refusal: all LLM configurations failed to retain the Partially Supported label in 83.3-100% of true Partially Supported cases, achieving high aggregate safety scores through label collapse rather than clinically calibrated reasoning. To address this limitation, we developed MTB-AuditAgent, a deterministic seven-module framework separating evidence verification, missing-information detection, safety classification, and abstention. It reduces over-refusal to 6.7% and achieves 91.2% accuracy (95% CI: 88.6-93.6%). A two-oncologist annotation study found disagreement concentrated at the boundary between information sufficiency and treatment optimization, underscoring the need to preserve clinically meaningful distinctions.

Figures & tables

Explore similar work

CardsList
  1. OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories

    Sep 26, 2026Anqi Li, Zhixuan Ge, Yixuan Duan +9Multidisciplinary Tumor BoardsMultimodal Clinical Data

  2. The Refusal--Compliance Tradeoff: A Large-Scale Safety Behavior Audit of Large Language Models

    May 6, 2026Alif Al Hasan, Sumon BiswasLarge Language Model SafetyRefusals

  3. Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context

    Jan 25, 2026Zhihao Zhang, Liting Huang, Guanghao Wu +3RefusalsReproducibility