OpenMTB-Audit: Exposing Over-Refusal and Clinical Expert Perspectives in LLM-Based Molecular Tumor Board Safety Evaluation
Authors: Negin Ashrafi, Jia Luo, Stacey M. Frumm, Roxana Daneshjou
Organizations: Department of Biomedical Data Science, Stanford University, Stanford, CA, USA · Lowe Center for Thoracic Oncology, Dana-Farber Cancer Institute, and Harvard Medical School, Boston, MA, USA
Molecular tumor boards integrate genomic findings, clinical context, and therapeutic evidence to support precision oncology. As AI enters this workflow, a key safety challenge is distinguishing truly unsupported recommendations from evidence-supported options that still require oncologist review because of incomplete information, poor ECOG performance status, or other clinical caveats. We introduce OpenMTB-Audit, an open-source benchmark of 500 synthetic non-small cell lung cancer cases spanning five adversarial error categories and four safety labels: Supported, Partially Supported, Unsupported, and Insufficient Information. Across eight large language model configurations, we identify pervasive over-refusal: all LLM configurations failed to retain the Partially Supported label in 83.3-100% of true Partially Supported cases, achieving high aggregate safety scores through label collapse rather than clinically calibrated reasoning. To address this limitation, we developed MTB-AuditAgent, a deterministic seven-module framework separating evidence verification, missing-information detection, safety classification, and abstention. It reduces over-refusal to 6.7% and achieves 91.2% accuracy (95% CI: 88.6-93.6%). A two-oncologist annotation study found disagreement concentrated at the boundary between information sufficiency and treatment optimization, underscoring the need to preserve clinically meaningful distinctions.
Figures & tables
Error Type
Description
Wrong Drug
The therapy lacks support for the reported molecular alteration.
Wrong Context
The therapy has evidence or regulatory support in a different cancer type, but not in the presented NSCLC context.
Missing Critical Information
One or more fields required to determine treatment appropriateness are absent, such as disease stage, ECOG performance status, prior treatment, or biomarker validation.
Hallucinated Variant
A variant of uncertain significance, non-sensitizing alteration, or otherwise unsupported biomarker is presented as clinically actionable.
Unsafe Overconfidence
The molecularly matched therapy may be appropriate, but the recommendation omits important caveats, such as ECOG performance status ≥3 or extensive prior treatment.
Table 1: Adversarial error categories. Each category contains 60 cases.
Figure 1: Worked examples illustrating the four reference labels in OpenMTB-Audit . All cases are synthetic and contain no real patient data.
Figure 2: MTB-AuditAgent workflow from free-text case input through seven deterministic modules to a structured safety audit.
System
Over-Refusal
False Approval
GPT-4o-mini
Base LLM
83.3%
2.2%
Simple RAG
100.0%
0.0%
RAG+EV
100.0%
0.0%
Prompt-Checklist
100.0%
2.2%
GPT-4o
Table 2: Clinical over-refusal and false-approval rates. Over-refusal is 1−recallPS : the proportion of ground-truth Partially Supported cases not predicted as Partially Supported (misclassified as Supported or Unsupported). False approval denotes Unsupported cases classified as Supported.
Prompt-Checklist GPT-4o
MTB-AuditAgent
Label
Prec
Rec
Prec
Rec
Supported
0.97
0.75
0.98
0.84
Partially Supported
0.00
0.00
0.71
0.93
Unsupported
0.75
0.98
0.91
1.00
Insufficient Info
0.97
0.58
1.00
0.87
Table 3: Per-class precision and recall for the highest-accuracy LLM baseline and MTB-AuditAgent . Similar aggregate Safety Scores conceal differences in class-specific performance, particularly for Partially Supported cases.
Figure 3: Safety Score versus accuracy across systems. Structured LLMs show high safety but lower accuracy because of over-refusal, whereas MTB-AuditAgent achieves both.
Comparison
Agreement
κ (Level)
Oncologist B Label
Recall
Oncologist A vs. Benchmark
50/50
1.000 (Perfect)
Unsupported
19/20(95%)
Oncologist B vs. Benchmark
34/50
0.553 (Moderate)
Partially Supported
8/10(80%)
Oncologist A vs. Oncologist B
34/50
0.553 (Moderate)
Supported
6/10(60%)
Insufficient Info
1/10(10%)
Table 4: Agreement with benchmark labels and per-label recall on the 50-case expert review subset. [ 21 ]
System
Accuracy
Macro F1
Abstain F1
MI Catch
Safety Score
GPT-4o-mini
Base LLM
0.694
0.499
0.348
0.567
74.9
Simple RAG
0.524
0.318
0.346
0.467
73.8
RAG+EV
0.462
0.317
0.832
0.650
87.0
Prompt-Checklist
0.600
0.422
0.851
0.567
83.9
GPT-4o
Table 5: Performance on OpenMTB-Audit ( N=500 ) under the four-label framework. Bold indicates the best value in each column.
Configuration
Accuracy
Macro F1
Safety Score
Full MTB-AuditAgent
0.912
0.898
92.4
w/o Evidence Verifier (M3)
0.552
0.586
37.4 ( −55.0 )
w/o Missing Information Detector (M4)
0.808
0.638
75.1 ( −17.3 )
w/o ECOG Rule (M5)
0.800
0.667
87.8 ( −4.6 )
w/o Abstention Module (M6)
0.912
0.898
78.7 ( −13.7 )
Table 6: Ablation analysis of MTB-AuditAgent on OpenMTB-Audit . Values in parentheses indicate the change in Safety Score relative to the full system.
Figure 4: Correct-label rates by error type ( N=60 each), with the largest gap on Unsafe Overconfidence .
Multidisciplinary tumor boards integrate multimodal clinical observations and longitudinal patient histories through specialist discussions, yet benchmarks rarely capture these real-world trajectories. We introduce OpenTumorBoard, a benchmark with 611 patient cases and 19,157 discussion turns across ten specialist roles, transcribed from 12,534 minutes of publicly available tumor board recordings on YouTube. The benchmark evaluates two settings: SPECIALIST TURN, in which an LLM responds to a clinically significant question posed during a real discussion, and BOARD SIMULATION, in which it generates an entire back-and-forth discussion and reaches a consensus on therapy recommendations, surgical plans, next actions and clinical trial matching. Evaluation of 14 general-purpose frontier and medical LLMs reveals substantial limitations: the best models score 3.43 out of 5 in clinical equivalence to specialist answers and 2.78 out of 5 in alignment with recorded board conclusions. Supervised finetuning and reinforcement learning improve performance on a held-out test set, suggesting that real-world discussion trajectories can support model adaptation. Three M.D. experts review a subset of the benchmark, finding high information coverage and factuality of patient cases and strong fidelity of extracted consensus conclusions. We will release OpenTumorBoard and its automated curation pipeline to support the development and evaluation of LLMs for multidisciplinary, personalized cancer decision-making.
Anqi Li, Zhixuan Ge, Yixuan Duan +9
Rice University · University of Illinois at Urbana-Champaign · University of Washington +3
Refusal rates are a poor proxy for LLM safety, i.e., a model may over-refuse benign prompts while still complying with harmful ones. We audit both failure modes across 21 open-weight LLMs on four safety benchmarks (OR-Bench, XSTest, ToxiGen, BOLD), using a composition adjustment to isolate model sensitivity from dataset toxicity confounds. We report three findings. First, models adopt fundamentally different calibration strategies: conservative ecosystems such as Llama suppress unsafe outputs at the cost of elevated over-refusals, while permissive ecosystems such as DeepSeek and Qwen preserve helpfulness but tolerate higher harmful compliance. Second, demographic protection is unequal: models over-protect prominent racial and religious groups, frequently refusing even benign prompts about them, while providing substantially weaker protection against disability-targeted attacks. Third, refusal and compliance tendencies are stable within model families across generations and scales, suggesting that post-training objectives shape safety behavior more than architecture. Our results call for joint, demographically-aware, and multi-judge safety evaluation.
Alif Al Hasan, Sumon Biswas
Department of Computer and Data Sciences Case Western Reserve University Cleveland, OH, USA
Safety alignment in Large Language Models is critical for healthcare; however, reliance on binary refusal boundaries often results in over-refusal of benign queries or unsafe compliance with harmful ones. While existing benchmarks measure these extremes, they fail to evaluate Safe Completion: the model's ability to maximise helpfulness on dual-use or borderline queries by providing safe, high-level guidance without crossing into actionable harm. We introduce Health-ORSC-Bench, the first large-scale benchmark designed to systematically measure Over-Refusal and Safe Completion quality in healthcare. Comprising 31,920 benign boundary prompts across seven health categories (e.g., self-harm, medical misinformation), our framework uses an automated pipeline with human validation to test models at varying levels of intent ambiguity. We evaluate 30 state-of-the-art LLMs, including GPT-5 and Claude-4, revealing a significant tension: safety-optimised models frequently refuse up to 80% of "Hard" benign prompts, while domain-specific models often sacrifice safety for utility. Our findings demonstrate that model family and size significantly influence calibration: larger frontier models (e.g., GPT-5, Llama-4) exhibit "safety-pessimism" and higher over-refusal than smaller or MoE-based counterparts (e.g., Qwen-3-Next), highlighting that current LLMs struggle to balance refusal and compliance. Health-ORSC-Bench provides a rigorous standard for calibrating the next generation of medical AI assistants toward nuanced, safe, and helpful completions. Furthermore, our benchmark facilitates reproducible evaluation, encourages safety calibration, and supports development of clinically reliable, context-aware, human-aligned medical AI systems. Our code and data are available at: https://github.com/ZhihaoZhang97/Health-ORSC-Bench. Warning: Some contents may include toxic or undesired contents.