OpenMTB-Audit: Exposing Over-Refusal and Clinical Expert Perspectives in LLM-Based Molecular Tumor Board Safety Evaluation
Authors: Negin Ashrafi, Jia Luo, Stacey M. Frumm, Roxana Daneshjou
Organizations: Department of Biomedical Data Science, Stanford University, Stanford, CA, USA · Lowe Center for Thoracic Oncology, Dana-Farber Cancer Institute, and Harvard Medical School, Boston, MA, USA
Molecular tumor boards integrate genomic findings, clinical context, and therapeutic evidence to support precision oncology. As AI enters this workflow, a key safety challenge is distinguishing truly unsupported recommendations from evidence-supported options that still require oncologist review because of incomplete information, poor ECOG performance status, or other clinical caveats. We introduce OpenMTB-Audit, an open-source benchmark of 500 synthetic non-small cell lung cancer cases spanning five adversarial error categories and four safety labels: Supported, Partially Supported, Unsupported, and Insufficient Information. Across eight large language model configurations, we identify pervasive over-refusal: all LLM configurations failed to retain the Partially Supported label in 83.3-100% of true Partially Supported cases, achieving high aggregate safety scores through label collapse rather than clinically calibrated reasoning. To address this limitation, we developed MTB-AuditAgent, a deterministic seven-module framework separating evidence verification, missing-information detection, safety classification, and abstention. It reduces over-refusal to 6.7% and achieves 91.2% accuracy (95% CI: 88.6-93.6%). A two-oncologist annotation study found disagreement concentrated at the boundary between information sufficiency and treatment optimization, underscoring the need to preserve clinically meaningful distinctions.
Figures & tables
Error Type
Description
Wrong Drug
The therapy lacks support for the reported molecular alteration.
Wrong Context
The therapy has evidence or regulatory support in a different cancer type, but not in the presented NSCLC context.
Missing Critical Information
One or more fields required to determine treatment appropriateness are absent, such as disease stage, ECOG performance status, prior treatment, or biomarker validation.
Hallucinated Variant
A variant of uncertain significance, non-sensitizing alteration, or otherwise unsupported biomarker is presented as clinically actionable.
Unsafe Overconfidence
The molecularly matched therapy may be appropriate, but the recommendation omits important caveats, such as ECOG performance status ≥3 or extensive prior treatment.
Table 1: Adversarial error categories. Each category contains 60 cases.
Figure 1: Worked examples illustrating the four reference labels in OpenMTB-Audit . All cases are synthetic and contain no real patient data.
Figure 2: MTB-AuditAgent workflow from free-text case input through seven deterministic modules to a structured safety audit.
System
Over-Refusal
False Approval
GPT-4o-mini
Base LLM
83.3%
2.2%
Simple RAG
100.0%
0.0%
RAG+EV
100.0%
0.0%
Prompt-Checklist
100.0%
2.2%
GPT-4o
Table 2: Clinical over-refusal and false-approval rates. Over-refusal is 1−recallPS : the proportion of ground-truth Partially Supported cases not predicted as Partially Supported (misclassified as Supported or Unsupported). False approval denotes Unsupported cases classified as Supported.
Prompt-Checklist GPT-4o
MTB-AuditAgent
Label
Prec
Rec
Prec
Rec
Supported
0.97
0.75
0.98
0.84
Partially Supported
0.00
0.00
0.71
0.93
Unsupported
0.75
0.98
0.91
1.00
Insufficient Info
0.97
0.58
1.00
0.87
Table 3: Per-class precision and recall for the highest-accuracy LLM baseline and MTB-AuditAgent . Similar aggregate Safety Scores conceal differences in class-specific performance, particularly for Partially Supported cases.
Figure 3: Safety Score versus accuracy across systems. Structured LLMs show high safety but lower accuracy because of over-refusal, whereas MTB-AuditAgent achieves both.
Comparison
Agreement
κ (Level)
Oncologist B Label
Recall
Oncologist A vs. Benchmark
50/50
1.000 (Perfect)
Unsupported
19/20(95%)
Oncologist B vs. Benchmark
34/50
0.553 (Moderate)
Partially Supported
8/10(80%)
Oncologist A vs. Oncologist B
34/50
0.553 (Moderate)
Supported
6/10(60%)
Insufficient Info
1/10(10%)
Table 4: Agreement with benchmark labels and per-label recall on the 50-case expert review subset. [ 21 ]
System
Accuracy
Macro F1
Abstain F1
MI Catch
Safety Score
GPT-4o-mini
Base LLM
0.694
0.499
0.348
0.567
74.9
Simple RAG
0.524
0.318
0.346
0.467
73.8
RAG+EV
0.462
0.317
0.832
0.650
87.0
Prompt-Checklist
0.600
0.422
0.851
0.567
83.9
GPT-4o
Table 5: Performance on OpenMTB-Audit ( N=500 ) under the four-label framework. Bold indicates the best value in each column.
Configuration
Accuracy
Macro F1
Safety Score
Full MTB-AuditAgent
0.912
0.898
92.4
w/o Evidence Verifier (M3)
0.552
0.586
37.4 ( −55.0 )
w/o Missing Information Detector (M4)
0.808
0.638
75.1 ( −17.3 )
w/o ECOG Rule (M5)
0.800
0.667
87.8 ( −4.6 )
w/o Abstention Module (M6)
0.912
0.898
78.7 ( −13.7 )
Table 6: Ablation analysis of MTB-AuditAgent on OpenMTB-Audit . Values in parentheses indicate the change in Safety Score relative to the full system.
Figure 4: Correct-label rates by error type ( N=60 each), with the largest gap on Unsafe Overconfidence .