OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories
Organizations: Rice University · University of Illinois at Urbana-Champaign · University of Washington · National Yang Ming Chiao Tung University · Microsoft · University of Pennsylvania
Abstract
Multidisciplinary tumor boards integrate multimodal clinical observations and longitudinal patient histories through specialist discussions, yet benchmarks rarely capture these real-world trajectories. We introduce OpenTumorBoard, a benchmark with 611 patient cases and 19,157 discussion turns across ten specialist roles, transcribed from 12,534 minutes of publicly available tumor board recordings on YouTube. The benchmark evaluates two settings: SPECIALIST TURN, in which an LLM responds to a clinically significant question posed during a real discussion, and BOARD SIMULATION, in which it generates an entire back-and-forth discussion and reaches a consensus on therapy recommendations, surgical plans, next actions and clinical trial matching. Evaluation of 14 general-purpose frontier and medical LLMs reveals substantial limitations: the best models score 3.43 out of 5 in clinical equivalence to specialist answers and 2.78 out of 5 in alignment with recorded board conclusions. Supervised finetuning and reinforcement learning improve performance on a held-out test set, suggesting that real-world discussion trajectories can support model adaptation. Three M.D. experts review a subset of the benchmark, finding high information coverage and factuality of patient cases and strong fidelity of extracted consensus conclusions. We will release OpenTumorBoard and its automated curation pipeline to support the development and evaluation of LLMs for multidisciplinary, personalized cancer decision-making.
Figures & tables
| Dataset | Real-world tumor board cases | Multi- disciplinary team | Ground- truth discussion | Multimodal data | No. of tumor board cases |
|---|---|---|---|---|---|
| MedQA ( Jin et al., 2020 ) | ✗ | ✗ | ✗ | ✗ | N/A |
| CRAFT-MD ( Johri et al., 2025 ) | ✗ | ✗ | ✗ | ✓ | N/A |
| MediQ ( Li et al., 2024 ) | ✗ | ✗ | ✗ | ✗ | N/A |
| AgentClinic ( Schmidgall et al., 2025 ) | ✗ | ✗ | ✗ | ✓ | N/A |
| MedAgentBench ( Jiang et al., 2025 ) | ✗ | ✗ | ✗ | ✗ | N/A |
| MedAgentGym ( Xu et al., 2025 ) | ✗ | ✗ | ✗ | ✗ | N/A |
| Specialist Turn | Board Simulation | |||||
|---|---|---|---|---|---|---|
| Question relevance | Response correctness | Discussion support | Case coverage | Case factuality | Consensus fidelity | |
| High | 76.5% | 92.6% | 94.4% | 100% | 100% | 100% |
| Medium | 18.3% | 4.9% | 3.8% | 0% | 0% | 0% |
| Low | 5.1% | 2.5% | 1.8% | 0% | 0% | 0% |
| Annotator agreement | 60.0% | 79.2% | 79.4% | 80.0% | 93.3% | 73.3% |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Stage | Metric | Value | Items |
|---|---|---|---|
| Case segmentation | Case count | 100.0% | 15 recordings |
| Case boundary | 98.1% | 106 boundaries | |
| Slide extraction | Precision | 93.0% | 313 extracted frames |
| Precision, duplicates excluded | 99.7% | 313 extracted frames | |
| Recall | 93.6% | 311 presented slides | |
| Role inference | Accuracy | 96.3% | 136 speakers |
| Model | Input | Equiv. | Crit. err. | Unsupp. | ROUGE-L | BERTScore | Tokens |
|---|---|---|---|---|---|---|---|
| DeepSeek-V4-Pro (R) | caption | 3.43 | 5.6% | 11.8% | 0.183 | 0.616 | 637 |
| DeepSeek-V4-Flash (R) | caption | 3.33 | 6.7% | 13.7% | 0.181 | 0.617 | 325 |
| Gemma 4 31B (R) | image | 3.08 | 6.5% | 10.0% | 0.198 | 0.633 | 771 |
| Ministral 3 14B | image | 3.02 | 12.1% | 28.7% | 0.169 | 0.608 | 91 |
| Meditron3-70B | caption | 3.01 | 11.2% | 20.9% | 0.217 | 0.643 | 75 |
| Llama 4 Scout | image | 2.97 | 10.2% | 15.0% | 0.209 | 0.642 | 63 |
| Model | Input | Align. | Malformed | ROUGE-L | BERTScore | Tokens |
| Gemini 3.7 Flash (R) | image | 2.78 | 0 | 0.177 | 0.641 | 2,759 |
| Grok 4.6 (R) | image | 2.58 | 0 | 0.159 | 0.610 | 4,932 |
| Qwen3.8-Max (R) | image | 2.57 | 0 | 0.150 | 0.603 | 11,276 |
| GPT-5.6 Sol (R) | image | 2.53 | 3 | 0.149 | 0.612 | 8,213 |
| Claude Opus 5 (R) | image | 2.53 | 0 | 0.140 | 0.601 | 5,540 |
| Gemma 4 31B (R) | image | 2.52 | 0 | 0.181 | 0.636 | 1,693 |
| Model | Conclusion alignment | Malformed | Four decisions | Therapy recommendation | Surgical plan | Next action | Clinical trial matching |
|---|---|---|---|---|---|---|---|
| Qwen2.5-VL-3B | 1.58 | 22 | 1.78 | 1.68 | 1.83 | 1.88 | 1.72 |
| + SFT | 1.67 | 3 | 2.11 | 1.78 | 2.42 | 2.37 | 1.89 |
| + SFT + RL | 1.86 | 4 | 2.23 | 1.95 | 2.69 | 2.44 | 1.84 |
| Specialist turn type | Mean of 9 | Best | Crit. err. | Unsupp. | |
|---|---|---|---|---|---|
| Clarification question | 178 | 3.30 | 3.65 | 5.1% | 8.4% |
| Findings interpretation | 1,284 | 3.17 | 3.65 | 2.9% | 10.3% |
| Agreement or support | 178 | 3.07 | 3.42 | 10.7% | 8.4% |
| Next action suggestion | 654 | 2.97 | 3.46 | 2.8% | 5.7% |
| Uncertainty | 339 | 2.95 | 3.38 | 2.1% | 9.7% |
| Eligibility assessment | 322 | 2.95 | 3.48 | 7.5% | 9.9% |