KlinikeBench: Evaluating Language Models Beyond Diagnostic Accuracy
Organizations: Zhejiang University · Imperial College London · Nanchang University · University of Toronto · Microsoft
Abstract
Most clinical benchmarks evaluate language models (LMs) on diagnosis using complete case descriptions. In clinical practice, however, patients present information in different ways, and clinicians must obtain relevant history and determine which examinations are needed before reaching a diagnosis. Diagnostic accuracy alone therefore cannot establish whether an agent gathered essential information or conducted an appropriate clinical assessment. Furthermore, existing benchmarks lack professional clinicians' verification. To address this gap, we introduce KlinikeBench, a benchmark of 333 clinician-authored tasks, each providing an isolated sandbox environment with a virtual patient, clinical tools, and task-specific success criteria. More than 35 clinicians contributed to case authoring and benchmark evaluation. In an empirical study, clinicians gave simulated dialogues higher mean quality ratings than reference conversations, which is adapted from real conversation. In each task, an LM has a fixed budget of turns to communicate with the patient, ask about relevant history, request examinations, follow action constraints, and record a final diagnosis. We score these steps separately as well as together. Across 31 models and seven model families, the best-performing models (e.g., GPT-6-astra and Claude Opus 5) succeed on less than 30% of tasks, even though their diagnosis accuracy reaches 90.7%. Some models benefit from talking with the patient; others diagnose well from a complete chart but perform much worse in conversation. Overall, KlinikeBench provides a testbed for evaluating the full clinical encounter and reveals a substantial gap between diagnostic accuracy and performance in interactive clinical assessment.
Figures & tables
| Benchmark | Metric | Interaction | Tools | Source | Clinician verification |
| MedQA ( 2021 ) | MCQA acc. | Single turn | None | Board exams | Exam keys |
| MMLU-med. ( 2020 ) | MCQA acc. | Single turn | None | Exam material | Source answers |
| MedMCQA ( 2022 ) | MCQA acc. | Single turn | None | Entry exams | Expert questions |
| HealthBench Pro. ( 2026 ) | Rubric | Next reply | Varies | Clinical chats | 3 raters/item |
| AgentClinic ( 2026 ) | Diag. acc. | 20 turns | Tests + aids | Clinical cases | Dialogue ratings |
| MedAgentBench ( 2025 ) | Task success | Multi-step | FHIR APIs | EHRs | Task authoring |
| Model | Strict pass@1 | Diagnosis | Tools | Must-ask |
|---|---|---|---|---|
| Claude Opus 5 ( Anthropic, 2026b ) | 29.4% | 90.7% | 44.1% | 89.9% |
| Claude Opus 5.5 ( Anthropic, 2026c ) | 28.8% | 90.7% | 49.8% | 84.4% |
| Claude Opus 4.8 ( Anthropic, 2026a ) | 16.8% | 89.2% | 27.9% | 85.5% |
| Claude Sonnet 5 ( Anthropic, 2026e ) | 13.8% | 87.4% | 24.0% | 85.1% |
| Claude Haiku 4.5 ( Anthropic, 2025a ) | 2.7% | 72.1% | 8.1% | 70.6% |
| GPT-6-astra ( OpenAI, 2026f ) | 24.3% | 87.4% | 39.3% | 88.0% |
| Dimension | Virtual patient | MTS-Dialog |
|---|---|---|
| Progressive history disclosure | ||
| Emotional appropriateness | ||
| Human-likeness |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Corpus | Total | Per task | Mean | Median | Min–Max |
|---|---|---|---|---|---|
| Clinical tasks | 333 | Must-ask topics | 4.5 | 4 | 0–26 |
| Distinct tool names | 140 | Required tools | 5.8 | 5 | 1–19 |
| Clinical subcategories | 17 | Exposed tools | 10.7 | 10 | 3–24 |
| Function | Names | Example tools | Execution behavior |
|---|---|---|---|
| Patient records | 6 | get_patient_info , get_patient_history , get_patient_vitals | Return demographics, history, medications, allergies, family history, or vital signs. |
| Tests and imaging | 3 | order_lab_test , get_lab_results , order_imaging | Store orders and retrieve available laboratory results. |
| Treatment and referral | 3 | prescribe_medication , refer_to_specialist , check_drug_interactions | Store prescriptions or referrals; return an interaction-check response. |
| Diagnosis recording | 2 | record_diagnosis , record_differential | Store the final diagnosis or differential; only the final diagnosis ends the encounter. |
| Clinical knowledge | 4 | get_treatment_guidance , get_disease_symptoms | Expose knowledge-query interfaces; return unavailable status in the current headless runtime. |
| Encounter memory | 3 | take_note , set_working_diagnosis , summarize_history_taken | Store notes and interim hypotheses, or summarize questions already asked. |
| Setting | SFT | RL (GRPO) |
|---|---|---|
| Initialization | Qwen3-4B | SFT checkpoint |
| Learning rate | ||
| Schedule | Cosine; 5% warm-up | Constant |
| Batch | 32 sequences (global) | 8 tasks 8 rollouts |
| Training duration | 3 epochs | 315 steps planned; results through 30 |
| Sequence limit | 20,480 tokens | 32k tokens |
| Model | Strict pass@1 | Diagnosis | Tools | Must-ask |
|---|---|---|---|---|
| Claude Opus 5 ( Anthropic, 2026b ) | 29.4% | 90.7% | 44.1% | 89.9% |
| Claude Opus 5.5 ( Anthropic, 2026c ) | 28.8% | 90.7% | 49.8% | 84.4% |
| Claude Opus 4.8 ( Anthropic, 2026a ) | 16.8% | 89.2% | 27.9% | 85.5% |
| Claude Sonnet 5 ( Anthropic, 2026e ) | 13.8% | 87.4% | 24.0% | 85.1% |
| Claude Sonnet 4.5 ( Anthropic, 2025b ) | 9.9% | 82.3% | 25.8% | 76.7% |
| Claude Sonnet 4.6 ( Anthropic, 2026d ) | 8.4% | 83.8% | 22.8% | 73.8% |
| Model | Step cap | pass@1 | Diagnosis | Tools | Must-ask coverage | [95% CI] |
|---|---|---|---|---|---|---|
| GPT-6-astra | Default | 28 | 86 | 46 | 88.9 | — |
| 45 | 34 | 88 | 60 | 90.4 | ||
| 60 | 32 | 86 | 56 | 91.2 | ||
| None | 36 | 88 | 52 | 89.0 | ||
| Claude Opus 5 | Default | 34 | 86 | 54 | 89.8 | — |
| 45 | 36 | 84 | 62 | 90.4 |
| Condition | OP5 | G6 | GR4.7 | GE3.8 | OP5.5 | G5.6 | S5 | GE3.7 |
|---|---|---|---|---|---|---|---|---|
| Judge changed; patient fixed to GPT-5.4-mini | ||||||||
| GPT-5.4-mini (original) | 24 | 30 | 20 | 20 | 26 | 18 | 8 | 8 |
| GPT-5.4-mini (repeat) | 26 | 28 | 24 | 20 | 26 | 18 | 6 | 6 |
| GPT-5.4 | 24 | 32 | 20 | 20 | 32 | 24 | 10 | 6 |
| Gemini-3.5-flash | 22 | 20 | 18 | 12 | 24 | 14 | 6 | 6 |
| Grok-4.5 | 22 | 28 | 18 | 14 | 28 | 16 | 8 | 6 |
| Stage | Strict pass@1 | Diagnosis | Required tools | Must-ask |
|---|---|---|---|---|
| Base | 0.0% | 26.0% | 3.0% | 39.7% |
| SFT | 8.0% | 39.0% | 33.0% | 77.5% |
| SFT + RL (diagnosis) | 10.0% | 44.0% | 32.0% | 76.8% |
| SFT + RL (multi-gate) | 6.0% | 42.0% | 32.0% | 73.1% |