A rubric landscape for evaluating clinical reasoning in large language models: what exists, what is missing, and what needs to be combined
Organizations: DRIVE-Health CDT Department of Biostatistics and Health Informatics Institute of Psychiatry, Psychology and Neuroscience King’s College London London, UK
Abstract
Exam-style accuracy does not establish whether large language models (LLMs) reason well over clinical records. We define clinical reasoning as integrating and updating evidence across time and sources to form, revise and justify a patient's problem representation and a defensible plan. This structured narrative review maps three literatures: medical education assessment instruments, clinical LLM benchmarks published from 2023 onwards, and general-domain methods for evaluating long-form generation. We examine six dimensions: problem representation, temporal synthesis, differential and management reasoning, counterfactual reasoning, calibrated uncertainty, and reasoning faithfulness. Preprints are included and flagged. No single instrument covers all six dimensions. Problem representation and differential or management reasoning are reasonably covered, although reliability varies by instrument and setting. TIMER-Eval targets temporal synthesis, and ER-Reason assesses sequential diagnostic belief updating. Dedicated uncertainty and counterfactual evaluations are emerging, but their applicability to longitudinal free-text reasoning remains limited. Factual completeness is well theorised in general-domain evaluation, with early clinical evidence of important omissions. Faithfulness remains the weakest dimension, with one identified clinical causal-ablation study on multiple-choice questions. Existing tools should be combined through binary rubric items, separate completeness and correctness scores, case-specific importance weighting with non-compensable safety caps, temporal order-consistency checks, and chance-corrected reliability reporting. Further design work is needed for calibrated uncertainty, counterfactual reasoning and faithfulness over longitudinal free-text records. This review provides a design rationale, not a validated instrument.
Figures & tables
| Dimension | Best-available instrument(s) | What is measured | What remains open |
|---|---|---|---|
| Problem representation | IDEA, R-IDEA, ART ( Baker et al., 2015 ; Schaye et al., 2022 ; Thammasitboon et al., 2018 ) | Structured write-up of interpretive summary and differential in clinical notes | Reliability varies by instrument and setting; transfer to LLM outputs requires validation; documentation quality does not establish causal faithfulness ( Baker et al., 2015 ; Schaye et al., 2022 ) |
| Differential and management reasoning | HealthBench, MedR-Bench, ER-Reason, Brodeur et al. ( Arora et al., 2025 ; Qiu et al., 2025 ; Mehandru et al., 2025 ; Brodeur et al., 2026 ) | Quality, safety and completeness of clinically relevant answers; correspondence to reference reasoning | Correspondence to reference text is not causal reasoning; text-only input, no bedside cues ( Brodeur et al., 2026 ) |
| Temporal / longitudinal synthesis | TIMER-Eval ( Cui et al., 2025 ) | Respect for temporal boundaries, trend identification, chronological order | Does not test whether a later record supersedes, rather than supplements, an earlier one |
| Counterfactual reasoning | MamaBench ( Adewuyi et al., 2026 ) | Bias Trap Rate on paired vignettes where one parameter changes the correct diagnosis | Single specialty; paired vignettes, not longitudinal free text |
| Calibrated uncertainty | Du et al. uncertainty-preservation benchmark, SCT-Bench ( Du et al., 2026 ; McCoy et al., 2025 ) | Preservation of five-level uncertainty labels; belief-shift patterns against expert panels | No rubric for uncertainty that must change as records accumulate |
| Reasoning faithfulness | Afolabi et al. causal-ablation probe ( Afolabi et al., 2026 ) | Causal necessity of individual chain-of-thought steps via redaction | Workshop-scale study on multiple-choice only; no free-text, longitudinal, rubric-integrated version |