cs.CLSep 29, 2026

A Proposed Rubric for Evaluating Expressed Clinical Reasoning in Large Language Model Responses

Authors: Zhangshu Joshua Jiang, Zina Ibrahim, James T. Teo

Organizations: DRIVE-Health CDT Department of Biostatistics and Health Informatics Institute of Psychiatry, Psychology and Neuroscience King’s College London, London, UK · Neurological Institute, Cleveland Clinic London, London, UK · King’s College London, London, UK · King’s College Hospital NHS Foundation Trust, London, UK

Abstract

We propose a rubric for assessing expressed clinical reasoning in model responses, drawing on three bodies of work: medical education assessment frameworks (ART, SCT, Key Feature Problems and OSCE); clinical LLM benchmarks (MedR-Bench, HealthBench, TIMER-Bench, DR. BENCH, PrIME-LLM and PatientSafeBench); and general LLM reasoning evaluation research, including the Factuality-Validity-Coherence-Utility taxonomy, FaithCoT-Bench and C2-Faith. We use groundedness as a clinically oriented adaptation of the taxonomy's factuality category. The rubric brings these concepts together in a multidimensional framework for scoring free-text responses to gold-standard clinical vignettes. It includes provisional behavioural anchors, applicability rules and a separate flag for case-specific safety-critical errors. General-domain frameworks inform its design but are not treated as validated clinical instruments. The rubric does not replace case-specific reference criteria or the task-specific metrics of existing benchmarks. It has not yet been tested for inter-rater reliability, construct validity or clinical utility. Its immediate purpose is to make evaluation decisions explicit and open to scrutiny before empirical testing.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. A rubric landscape for evaluating clinical reasoning in large language models: what exists, what is missing, and what needs to be combined

    Oct 1, 2026Zhangshu Joshua Jiang, Zina Ibrahim, James T. TeoClinical Reasoning TrainingLLM Reasoning Strategies

  2. A rubric-based controlled comparison of frontier language models on expert-authored clinical reasoning tasks

    Jul 2, 2026Samiha A. Ismail, Fan X. Chen, Ali MeraliRubric-Based ScoringComparison

  3. When Rubrics Fail: Hallucinations Reveal Blind Spots in Medical AI Evaluation

    Sep 14, 2026Griffin Farrow, Lily Sijia Li, Jack Johnson +4Clinician TrustVisual Hallucinations