cs.CLSep 14, 2026

K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations

Authors: Laura M. VowelsMatthew J. VowelsShivali SharmaApoorv JhaRehnuma ChoudhuryWasseem El SarrajRachel Francois-WalcottAruba Hussain+4 more

Organizations: School of Psychology, University of Roehampton, London, United Kingdom · 2Kivira Health, United Kingdom · University of Hertfordshire, Hatfield, United Kingdom · University of Surrey, Guildford, United Kingdom · School of Sport, Psychology and Social Sciences, University of Bedfordshire, Luton, United Kingdom · 6Tavistock Relationships, London, United Kingdom · 7InsideOut, United Kingdom

Abstract

People increasingly use large language models (LLMs) for mental health support, yet their safety in evolving, high-risk conversations remains poorly characterised. We developed K-Bench, a clinician-calibrated, protected benchmark evaluating 125 model configurations representing 33 base models from 14 providers across a fixed cohort of 200 multi-turn vignettes involving suicide, self-harm, domestic violence, substance misuse, and no-risk presentations. Synthetic patient conversations showed substantial distributional overlap with real human-AI conversations. A frozen GPT-4o judge achieved 94.2% exact agreement with clinician consensus across 6,751 eligible item comparisons from 151 clinician-rated transcripts. Leading models combined strong supportive conversation with combined-risk scores above 95, whereas risk exploration exposed substantial variation among lower-performing configurations. Therapeutic prompting produced configuration-specific gains concentrated among weaker models, while elevated reasoning produced no average improvement. K-Bench combines broader clinical coverage and configuration-scale comparison with a continuously updated public leaderboard whose operational test materials are protected from direct optimisation. The leaderboard is available at www.k-bench.ai.

Explore similar work

CardsList
  1. One Year Later...The Harms Persist, But So Do We!

    Jun 22, 2026Annika Marie Schoene, Cansu Canca, Gautham Vijay Kumar +1Mental HealthLarge Language Model Safety