cs.AIJun 3, 2026

PSEBench: A Controllable and Verifiable Benchmark for Evaluating LLMs in Patient Safety Event Triage

Authors: Keqi Han, Ryan Young, Annabel Strauss, Lindsey Hughes, Katharine M. Nesbitt, Nicole Schueler, Che Ngufor, Carl Yang, +2 more

Organizations: 1Emory University · 2Scale AI · 3Mayo Clinic · 4Vanderbilt University Medical Center

Abstract

Patient safety event triage, determining whether a clinical event is reportable under jurisdiction-specific policy, is a high-stakes task typically performed manually by patient safety experts. Although LLMs may support this workflow, reliable evaluation is limited by the lack of benchmarks to capture evidence-grounded policy reasoning, proactive information seeking for incomplete reports, and principled abstention in irreducibly ambiguous cases. We address this gap with a policy-grounded construction methodology centered on the clause card, a structured representation that factorizes regulatory text into auditable decision specifications. Combining clause cards with anchor-driven instantiation and closed-loop verification, our scalable pipeline produces narratives with by-construction ground truth and naturally supports generating missing information and uncertain variants. We instantiate this method on Minnesota's 29 Reportable Adverse Health Events, producing PSEBench, a 5,074-case benchmark with an agentic evaluation environment. Evaluation on 15 representative LLMs reveals consistent capability trends, demonstrates the benchmark's utility, and identifies actionable gaps toward reliable LLM-based patient safety event triage.

Explore similar work

CardsList
  1. CARE-Bench: Benchmarking Patient-Facing LLM Triage

    Aug 4, 2026Yining Hua, Hongbin Na, Cyrus AyubchaTriageCare

  2. Auditable Emergency Triage for Maternal and Newborn Care in India

    Sep 10, 2026Shobhit Jagga, Aman Dalmia, Niharika Priyadarshini +7TriageModel Interpretability