cs.CLSep 5, 2026

Beyond the Flag: Clinical Framing Closes the Moderation Gap in Suicide Risk Measurement

Authors: Shreyas Krishnan, Gun Ahn, Jungjin Kim

Organizations: University of California, Berkeley · Wondi AI · MIT · Harvard Medical School · McLean Hospital

Abstract

Moderation APIs are built to flag policy-violating content, not to measure graded clinical risk. But a platform's duty does not end at detection: the response owed to passive distress differs sharply from the response owed to active planning with means access, and emerging regulation (e.g., California Senate Bill 243) is turning that distinction into a compliance requirement. We therefore ask how well deployed safety signals recover clinically meaningful severity. We release a benchmark of 516 r/SuicideWatch posts rated by a licensed psychiatrist on a four-level ordinal schema (Indicator, Ideation, Behavior, Attempt) grounded in the Columbia Suicide Severity Rating Scale, and evaluate moderation APIs, prompted LLMs, and supervised baselines under seven ordinal-aware metrics. Three findings. Vendor moderation APIs separate low- from high-severity posts well (0.860 high-risk F1) but measure severity poorly (0.395 macro F1), systematically over-predicting the most severe category. Clinically grounded zero-shot prompting recovers much of that gap (0.562 macro F1), and expert-authored framing (not fine-tuning, added reasoning, or naive multi-agent aggregation) is the effective lever. The value of reasoning depends on register: it hurts on long, noisy Reddit posts and helps on short, clinician-authored statements. We argue graded severity, not a binary flag, is what a proportionate duty of care requires, and release our evaluation framework to support that measurement.

Explore similar work

CardsList
  1. Explainable Suicide Risk Assessment on Social Media with Multi-Task QLoRA

    Sep 30, 2026Xuan Zhong Feng, Geoffrey Martin, Hexin Dong +1Social Media AnalysisMulti-Task Learning

  2. Bag of Tricks or Bag of Myths? Reducing Modeling Complexity with Task Knowledge in Explainable Suicide Risk Assessment

    Sep 7, 2026Shlok Shelat, Shrey Salvi, Souvik Roy +2Language Model EnsemblesSuicide Risk Modeling

  3. First, do no harm: Breaking suicidogenic echo chambers in media recommendation

    May 24, 2026Alberto Díaz-Álvarez, Raúl Lara-Cabrera, Fernando Ortega-Requena +1Mental HealthRecommender Systems