Explainable Suicide Risk Assessment on Social Media with Multi-Task QLoRA
Authors: Xuan Zhong Feng, Geoffrey Martin, Hexin Dong, Yifan Peng
Organizations: Department of Population Health Sciences, Weill Cornell Medicine, New York, NY, USA · Systems Engineering, Cornell University, New York, NY, USA
Explainable suicide-risk assessment requires models not only to estimate risk severity, but also to identify supporting language and the risk and protective factors expressed in a post. We present our system for the IEEE BigData 2026 Cup on Explainable Suicide Risk Assessment on Social Media, which addresses three tasks: risk-level classification, evidence phrase extraction, and multi-label factor identification. Our approach adapts Qwen2.5-Instruct models using quantized low-rank adaptation (QLoRA) and an answer-masked causal language-model objective. We jointly train across all three tasks for risk classification, jointly train on Tasks1a and 1b for evidence extraction, and adapt Task2 separately for factor identification. We also tailor aggregation to each output: we average risk-level probabilities from the 32B and 72B models, combine evidence phrases through cross-fold consensus, and calibrate factor-specific decisions through rate matching based on out-of-fold operating points. On the official leaderboard, the final system achieved a composite score of 0.7738, with 0.8089 on Task1 and 0.6919 on Task2. Across the evaluated configurations, three-task training performed best for Task1a, joint training on Tasks1a and 1b performed best for Task1b, and task-specific training performed best for Task2. Probability averaging further improved Task~1a when component models had complementary errors. These findings highlight the value of tailoring both training objectives and aggregation strategies to the output structure of each task within a unified language-model framework.
Figures & tables
n
%
Task 1a: Risk level
Indicator
611
37.4
Ideation
519
31.7
Behavior
391
23.9
Attempt
114
7.0
Task 1b: Evidence phrases
Table 1: Training-set distributions across the three shared tasks.
Figure 1: System architecture for the three prediction tasks. Task 1a applies cross-model probability averaging. Task 1b uses cross-fold phrase consensus. Task 2 makes factor-level decisions using rate-matched operating points estimated from out-of-fold predictions.
Model
Strategy
Weighted F1
MentalRoBERTa Classifier
-
0.7636
Qwen2.5-32B
1a
0.8209
1a + 1b
0.8205
1a + 1b + 2
0.8259
Qwen2.5-72B
1a
0.8272
1a + 1b
0.8225
Table 2: Results for Task 1a. Weighted F1 on pooled five-fold out-of-fold predictions ( n=1635 ).
Model
Strategy
Phrase F1
MentalRoBERTa + CRF
–
0.7044
Qwen2.5-32B
1b
0.7676
1a + 1b
0.7792
1a + 1b + 2
0.7821
Qwen2.5-72B
1b
0.7761
1a + 1b
0.7874
Table 3: Results for Task 1b. Phrase F1, defined as the mean post-level score, on pooled five-fold out-of-fold predictions ( n=1635 ).
Model
Strategy
Macro-F1
MentalRoBERTa
Asymmetric Loss
0.4507
Zero-shot
Qwen2.5-32B
0.6384
Qwen2.5-72B
0.6507
Qwen3-32B
0.6345
Llama-3.3-70B
0.6225
GPT-5.4 *
0.6191
Table 4: Pooled five-fold Task 2 results over n=1635 posts.
Assessing suicide risk from social media text is a small-data, high-stakes setting requiring not only severity prediction but also supporting evidence and clinically relevant risk and protective factors. Yet common NLP techniques, including model scaling, synthetic data, loss reweighting, ensembling, and threshold tuning, are often applied without testing whether their gains hold up under severe class imbalance, coupled outputs, and limited author-level data. We study 1,635 clinician-annotated posts and audit 31 pre-specified techniques from 7 methodological families through roughly 300 controlled experiments on author-disjoint partitions. We found no prior audit of this playbook in this regime. The findings guide a task-grounded system for three outputs: 4-level suicide risk, evidence spans, and 24 clinical risk and protective factors. Only 5 of 31 comparisons produced reliable gains. We reformulate factor prediction as entailment between each post and its codebook definitions, using an architecturally diverse ensemble with class-balanced training and score rescaling. Risk predictions condition a 7-model evidence tagger ensemble; evidence restricts symbolic risk rules; and a difficult risk class is routed separately. The factor predictor remains independent because risk evidence provides no additional factor signal. We also correct a mismatch between validation scores used for threshold fitting and test-time ensemble scores through deployment-consistent calibration, yielding the largest improvement to the factor system. The final system achieves 0.8203 for risk, 0.7953 for evidence, and 0.7045 macro-F1 for factors, with a 0.7781 composite, ranking third among 53 teams. We call the underlying principle task-conditioned technique selection: retain techniques only when task-specific knowledge, structure, or empirical evidence justifies them.
Shlok Shelat, Shrey Salvi, Souvik Roy +2
Indian AI Research Organisation · Ahmedabad University · University of Maryland, Baltimore County
Suicide is a critical global public health challenge, causing approximately 720,000 deaths each year and calling for timely, effective prevention strategies. Existing computational studies primarily focus on post-based social media platforms such as Twitter and Weibo, leaving instant messaging environments such as Telegram underexplored. Yet group chats pose distinct challenges: messages are short, fragmented, multi-party, and often rely on implicit or culturally specific expressions, making isolated post-level analysis insufficient. We introduce SuiChat-CN, a Chinese group-chat benchmark for contextual suicide risk assessment. We collect public Telegram group-chat data, construct coherent conversational segments through signal-word extraction and bidirectional context expansion, and annotate user risk levels with an expert-validated, LLM-assisted paradigm. SuiChat-CN contains 13,312 contextual segments from 1,406 users, covering 258,228 raw chat messages. Extensive experiments with PLMs and more than 40 LLMs demonstrate that contextual information is essential for reliable risk assessment, while fine-tuning and partial-context evaluation further reveal the challenges of early detection in multi-party conversations. Due to ethical and sensitivity concerns, the dataset is not publicly released but will be shared with accredited mental health and suicide-prevention research institutions upon reasonable request.
Xiangyu Wang, Zhiwei Yu, Chengze Du +3
University of Chinese Academy of Sciences · 2Tsinghua University · 3Beijing University of Posts and Telecommunications +1
Crisis helplines assess suicide risk through structured interviews, a process that is slow and dependent on operator training and workload. Natural language processing could support risk assessment and call prioritization, but almost no work addresses Arabic-language helpline calls or operates within the privacy constraints of real helpline data. We analysed de-identified transcripts from Lebanon's National Lifeline for Emotional Support and Suicide Prevention. Audio never left the helpline: calls were transcribed on site with a speech recognition model for Levantine Arabic, and an Arabic named-entity recognition model removed identifying information locally. Only the de-identified transcripts were shared with the research team. Operators recorded the five suicidal ideation items of the Columbia Suicide Severity Rating Scale, which we combined into two binary outcomes: at-risk and high-risk. We also machine-translated the transcripts into English, giving a paired Arabic/English comparison. On each corpus, we fine-tuned five instruction-tuned large language models alongside six transformer encoder baselines (four Arabic, two English) and evaluated all models on a held-out test set. We included 383 calls: 373 for the at-risk task (52.3% positive) and 297 for the high-risk task (30.0% positive). The best Arabic model reached a macro-F1 of 81.19 and a ROC-AUC of 90.61 on high-risk; the best English model reached 85.00 and 92.59, identifying 88.9% of high-risk calls. In both languages, high-risk calls separated more cleanly than at-risk calls, and translation to English did not reduce the best observed performance. Suicide risk can be classified from de-identified Arabic transcripts without sending audio outside the helpline. The high-risk results support further testing as an operator-facing tool; lower-severity ideation proved the harder case.
Linhai Ma, Rita El Hachem, Mahatab El Hajj +2
Department of Emergency Medicine, Yale University, New Haven, CT, USA. · Department of Epidemiology and Population Health, American University of Beirut, Beirut, Lebanon. · Embrace, Mental Health Center, Beirut, Lebanon.