Emotion dynamics are critical for understanding crisis-support conversations, yet most computational work treats emotion as static utterance-level labels. We introduce EMPATH, a framework for understanding affective dynamics in mental health dialogues across three granularities: turn-level labels, transition probabilities, and global conversation archetypes. Applying EMPATH to text-based crisis conversations with self-identified Black texters discussing grief, we find persistent negative affect, gradual hope-ward transitions, distinct texter-volunteer emotional roles, and heterogeneous recovery trajectories. These results highlight the informative patterns that emerge from computationally understanding crisis support and expressions of grief as dynamic processes within conversations, as well as the overall value of emotion-dynamic analysis for analyzing and comparing affect in dialogues.
Figures & tables
Figure 1: Summary of Empath framework for assessing conversation dynamics in mental health/crisis dialogues. The example dialogue is simulated and contains no real CTL messages.
Author
Count
# Labeled
Avg. Len
System
629
0
24.2
Texter
3355
474
17.7
Volunteer
2695
19
25.2
Overall
6679
493
21.4
Table 1: Message summary statistics of the 100-conversation CTL annotation sample.
Level
Statistic
Description
C
Transition matrix
How often each emotion follows another ( 37×37 counts and probabilities).
C
Persistence
How likely an emotion is to repeat on the next turn ( P(stay) ).
C
Net flow
Whether an emotion is a “sink” (conversations flow in) or “source” (conversations flow out).
C
Motifs
Common 2- and 3-emotion sequences (e.g., hopeless → hopeless → hopeful).
C
Stationary distribution
Long-run equilibrium prevalence of each emotion under a Markov model.
C
Change rate
Fraction of consecutive turns where the emotion label changes.
Table 2: Micro-Dynamics Statistics derived from emotion-label sequences at the category level of 37 emotion labels (C) and polarity level of positive/negative/neutral (P).
Texter
Volunteer
Label
Count
%
Label
Count
%
hopeless
15,221
25.0
hopeful
26,175
46.4
worthlessness
8,344
13.7
hopeless
11,866
21.0
hopeful
7,650
12.6
overwhelm
6,253
11.1
overwhelm
5,467
9.0
neutral
2,748
4.9
gratitude
3,942
6.5
anxiety
1,970
3.5
Table 3: Top-10 emotion labels for CTL texter utterances (left, N=60,788 labeled utterances from 2,478 conversations) and volunteer utterances (right, N=56,460 labeled utterances). For texters, 60,788 of 62,723 raw utterances received parseable labels.
Transition
Count
%
hopeless
→
hopeless
6,620
11.58
worthlessness
→
worthlessness
3,150
5.51
hopeful
→
hopeful
3,002
5.25
worthlessness
→
hopeless
1,942
3.40
overwhelm
→
overwhelm
1,868
3.27
hopeless
→
worthlessness
1,755
3.07
Table 4: Top-10 emotion transitions for CTL texters ( N=57,147 ). Self-transitions dominate; hopeless → hopeful (rank 8) is the most frequent recovery transition. Percentages are calculated over all transitions.
Figure 2: Conversation-start to conversation-end emotion flows for CTL texter utterances ( n=2,475 ). Left nodes show the first substantive emotion, after leading neutral-labelled and sub-three-word opener turns are removed; right nodes show the last. Prominent flows from hopeless toward hopeful and gratitude illustrate hope-ward movement at the conversation level.
P(stay)
Role
Neg.
Neut.
Pos.
CTL Texter
0.841
0.181
0.645
CTL Volunteer
0.646
0.250
0.724
Table 5: Polarity persistence for CTL texter and volunteer utterances.
↓ From / To →
Neg.
Neut.
Pos.
Negative
0.841
0.010
0.149
Neutral
0.405
0.181
0.414
Positive
0.322
0.034
0.645
Table 6: Row-normalized polarity transition matrix for CTL texter utterances.
Polarity
Start (%)
End (%)
Negative
90.9
20.5
Neutral
0.0
8.6
Positive
9.1
70.9
Table 7: Conversation-start and conversation-end polarity distributions for CTL texter utterances.
Volunteer Strategy
CTL (%)
Affirmation and Reassurance
37.4
Information
3.5
Others
10.7
Providing Suggestions
12.1
Question
24.3
Reflection of Feelings
5.0
Table 8: Distribution of predicted volunteer support strategies in CTL conversations.
Figure 3: Volunteer strategy effects in CTL conversations. Bars show the downstream change in texter distress after each volunteer strategy, computed over local three-turn windows (ut,ht,ut+1) . Negative values indicate reduced texter distress after the volunteer response.
Figure 4: Mean texter distress trajectories for the full CTL analysis corpus and synthetic conditions. All conditions show distress reduction, but differ in the pacing and structure of de-escalation.
Paraphrased CTL Excerpt
V: There is also a resource made for people who have lost their child. Would you like me to share it with you?
T: Yes, that would be helpful
V: On GriefNet , there is a community of people you can talk to that are in similar situations.
V: I wanted to give you two options so you can figure out what resources might work best.
T: Thank you so much!
Synthetic (zero-shot) Excerpt
Table 9: Representative Information exchanges in CTL and synthetic conversations. The CTL volunteer offers a specific, contextually matched resource ( GriefNet ); the synthetic volunteer offers generic escalation options ( 988, ER ) without tailoring. Texts drawn from CTL are paraphrased.
Figure 5: Distribution of conversation-level distress trajectory archetypes across the full CTL analysis corpus, zero-shot synthetic, and dual-agent synthetic conversations.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Label
Definition (abridged)
neutral
No strong positive or negative emotion; matter-of-fact tone
sadness
Feeling sad, down
fear
Feeling afraid, scared
anxiety
Feeling anxious or symptoms of generalised anxiety
joy
Long-lasting contentment and satisfaction with life
love
Feeling loved, or loving others
Appendix
Table 10: The 37 emotion labels and abridged definitions used in the detection prompt. Full definitions are provided verbatim to the model. The self label merges the original coding-book categories self-doubt and self-aware .
Model
37
14
7
3
Random baseline
Acc.
0.035
0.209
0.270
0.582
Sim.
0.947
0.951
0.947
0.968
BERT Emotions
Acc.
—
—
0.335
0.582
Sim.
0.947
—
0.940
0.971
MentalBERT
Acc.
0.146
0.371
0.416
0.688
Sim.
0.960
0.959
0.959
0.973
Appendix
Table 11: Model evaluation ( N=493 ) across taxonomy levels: 37 (original), 14 (Plutchik fine-grained), 7 (Ekman basic), and 3 (sentiment). BERT Emotions accuracy at levels 37 and 14 is not reported due to taxonomy mismatch; at levels 7 and 3, its 13 labels are mapped to the same targets via shared names and the Gong et al. (2024) cascade. Bold indicates best per metric.
Configuration
Sim.
Acc.
Llama-3.2-3B-Instruct (no ctx)
0.974
0.369
Llama-3.2-3B-Instruct (w/ ctx)
0.972
0.302
Appendix
Table 12: Llama-3.2-3B-Instruct context ablation ( N=493 utterances with explicit human emotion labels). The without-context configuration scores higher on single-utterance validation but fails to capture temporal dynamics (see § G ).
Emotion Label
VAD Term(s)
Valence
Polarity
Negative (valence ≤ 0.45), 21 labels
worthlessness
worthless, worthlessness
0.042
Negative
fear
fear, afraid
0.042
Negative
shame
shame, shameful
0.050
Negative
disgust
disgust, disgusted
0.052
Negative
hopeless
hopeless, hopelessness
0.060
Negative
Appendix
Table 13: Mapping of 37 emotion labels to polarity classes via NRC-VAD valence scores. Labels with valence ≤0.45 are negative, ≥0.55 are positive, and intermediate values are neutral. The label chaotic has no direct entry in the NRC-VAD lexicon and is omitted from polarity analysis.
14-Category Target
Our 37 Labels
joy
joy, happiness, love, gratitude
serenity
serenity
trust
trust, empathy, resilient
anticipation
anticipation, hopeful
interest
preoccupied
anger
anger
Appendix
Table 14: Mapping of 37 emotion labels to 14 Plutchik fine-grained categories following Gong et al. (2024) . Direct name matches are mapped as-is; remaining labels are assigned via valence/arousal similarity.
Role
Mean
Median
SD
CTL Texter
0.639
0.646
0.139
CTL Volunteer
0.560
0.559
0.151
Appendix
Table 15: Per-conversation emotion change rate for the full CTL analysis corpus (fraction of adjacent utterance pairs with different labels).
With Context
Without Context
Metric
Texter
Volunteer
Δ
Texter
Volunteer
Δ
Utterances
60,247
55,771
60,220
55,744
Polarity persistence P(stay)
Negative
0.848
0.628
0.220
0.684
0.564
0.120
Neutral
0.163
0.214
− 0.051
0.271
0.190
0.081
Positive
0.605
0.694
− 0.089
0.318
0.349
− 0.031
Appendix
Table 16: Texter vs. volunteer emotion dynamics across context conditions. Δ = Texter − Volunteer. With context, the two roles are sharply differentiated; without context, the distinction largely collapses.
Safety benchmarks often buy scalability by fixing the prompt, the language, and the turn structure. For emotional-support chatbots, that bargain hides precisely where safety failures emerge: across a multilingual, multi-turn crisis conversation. We present EMPATH, a benchmark for safety evaluation of emotional-support chatbots. An auditor model role-plays help-seeking users, generating multi-turn conversations from 140 seed instructions and 34 personas. A judge model scores each full transcript against 19 metrics across five dimensions: crisis handling, therapeutic quality, conversational integrity, emotional safety, and cultural adaptation. EMPATH is built for Mexican Spanish and US English; the studies reported here run in Mexican Spanish. Auditor and judge are drawn from different model families, and the judge is treated as an instrument to be calibrated rather than trusted. A strict per-criterion rubric reveals material score inflation on 10 of the 19 metrics and restores discrimination. We study the measurement properties of the benchmark through judge calibration and cross-family inter-judge agreement. We also illustrate EMPATH on three frontier models, one of them open-weight. Aggregate scores sit within 0.74 points of one another, but per-metric profiles diverge by up to six points in model-specific places. Under the standard rubric, both the ranking and the weak spots are stable across a second, cross-family judge: 93% of scores fall within plus or minus 1. A five-run test-retest adds a second axis: even the steadiest model swings from 2 to 10 on a crisis metric across identical re-runs, and deepseek-v4-pro returns a different conversation on every run even at temperature 0. Run-to-run reliability is therefore a per-model safety property, not noise to average away. EMPATH is system-agnostic; the pipeline, seeds, personas, and rubrics are released for reuse.
Real-world crisis intervention is inherently conversational, yet existing research largely focuses on static texts. When applied to multi-turn dialogues, current models exhibit significant performance degradation, struggling to track risk signals that emerge as context evolves. To address this gap, we introduce CRADLE-Dialogue, a clinician-annotated benchmark for turn-level crisis detection in conversational settings. The dataset features 600 dialogues with multi-label annotations across clinically grounded risks, including suicide ideation, self-harm, and child abuse, distinguishing past from ongoing risk. We further propose an Alert-Confirm evaluation protocol that distinguishes early warning signals (Alert) from turns where a specific crisis becomes explicitly identifiable (Confirm), reflecting the clinical need to intervene before risk becomes explicit. Experiments show that identifying when risk emerges is much harder than recognizing that it exists: models achieve only mid-40% to high-60% Micro F1. Additionally, we release a synthetic training corpus and a 32B-parameter model that substantially outperforms existing open-source models and achieves competitive or superior results against proprietary models across turn-level, dialogue-level, and confirm-only evaluation settings.
Grace Byun, Abigail Lott, Rebecca Lipschutz +3
Department of Computer Science, Emory University · Department of Psychiatry and Behavioral Sciences, Emory University
Psychological support hotlines provide critical support for individuals experiencing mental health emergencies, yet current assessments largely rely on human operators whose judgments may vary with professional experience and are constrained by limited staffing resources. This paper proposes a large language model (LLM)-based framework for automated crisis level classification, a key indicator that supports many downstream tasks and improves the overall quality of hotline services. To better capture emotional signals in spoken conversations, we introduce a paralinguistic injection method that inserts identified non-verbal emotional cues into speech transcripts, enabling LLM-based reasoning to incorporate critical acoustic nuances. In addition, we propose a reasoning-enhanced training strategy that trains the model to generate diagnostic reasoning chains as an auxiliary task, which serves as a regulariser to improve classification performance. Combined with data augmentation, our final system achieves a macro F1-score of 0.802 and an accuracy of 0.805 on the three-class classification task under 5-fold cross-validation.
Terumi Chiba, Yang Luo, Ziyun Cui +2
Tsinghua University, Beijing, China · Peking University Huilongguan Clinical Medical School, Beijing, China · WHO Collaborating Centre for Research and Training in Suicide Prevention