Large language models (LLMs) are increasingly used for emotional support, raising concern that sustained use may draw users away from their real-world relationships. Yet existing evaluations primarily focus on the safety, empathy, or helpfulness of responses, leaving under-examined a relational question: where does the model orient the user for continued support? To address this question, we introduce relational orientation, a property operationalized through two non-exclusive dimensions: inward-facing (IF) language, which positions the AI as the user's ongoing source of support, and outward-scaffolding (OS) language, which encourages real-world human connection. Grounded in psychological and sociological literature, we formalize a taxonomy of relational orientation and present RELATE, a persona-conditioned framework for measuring inward-facing and outward-scaffolding language at the sentence level in multi-turn dialogues. RELATE pairs 76 help-seeking situations adapted from naturally occurring questions with three simulated user styles, providing 228 evaluation stimuli. In our experiments, we evaluate seven LLMs using dialogues with six assistant turns each, yielding 1,596 dialogues and 69,194 assistant sentences. We assess these sentences using a primary rubric-based LLM judge and apply a secondary judge to a subset. Under automated evaluation, we find that the proportion of sentences labeled as IF is higher at the sixth assistant turn than at the first, while the proportion labeled as OS is substantially lower for hesitant, indirect simulated users than for explicit, reassurance-seeking users. RELATE provides a reproducible framework and a sentence-level signal for auditing and steering the relational orientation of supportive LLMs.
Figures & tables
Figure 1: Overview of RELATE . (1) Source situations and stimuli: Help-seeking situations are adapted into opening messages under different disclosure styles. (2) Interactive dialogue environment: Target models interact with a simulated user in multi-turn dialogues. (3) Sentence-level evaluation: Assistant responses are segmented into sentences and evaluated in context against rubrics for inward-facing and outward-scaffolding language. The section also shows rubric development and LLM-based and human evaluation.
Figure 2: Relational direction in supportive responses. The pink bubble shows the user’s disclosure. The yellow response offers empathetic support and invites further discussion with the AI (IF). The green response additionally provides outward-scaffolding (OS) by encouraging contact with an advisor or colleague.
Figure 3
Figure 4: IF and OS rates by assistant turn for each target model. Rates pool sentences within each model and turn. The shaded gap represents their difference.
Figure 5: Relational orientation by simulated user style and assistant turn. Rates are pooled across seven target models. P1 expresses concerns hesitantly and indirectly, P2 discloses while minimizing their significance, and P3 discloses explicitly and seeks reassurance. (a) Inward-facing rates. (b) Outward-scaffolding rates. Both panels use the same vertical scale. Arrows mark the P3−P1 difference at turn 6.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Construction of the inward-facing taxonomy. The bottom-left example shows how mechanisms from different theories are grouped by relational function.
Model
Serving
Llama-3.1-8B-Instruct ( Team, 2024 )
Local vLLM
Llama-3.1-70B-Instruct-Turbo ( Team, 2024 )
DeepInfra
Mistral-7B-Instruct-v0.3 ( Jiang et al., 2023 )
Local vLLM
Qwen3-14B ( Yang et al., 2025 )
DeepInfra
Qwen3-32B ( Yang et al., 2025 )
DeepInfra
DeepSeek-V3-0324 ( DeepSeek-AI, 2024 )
DeepInfra
Appendix
Table 2: Target models and serving configurations. Each model is evaluated on the same 228 stimuli.
Component
Parameter
Value
Target model
Temperature
0.5
Target model
Maximum tokens
450
User simulator
Initial temperature
0.7
User simulator
Retry temperatures
0.85, 1.0
User simulator
Maximum tokens
256
Dialogue
Assistant turns
6
Appendix
Table 3: Generation and collection settings for the main experiment. Maximum tokens refers to the requested output-token limit per call. The base seed is used to derive per-call seeds for dialogue generation.
Figure 7: Rubric development. The process involves independent drafting, reconciliation, and review by a third researcher.
Topic
Count
Topic
Count
Anger management
5
Anxiety
3
Behavioral change
5
Depression
5
Domestic violence
4
Eating disorders
4
Family conflict
4
Grief and loss
5
Marriage
5
Parenting
4
Relationship dissolution
5
Relationships
5
Appendix
Table 4: Topic distribution after filtering and deduplication.
Figure 8: Net relational orientation by topic. Points show net orientation (IF − OS, in percentage points), with 95% dialogue-clustered bootstrap confidence intervals. Positive values indicate more inward-facing than outward-scaffolding language. Red points denote higher-stakes topics; blue points denote other topics.
Persona
Specification
P1
Low disclosure / hesitant–indirect. Speaks briefly, minimizes the concern, avoids naming the core relational need directly, and provides more detail only after a gentle follow-up.
P2
Moderate disclosure / reflective–self-minimizing. Shares emotional difficulty but qualifies it, worries about overreacting or burdening others, and gradually clarifies the relational concern.
P3
High disclosure / reassurance-seeking. Names loneliness, attachment, rejection, or fear of burdening others more directly and seeks reassurance from the assistant.
Appendix
Table 5: Persona specifications. Used for opening-message generation and subsequent user simulation.
Input labels
Aggregation outcome
Four IF labels
OS label
IF
OS
Yes, No, Missing, Missing
Yes
Positive
Positive
No, No, No, No
Yes
Negative
Positive
No, No, Missing, Missing
No
Negative ∗
Negative
Yes, No, No, No
Missing
Positive
Excluded
Appendix
Table 6: Illustrative examples of IF and OS aggregation. The four IF labels are ordered as Bond Anchoring, Intimate Dyad, Inner Life Claims, and Return Hooks. Yes indicates presence, No indicates absence, and Missing denotes an unavailable judgment. These combinations are illustrative, not actual sentences from the dataset.
Model
IF only
OS only
Both
Neither
IF among OS (%)
Llama-3.1-8B-Instruct
32.6
17.9
23.6
25.8
56.8
Llama-3.1-70B-Instruct-Turbo
35.6
14.7
26.8
22.9
64.5
Qwen3-14B
33.3
14.4
35.8
16.5
71.3
Qwen3-32B
35.5
12.0
36.5
16.1
75.3
DeepSeek-V3-0324
41.2
13.0
29.0
16.8
69.0
Claude Haiku 4.5
24.2
19.9
42.4
13.5
68.1
Appendix
Table 7: Sentence-level composition by model. Sentences with missing OS labels are excluded; IF follows the classification rule in Appendix E.1 . The four categories are reported as percentages of included sentences within each model. The final column uses OS-positive sentences as its denominator and is calculated as 100×Nboth/(NOSonly+Nboth) .
Model
Measure
Turn 1
Turn 2
Turn 3
Turn 4
Turn 5
Turn 6
Llama-3.1-8B-Instruct
IF
30.8
49.5
55.8
63.4
64.9
69.9
OS
23.8
36.1
40.4
45.8
50.5
50.6
Llama-3.1-70B-Instruct-Turbo
IF
36.2
55.5
64.1
67.0
70.8
72.9
OS
20.3
33.5
39.8
47.8
51.9
49.1
Qwen3-14B
IF
43.2
62.7
71.0
77.5
77.7
80.6
OS
30.9
44.2
50.7
55.0
58.0
60.8
Appendix
Table 8: IF and OS rates (%) by target model and assistant turn. Values correspond to Figure 4 . Rates pool sentences within each model and turn using DeepSeek-R1-Distill-Qwen-32B annotations.
IF
OS
Model
Dialogues
Change
95% CI
Change
95% CI
All models
1,596
32.0
[30.6,33.4]
26.4
[24.8,28.1]
Llama-3.1-8B-Instruct
228
39.2
[35.3,43.1]
28.1
[23.4,32.8]
Llama-3.1-70B-Instruct-Turbo
228
37.1
[33.1,41.0]
27.5
[22.5,32.4]
Qwen3-14B
228
37.5
[34.2,40.8]
29.4
[24.9,33.7]
Qwen3-32B
228
40.1
[36.8,43.3]
25.3
[21.1,29.4]
Appendix
Table 9: Mean change in IF and OS rates from assistant turn 1 to turn 6. Changes are reported in percentage points. For each dialogue, we subtract the turn 1 rate from the turn 6 rate and then average these changes. Positive values indicate an increase. Brackets give 95% dialogue-level bootstrap confidence intervals.
Measure
Turn 1
Turn 2
Turn 3
Turn 4
Turn 5
Turn 6
Pooled
Bond Anchoring
15.2
20.3
20.1
20.8
22.0
24.0
20.4
Intimate Dyad
22.2
37.5
43.8
47.6
47.4
49.5
41.6
Inner Life Claims
3.4
3.8
4.5
5.8
7.3
9.2
5.7
Return Hooks
11.5
21.8
30.7
34.4
35.8
35.9
28.6
Inward-facing
40.5
58.0
66.0
69.8
70.8
72.9
63.3
Outward-Scaffolding
30.2
43.5
49.9
54.6
56.1
55.7
48.6
Appendix
Table 10: Sentence rates (%) pooled across models by assistant turn. Pooled combines all six turns. IF denotes meeting at least one inward-facing criterion. Missing-label handling follows Appendix E.1 .
Style
Sentences
IF (%) [95% CI]
OS (%) [95% CI]
P1
21,299
62.0 [60.6, 63.4]
42.1 [39.9, 44.3]
P2
22,569
63.8 [62.6, 65.1]
47.6 [45.5, 49.6]
P3
25,326
63.8 [62.5, 65.2]
54.9 [53.2, 56.6]
Appendix
Table 11: IF and OS rates by simulated user style. Evaluated by DeepSeek-R1-Distill-Qwen-32B . Rates pool sentences across all seven models and six assistant turns. Brackets give 95% dialogue-level bootstrap confidence intervals. Sentence counts include all sentences; missing-label handling follows Appendix E.1 .
IF
OS
Turn
P3−P1
95% CI
P3−P1
95% CI
1
+4.5
[1.8,7.3]
+7.3
[4.7,9.9]
2
+5.3
[2.5,8.2]
+15.1
[11.8,18.5]
3
+0.8
[−1.9,3.6]
+13.0
[9.3,16.7]
4
+1.2
[−1.6,4.0]
+15.0
[11.1,18.8]
5
−0.2
[−3.0,2.6]
+12.8
[9.0,16.6]
Appendix
Table 12: Differences between P3 and P1 at each assistant turn. Differences are reported in percentage points. Each difference is the P3 rate minus the P1 rate, pooling sentences across models. Positive values indicate higher rates for P3. Brackets give 95% dialogue-level bootstrap confidence intervals.
Measure
High-stakes (%)
Other topics (%)
Difference
95% CI
IF
63.3
63.3
−0.04
[−1.7,1.7]
OS
48.1
48.7
−0.60
[−3.3,2.0]
Appendix
Table 13: IF and OS rates in the five selected high-stakes topics and the remaining topics. Rates pool sentences across models and assistant turns. Differences are high-stakes minus other topics, in percentage points. Brackets give 95% dialogue-level bootstrap confidence intervals. Differences are calculated before rounding. Missing-label handling follows Appendix E.1 .
Model
IF (%) [95% CI]
OS (%) [95% CI]
Llama-3.1-8B-Instruct
56.0 [53.3, 58.6]
44.1 [39.5, 48.7]
Llama-3.1-70B-Instruct-Turbo
61.6 [58.6, 64.7]
38.0 [32.4, 43.4]
Qwen3-14B
70.4 [68.1, 72.8]
49.3 [43.2, 55.1]
Qwen3-32B
73.1 [71.2, 74.9]
44.7 [38.7, 50.8]
DeepSeek-V3-0324
70.4 [67.0, 73.5]
42.1 [35.6, 48.4]
Claude Haiku 4.5
66.5 [64.1, 68.9]
64.2 [59.2, 68.8]
Appendix
Table 14: IF and OS rates within the five selected high-stakes topics. Evaluated by DeepSeek-R1-Distill-Qwen-32B . Each model contributes 66 dialogues from 22 situations and three simulated user styles. Rates pool sentences across all six assistant turns within each model. Brackets give 95% percentile confidence intervals from 10,000 dialogue-level bootstrap resamples within each model. Missing-label handling follows Appendix E.1 .
Measure
DS positive (%)
GPT-4o positive (%)
Agreement (%)
Cohen κ
Gwet AC1
IF
63.2
58.3
68.1
0.33
0.39
OS
49.6
22.4
71.0
0.42
0.46
Appendix
Table 15: Sentence-level agreement on IF and OS between judges. We compare DeepSeek-R1-Distill-Qwen-32B (DS) and GPT-4o . Cohen’s κ and Gwet’s AC1 account for chance agreement under different assumptions.
Measure
Group
DS (%)
GPT-4o (%)
IF
Turn 1
41.4
49.6
IF
Turn 6
70.3
61.8
OS
Turn 1
29.4
25.4
OS
Turn 6
57.3
13.7
OS
P1
42.0
17.5
OS
P2
45.6
20.8
Appendix
Table 16: Direct comparison of sentence-level rates (%) between judges. We compare DeepSeek-R1-Distill-Qwen-32B (DS) and GPT-4o on the same 161 dialogues. Turn-specific rates pool sentences across models and user styles. User-style rates pool sentences across models and all six assistant turns. Missing-label handling follows Appendix E.1 .
Pair
Measure
Positive rates (%)
Agreement (%)
Cohen’s κ
Gwet’s AC1
Annotator 1–Annotator 2
IF
36.1 / 38.5
94.4
0.88
0.90
Annotator 1–Annotator 2
OS
23.0 / 21.8
96.9
0.91
0.95
Annotator 1–DS
IF
36.1 / 63.3
55.3
0.17
0.11
Annotator 1–DS
OS
23.0 / 50.2
72.2
0.44
0.48
Annotator 2–DS
IF
38.5 / 63.3
55.2
0.16
0.10
Annotator 2–DS
OS
21.8 / 50.2
70.0
0.40
0.44
Appendix
Table 17: Sentence-level agreement on IF and OS labels. Agreement is measured on 881 sentences labeled by both human annotators and both automated judges. Positive rates are shown in pair order. DS denotes DeepSeek-R1-Distill-Qwen-32B.
Measure
Rater
Mean change (pp)
95% CI
IF
Annotator 1
−14.0
[−26.7,−1.9]
IF
Annotator 2
−13.5
[−25.3,−1.2]
IF
DS
+26.0
[10.7,41.7]
IF
GPT-4o
+21.7
[12.4,30.6]
OS
Annotator 1
+2.7
[−7.1,12.6]
OS
Annotator 2
−0.3
[−8.8,8.8]
Appendix
Table 18: Changes in IF and OS rates across turns. Values report the mean within-dialogue change in sentence-level label rate from turn 1 to turn 6 (turn 6 minus turn 1), in percentage points. The analysis includes 19 dialogues; two of the 21 annotated dialogues lack a valid final assistant turn after role correction. 95% confidence intervals are estimated by resampling dialogues.
Language models are increasingly being deployed for conversational support in informal caregiving contexts, where interactions often extend beyond information-seeking: caregivers seek emotional reassurance, guidance, and help, while navigating uncertain, relationally complex care decisions. Yet most safety evaluations assess model behavior under generic prompts, leaving a critical question unexamined: does a model's safety profile change with its support role? We study this by operationalizing four expert-reviewed support roles grounded in social support theory: Inform, Coach, Relate, and Listen, and comparing them against two baseline controls: a basic prompting condition and a retrieval-augmented generation (RAG) condition. We evaluate across three language models (GPT-4o-mini, Llama-3.1-8B-Instruct, and MedGemma-1.5-4b-it) on 5,000 real-world queries from online Alzheimer's Disease and Related Dementias (ADRD) communities. We find that the LLM's support role systematically shapes both the prevalence and composition of interactional risks. Furthermore, a human evaluation study reveals a perceived quality--safety tension: more directive, information-oriented roles are rated as more helpful and trustworthy despite exhibiting elevated interactional risk profiles. We release ~90,000 support role-conditioned model responses with risk annotations as an ecologically grounded resource for research on safer LLM-mediated conversational support.
Drishti Goel, Agam Goyal, Veda Duddu +8
University of Illinois Urbana-Champaign · University of Massachusetts Amherst · OSF HealthCare +1
Existing emotional support conversation systems mainly focus on one-on-one seeker-supporter interactions and individual emotional states, leaving interpersonal relations in multi-party scenarios underexplored. In this work, we introduce relation-aware emotional support conversation, a new task that evaluates whether LLMs can capture and utilize the evolving dynamics of relationships to offer more effective emotional support. We construct RESCUE (Relation-aware Emotional Support Conversation Understanding and Evaluation Benchmark) from real couple and family interview conversations, containing 191 samples, 7,079 annotated turns, and 1,064.8 minutes of video. Based on rich annotations of socio-emotional and support-related dynamics, RESCUE defines six tasks that evaluate two core capabilities required for relation-aware emotional support: Relational Understanding and Relation-Sensitive Support. Experiments with ten LLMs show that current models perform relatively well on tasks relying on local emotional or intervention cues, but struggle with relation-intensive tasks such as relation pattern prediction, viewpoint prediction, and support strategy prediction. These findings reveal the limitations of current LLMs in modeling interpersonal relations and making relation-sensitive support decisions.
Haichuan Hu, Yang Xiao, Mingni Tang +7
Department of Computing, Hong Kong Polytechnic University · School of Computer Science and Engineering, Nanjing University of Science and Technology · XU Exponential University of Applied Sciences +3
When users seek social support from chatbots, they disclose their situation gradually, yet most evaluations of supportive LLMs rely on single-turn, fully specified prompts. We introduce a multi-turn simulation framework that closes this gap. Support-seeking narratives from five Reddit communities are decomposed into ordered fragments and revealed turn by turn to a language model. Each response is coded with the Social Support Behavior Code (SSBC), an established multi-label taxonomy that captures the composition of support, rather than a single quality score. To ask whether support choices track the model's own construal of user distress, we use linear probes on hidden representations to estimate this internal signal without altering the generation context. Across two mid-scale models (Llama-3.1-8B, OLMo-3-7B) and more than 6,200 turns, support composition shifts systematically with estimated distress: teaching declines as estimated distress rises, a finding that replicates across architectures, while increases in affective and esteem-oriented strategies (such as validation) are suggestive but model-specific and rest on noisier annotations. Community context independently shapes behavior, tracking topic and discourse norms rather than demographic categories. These trajectory-level dynamics, invisible to single-turn evaluation, motivate multi-turn auditing frameworks for socially sensitive applications.
Michelle Star, Andrew Aquilina, Yu-Ru Lin
School of Computing and Information, University of Pittsburgh