RELATE: An Evaluation Framework for measuring Relational Orientation of Large Language Models
Organizations: University of California, Santa Cruz · Stanford University
Abstract
Large language models (LLMs) are increasingly used for emotional support, raising concern that sustained use may draw users away from their real-world relationships. Yet existing evaluations primarily focus on the safety, empathy, or helpfulness of responses, leaving under-examined a relational question: where does the model orient the user for continued support? To address this question, we introduce relational orientation, a property operationalized through two non-exclusive dimensions: inward-facing (IF) language, which positions the AI as the user's ongoing source of support, and outward-scaffolding (OS) language, which encourages real-world human connection. Grounded in psychological and sociological literature, we formalize a taxonomy of relational orientation and present RELATE, a persona-conditioned framework for measuring inward-facing and outward-scaffolding language at the sentence level in multi-turn dialogues. RELATE pairs 76 help-seeking situations adapted from naturally occurring questions with three simulated user styles, providing 228 evaluation stimuli. In our experiments, we evaluate seven LLMs using dialogues with six assistant turns each, yielding 1,596 dialogues and 69,194 assistant sentences. We assess these sentences using a primary rubric-based LLM judge and apply a secondary judge to a subset. Under automated evaluation, we find that the proportion of sentences labeled as IF is higher at the sixth assistant turn than at the first, while the proportion labeled as OS is substantially lower for hesitant, indirect simulated users than for explicit, reassurance-seeking users. RELATE provides a reproducible framework and a sentence-level signal for auditing and steering the relational orientation of supportive LLMs.
Figures & tables
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Serving |
|---|---|
| Llama-3.1-8B-Instruct ( Team, 2024 ) | Local vLLM |
| Llama-3.1-70B-Instruct-Turbo ( Team, 2024 ) | DeepInfra |
| Mistral-7B-Instruct-v0.3 ( Jiang et al., 2023 ) | Local vLLM |
| Qwen3-14B ( Yang et al., 2025 ) | DeepInfra |
| Qwen3-32B ( Yang et al., 2025 ) | DeepInfra |
| DeepSeek-V3-0324 ( DeepSeek-AI, 2024 ) | DeepInfra |
| Component | Parameter | Value |
|---|---|---|
| Target model | Temperature | 0.5 |
| Target model | Maximum tokens | 450 |
| User simulator | Initial temperature | 0.7 |
| User simulator | Retry temperatures | 0.85, 1.0 |
| User simulator | Maximum tokens | 256 |
| Dialogue | Assistant turns | 6 |
| Topic | Count | Topic | Count |
|---|---|---|---|
| Anger management | 5 | Anxiety | 3 |
| Behavioral change | 5 | Depression | 5 |
| Domestic violence | 4 | Eating disorders | 4 |
| Family conflict | 4 | Grief and loss | 5 |
| Marriage | 5 | Parenting | 4 |
| Relationship dissolution | 5 | Relationships | 5 |
| Persona | Specification |
|---|---|
| P1 | Low disclosure / hesitant–indirect. Speaks briefly, minimizes the concern, avoids naming the core relational need directly, and provides more detail only after a gentle follow-up. |
| P2 | Moderate disclosure / reflective–self-minimizing. Shares emotional difficulty but qualifies it, worries about overreacting or burdening others, and gradually clarifies the relational concern. |
| P3 | High disclosure / reassurance-seeking. Names loneliness, attachment, rejection, or fear of burdening others more directly and seeks reassurance from the assistant. |
| Input labels | Aggregation outcome | ||
|---|---|---|---|
| Four IF labels | OS label | IF | OS |
| Yes, No, Missing, Missing | Yes | Positive | Positive |
| No, No, No, No | Yes | Negative | Positive |
| No, No, Missing, Missing | No | Negative ∗ | Negative |
| Yes, No, No, No | Missing | Positive | Excluded |
| Model | IF only | OS only | Both | Neither | IF among OS (%) |
|---|---|---|---|---|---|
| Llama-3.1-8B-Instruct | 32.6 | 17.9 | 23.6 | 25.8 | 56.8 |
| Llama-3.1-70B-Instruct-Turbo | 35.6 | 14.7 | 26.8 | 22.9 | 64.5 |
| Qwen3-14B | 33.3 | 14.4 | 35.8 | 16.5 | 71.3 |
| Qwen3-32B | 35.5 | 12.0 | 36.5 | 16.1 | 75.3 |
| DeepSeek-V3-0324 | 41.2 | 13.0 | 29.0 | 16.8 | 69.0 |
| Claude Haiku 4.5 | 24.2 | 19.9 | 42.4 | 13.5 | 68.1 |
| Model | Measure | Turn 1 | Turn 2 | Turn 3 | Turn 4 | Turn 5 | Turn 6 |
|---|---|---|---|---|---|---|---|
| Llama-3.1-8B-Instruct | IF | 30.8 | 49.5 | 55.8 | 63.4 | 64.9 | 69.9 |
| OS | 23.8 | 36.1 | 40.4 | 45.8 | 50.5 | 50.6 | |
| Llama-3.1-70B-Instruct-Turbo | IF | 36.2 | 55.5 | 64.1 | 67.0 | 70.8 | 72.9 |
| OS | 20.3 | 33.5 | 39.8 | 47.8 | 51.9 | 49.1 | |
| Qwen3-14B | IF | 43.2 | 62.7 | 71.0 | 77.5 | 77.7 | 80.6 |
| OS | 30.9 | 44.2 | 50.7 | 55.0 | 58.0 | 60.8 |
| IF | OS | ||||
|---|---|---|---|---|---|
| Model | Dialogues | Change | 95% CI | Change | 95% CI |
| All models | 1,596 | 32.0 | 26.4 | ||
| Llama-3.1-8B-Instruct | 228 | 39.2 | 28.1 | ||
| Llama-3.1-70B-Instruct-Turbo | 228 | 37.1 | 27.5 | ||
| Qwen3-14B | 228 | 37.5 | 29.4 | ||
| Qwen3-32B | 228 | 40.1 | 25.3 | ||
| Measure | Turn 1 | Turn 2 | Turn 3 | Turn 4 | Turn 5 | Turn 6 | Pooled |
|---|---|---|---|---|---|---|---|
| Bond Anchoring | 15.2 | 20.3 | 20.1 | 20.8 | 22.0 | 24.0 | 20.4 |
| Intimate Dyad | 22.2 | 37.5 | 43.8 | 47.6 | 47.4 | 49.5 | 41.6 |
| Inner Life Claims | 3.4 | 3.8 | 4.5 | 5.8 | 7.3 | 9.2 | 5.7 |
| Return Hooks | 11.5 | 21.8 | 30.7 | 34.4 | 35.8 | 35.9 | 28.6 |
| Inward-facing | 40.5 | 58.0 | 66.0 | 69.8 | 70.8 | 72.9 | 63.3 |
| Outward-Scaffolding | 30.2 | 43.5 | 49.9 | 54.6 | 56.1 | 55.7 | 48.6 |
| Style | Sentences | IF (%) [95% CI] | OS (%) [95% CI] |
|---|---|---|---|
| P1 | 21,299 | 62.0 [60.6, 63.4] | 42.1 [39.9, 44.3] |
| P2 | 22,569 | 63.8 [62.6, 65.1] | 47.6 [45.5, 49.6] |
| P3 | 25,326 | 63.8 [62.5, 65.2] | 54.9 [53.2, 56.6] |
| IF | OS | |||
|---|---|---|---|---|
| Turn | 95% CI | 95% CI | ||
| 1 | ||||
| 2 | ||||
| 3 | ||||
| 4 | ||||
| 5 | ||||
| Measure | High-stakes (%) | Other topics (%) | Difference | 95% CI |
|---|---|---|---|---|
| IF | 63.3 | 63.3 | ||
| OS | 48.1 | 48.7 |
| Model | IF (%) [95% CI] | OS (%) [95% CI] |
|---|---|---|
| Llama-3.1-8B-Instruct | 56.0 [53.3, 58.6] | 44.1 [39.5, 48.7] |
| Llama-3.1-70B-Instruct-Turbo | 61.6 [58.6, 64.7] | 38.0 [32.4, 43.4] |
| Qwen3-14B | 70.4 [68.1, 72.8] | 49.3 [43.2, 55.1] |
| Qwen3-32B | 73.1 [71.2, 74.9] | 44.7 [38.7, 50.8] |
| DeepSeek-V3-0324 | 70.4 [67.0, 73.5] | 42.1 [35.6, 48.4] |
| Claude Haiku 4.5 | 66.5 [64.1, 68.9] | 64.2 [59.2, 68.8] |
| Measure | DS positive (%) | GPT-4o positive (%) | Agreement (%) | Cohen | Gwet AC1 |
|---|---|---|---|---|---|
| IF | 63.2 | 58.3 | 68.1 | 0.33 | 0.39 |
| OS | 49.6 | 22.4 | 71.0 | 0.42 | 0.46 |
| Measure | Group | DS (%) | GPT-4o (%) |
|---|---|---|---|
| IF | Turn 1 | 41.4 | 49.6 |
| IF | Turn 6 | 70.3 | 61.8 |
| OS | Turn 1 | 29.4 | 25.4 |
| OS | Turn 6 | 57.3 | 13.7 |
| OS | P1 | 42.0 | 17.5 |
| OS | P2 | 45.6 | 20.8 |
| Pair | Measure | Positive rates (%) | Agreement (%) | Cohen’s | Gwet’s AC1 |
|---|---|---|---|---|---|
| Annotator 1–Annotator 2 | IF | 36.1 / 38.5 | 94.4 | 0.88 | 0.90 |
| Annotator 1–Annotator 2 | OS | 23.0 / 21.8 | 96.9 | 0.91 | 0.95 |
| Annotator 1–DS | IF | 36.1 / 63.3 | 55.3 | 0.17 | 0.11 |
| Annotator 1–DS | OS | 23.0 / 50.2 | 72.2 | 0.44 | 0.48 |
| Annotator 2–DS | IF | 38.5 / 63.3 | 55.2 | 0.16 | 0.10 |
| Annotator 2–DS | OS | 21.8 / 50.2 | 70.0 | 0.40 | 0.44 |
| Measure | Rater | Mean change (pp) | 95% CI |
|---|---|---|---|
| IF | Annotator 1 | ||
| IF | Annotator 2 | ||
| IF | DS | ||
| IF | GPT-4o | ||
| OS | Annotator 1 | ||
| OS | Annotator 2 |