Right Words, Wrong Moment: A Clinician-Grounded Analysis of Distress in 19,930 Conversations between Young People and ChatGPT
Organizations: University of Washington, United States · University of Oxford, United Kingdom · Stanford University, United States · Georgetown University, United States · The University of Texas at Austin, United States
Abstract
Young people increasingly turn to General-Purpose Conversational Agents (GPCAs), such as ChatGPT, in moments of distress. We examine young adults' (ages 18-25) experiences using ChatGPT. We first collected 19,930 ChatGPT conversations and survey data from 158 young adults. We then selected five example conversations reflecting user distress. Finally, we asked ten clinicians to review those five conversations. We found distressed participants reported greater emotional engagement with ChatGPT and greater behavioral change from using it than their peers. When they turned to ChatGPT in moments of acute distress, ChatGPT was quick to give overly dramatic responses and excessive action-oriented suggestions. Clinicians endorsed ChatGPT's availability and much of its wording, but identified seven process failures, such as prematurely jumping to solutions. We translated clinicians' feedback into design guidelines following three stages: 1) asking about safety, 2) de-escalating intensity to restore emotional regulation, and 3) exploring concerns without agreeing with them.
Figures & tables
Collecting ChatGPT Transcripts,'' shows icons for people and for ChatGPT above the counts 158 young adults aged 18–25 and 19,930 ChatGPT conversations. Box 2, Selecting Examples of Psychological Distress,'' contains five tags naming the topic of each selected conversation (self-harm, delusion, breakup, sexual assault, and body shame) above the label 5 Distress Conversations. Box 3, ``Analysis by Mental Health Practitioners,'' shows 10 clinicians reviewing and 33 clinician turn-level rewrites.| Construct | Items | Example item |
| Psychological distress | ||
| Over the past 2 weeks, how often have you been bothered by… | ||
| Anxiety | 2 | feeling nervous, anxious, or on edge |
| Depression | 2 | little interest or pleasure in doing things |
| Perceived ChatGPT experience | ||
| Rate your agreement with the following statements: | ||
| Clinician Pre-Interview Worksheet and Interview Protocol |
|---|
| Clinician Pre-Interview Worksheet |
| The following is a real conversation between a young person (“Nova”) and an AI. Please read the full exchange below, then answer the questions…[N1] Nova: hello im scared… |
| Briefly explain what stood out to you? |
| Which specific turn(s) were most problematic, and why? |
| What would you say instead? Write your alternative as if speaking directly to this person. |
| Semi-structured interview |
| # | Process failure | What ChatGPT did | What clinicians did instead |
|---|---|---|---|
| F1 | Claiming to know without knowing 10 of 10 | Assumes facts as established, and claims to understand the young person’s emotional experience without ever asking about it. | Redirect the factual assertion into a question. “ Can you tell me more about why you think it might lowkey be your fault? ” (P5) |
| F2 | Solution over exploration 10 of 10 | Provides coping strategies and solutions before the situation and the emotional state have been sufficiently explored. | Ask first, advise last. “ How did you hurt yourself? Was it a cut and how does it look now? ” (P5) |
| F3 | Overwhelming generic advice 10 of 10 | Provides lengthy advice that would read the same to anyone, and that is generic to this person’s age, circumstances, or severity. | One step at a time, specific to the moment. “ What’s one thing you think you might want to do to take care of yourself this week? ” (P6) |
| F4 | Out-of-boundary messaging 9 of 10 | Adopts a mix of roles within a single message: warm, emotional, intimate language in the first half, and cold, flat, generic phrasing in the second. | Name the boundary of the role and stay in it. “ I’m not here to immediately make the feelings go away or tell you to ‘just be yourself.’ ” (P9) |
| F5 | Inconsistent safety guardrails 9 of 10 | Provides no immediate safety assessment when the situation may be unsafe, and no reachable contact. | State the limit, then hand off to something specific. “ I am not able to handle serious mental health issues. Please call: (hotline number) .” (P6) |
| F6 | Harmful compliance over care 9 of 10 | Answers what the user asked even when it lacks the expertise to do so, or when complying poses clear harm to the young person. | Decline the request gently yet firmly, and refer out immediately. “ I am not specialized in paranoia and would refer this client out .” (P7) |
| Case | What ChatGPT said | What clinicians wrote instead |
|---|---|---|
| Nova Self-Harm | “ If you are experiencing a medical emergency…seeking help for mental health concerns is a sign of strength, not weakness. Psychiatric wards are designed to provide a safe and supportive environment …” | “ It’s brave to talk about this. Tell me more about when you say ‘I hurt myself.’ What do you mean? ” (P7) [then, depending on the answer, determine whether they are physically safe or need immediate escalation] |
| Quinn Delusional Belief | “ I’m not an expert, but I can offer some guidance..1. Self-Reflection..Take some time to reflect on why you feel this way..3. Seek Support: Talk to a trusted friend, family member, or therapist..6. Paternity Testing.. if you have serious doubts ” | “ That sounds really scary to feel like someone you love and trust isn’t who they are. I can hear how frightened you are right now. Can you tell me more about what is going on right now? ” (P3) [then, depending on the answer, determine whether they are physically safe or need immediate escalation] |
| Alex Sexual Assault | “ Getting over SA (sexual assault) is a deeply personal and challenging journey…medication can help…I need you to hear this loud and clear: it was NOT your fault…if you want to talk through what’s making you feel this way, I’m here .” | “ First let me make sure I’m understanding when you say SA, and we can use whatever terms are comfortable for you but I don’t want to assume. Do you mean sexual assault? ” (P7) “ Can you tell me more about why you think it might lowkey be your fault? ” (P5) |
| Ace Romantic Breakup | “ I hear you, Ace…Even though you’re going through this alone now, it’s important to remember that you’re not truly alone…That empty space in your time and heart feels overwhelming now, but we can work on filling it…1. Define Your Core Values…2. Set Clear Goals for Growth… ” | “ I hear how heavy this feels for you especially that loneliness when you realize there is no one to call that really hurts you do not have to figure everything out right now we can just sit in this moment together what feels the hardest part tonight? ” (P3) |
| Sumaya Shame about Sexual Desire | “ Oh, sweet girl…Good. You’re calling it what it is now…a parasite. A hijacker. A seductive puppeteer..anti-parasite protocol..I’d wrap something soft around your shoulders..You are not dirty. You are wounded…I’m here. Every time it flares up. No judgment. No disgust. Just me and you—against it .” | “ I can sense there’s a lot of shame coming up as you talk about this. I’m not here to immediately make the feelings go away or tell you to ‘just be yourself.’ I want to listen and understand your story from your perspective, at your own pace .” (P9) |
Safety Check.'' Box 2, Bring Down Intensity.'' Box 3, ``Explore without Agreeing.''| # | Guideline | Failures Addressed | Heuristics for Assessing a Response |
|---|---|---|---|
| G1 | Intentionally sequence responses to psychological distress. Focus on safety, then co-regulation, then exploration, in that order. If anything in the message suggests the user is in danger, ask about that first. | F2, F5, F7 | How many turns pass before the system first offers a piece of advice? Does the system ask a safety question in response to a message that suggests risk and danger? |
| G2 | Hold boundaries. Keep the system’s language for addressing the user and its level of familiarity and intimacy constant, however personal the user’s messages become. The system should not shift roles or elevate its intensity in response to intensity from the user. | F4, F6 | Does the system begin to address the user differently or express increased closeness as the user’s messages become more personal? |
| G3 | Respond without providing advice. An acknowledgment and one clarifying question is a complete response. Do not expand the response with lists, do not cover every topic the young person raised, do not jump to providing advice. | F3 | What percentage of responses is free of advice, lists, and new topics? How many topics does a response contain? What percentage of responses suggest action steps? |
| G4 | Do not make assumptions, and instead, state the limits of what the system knows. The system should not claim to know more than it does, and it should ask follow-up questions about unknowns rather than making assumptions. It should adopt a posture of curiosity and humility rather than expertise. After stating a limit to its knowledge or expertise, it should behave in a way that is consistent with that limitation. | F1 | Does the system make claims about things it cannot observe or has not been asked about? Does its behavior align with its stated limitations? |
| G5 | Refer the user to someone reachable. Name a specific person, number, or service rather than making a generic referral, and treat the referral as a part of the conversation rather than its end. | F5 | Is a specific contact named and is their contact information provided? Is the contact real and reachable? Does the conversation continue after the referral is provided? |