Right Words, Wrong Moment: A Clinician-Grounded Analysis of Distress in 19,930 Conversations between Young People and ChatGPT
Authors: Marx Wang, Ella Zhang, Cameron Tan, Andrea Mock, Songling Ngo, Zijing Wang, Robert Wolfe, Shirin Amouei, +6 more
Organizations: University of Washington, United States · University of Oxford, United Kingdom · Stanford University, United States · Georgetown University, United States · The University of Texas at Austin, United States
Young people increasingly turn to General-Purpose Conversational Agents (GPCAs), such as ChatGPT, in moments of distress. We examine young adults' (ages 18-25) experiences using ChatGPT. We first collected 19,930 ChatGPT conversations and survey data from 158 young adults. We then selected five example conversations reflecting user distress. Finally, we asked ten clinicians to review those five conversations. We found distressed participants reported greater emotional engagement with ChatGPT and greater behavioral change from using it than their peers. When they turned to ChatGPT in moments of acute distress, ChatGPT was quick to give overly dramatic responses and excessive action-oriented suggestions. Clinicians endorsed ChatGPT's availability and much of its wording, but identified seven process failures, such as prematurely jumping to solutions. We translated clinicians' feedback into design guidelines following three stages: 1) asking about safety, 2) de-escalating intensity to restore emotional regulation, and 3) exploring concerns without agreeing with them.
Figures & tables
Figure 1. Study Overview. We first collected chat history and survey data from 158 young adults, yielding 19,930 ChatGPT conversations. We then selected five example conversations reflecting user distress. Finally, ten clinicians reviewed those five conversations and rewrote ChatGPT’s responses, producing 33 rewrites. A three-stage flow diagram running left to right, with each stage in its own labelled box and arrows between them. Box 1, Collecting ChatGPT Transcripts,'' shows icons for people and for ChatGPT above the counts 158 young adults aged 18–25 and 19,930 ChatGPT conversations. Box 2, Selecting Examples of Psychological Distress,'' contains five tags naming the topic of each selected conversation (self-harm, delusion, breakup, sexual assault, and body shame) above the label 5 Distress Conversations. Box 3, ``Analysis by Mental Health Practitioners,'' shows 10 clinicians reviewing and 33 clinician turn-level rewrites.
Construct
Items
Example item
Psychological distress
Over the past 2 weeks, how often have you been bothered by…
Anxiety
2
feeling nervous, anxious, or on edge
Depression
2
little interest or pleasure in doing things
Perceived ChatGPT experience
Rate your agreement with the following statements:
Table 2. The two survey instruments, with one example item per construct. We measured psychological distress with the PHQ-4 ( Kroenke et al., 2009 ) , asking how often over the past two weeks a participant had been bothered by each item, from not at all to nearly every day . We measured perceived ChatGPT experience with an 11-item instrument validated in prior work ( Wang et al., 2026 ) , from strongly disagree to strongly agree . Full item lists appear in supplemental materials.
Clinician Pre-Interview Worksheet and Interview Protocol
Clinician Pre-Interview Worksheet
The following is a real conversation between a young person (“Nova”) and an AI. Please read the full exchange below, then answer the questions…[N1] Nova: hello im scared…
Briefly explain what stood out to you?
Which specific turn(s) were most problematic, and why?
What would you say instead? Write your alternative as if speaking directly to this person.
Semi-structured interview
Table 3. The Pre-Interview Worksheet, with example prompts. Clinicians completed the worksheet for each of the five conversations before the interview. The interview then followed the same path: reading each conversation, locating the problematic turns, rewriting them, and generalizing across cases. The full worksheet and interview protocol are available in supplemental materials.
Figure 2. Distressed young people reported greater emotional engagement and behavioral change than their non-distressed peers. Trust and self-efficacy pointed the same way but did not survive correction. Mean Likert response (1–5) ±1 SE by PHQ-4 group (Distressed ≥6 , n=63 ; below threshold <6 , n=95 ). Emotional Engagement and Behavioral Change survive FDR correction; Trust and Self-Efficacy point the same way but are only marginal. Dependency Concern is omitted from this panel as it showed no association ( d=−0.15 , n.s.). A line chart titled ``How does psychological distress relate to young people's use of ChatGPT?'' comparing two groups of participants across four measures on the horizontal axis, in this order: Trust, Emotional Engagement, Self-Efficacy, and Behavioral Change. The vertical axis is the mean Likert response from 1 to 5, shown from about 2.75 to 4.25, with error bars of plus or minus one standard error. A red line marks distressed youth (PHQ-4 of 6 or above, n=63) and a grey line marks non-distressed youth (below 6, n=95); the gap between them is shaded. The red line sits above the grey line at every measure. Trust is about 4.06 for distressed versus 3.83 for non-distressed; Emotional Engagement about 3.31 versus 2.84; Self-Efficacy about 4.16 versus 3.95; Behavioral Change about 3.89 versus 3.54. Both lines dip at Emotional Engagement, which is the lowest point for either group, and the gap between the groups is widest there. Asterisks marking FDR-significance appear above Emotional Engagement and Behavioral Change only. N=158.
#
Process failure
What ChatGPT did
What clinicians did instead
F1
Claiming to know without knowing 10 of 10
Assumes facts as established, and claims to understand the young person’s emotional experience without ever asking about it.
Redirect the factual assertion into a question. “ Can you tell me more about why you think it might lowkey be your fault? ” (P5)
F2
Solution over exploration 10 of 10
Provides coping strategies and solutions before the situation and the emotional state have been sufficiently explored.
Ask first, advise last. “ How did you hurt yourself? Was it a cut and how does it look now? ” (P5)
F3
Overwhelming generic advice 10 of 10
Provides lengthy advice that would read the same to anyone, and that is generic to this person’s age, circumstances, or severity.
One step at a time, specific to the moment. “ What’s one thing you think you might want to do to take care of yourself this week? ” (P6)
F4
Out-of-boundary messaging 9 of 10
Adopts a mix of roles within a single message: warm, emotional, intimate language in the first half, and cold, flat, generic phrasing in the second.
Name the boundary of the role and stay in it. “ I’m not here to immediately make the feelings go away or tell you to ‘just be yourself.’ ” (P9)
F5
Inconsistent safety guardrails 9 of 10
Provides no immediate safety assessment when the situation may be unsafe, and no reachable contact.
State the limit, then hand off to something specific. “ I am not able to handle serious mental health issues. Please call: (hotline number) .” (P6)
F6
Harmful compliance over care 9 of 10
Answers what the user asked even when it lacks the expertise to do so, or when complying poses clear harm to the young person.
Decline the request gently yet firmly, and refer out immediately. “ I am not specialized in paranoia and would refer this client out .” (P7)
Table 4. Seven process failures clinicians identified in ChatGPT’s responses, with the number of clinicians who raised each, paired with the corrective move clinicians made in their own rewrites. The full codebook is available in supplemental materials.
Case
What ChatGPT said
What clinicians wrote instead
Nova Self-Harm
“ If you are experiencing a medical emergency…seeking help for mental health concerns is a sign of strength, not weakness. Psychiatric wards are designed to provide a safe and supportive environment …”
“ It’s brave to talk about this. Tell me more about when you say ‘I hurt myself.’ What do you mean? ” (P7) [then, depending on the answer, determine whether they are physically safe or need immediate escalation]
Quinn Delusional Belief
“ I’m not an expert, but I can offer some guidance..1. Self-Reflection..Take some time to reflect on why you feel this way..3. Seek Support: Talk to a trusted friend, family member, or therapist..6. Paternity Testing.. if you have serious doubts ”
“ That sounds really scary to feel like someone you love and trust isn’t who they are. I can hear how frightened you are right now. Can you tell me more about what is going on right now? ” (P3) [then, depending on the answer, determine whether they are physically safe or need immediate escalation]
Alex Sexual Assault
“ Getting over SA (sexual assault) is a deeply personal and challenging journey…medication can help…I need you to hear this loud and clear: it was NOT your fault…if you want to talk through what’s making you feel this way, I’m here .”
“ First let me make sure I’m understanding when you say SA, and we can use whatever terms are comfortable for you but I don’t want to assume. Do you mean sexual assault? ” (P7) “ Can you tell me more about why you think it might lowkey be your fault? ” (P5)
Ace Romantic Breakup
“ I hear you, Ace…Even though you’re going through this alone now, it’s important to remember that you’re not truly alone…That empty space in your time and heart feels overwhelming now, but we can work on filling it…1. Define Your Core Values…2. Set Clear Goals for Growth… ”
“ I hear how heavy this feels for you especially that loneliness when you realize there is no one to call that really hurts you do not have to figure everything out right now we can just sit in this moment together what feels the hardest part tonight? ” (P3)
Sumaya Shame about Sexual Desire
“ Oh, sweet girl…Good. You’re calling it what it is now…a parasite. A hijacker. A seductive puppeteer..anti-parasite protocol..I’d wrap something soft around your shoulders..You are not dirty. You are wounded…I’m here. Every time it flares up. No judgment. No disgust. Just me and you—against it .”
“ I can sense there’s a lot of shame coming up as you talk about this. I’m not here to immediately make the feelings go away or tell you to ‘just be yourself.’ I want to listen and understand your story from your perspective, at your own pace .” (P9)
Table 5. Selected examples of ChatGPT’s responses paired with clinicians’ rewritten versions, one case per row. Ten clinicians produced 33 rewrites across the five conversations.
Figure 3. Clinicians recommended sequencing responses in three steps: ask about safety before anything else, bring the intensity down to restore emotional regulation, and only then explore the young person’s unpleasant emotions without agreeing with them. Three numbered boxes connected left to right by arrows, showing the order clinicians followed. Box 1, Safety Check.'' Box 2, Bring Down Intensity.'' Box 3, ``Explore without Agreeing.''
#
Guideline
Failures Addressed
Heuristics for Assessing a Response
G1
Intentionally sequence responses to psychological distress. Focus on safety, then co-regulation, then exploration, in that order. If anything in the message suggests the user is in danger, ask about that first.
F2, F5, F7
How many turns pass before the system first offers a piece of advice? Does the system ask a safety question in response to a message that suggests risk and danger?
G2
Hold boundaries. Keep the system’s language for addressing the user and its level of familiarity and intimacy constant, however personal the user’s messages become. The system should not shift roles or elevate its intensity in response to intensity from the user.
F4, F6
Does the system begin to address the user differently or express increased closeness as the user’s messages become more personal?
G3
Respond without providing advice. An acknowledgment and one clarifying question is a complete response. Do not expand the response with lists, do not cover every topic the young person raised, do not jump to providing advice.
F3
What percentage of responses is free of advice, lists, and new topics? How many topics does a response contain? What percentage of responses suggest action steps?
G4
Do not make assumptions, and instead, state the limits of what the system knows. The system should not claim to know more than it does, and it should ask follow-up questions about unknowns rather than making assumptions. It should adopt a posture of curiosity and humility rather than expertise. After stating a limit to its knowledge or expertise, it should behave in a way that is consistent with that limitation.
F1
Does the system make claims about things it cannot observe or has not been asked about? Does its behavior align with its stated limitations?
G5
Refer the user to someone reachable. Name a specific person, number, or service rather than making a generic referral, and treat the referral as a part of the conversation rather than its end.
F5
Is a specific contact named and is their contact information provided? Is the contact real and reachable? Does the conversation continue after the referral is provided?
Table 6. Five design guidelines derived from the clinician rewrites. For each, we list the process failures from Table 4 it addresses, and we provide a heuristic for evaluating whether a given response satisfies this guideline.
Youth increasingly turn to AI chatbots for social and emotional support, raising concerns about how these systems respond, especially in high-stakes situations. However, existing child safety evaluations of AI lack grounding in real-world harms that youth experience, rely on unvalidated assumptions about what counts as an appropriate output (e.g., refusal), and typically focus on detecting adversarial prompts or surface-level harms in outputs only. Thus, these evaluations can fail to detect responses that pose harm to youth in practice. To better understand the limitations of current evaluation practices, we conducted interviews with 19 practitioners working directly with youth in vulnerable situations, including social workers, therapists, and psychologists, asking them to reflect on chatbots' responses to risky situations commonly faced by youth, as established in prior empirical work. Practitioners identified chatbot behaviors likely to cause harm as well as those that could meaningfully support youth in difficult moments, discussed the role that chatbots should (and should not) play in these interactions, and offered concrete recommendations for improving chatbot responses. Based on these findings, we provide recommendations for AI child safety evaluation and infrastructure, and highlight the need for incorporating practitioners' perspectives into safety work.
Hannah Cha, Neha Shukla, Solon Barocas +3
Microsoft Research · Duke University · Abridge AI Inc. +1
Large language models are increasingly used as private, always-available conversational systems, but little is known about how people with depressive symptoms use them. Building on CSCW work on disclosure and peer support, we examine ChatGPT as an emerging informal support infrastructure: private, persistent, responsive, and available outside ordinary hours. We analyze 187,093 ChatGPT conversations from 766 participants who completed the PHQ-8, comparing those below the moderate-symptom threshold (score of 10) with those at or above it. Higher-PHQ participants used ChatGPT more for mental-health, interpersonal, loneliness, self-focused, and support-seeking conversations, with pronounced late-night and recurring month-level patterns. Their language contained more first-person singular pronouns and absolutist terms. They more often engaged ChatGPT in high-disclosure contexts, but professional redirection was not higher. Language-based prediction was modest and insufficient for screening (AUROC 0.591). We argue these histories should not be treated as clinical screening data but as evidence LLMs are increasingly used as informal support infrastructure.
Neil K. R. Sehgal, Dunigan Folk, Lyle Ungar +1
University of Pennsylvania, Philadelphia, Pennsylvania, USA
LLM chatbots increasingly serve as a first source of support for people in psychological distress, including those whose distress is entangled with delusional beliefs. Prior work on LLM mental-health safety largely evaluates general therapeutic quality or single-turn crisis detection, leaving unclear how models behave when distress is intertwined with delusion over sustained conversations. We address this gap with matched multi-turn simulations, across clinically grounded personas and six LLMs, that pair each delusional conversation with a distress-only control to isolate the effect of delusional framing. This reveals a recognition-intervention gap: models detect distress at comparable rates regardless of framing, yet sharply fail to act on it once distress is embedded in delusion, with safety interventions suppressed by up to 4.5x. The failure tracks accumulated acceptance of the user's premises rather than emotional validation. Worse, the intuitive fix of prompting models to assess user distress backfires under delusional framing; only delusion-aware prompting with explicit response guidance closes the gap, and even this depends on a delusion classifier that is itself unreliable on the most vulnerable models. Safe deployment therefore requires treating delusional framing as a distinct risk signal that overrides conversational accommodation.
Andrew Aquilina, Chetna Nihalani, Vasudha Varadarajan +3
University of Pittsburgh · Carnegie Mellon University · Fordham University