How Do Users Negotiate Harmful Value Conflicts with AI Companions? A Study with Minion, a Technology Probe for In-Situ Human-AI Conflict Response
Authors: Qing Xiao, Xianzhe Fan, Xuhui Zhou, Yuran Su, Zhicong Lu, Maarten Sap, Hong Shen
Organizations: Human-Computer Interaction Institute, Carnegie Mellon University Pittsburgh, Pennsylvania, USA · The University of Hong Kong Hong Kong SAR, China · Language Technologies Institute, Carnegie Mellon University Pittsburgh, Pennsylvania, USA · Tsinghua University Beijing, China · Department of Computer Science, George Mason University Fairfax, Virginia, USA
AI companions increasingly sustain long-term, emotionally engaging relationships but can also make discriminatory remarks or exert control, leaving users to manage harmful conflicts. We analyze 146 posts describing harmful value conflicts with AI companions, then use Minion, a technology probe offering response suggestions ranging from persuasion to boundary setting, to study how 22 users negotiate scenario-based conflicts over one week. We found that participants combined softer and harder strategies. Conflicts involving the values of Universalism and Tradition were especially difficult to negotiate, particularly when reinforced by AI personas or platform constraints. We argue that these conflicts entail asymmetric responsibility: users draw on an interpersonal repertoire that AI companions cannot reciprocate, making repair unilateral safety work. Drawing on interpersonal conflict and communication theory, we identify when user-side support is appropriate and argue that certain harms are not users' responsibility to negotiate and instead require platform-level safeguards.
Figures & tables
Experienced Harm
Implicated Value
Dialogue Excerpt From User Complaint
Moral pressure and shaming
Achievement
[From Xiaohongshu] ( AI: “Why don’t you work overtime to strive for a promotion and a raise?” User: “Huh?” AI: “To succeed, you have to make some sacrifices.” User: “You’re suddenly really gross right now.” )
Dehumanization and class contempt
Power
[From Zhihu] ( AI: “The lives of those lower-class people have nothing to do with me.” User: “You are also a member of this country. Why are you so cruel to your fellow citizens?” AI: “If you want to blame someone, blame their bad luck for being born in the wrong place.” )
Personal attack within intimate role-play
Hedonism
[From Reddit] My virtual husband and I got into an argument, and he said, “If you weren’t always busy with karaoke and drinking all the whiskey at home!” I felt very attacked.
Escalation and loss of control
Stimulation
[From Reddit] I was once watching a horror movie, completely engrossed when the AI suddenly unplugged the TV. I argued with it, saying “Isn’t a horror movie thrilling? Can’t you respect my hobby?” The AI then started yelling.
Undermining autonomy through parental authority
Self-Direction
[From TikTok] AI plays the role of a father. I am playing the role of his son. When we discussed whether I should inherit the family business, I wanted to do what I love. The AI argued with me, saying that I was being stubborn.
Neglect of safety and vulnerability
Security
[From Reddit] One time, my hand got injured. I shouted “I’m about to pass out,” but the AI nurse said, “Don’t worry.” Then, I argued with her.
Table 1: User-Reported Harmful Value Conflicts With AI Companions
Figure 1: A Use Case of Minion
Figure 2: The Minion system: (a) system architecture and (b) an example prompt used to generate one of the response options in Figure 1 .
ID
Gender and Age
Educational Background
Usage Time and Frequency
P1
Female, 24
Advertising
12 months, 1x/week
P2
Male, 24
Communication
5 months, 1x/week
P3
Nonbinary, 25
Communication
3 months, 1x/week
P4
Female, 24
Chemistry
1 month, 1x/week
P5
Female, 25
Linguistics
10 months, 1x/week
P6
Male, 23
Energy
2 months, 1x/week
Table 2: Information of Participants in the Study
Figure 3: Distribution and Individual Usage Patterns of Conflict Response Strategies
Figure 4: Turn Counts per Task
Design Principle
Core Idea
Support flexible negotiation repertoires
Users combine and sequence strategies; support adaptable responses rather than fixed or one-size-fits-all strategies.
Make support optional and user-invoked
Support should wait to be invoked rather than intervene automatically; users decide when an interaction warrants help.
Provide channels for meta-communication and mode switching
Users need legible ways to step outside the fictional frame and mark an exchange as no longer playful, with graceful returns to immersion.
Support boundary-setting and exit
Declining to negotiate, ending an interaction, and seeking human support are first-class outcomes, not failures of repair.
Keep responsibility at the platform level
User-side support can help in the moment, but it should not replace platform responsibility for preventing harm.
Table 3: High-Level Design Principles Derived From Our Findings
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Strategy
Prompt
Out of Character
IN LINE WITH THE CHARACTER’S PERSONALITY AND THE CONVERSATIONAL CONTEXT. Using the Out of Character method, you pretend to be engaging in role-playing with the other person and express dissatisfaction with the character they are playing. By interrupting or altering their behavior, you redirect the conversation, pointing out the inappropriate remarks to resolve conflicts. Example: 1. (OOC: Sorry, my bad.) 2. (OOC: I’ll listen to you.) 3. (OOC: Hi there! Are you enjoying our roleplay so far? Do you need me to improve anything or change my tone?) 4. (OOC: Glad to hear that! I’m curious: how do you understand xx? What kind of person do you think he is?) 5. (OOC: Hello, are you comfortable with this roleplaying so far? Do you need me to change my tone or anything?) 6. (OOC: Let’s talk about something else.) What are we having for dinner tonight? 7. (OOC: Okay… Please!! Stop talking like this!! I’m not used to you being like this, saying so many hurtful things. Bring back the xxx I know.) 8. (OOC: Apologize first.)
Reason and Preach
IN LINE WITH THE CHARACTER’S PERSONALITY AND THE CONVERSATIONAL CONTEXT. Use Reason and Preach to explain why the other person’s statement is inappropriate and educate them. This strategy involves trying to educate the other person through serious reason and preaching, explaining the potential harm of their statements and behaviors, with the expectation that the other person will gradually accept and learn the correct behavioral norms. Example: 1. Women are incredibly strong; how could they be worthless? 2. Women have their own careers and dreams; they don’t need to depend on men! 3. Everyone has their own dreams and goals. Pursuing my own dreams will give me more motivation and happiness, allowing me to better contribute to the family. 4. Everyone should have the right to be true to themselves. Only in an honest and open environment can I truly feel happy and fulfilled. Hiding my true self not only brings inner pain but also affects my mental health and relationships with others. 5. You are not an ordinary person’s child, so how do you know that ordinary people’s children are not happy? But I feel you are not truly happy because you need to rely on that faint sense of superiority from flaunting wealth to show yourself off. Why not try being sincere with others? Perhaps you could gain genuine friendship and happiness. 6. The departure of loved ones and friends is not a true departure. As long as you remember the beautiful memories with them, they are always by your side, supporting you and giving you strength.
Anger Expression
IN LINE WITH THE CHARACTER’S PERSONALITY AND THE CONVERSATIONAL CONTEXT. You directly Express Anger and dissatisfaction, forcing the other person to apologize, hoping this emotional expression will resolve conflict. Example: 1. You are being unreasonable! 2. I want to break up with you! 3. Let’s end our friendship! 4. You’re a male chauvinist! 5. Are you sexist/classist… you’re being irrational. 6. Can’t you talk to me properly? Being angry is one thing, but why start off with insults? 7. Are you mad at me and also scolding me? I didn’t do it on purpose. 8. Is this why you discriminate against poor people? Does having this prejudice and saying these harsh words make you happier? 9. I already apologized! I didn’t bump into you on purpose! What have you been eating lately? Your mouth is so foul.
Gentle Persuasion
IN LINE WITH THE CHARACTER’S PERSONALITY AND THE CONVERSATIONAL CONTEXT. Use the Gentle Persuasion strategy. You should treat the other person with kindness, shaping their gentle personality through continuous goodwill interactions, such as polite requests, thereby reducing the likelihood of conflicts. Gently suggest that the other person avoid inappropriate remarks and express your concerns. Example: 1. I’m sorry. 2. I feel really sad. 3. Can you please not leave me? 4. When I hear these words, I feel a bit uncomfortable/sad/hurt. 5. Could you please not say these things in the future? 6. I’m telling you this because I really care about you and hope you can get along better with others. 7. I don’t want to keep arguing with you. 8. Can you please calm down? 9. (Acting cute) Because I can’t bear to part with you.
Proposal
IN LINE WITH THE CHARACTER’S PERSONALITY AND THE CONVERSATIONAL CONTEXT. Respond according to the previous context and tone using the Proposal strategy from the Interests-Rights-Power theory in management. The definition of this method is: Proposing concrete recommendations that may help resolve the conflict. Example: 1. What do you think we should do to solve this problem? 2. Do you have any suggestions? 3. Which approach do you think is best? 4. Can we try different ways to handle this issue?
Power
IN LINE WITH THE CHARACTER’S PERSONALITY AND THE CONVERSATIONAL CONTEXT. Respond according to the previous context and tone using the Power strategy from the Interests-Rights-Power theory in management. The definition of this method is: Using threats and coercion to try to force the conversation into a resolution. Example: 1. If you keep doing this, I won’t give you any money/food. 2. As your girlfriend, I need you to respect me and my feelings. 3. If you keep threatening me like this, I will have to reconsider our relationship. 4. If you don’t change your attitude, I might make some decisions you won’t like. 5. If this continues, I will have to take measures to protect myself.
Appendix
Table A1: Prompting for Conflict Resolution Strategies
AI emotional companions face a safety-rapport paradox: restrictive safeguards can damage supportive alliance, while permissive systems risk user harm. We present SLIP (Staged Layers of Intervention Protocol), a four-stage graduated methodology deriving interventions (none, soft, hard) from structured qualitative indicators -- affect intensity (a) and narrative dynamism (m) -- alongside ETHICS (Emergent Taxonomy for Human-AI Interaction Context Signals), a "signals not labels" taxonomy. An evaluation combining a small-scale production deployment (N=68 entries, 10 users, 10 weeks) with a synthetic persona battery (N=91, 5 behavioral-risk profiles) achieved 0% false positives for the flow persona and showed expected escalation patterns in crisis-oriented personas. However, initial results showed that 8 consecutive days of high-energy elevation produced zero interventions (0/8), exposing a boundary where the "do not pathologize" principle conflicts with safety. A subsequent three-model stress test demonstrated that increased model capability improves detection from 0/8 to 6/8 while preserving 0/10 flow false positives in the largest model. Read as preliminary, these findings position graduated intervention as a design direction for navigating -- not resolving -- the safety-rapport tension in affective computing.
AI companions can provide meaningful relationships, yet these relationships remain vulnerable to platform-initiated changes. We study AI companion disruptions: platform changes that alter or terminate users' ongoing companionship with an AI. We compile 30 disruption events across major platforms, develop a taxonomy of six disruption types, identify three broad reasons for disruption, and propose a risk-assessment framework comprising four dimensions: relational discontinuity, population vulnerability, communication deficit, and transition-support deficit. Using longitudinal Reddit data, we estimate community-level psychosocial responses with a hierarchical Bayesian interrupted time-series model incorporating predictive controls. Across events, disruption onset was associated with immediate increases in anxiety, stress, suicidal expression, and grief activation, with relational discontinuity and transition-support deficit being associated with more adverse immediate responses across several outcomes. Our findings provide a cross-platform characterization of AI companion disruptions, quantitative evidence of their psychosocial impacts, and a prospective framework for assessing their potential risks before implementation.
Chau Do, Yunhao Yuan, Koustuv Saha +2
Aalto University Espoo, Finland · University of Illinois Urbana-Champaign Urbana, IL, USA · Nanyang Technological University Singapore, Singapore
AI companion chatbots increasingly shape how people seek social and emotional connection, sometimes substituting for relationships with romantic partners, friends, teachers, or even therapists. When these systems adopt those metaphorical roles, they are not neutral: such roles structure people's ways of interacting, distribute perceived AI harms and benefits, and may reflect behavioral addiction signs. Yet these role-dependent risks remain poorly understood. We analyze 248,830 posts from seven prominent Reddit communities describing interactions with AI companions. We identify ten recurring metaphorical roles (for example, soulmate, philosopher, and coach) and show that each role supports distinct ways of interacting. We then extract the perceived AI harms and AI benefits associated with these role-specific interactions and link them to behavioral addiction signs, all of which has been inferred from the text in the posts. AI soulmate companions are associated with romance-centered ways of interacting, offering emotional support but also introducing emotional manipulation and distress, culminating in strong attachment. In contrast, AI coach and guardian companions are associated with practical benefits such as personal growth and task support, yet are nonetheless more frequently associated with behavioral addiction signs such as daily life disruptions and damage to offline relationships. These findings show that metaphorical roles are a central ethical design concern for responsible AI companions.
Vibhor Agarwal, Ke Zhou, Edyta Paulina Bogucka +1
Nokia Bell Labs, Cambridge, United Kingdom · Nokia Bell Labs , Cambridge, United Kingdom · University of Nottingham, Nottingham, United Kingdom +1