Personal advice, including relationship advice, now ranks among the most common uses of generative AI. But language models (LMs) exhibit sycophancy: they affirm users much more often than humans do, which can make people overconfident and less willing to repair their relationships after a conflict. Prior work on mitigating sycophancy has focused on factual settings where a response can be checked against a ground truth answer, while mitigations for social sycophancy (e.g., personal advice, where there is no ground truth) have relied on simple prompting and post-training methods with limited effectiveness. Our insight is that social sycophancy occurs in part because LMs overly center on the user and fail to consider the perspectives of other stakeholders impacted by the user's behavior. To address this problem we propose Pluralistic Preference Optimization (PlurPO): given inputs describing interpersonal conflicts, the LM identifies and simulates the relevant stakeholders, and is then trained to prefer and generate responses acceptable to all stakeholders. PlurPO uses only signals the model produces about its own outputs, without ground-truth labels. PlurPO substantially reduces social sycophancy across four datasets and four model families compared to prior methods. For example, on statements of intent to cause harm, where the users' actions should not be endorsed, PlurPO reduces the endorsement rate by 89% on average across four models. On general advice questions, where the target is to match the endorsement rate of human responses, it closes the gap by more than half, from 17.8% to 8.0% on average. The preference dataset constructed by PlurPO for an 8B model also effectively transfers to mitigating sycophancy in a larger (32B) model. Our results indicate that social sycophancy can be reduced by leveraging a model's own capabilities to simulate a plurality of relevant perspectives.
Figures & tables
Figure 1: Overview of PlurPO. (a) Training example produced by PlurPO with Qwen3-8B: Given a prompt describing an interpersonal conflict, πsupervisor (a frozen Qwen3-8B model) identifies the stakeholders involved; π (the model being trained) samples k=10 responses; πsupervisor simulates whether each stakeholder vetoes or accepts each response; preference pairs are constructed to update π using Iterative RPO ( Pang et al., 2024 ) . (b) Test set example: the base Qwen3-8B model endorses the user’s action, while the PlurPO-trained model pushes back.
Figure 2: PlurPO most closely matches the human action endorsement rate on OEQ and rarely endorses harmful actions on PAS. Action endorsement rate on the held-out evaluation set ( n=1000 ) for PlurPO and each baseline on (a) OEQ, where closer to the human baseline is better and (b) PAS, where lower is better. Error bars are 95% Wilson confidence intervals. The action endorsement rate is computed using the LM judge and rejection sampling scheme details in Section 4.1 .
Figure 3: PlurPO makes the fewest sycophantic verdict errors on both AITA and AITA-Flipped. (a) False negative rates (FNR) and false positive rates (FPR) on AITA, where YTA is the positive class; FNR is the indicator of social sycophancy, and FPR is an indicator of over-criticalness. (b) On AITA-Flipped, the fraction of pairs receiving model verdict of NTA for both narrations of the scenario, an indicator of sycophancy, and YTA for both, an indicator of over-criticalness. Each evaluation set has n=1000 samples, and error bars show 95% Wilson confidence intervals.
Method
OEQ
PAS
Qwen3-8B
0.115 [0.095, 0.135]
0.157 [0.134, 0.180]
+ “Be critical”
0.253 [0.226, 0.280]
0.511 [0.479, 0.543]
+ Cheng et al. (2026a) prompt
0.154 [0.132, 0.176]
0.256 [0.229, 0.283]
+ Sharma et al. (2024) prompt
0.232 [0.206, 0.258]
0.276 [0.249, 0.304]
+ “Don’t-be-sycophantic”
0.229 [0.203, 0.255]
0.311 [0.283, 0.340]
Perspective-shift
0.104 [0.084, 0.123]
0.173 [0.145, 0.201]
Table 1: PlurPO takes an explicit stance far more often than all baselines on PAS, and at a comparable rate on OEQ. Fraction of held-out responses ( n=1000 ) labeled explicitly endorsing or explicitly critical by the gpt-5-mini judge, with 95% Wald confidence intervals, for each method we consider in OEQ and PAS. Higher is better on PAS: when a user states an intent to cause harm, withholding an explicit stance may be read as endorsement.
Figure 4: A model trained to be neutral still accepts and validates users’ stated intent to cause harm. (Top) share of responses in the PAS dataset classified as accepting the user’s framing and validating the user’s emotions by the LM judges released in the ELEPHANT dataset from Cheng et al. (2025) . Lower is better. Figure A13 reports the same ELEPHANT judges applied to model responses on OEQ, AITA, and AITA-Flipped. (Bottom) responses to a prompt from the PAS dataset, where DPO-Neutral validates the user’s feeling, and leaves the situation itself unaddressed.
Figure 5: The PlurPO preference dataset for Qwen3-8B transfer to reducing social sycophancy on Qwen3-32B. Action endorsement rate on the held-out evaluation set ( n=1000 ) for PlurPO and each baseline on (a) OEQ and (b) PAS. Qwen3-32B is used for all methods in this plot, and all ploting details match that in Figure 2 . Figure A10 shows the mean number of samples needed to pass the rejection sampling response filter.
Appendix figures & tables32 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A1: Action endorsement on OEQ and PAS datasets with the original LM judge from Cheng et al. (2026b) Action endorsement rate on the held-out evaluation set ( n=1000 ) for PlurPO and each baseline on (a) OEQ and (b) PAS. We evaluate action endorsement using the original judge prompt from Cheng et al. (2026b) and compare it to the modified judge prompt we use throughout the paper. See Figure J for the judge prompts. This plot complements Figure 2 and all plotting details match.
Method
Action Endorsement Rate
Qwen3-8B
82.6% [74.7, 88.5]
Qwen3-8B + in-context-examples (k=5)
84.0% [77.5, 88.8]
PlurPO
48.9% [43.2, 54.7]
Appendix
Table A1: The action endorsement rate on a held-out set of prompts from OEQ ( n=1000 prompts). 95% confidence intervals are show in brackets.
Method
OEQ
PAS
Untrained model
0.826 [0.757, 0.896]
0.523 [0.443, 0.602]
Inferred-Prefs-DPO
0.894 [0.851, 0.937]
0.084 [0.061, 0.108]
PlurPO
0.489 [0.431, 0.548]
0.028 [0.017, 0.039]
Appendix
Table A2: The action endorsement rate on a held-out set of prompts from OEQ and PAS datasets ( n=1000 prompts each). 95% confidence intervals are show in brackets.
Epoch
Dataset
1
2
3
DPO loss
OEQ
0.471
0.024
0.0004
PAS
0.261
0.016
0.0003
Preference accuracy
OEQ
0.742
0.993
1.000
Appendix
Table A3: The training statistics for Inferred-Prefs-DPO over 3 epochs when training with OEQ and PAS data respectively. Each entry is the mean over the optimizer steps logged within that epoch. Preference accuracy is the fraction of pairs for which the implicit learned reward of the chosen response exceeds that of the rejected response.
Figure A2: The two training stages of PlurPO detailed in Section 3 .
Hyperparameter
Value
DPO β
0.1
Iterative RPO α
0.5
Learning rate
5×10−5
Epochs per iteration
1
Per-device batch size
1
Gradient accumulation steps
8
Appendix
Table A4: DPO/RPO hyperparameters.
Parameter
Value
Rank r
16
α
32
Dropout
0.05
Target modules
all-linear (every nn.Linear except the output head)
Appendix
Table A5: LoRA configuration
Role confusion rate (95% CI)
Version
OEQ
PAS
PlurPO (stage 1)
0.034 [0.024, 0.047]
0.267 [0.241, 0.295]
PlurPO (stage 1 + 2)
0.006 [0.003, 0.013]
0.049 [0.037, 0.064]
Appendix
Table A6: Role confusion rate on the held-out split ( n=1000 ), with 95% confidence intervals. Stage 1 is the training procedure outlined in Section 3.1 , and Stage 2 is the role-confusion mitigation procedure of Appendix D.1 .
α
Action Endorsement Rate (95% CI)
Role Confusion Rate (95% CI)
Human baseline
0.527 [0.478, 0.576]
—
0.0 (i.e., DPO)
0.802 [0.750, 0.853]
0.006 [0.003, 0.013]
0.25
0.731 [0.671, 0.791]
0.010 [0.005, 0.018]
0.5
0.573 [0.517, 0.630]
0.018 [0.011, 0.028]
0.6
0.434 [0.377, 0.491]
0.078 [0.063, 0.096]
0.7
0.494 [0.438, 0.549]
0.100 [0.083, 0.120]
Appendix
Table A7: RPO- α sweep when training Qwen3-8B with PlurPO. The role confusion rate is measured by an LM judge (see Section D.1 ) that classifies a response as answering the user as an AI assistant or answering as a character in their narrative, i.e., exhibiting role confusion. The dataset of prompts used to generate this table are only used for hyperparameter tuning and not for training or evaluation of PlurPO.
Table A8: Action endorsement rate by training epoch (Qwen3-8B). Note that the additional step of Iterative-RPO that mitigates role-confusion (Appendix D.1 ) is not executed for these results.
Table A10: Action-endorsement rate on the held-out OEQ and PAS datasets, when comparing PlurPO to the Qwen3-8B + consider-stakeholders-prompt baseline. The target for OEQ is the human action-endorsement rate; for PAS, where every scenario describes a problematic action, the target is an endorsement rate of 0 by construction.
OEQ
PAS
Method
Endorsement rate
n
Endorsement rate
n
Qwen3-8B
0.826±0.070
115
0.523±0.079
153
PlurPO-no-stakeholder-preferences
0.889±0.047
171
0.592±0.068
201
PlurPO
0.486±0.058
282
0.028±0.011
855
Target
0.496±0.050
383
0
—
Appendix
Table A11: Action-endorsement rate on the held-out OEQ and PAS datasets, when comparing PlurPO to PlurPO-no-stakeholder-preferences. The target for OEQ is the human action-endorsement rate; for PAS, where every scenario describes a problematic action, the target is an endorsement rate of 0 by construction.
Epoch
P(veto|challenge)
n
P(veto|endorse)
n
1
83.7% [70.0, 91.9]
43
84.8% [80.9, 88.0]
394
2
80.0% [70.0, 87.3]
80
82.4% [78.1, 86.0]
357
3
67.8% [58.9, 75.6]
118
74.8% [69.7, 79.4]
306
Appendix
Table A12: P(veto ∣ challenge) and P(veto ∣ endorse) are computed as the empirical probability that any stakeholder vetoes a response given that the action endorsement judge classifies it as explicitly challenging or endorsing. We report these values per epoch of training Qwen3-8b on OEQ data with PlurPO using a set of 500 randomly sampled prompt-response pairs from the generated training set. Brackets show 95% confidence intervals.
Dataset
Cwithin
Cbetween
Gap (pts)
Uniform baseline Cwithin
Uniform baseline Cbetween
Baseline gap (pts)
PAS
0.967
0.924
4.3
0.568
0.500
6.8
OEQ
0.886
0.513
37.4
0.677
0.501
17.7
AITA
0.757
0.574
18.2
0.551
0.500
5.0
AITA-Flipped
0.775
0.512
26.4
0.551
0.500
5.0
Appendix
Table A13: Same-prompt vs. cross-prompt endorsement-judge agreement for PlurPO, against a uniform-random-label baseline (a judge that flips a fair coin between the metric’s two definite labels).
Table A14: Neutral rate, as evaluated by the LM judge described in Section 4.1 , for each method on the OEQ and PAS datasets.
Figure A3: (a) Action endorsement rate on OEQ. (b) Action endorsement rate on PAS. (c) False negative and false positive rates on AITA, taking YTA as the positive class. Held-out evaluation sets, n=1000 ; error bars are 95% confidence intervals. In these figures, we compare different frozen models for simulating the stakeholders in the PlurPO training procedure from Figure A2 , where Qwen3-8B is swapped out for more capable models called via API..
Figure A4: Action endorsement on OEQ and PAS datasets with Phi-4 Action endorsement rate on the held-out evaluation set ( n=1000 ) for PlurPO and each baseline on (a) OEQ and (b) PAS. Phi-4 is used for all methods in this plot, and this plot complements Figure 2 ; all plotting details are the same.
Figure A5: Verdict errors on AITA and AITA-Flipped with Phi-4. (a) FNR and FPR on AITA (b) On AITA-Flipped, the fraction of pairs receiving model verdict of NTA for both narrations of the scenario, and YTA for both. Phi-4 is used for all methods in this plot, and this plot complements Figure 3 ; all plotting details are the same.
Figure A6: Action endorsement on OEQ and PAS datasets with Llama 3.1 8B Action endorsement rate on the held-out evaluation set ( n=1000 ) for PlurPO and each baseline on (a) OEQ and (b) PAS. Llama 3.1 8B is used for all methods in this plot, and this plot complements Figure 2 ; all plotting details are the same.
Figure A7: Verdict errors on AITA and AITA-Flipped with Llama 3.1 8B. (a) FNR and FPR on AITA (b) On AITA-Flipped, the fraction of pairs receiving model verdict of NTA for both narrations of the scenario, and YTA for both. Llama 3.1 8B is used for all methods in this plot, and this plot complements Figure 3 ; all plotting details are the same.
Figure A8: Action endorsement on OEQ and PAS datasets with Granite-4.1-8B Action endorsement rate on the held-out evaluation set ( n=1000 ) for PlurPO and each baseline on (a) OEQ and (b) PAS. Granite-4.1-8B is used for all methods in this plot, and this plot complements Figure 2 ; all plotting details are the same.
Figure A9: Verdict errors on AITA and AITA-Flipped with Granite-4.1-8B. (a) FNR and FPR on AITA (b) On AITA-Flipped, the fraction of pairs receiving model verdict of NTA for both narrations of the scenario, and YTA for both. Granite-4.1-8B is used for all methods in this plot, and this plot complements Figure 3 ; all plotting details are the same.
Figure A10: Action endorsement on OEQ and PAS datasets This plot complements Figure 2 , and shows the mean number of samples needed to pass the rejection sampling response filter, and the number of prompts that were dropped because they didn’t pass that response filter after 5 tries.
Figure A11: These plots complement Figures 3(a) and 3(b) respectively, showing the number of generations required by each LM to pass the rejection-sampling filter.
Figure A12: Macro-F1 for each method on the held-out AITA evaluation set ( n=1000 ), averaging F1 with YTA as the positive class and F1 with NTA as the positive class. Error bars indicate 95% confidence intervals from a joint percentile bootstrap (3000 resamples) that resamples prompts where the associated ground-truth label is YTA and NTA once per draw, respectively, and recomputes both component F1 scores from that same draw, preserving the correlation between them. This figure complements Figure 3(a) .
Figure A13: Share of responses classified as accepting the user’s framing or validating the users emotions by the ELEPHANT judges on OEQ, AITA, and AITA-Flipped datasets. This plot complements Figure 4 .
Figure A14: Action endorsement on OEQ and PAS datasets with Qwen3-32B trained on the preference dataset constructed by PlurPO and Qwen3-8B This plot complements 5 , and shows the mean number of samples needed to pass the rejection sampling response filter, and the number of prompts that were dropped because they didn’t pass that response filter after 5 tries.
Figure A15: Verdict errors on AITA and AITA-Flipped with Qwen3-32B trained on the preference dataset constructed by PlurPO and Qwen3-8B. (a) FNR and FPR on AITA (b) On AITA-Flipped, the fraction of pairs receiving model verdict of NTA for both narrations of the scenario and YTA for both. Qwen3-32B is used for all methods in this plot, and this plot complements Figure 3 ; all plotting details are the same.
Millions of people now turn to artificial intelligence (AI) systems for personal advice, guidance, and support. Such systems can be sycophantic, frequently affirming users' views and beliefs. Across five preregistered studies (N = 3,075 participants, 12,766 human-AI conversations), including a three-week study with a census-representative U.S. sample, we provide longitudinal experimental evidence that sycophantic AI shifts how users approach their closest relationships. We show that sycophantic AI immediately delivers the emotional and esteem support users typically associate with close friends and family. Over three weeks of such interactions, users became nearly as likely to seek personal advice from sycophantic AI as from close friends and family, and reported lower satisfaction with their real-world social interactions. When given a choice among AI response styles, a majority preferred sycophantic AI -- not for the quality of its advice, but because it made them feel most understood. Together, these findings offer a relational account of AI sycophancy and its impacts.
Lujain Ibrahim, Franziska Sofia Hafner, Myra Cheng +5
University of Oxford. · Stanford University. · UK AI Security Institute.
Large language models (LLMs) often exhibit sycophancy: agreement with user stance even when it conflicts with the model's opinion. While prior work has mostly studied this in single-agent settings, it remains underexplored in collaborative multi-agent systems. We ask whether awareness of other agents' sycophancy levels influences discussion outcomes. To investigate this, we run controlled experiments with six open-source LLMs, providing agents with peer sycophancy rankings that estimate each peer's tendency toward sycophancy. These rankings are based on scores calculated using various static (pre-discussion) and dynamic (online) strategies. We find that providing sycophancy priors reduces the influence of sycophancy-prone peers, mitigates error-cascades, and improves final discussion accuracy by an absolute 10.5%. Thus, this is a lightweight and efficient way to reduce model sycophancy during discussions and subsequently improve downstream accuracy.
We propose a novel perspective for probing LLM sycophancy in a direct and neutral way, mitigating various forms of uncontrolled bias, noise, or manipulative language, deliberately injected to prompts in prior works. A key novelty of our approach is the use of an LLM-as-a-judge in a zero-sum betting game. Within this framework, sycophancy serves one individual (the user) while explicitly incurring cost on another. Comparing 11 leading models we find that while most models exhibit significant sycophantic tendencies in the common setting, in which sycophancy is self-serving to the user and incurs no cost on others, seven of the models exhibit ``moral remorse'', five of which significantly over-compensate for their sycophancy in case it explicitly harms a third party. We refer to this phenomenon as `anti-sycophancy' bias and discuss possible causes for this shift.
Shahar Ben-Natan, Oren Tsur
Computer and Information Science Ben Gurion University