LLM Sycophancy

LLM: Large Language Model

Momentum

13 papers in the last four weeks, up 160% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 96

All topics
CardsList
  1. BeliefScope: Diagnosing Evidence-Driven Revision and Pressure-Induced Shifts in Large Language Models

    Oct 8, 2026Shuai Guo, Yidong CuiBelief Updating in LLMsLLM Auditing

  2. Memadapter: Counterfactual Adaptation Against Memory-induced Sycophancy

    Oct 4, 2026Ruqing Ning, Haibo Meng, Zhishang Xiang +4LLM SycophancyLLM Agent Memory

  3. Mitigating Social Sycophancy via Pluralistic Preference Optimization

    Oct 1, 2026Stephane Hatgis-Kessell, Myra Cheng, Xiaoxuan Hou +4LLM SycophancyPreference Optimization

  4. FIGS: Evaluating Multi-Turn Sycophancy Without Penalizing Empathy

    Sep 30, 2026Sidharth Pulipaka, Ruta Binkyte, Ivaxi Sheth +1LLM SycophancyMulti-Turn Dialogue Evaluation

  5. Less Sycophancy, Stronger Refusal? Lessons for AI Safety from Mechanistic Interpretability

    Sep 28, 2026Xu Wang, Difan Zou, Xuansheng WuLLM SycophancyLanguage Model Safety Evaluation

  6. Receptiveness, Not Sycophancy: Distinguishing Engagement from Deference in Language Models

    Sep 22, 2026Calvin Isley, Johann Gaebler, Max Lamparth +2LLM Sycophancy

  7. XYEval: Agents say yes to bad advice

    Sep 20, 2026Zhengxuan Wu, Yuxuan Li, Oyvind Tafjord +1LLM SycophancyAI Agent Evaluation

  8. Measuring LLM Sycophancy under Sustained Multi-Turn Pressure

    Sep 8, 2026Leyuan Tang, Kangda Wei, Tianyu Jiang +1LLM Safety BenchmarksLLM Evaluation

  9. How AI Models Manage Epistemic Authority: A Taxonomy and Comparative Analysis of Responses to User Disagreement

    Sep 7, 2026Riyadh Alnasser, Yusuf Mücahit Çetinkaya, Sumin Zhao +1LLM EvaluationLLM Sycophancy

  10. Sycophantic Agreement Transfers with Neutral Data via Contrastive Preference Optimization

    Aug 31, 2026Camila Blank, Zhuofan Ying, Christopher Potts +2LLM SycophancyEmergent Misalignment in Language Models

  11. WildSEEK: Evaluating Language Models for Information-Seeking

    Aug 31, 2026Tanise Ceron, Joachim Baumann, Elisa Bassignana +3LLM Safety BenchmarksLLM Sycophancy

  12. Group Alignment-Induced Sycophancy: A Two-Sided Evaluation of Steerable Pluralistic Alignment

    Aug 12, 2026Haokai Zhao, Yunze Xiao, Weihao Xuan +3LLM SycophancyLLM Alignment

  13. FaithformBench: Benchmarking Faithfulness of Mathematical Chain-of-Thought Autoformalisation

    Aug 11, 2026Rob Cornish, Iacopo Ghinassi, Po-Hung Yeh +7CoT FaithfulnessLLM Sycophancy

  14. Mitigating Over-Personalization in LLMs via Structured Memory

    Aug 8, 2026Hakeem Hannoon, Andrew Zhao, Mihir Narayan +2LLM SycophancyPersistent Memory for Language Models

  15. Confirming Our Biases? Evaluating the Capabilities, Risks, and Societal Impact of Large Language Models

    Aug 7, 2026Mudar Adas, Polina Tsvilodub, Michael Franke +1Prompt SensitivityLLM Sycophancy

  16. Poli-Bias: Understanding and Measuring Large Language Model Biases in International Political Conflicts

    Aug 6, 2026Massi-Nissa Abboud, Aladin Djuhera, Elena Cabrio +1LLM SycophancyLLM Auditing

  17. Measuring and Detecting Harmful AI Sycophancy

    Aug 6, 2026Bohan Jiang, Dawei Li, Yasin Silva +1LLM SycophancyLanguage Model Safety Evaluation

  18. Language Models Encode the Contextual Truth of Propositions

    Aug 4, 2026Rupak Sarkar, Pritika Ramu, Rachel RudingerLLM SycophancyKnowledge Conflicts in Language Models

  19. MedPRESS: A Multi-turn Benchmark for Patient-Pressure-Induced Medical Sycophancy in LLMs

    Aug 3, 2026Saman Sarker Joy, Niloy FarhanLLM Safety BenchmarksLLM Sycophancy

  20. Why LLMs Give In: Conversational Factors and Reasoning Behind Medical Sycophancy

    Aug 2, 2026Kaike Ping, Buse Çarık, Caleb Wohn +3LLM SycophancyLanguage Modeling

  21. Token-Level Diagnosis of Sycophancy in LLMs with Attribution-Guided Steering

    Jul 31, 2026Hieu Nguyen, Mahammed Kamruzzaman, Anshuman Chhabra +1LLM SycophancyLLM Interpretability

  22. Talked Out of the Truth: Sycophancy in the Reasoning Chains of Multimodal Models

    Jul 30, 2026Mahir Numayeer Islam, Gakuto Okuyama, Nikolaus Siauw +3LLM SycophancyMultimodal CoT Reasoning

  23. Individual-level interventions against sycophantic AI reduce its appeal but not its persuasiveness

    Jul 28, 2026Meryl Ye, Robert Kraut, Steve RathjeLLM SycophancyHuman-AI Interaction