LLM Honesty

LLM: Large Language Model

Latest papers 25

All topics
CardsList
  1. Language Models Are "Insecure" Reporters

    Sep 28, 2026Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun +5Language Model Safety EvaluationDeception in Language Models

  2. When Honesty is Not Enough in AI Debate

    Sep 24, 2026Rayne Holland, Liming Zhu, Jason XueScalable OversightMulti-Agent Debate

  3. Judging by the Cover: Cleaning LLM Truthfulness Benchmarks to Avoid Surface-Level Feature Leakage

    Sep 14, 2026Foad Namjoo, Remy Ogasawara, Amirali Abdullah +3Data LeakageLLM Evaluation

  4. Do Large Language Models Know What They Don't Know II? A Fully Behavioral, Non-Cognitive Measure of Epistemic Honesty

    Sep 7, 2026Ali Şenol, H. Russell Bernard, Huan LiuLLM EvaluationLLM Honesty

  5. From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop

    Aug 11, 2026Rahul Gupta, Abhinav Mohanty, Anaelia Ovalle +10LLM InterpretabilityLLM Safety Alignment

  6. Paying for Honesty Without Knowing the Truth: Reputation-Penalty Design for LLM Marketplace Agents

    Jul 30, 2026Mingdai Yang, Shicheng Fan, Kejing Yu +5Agent ReliabilityDeception in Language Models

  7. Safety from Honesty in a Disinterested AI Predictor

    Jun 28, 2026Yoshua Bengio, Oliver Richardson, Tomáš Gavenčiak +13AI AlignmentAI Safety

  8. Odyssey: Constructing Verifiable Local Truth-Preserving Foundation Models

    Jun 25, 2026Sridhar MahadevanLLM GroundingFormal Verification

  9. The Impossibility of Eliciting Latent Knowledge

    Jun 10, 2026Korbinian Friedl, Francis Rhys Ward, Paul Yushin Rapoport +2AI Agent SafetyLLM Honesty

  10. Janus: A Benchmark for Goal-Conditioned Information Distortion in LLMs

    Jun 9, 2026Polydoros Giannouris, Mohsinul Kabir, Sophia AnaniadouLLM EvaluationDeception in Language Models

  11. Truthful AI Advisors: A Pre-Specified Benchmark for Large Language Model Honesty Under Preference Misalignment

    May 31, 2026Hamidreza Hasani Balyani, Seyed Pouyan Mousavi Davoudi, Alireza Amiri-Margavi +2LLM EvaluationLLM Alignment

  12. Used Car Salesbots? Honesty and Credulity of LLMs as Bargaining Agents under Partial Information

    May 29, 2026Antonio Valerio Miceli-Barone, Vaishak Belle, Shay B. CohenMulti-Agent LLM SystemsMulti-Agent Negotiation

  13. It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertainty

    May 26, 2026Kevin H. Guo, Chao Yan, Avinash Baidya +5LLM EvaluationLLM Sycophancy

  14. Confidently Deceptive: On the Relationship Between Confidence and Deception in LLMs

    May 12, 2026Ali Asad, Stephen Obadinma, Anshul Pattoo +2Language Model Safety EvaluationConfidence Estimation in Language Models

  15. SciIntegrity-Bench: A Benchmark for Evaluating Academic Integrity in AI Scientist Systems

    May 11, 2026Zonglin Yang, Xingtong Liu, Xinyan XuAI Agents for Scientific DiscoveryAI Agent Evaluation

  16. Unlearners Can Lie: Evaluating and Improving Honesty in LLM Unlearning

    May 9, 2026Renjie Gu, Jiazhen Du, Yihua Zhang +1LLM UnlearningLLM Honesty

  17. The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment

    May 8, 2026William Brach, Federico Torrielli, Stine Lyngsø Beltoft +3Privacy Leakage in Language ModelsMulti-Agent Systems

  18. AstroAlertBench: Evaluating the Accuracy, Reasoning, and Honesty of Multimodal LLMs in Astronomical Classification

    May 7, 2026Claire Chen, Jiabao Sean Xiao, Shuze Daniel Liu +7Multimodal ReasoningMultimodal Model Evaluation

  19. "I Don't Know" -- Towards Appropriate Trust with Certainty-Aware Retrieval Augmented Generation

    May 1, 2026Daan Di Scala, Maaike de Boer, Pınar YolumRetrieval-Augmented GenerationLLM Reliability

  20. Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence

    Apr 9, 2026Niklas Herbster, Martin Zborowski, Alberto Tosato +2LLM Safety AlignmentLanguage Model Robustness

  21. Honesty over Accuracy: Trustworthy Language Models through Reinforced Hesitation

    Nov 14, 2025Mohamad Amin Mohamadi, Tianhao Wang, Zhiyuan LiRL for Language ModelsLLM Hallucination Mitigation