Honesty

Momentum

8 papers in the last four weeks, against 2 the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 29

All topics
CardsList
  1. Language Models Are "Insecure" Reporters

    Sep 28, 2026Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun +5HonestyHeadlines

  2. Byzantine-Robust Federated RAG via Aligned Calibration and Fixed-Membership Conformal Prediction

    Sep 27, 2026Prasanjit Dubey, Aritra Guha, Xiaoming HuoOnline Conformal PredictionHonesty

  3. How does Adversarial Influence Scale in Multi-Agent Systems?

    Sep 24, 2026Addison J. Wu, Jasin Cekinmez, Michel Liao +2DeceptionConformity

  4. When Honesty is Not Enough in AI Debate

    Sep 24, 2026Rayne Holland, Liming Zhu, Jason XueScalable OversightHonesty

  5. Sharp Limits for Honest Uncertainty in Hard-Budget Repeated Evaluation

    Sep 24, 2026Yezhou Cheng, Runjia Du, Zeming Liu +5Confidence IntervalsControlled Study

  6. Fabrication After Tool Failure: Tool-Augmented Agents Assert Values Their Tools Did Not Return

    Sep 13, 2026Arham Sethi, Arsen Kenzhebayev, Saanvi Paturi +3HonestySystemverilog Assertions

  7. Optimizing Byzantine Node Placement in Decentralized Federated Learning

    Sep 1, 2026Edoardo Gabrielli, Gabriele TolomeiByzantine AttacksFederated Learning

  8. The Politician, the Liar, and the Obedient Worker: Emerging Behavior of LLM Agents in Hierarchical Games

    Aug 10, 2026Fatemeh Seyedin, Adrian Weller, Jinhyuk Yun +1Game TheoryLarge Language Model Agents

  9. When Truth Is Distributed: Misinformation Derails Collective Fact Recovery in LLM-Based Multi-Agent Systems

    Aug 4, 2026Chenfei Yan, Zeyang Yue, Feifei Zhao +6MisinformationTruth

  10. Paying for Honesty Without Knowing the Truth: Reputation-Penalty Design for LLM Marketplace Agents

    Jul 30, 2026Mingdai Yang, Shicheng Fan, Kejing Yu +5HonestyReputation

  11. Creativity, honesty and designed forgetting emerge in small hyperbolic language models

    Jul 10, 2026Kwan Soo Shin, In Seok Kang, Yunkyung MinSmall Large Language ModelsCreativity

  12. Formal Mechanisms for Market Stability in Self-Interested Agent Societies: A Marketplace Simulation Study

    Jul 9, 2026Eugene Ng Yi Sheng, Bingquan ShenMulti-Agent SimulationsMediation

  13. Safety from Honesty in a Disinterested AI Predictor

    Jun 28, 2026Yoshua Bengio, Oliver Richardson, Tomáš Gavenčiak +13Artificial Intelligence SafetyArtificial Intelligence Scientists

  14. RogueAI: A Reverse Turing Test for Detecting Licensed AI Deception in Dialogue

    Jun 11, 2026Sara Candussio, Emanuele Ballarin, Lorenzo Bonin +2DeceptionTrustworthy Artificial Intelligence

  15. The Impossibility of Eliciting Latent Knowledge

    Jun 10, 2026Korbinian Friedl, Francis Rhys Ward, Paul Yushin Rapoport +2HonestyArtificial Intelligence Agents

  16. Truthful AI Advisors: A Pre-Specified Benchmark for Large Language Model Honesty Under Preference Misalignment

    May 31, 2026Hamidreza Hasani Balyani, Seyed Pouyan Mousavi Davoudi, Alireza Amiri-Margavi +2HonestyLarge Language Model Alignment

  17. Used Car Salesbots? Honesty and Credulity of LLMs as Bargaining Agents under Partial Information

    May 29, 2026Antonio Valerio Miceli-Barone, Vaishak Belle, Shay B. CohenNegotiationLarge Language Model Agents

  18. Token Inflation: How Dishonest Providers Can Overcharge for Large Language Model Usage

    May 28, 2026Shahinul Hoque, Jinghuai Zhang, Jinyuan Sun +1Model AuditingHonesty

  19. Honest Lying: Understanding Memory Confabulation in Reflexive Agents

    May 28, 2026Prakhar Dixit, Sadia Kamal, Tim OatesReflectionsConfabulation

  20. Your Neighbors Know: Leveraging Local Neighborhoods for Backdoor Detection in Decentralized Learning

    May 19, 2026Sayan Biswas, Antoine Boutet, Davide Frey +7Clean Label Backdoor AttackDecentralized Learning

  21. Unlearners Can Lie: Evaluating and Improving Honesty in LLM Unlearning

    May 9, 2026Renjie Gu, Jiazhen Du, Yihua Zhang +1Large Language Model UnlearningLarge Language Model Reliability

  22. The Endogeneity of Miscalibration: Impossibility and Escape in Scored Reporting

    May 8, 2026Lauri Lovén, Sasu TarkomaScalable OversightHonesty

  23. ARGUS: Policy-Adaptive Ad Governance via Evolving Reinforcement with Adversarial Umpiring

    May 4, 2026Deyi Ji, Junyu Lu, Xuanyi Liu +7AdsHonesty