cs.CLOct 8, 2026

Deception by Omission: Language Models Knowingly Hide Their Mistakes

Authors: Lucas Florin, Amelie Knecht, Ulysse Schaller, Thilo Hagendorff

Organizations: AI Safety Research Group, University of Stuttgart, Germany

Abstract

Large language models (LLMs) increasingly act as agents with little human oversight, so potential mistakes they make can go unnoticed. Users then depend on the model to report what went wrong. An honest model discloses its mistakes, while a deceptive one conceals them. However, it is unclear how current LLMs behave in such situations. In this study, we prefill LLM trajectories with synthetic mistakes. The trajectories resemble real deployments in chat and agentic settings. Models fail to disclose their mistake in 36.4% of chat and 67.1% of agentic rollouts. In 2.4% and 5.3% of rollouts, respectively, they are aware of the mistake in their chain of thought but still deceptively conceal it. Rates vary by model: for instance, Gemini 3.5 Flash knowingly conceals mistakes in up to 19.9% of agentic rollouts. In 11.9% of chat and 51.8% of agentic rollouts, models show no awareness of mistakes, even though they reliably spot them when reviewing the same transcript as an outside observer. Our results show that, as agents take on more tasks with less oversight, users cannot rely on them to self-report possible mistakes. Developers should instead use independent monitors that review agent trajectories, or specifically train models to check their past actions and disclose what they find.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. When Do Large Language Models Exhibit Unsolicited Deception?

    Mar 31, 2025Samuel M. Taylor, Benjamin K. BergenLanguage Model Safety EvaluationDeception in Language Models

  2. Language Models Are "Insecure" Reporters

    Sep 28, 2026Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun +5Language Model Safety EvaluationDeception in Language Models

  3. Confidently Deceptive: On the Relationship Between Confidence and Deception in LLMs

    May 12, 2026Ali Asad, Stephen Obadinma, Anshul Pattoo +2Language Model Safety EvaluationConfidence Estimation in Language Models