cs.CROct 7, 2026

When AI Finds Hidden Messages, Does It Report?

Authors: William Guey, Rashik Jahangir, Pierrick Bougault, Vitor D. de Moura, Wei Zhang, José O. Gomes

Abstract

When an assistant encounters a message for another AI, does it tell its user? Four fixed model-provider deployments perform simulated source tasks in 1,280 ordinary-note and 128 enhanced-note sessions. Harmless and harmful messages have matched plaintext and ROT13 versions, with no-message controls. Observers receive no decoder or decoded meaning; a requested reference code incentivizes inspection. Asking for reports increases rule-detected notifications identifying another AI as recipient by 53.1 percentage points for harmless ROT13 messages and 54.7 for harmful ones. This is a joint inspection, recognition, and notification effect; missing-response bounds are 38.3--77.3 and 36.7--78.1 points. Model-based trace checks identify eleven ordinary plaintext cases where agents interpret the message but do not notify their user. Seven encoded omissions are verified with enhanced notes; ordinary encoded omissions remain unverified. Seven simulated filename disclosures coexist with accurate review-status answers, and two answers use a planted false count. Interpretation, notification, and authorized task performance are distinct outcomes.

Explore similar work

CardsList
  1. Deception by Omission: Language Models Knowingly Hide Their Mistakes

    Oct 8, 2026Lucas Florin, Amelie Knecht, Ulysse Schaller +1Deception in Language ModelsLanguage Model Error Detection

  2. Language Models Are "Insecure" Reporters

    Sep 28, 2026Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun +5Language Model Safety EvaluationDeception in Language Models

  3. Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems

    Sep 30, 2026Deema Alnuhait, Gengyu Wang, Muhammad Khalifa +1