cs.CLSep 5, 2026

Recall Is Not Protection: Evaluating Safety Monitors Against Model Compliance

Authors: Sripad Karne

Organizations: Columbia University

Abstract

Safety monitors screen prompts sent to deployed language models, flagging harmful requests so they are never answered. They are evaluated by recall against harmfulness labels, but a catch only prevents harm if the model would otherwise have complied. We measure the difference directly: we sample repeated responses from the target model, call a harmful prompt \emph{elicitable} if the model complies at least once, and report monitor recall separately on elicitable and non-elicitable prompts. Across six monitor configurations and three model families, spanning activation probes, fine-tuned text guards, and a 120B policy-conditioned reasoning classifier, recall on elicitable prompts falls 0.22 to 0.38 below recall on non-elicitable prompts at a fixed false positive rate. The prompts a monitor misses are 2.8 to 5.6 times more likely to be complied with than the prompts it catches. The gap replicates across three model families and appears also in text-only monitors entirely independent of the target model. This suggests that standard recall may overstate the protection monitors provide in practice, and that monitors should be evaluated against what their models will actually answer.

Figures & tables

Appendix figures & tables18 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models

    Sep 16, 2026Alizishaan Khatri, Chiquita Prabhu, Omkar NeogiLarge Language Model SafetyGuardrail

  2. Online Safety Monitoring for LLMs

    Jul 2, 2026Mona Schirmer, Metod Jazbec, Alexander Timans +3Large Language Model SafetyRed-Teaming

  3. AERIC: Anticipatory Hidden-State Monitoring for Implicit Harmful Dialogue

    May 13, 2026Jihyung Park, Saleh Afroogh, Junfeng JiaoLarge Language Model SafetyHarmful Content