cs.LGSep 29, 2026

Guard Models Are Overconfident Where Base Models Are Uncertain

Authors: Jonghyun Hong, MinJae Jung, Minwoo Kim

Organizations: DATUMO INC.

Abstract

Guard models are used as safety classifiers, with confidence scores driving downstream moderation decisions. We evaluate five guard models for prompt classification and find that although several are nearly calibrated on clean inputs, adversarial attacks degrade their calibration by an order of magnitude, turning false negatives into high-confidence errors indistinguishable from correct detections. Comparing each guard with its corresponding base LM, we find that uncertainty signals often remain available, with the base model typically expressing uncertainty on the same inputs where the guard fails. Layer-wise analyses localize this guard-base divergence to later layers, where guard models exhibit sharper safe/unsafe separation and lower-rank representations, while adversarial harmful inputs lie closer to the clean-safe region. These findings highlight a mismatch between guard confidence and base model uncertainty under attack.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Beyond Red-Teaming: Formal Guarantees of LLM Guardrail Classifiers

    May 11, 2026Nikita Kezins, Urbas Ekka, Pascal Berrang +1Large Language Model SafetyToxicity

  2. Safety Reconstructed: Generative Modeling via Masked Diffusion Builds Strong Safety Guardrails

    Sep 27, 2026Gert Lek, Abele Malan, Chaoyi Zhu +3Large Language Model SafetyMasked Diffusion Language Models