cs.CLSep 27, 2026

Safety Reconstructed: Generative Modeling via Masked Diffusion Builds Strong Safety Guardrails

Authors: Gert Lek, Abele Malan, Chaoyi Zhu, Pin-Yu Chen, Robert Birke, Lydia Chen

Organizations: University of Neuchâtel · Delft University of Technology · IBM Research · University of Turin

Abstract

Guard models are the last line of defense between a language model and a harmful output, yet their training objective is surprisingly narrow. Existing guards learn to predict a single verdict token from a conversational context, concentrating supervision on a single target. The consequences are structural: models latch onto shortcut features, are overconfident, and remain sensitive to where safety evidence appears in the sequence rather than its role in the full context. We propose a different framing. Rather than predicting a label from text, our LLaDA-Guard asks which label better explains the text: scoring the prompt or response under each label hypothesis and classifying based on their difference. This shifts supervision to every token in the moderated region, forcing the model to account for full content rather than its most discriminative fragments. We instantiate this idea with a masked diffusion language model, fine-tuning LLaDA-8B-Instruct with a class-conditional reconstruction objective using LoRA and requiring no architectural changes beyond the base model. LLaDA-Guard leads on average rank against discriminative baselines trained on stronger backbones across seven held-out safety benchmarks, while exhibiting substantially better confidence calibration (ECE 0.0875 vs. 0.1384 for Qwen3Guard), less over-defense on benign prompts with unsafe-looking cues, and less prompt leakage when moderating responses. Its generative nature further enables token-level risk localization as a natural byproduct, yielding a pipeline for rewriting unsafe prompts into safe equivalents without additional training and achieving a 60.7% average conversion-to-safe rate.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions

    Aug 11, 2026Xinzhe Huang, Biwu Yao, Kedong Xiu +4Large Language Model SafetyGuardrail

  2. SentGuard: Sentence-Level Streaming Guardrails for Large Language Models

    Jun 1, 2026Jiaqi Yu, Xin Wang, Yixu Wang +4Large Language Model SafetyStreaming Guardrails

  3. Robust and Efficient Guardrails with Latent Reasoning

    May 27, 2026Siddharth Sai, Xiaofei Wen, Muhao ChenLarge Language Model SafetyGuardrail