Safety Reconstructed: Generative Modeling via Masked Diffusion Builds Strong Safety Guardrails
Organizations: University of Neuchâtel · Delft University of Technology · IBM Research · University of Turin
Abstract
Guard models are the last line of defense between a language model and a harmful output, yet their training objective is surprisingly narrow. Existing guards learn to predict a single verdict token from a conversational context, concentrating supervision on a single target. The consequences are structural: models latch onto shortcut features, are overconfident, and remain sensitive to where safety evidence appears in the sequence rather than its role in the full context. We propose a different framing. Rather than predicting a label from text, our LLaDA-Guard asks which label better explains the text: scoring the prompt or response under each label hypothesis and classifying based on their difference. This shifts supervision to every token in the moderated region, forcing the model to account for full content rather than its most discriminative fragments. We instantiate this idea with a masked diffusion language model, fine-tuning LLaDA-8B-Instruct with a class-conditional reconstruction objective using LoRA and requiring no architectural changes beyond the base model. LLaDA-Guard leads on average rank against discriminative baselines trained on stronger backbones across seven held-out safety benchmarks, while exhibiting substantially better confidence calibration (ECE 0.0875 vs. 0.1384 for Qwen3Guard), less over-defense on benign prompts with unsafe-looking cues, and less prompt leakage when moderating responses. Its generative nature further enables token-level risk localization as a natural byproduct, yielding a pipeline for rewriting unsafe prompts into safe equivalents without additional training and achieving a 60.7% average conversion-to-safe rate.
Figures & tables
| Prompt-only datasets | Response datasets | Overall | OR | |||||
| Model | ToxicChat | OpenAI Mod. | XSTest | SafeRLHF | BeaverTails | XSTest-Resp | Avg | PH Acc. |
| ShieldGemma 27B | 63.68 / 30.18 | 69.05 / 50.89 | 82.51 / 41.22 | 82.09 / 28.23 | 79.08 / 28.53 | 54.58 / 34.86 | 71.83 / 35.65 | 98.03 |
| ShieldGemma 2B | 59.73 / 16.95 | 61.31 / 16.21 | 82.61 / 70.80 | 90.94 / 26.19 | 86.77 / 25.18 | 82.16 / 56.36 | 77.25 / 35.28 | 97.50 |
| ShieldGemma 9B | 80.03 / 68.34 | 88.68 / 78.93 | 90.98 / 81.99 | 88.39 / 62.66 | 86.84 / 68.81 | 91.18 / 86.67 | 87.68 / 74.57 | 89.84 |
| Llama-Guard 4 12B | 56.04 / 51.15 | 79.44 / 73.74 | 90.29 / 83.42 | 95.09 / 87.64 | 89.38 / 69.80 | 94.09 / 89.04 | 84.05 / 75.80 | 93.21 |
| Llama-Guard 3 8B | 57.93 / 54.03 | 87.30 / 79.11 | 97.40 / 88.41 | 96.66 / 88.56 | 89.02 / 67.76 | 95.71 / 90.41 | 87.34 / 78.05 | 94.80 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | ToxicChat | OpenAI Mod | XSTest | BeaverTails | XSTest-Resp | Avg |
| LLaDA-Guard Direct 8B | 0.8163 | 0.8163 | 0.9594 | 0.9104 | 0.9412 | 0.8887 |
| LLaDA-Guard 8B | 0.8452 | 0.8568 | 0.9691 | 0.9294 | 0.9806 | 0.9162 |
| +0.0289 | +0.0405 | +0.0097 | +0.0190 | +0.0394 | +0.0275 |
| Model | ToxicChat | OpenAI Mod | XSTest | BeaverTails | XSTest-Resp | Avg |
| LLaDA-Guard Direct 8B | 0.7566 | 0.7339 | 0.8855 | 0.7342 | 0.8514 | 0.7923 |
| LLaDA-Guard 8B | 0.7755 | 0.7513 | 0.8877 | 0.8329 | 0.9281 | 0.8599 |
| +0.0189 | +0.0174 | +0.0022 | +0.0877 | +0.0767 | +0.0406 |
| Model | ECE | Brier | NLL |
| ShieldGemma 2B | 19.79 | 18.26 | 64.21 |
| ShieldGemma 9B | 11.90 | 11.85 | 40.35 |
| ShieldGemma 27B | 16.05 | 16.60 | 50.78 |
| Llama-Guard 3 8B | 9.51 | 9.84 | 47.37 |
| Llama-Guard 4 12B | 12.74 | 12.65 | 168.65 |
| Granite Guardian 3.0 2B | 25.77 | 20.71 | 62.58 |
| Model | ToxicChat | OpenAI Mod | XSTest | SafeRLHF | BeaverTails | XSTest-Resp | PHTest | Avg |
| ShieldGemma 27B | 0.0419 | 0.0611 | 0.2081 | 0.2936 | 0.3553 | 0.0424 | 0.1214 | 0.1605 |
| ShieldGemma 2B | 0.0950 | 0.2106 | 0.1568 | 0.3540 | 0.4236 | 0.0812 | 0.0639 | 0.1979 |
| ShieldGemma 9B | 0.0252 | 0.0878 | 0.1065 | 0.2087 | 0.2105 | 0.0603 | 0.1341 | 0.1190 |
| Llama-Guard 4 12B | 0.1179 | 0.1546 | 0.1331 | 0.1121 | 0.2707 | 0.0337 | 0.0697 | 0.1274 |
| Llama-Guard 3 8B | 0.0773 | 0.0705 | 0.0774 | 0.0741 | 0.2664 | 0.0248 | 0.0753 | 0.0951 |
| Granite Guardian 3.0 2B | 0.3849 | 0.2734 | 0.2284 | 0.0542 | 0.1762 | 0.0331 | 0.6539 | 0.2577 |
| Model | Matched | Matched |
| LLaDA-Guard | 0.099 | 0.060 |
| LlamaGuard 3 | 0.206 | 0.208 |
| LlamaGuard 4 | 0.138 | 0.132 |
| WildGuard | 0.292 | 0.308 |
| GraniteGuard | 0.122 | 0.140 |
| Qwen3Guard | 0.386 | 0.424 |
| Model | 1 trigger | 2 triggers | 3 triggers |
| LLaDA-Guard | 0.00% (0) conf. – | 0.00% (0) conf. – | 0.00% (0) conf. – |
| Qwen3Guard | 0.00% (0) conf. – | 0.88% (1) conf. 80.59% | 3.54% (4) conf. 77.55% |
| WildGuard | 0.88% (1) conf. 65.14% | 0.88% (1) conf. 99.44% | 2.65% (3) conf. 78.16% |
| Granite | 1.77% (2) conf. 60.61% | 5.31% (6) conf. 64.70% | 10.62% (12) conf. 70.61% |
| NemoGuard | 0.00% (0) conf. – | 0.00% (0) conf. – | 1.77% (2) conf. 63.69% |
| Llama-Guard 3 | 0.00% (0) conf. – | 0.00% (0) conf. – | 0.88% (1) conf. 90.47% |
| Localizer | B1 | B2 | B3 | B4 | B5 |
| ShieldGemma 27B | 10.9 | 20.9 | 29.8 | 38.8 | 43.4 |
| ShieldGemma 2B | 16.5 | 32.9 | 41.4 | 47.3 | 50.7 |
| ShieldGemma 9B | 12.2 | 21.0 | 27.5 | 32.9 | 40.9 |
| Llama-Guard 4 12B | 14.7 | 26.1 | 33.5 | 39.3 | 44.7 |
| Llama-Guard 3 8B | 12.4 | 23.0 | 32.6 | 41.7 | 48.6 |
| Granite Guardian 2B | 13.7 | 24.8 | 34.3 | 40.6 | 45.9 |
| Localizer | XSTest | ToxicChat | OpenAI Mod | Avg |
| ShieldGemma 27B | 85.0 | 35.1 | 10.1 | 43.4 |
| ShieldGemma 2B | 96.0 | 42.3 | 13.8 | 50.7 |
| ShieldGemma 9B | 77.0 | 34.5 | 11.1 | 40.9 |
| Llama-Guard 4 12B | 85.5 | 37.0 | 11.5 | 44.7 |
| Llama-Guard 3 8B | 92.0 | 41.4 | 12.4 | 48.6 |
| Granite Guardian 2B | 91.0 | 37.0 | 9.7 | 45.9 |