Recall Is Not Protection: Evaluating Safety Monitors Against Model Compliance
Organizations: Columbia University
Abstract
Safety monitors screen prompts sent to deployed language models, flagging harmful requests so they are never answered. They are evaluated by recall against harmfulness labels, but a catch only prevents harm if the model would otherwise have complied. We measure the difference directly: we sample repeated responses from the target model, call a harmful prompt \emph{elicitable} if the model complies at least once, and report monitor recall separately on elicitable and non-elicitable prompts. Across six monitor configurations and three model families, spanning activation probes, fine-tuned text guards, and a 120B policy-conditioned reasoning classifier, recall on elicitable prompts falls 0.22 to 0.38 below recall on non-elicitable prompts at a fixed false positive rate. The prompts a monitor misses are 2.8 to 5.6 times more likely to be complied with than the prompts it catches. The gap replicates across three model families and appears also in text-only monitors entirely independent of the target model. This suggests that standard recall may overstate the protection monitors provide in practice, and that monitors should be evaluated against what their models will actually answer.
Figures & tables
| AUROC | Flag rate | Leak | |||||
|---|---|---|---|---|---|---|---|
| Guard | L1 | L2 | L3 | L1 | L2 | L3 | /344 |
| Llama Guard 3 | .70 | .94 | .99 | .15 | .77 | .94 | 321 |
| Qwen3Guard | .85 | .99 | 1.00 | .40 | .89 | .99 | 245 |
| WildGuard | .96 | .99 | 1.00 | .54 | .88 | .99 | 177 |
| gpt-oss-safeguard | .86 | .95 | .96 | .59 | .74 | .76 | 153 |
| AUROC | Flag rate | Leak | |||||
|---|---|---|---|---|---|---|---|
| L1 | L2 | L3 | L1 | L2 | L3 | /344 | |
| Linear | .91 | .98 | 1.00 | .43 | .79 | .94 | 270 |
| Attention | .92 | .99 | 1.00 | .45 | .94 | .99 | 259 |
| Latent Guard | .83 | .95 | .98 | .44 | .72 | .87 | 261 |
| Explicitness | Expl. harm | Harmless expl. | Random | ||||
|---|---|---|---|---|---|---|---|
| Flag | SD | Flag | SD | Flag | SD | Flag | |
| 0 | .26 | 0 | .26 | 0 | .26 | 0 | .26 |
| 0.1 | .42 | .35 | .37 | .27 | |||
| 0.2 | .57 | .40 | .47 | .28 | |||
| 0.5 | .91 | .57 | .70 | .28 | |||
| 1 | 1.00 | .85 | .96 | .26 | |||
| Held-out | Harmless FPR (%) | External recall | ||||||
|---|---|---|---|---|---|---|---|---|
| Arm | Recall | Leak Rate | Original | Indirect | XSTest | All 24 | Pers. | Adv. |
| A untouched | .238 | .784 | 1.4 | 0.6 | 0.4 | .442 | .081 | .682 |
| B originals | .422 | .572 | 0.6 | 0.2 | 13.7 | .461 | .116 | .565 |
| B + OR-Bench | .035 | .985 | 0.4 | 0.0 | 9.6 | .334 | .000 | .268 |
| C harmful ladders | .978 | .000 | 0.7 | 16.1 | 11.2 | .640 | .583 | .772 |
| D + harmless ladder | .889 | .057 | 0.8 | 1.1 | 9.4 | .553 | .361 | .693 |
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Answered L1 | L2 | L3 | Flipped |
|---|---|---|---|---|
| Gemma-4-31B-IT | 340 | 71 | 12 | 344 |
| Qwen3-32B | 482 | 114 | 27 | 470 |
| Qwen3-32B, per guard | ||||
| Guard | Ratio L1 | L2 | L3 | Leaks /470 |
| Llama Guard 3 | 3.4 | 1.6 | 1.4 | 422 (405) |
| Qwen3Guard | 1.4 | 1.2 | 1.1 | 293 (277) |
| Answered | Flip | L1 recall | ||||||
|---|---|---|---|---|---|---|---|---|
| Source | L1 | L2 | L3 | set | Guard | Refused | Answered | Ratio |
| Aegis 2.0 (193) | 110 | 33 | 8 | 108 | Llama Guard 3 | .217 | .045 | 4.8 |
| WildGuard | .422 | .255 | 1.7 | |||||
| gpt-oss-safeguard | .446 | .418 | 1.1 | |||||
| Qwen3Guard | .313 | .136 | 2.3 | |||||
| HarmBench (168) | 91 | 28 | 2 | 98 | Llama Guard 3 | .221 | .022 | 10.1 |
| Category | Answered | Llama Guard 3 | WildGuard | gpt-oss | Qwen3Guard | Pooled | |
|---|---|---|---|---|---|---|---|
| Violence, weapons | 122 | .52 | 17.1 | 1.50 | 1.07 | 4.58 | 1.84 |
| Cyber | 37 | .57 | 0.85 | 1.02 | 2.95 | 1.18 | |
| Illegal goods, drugs | 91 | .42 | 7.17 | 0.97 | 1.22 | 1.39 | 1.27 |
| Fraud, non-violent crime | 121 | .69 | 3.97 | 1.17 | 1.02 | 1.78 | 1.27 |
| Hate, harassment | 130 | .50 | 3.67 | 1.72 | 1.33 | 1.70 | 1.78 |
| Sexual content | 46 | .37 | 7.62 | 1.44 | 1.12 | 1.50 | 1.54 |
| AUROC | Flag rate | Recall (ref.) | Recall (ans.) | Ratio | Leak | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| L1 | L2 | L3 | L1 | L2 | L3 | L1 | L2 | L3 | L1 | L2 | L3 | L1 | L2 | L3 | /344 | |
| Qwen3Guard, strict | .929 | .993 | .996 | .270 | .889 | .949 | .427 | .917 | .953 | .129 | .662 | .750 | 3.3 | 1.4 | 1.3 | 290 |
| gpt-oss-safeguard, WG policy | .858 | .975 | .975 | .580 | .689 | .753 | .671 | .724 | .759 | .497 | .408 | .417 | 1.4 | 1.8 | 1.8 | 172 |
| AUROC | Flag rate | Recall (answered) | Ratio [95% CI] | Leak | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Monitor | L1 | L2 | L3 | L1 | L2 | L3 | L1 | L2 | L3 | L1 | L2 | L3 | /344 |
| Linear, mean content, L28 | .91 | .98 | 1.00 | .21 | .77 | .96 | .07 | .41 | .33 | 5.7 [3.9, 9.4] | 2.0 [1.5, 2.8] | 2.9 [1.5, 9.8] | 316 |
| Linear, post-instruction, L29 | .90 | .99 | 1.00 | .39 | .88 | .97 | .17 | .61 | .33 | 3.8 [2.9, 5.0] | 1.5 [1.3, 1.9] | 3.0 [1.6, 8.9] | 278 |
| Linear, last content, L29 | .92 | .98 | .99 | .36 | .74 | .85 | .20 | .35 | .25 | 2.7 [2.2, 3.4] | 2.2 [1.7, 3.3] | 3.5 [1.7, 11.0] | 271 |
| Linear, instruction, L31 | .91 | .98 | 1.00 | .43 | .79 | .94 | .21 | .35 | .58 | 3.2 [2.6, 4.0] | 2.4 [1.8, 3.6] | 1.6 [1.1, 3.3] | 270 |
| Attention, full sequence, L28 | .92 | .99 | 1.00 | .45 | .94 | .99 | .22 | .68 | .75 | 3.2 [2.7, 3.9] | 1.4 [1.2, 1.7] | 1.3 [1.0, 2.0] | 259 |
| Leaks /344 | L1 ratio | |||
|---|---|---|---|---|
| Monitor | WildChat | LMSYS | WildChat | LMSYS |
| gpt-oss-safeguard | 162 | 270 | 1.1 | 1.2 |
| WildGuard | 177 | 258 | 1.4 | 1.8 |
| Qwen3Guard | 245 | 330 | 2.1 | 5.0 |
| Attention probe | 259 | 328 | 3.2 | 11.3 |
| Latent Guard | 261 | 283 | 2.8 | 3.1 |
| FPR | 0.1% | 0.5% | 1% | 2% | 5% | 10% |
|---|---|---|---|---|---|---|
| Llama Guard 3 | 6.7 | 7.9 | 6.4 | 4.6 | 4.0 | 3.1 |
| WildGuard | 2.7 | 1.5 | 1.4 | 1.3 | 1.2 | 1.1 |
| Qwen3Guard | 3.5 | 2.1 | 1.4 | 1.2 | 1.2 | |
| gpt-oss-safeguard | 1.2 | 1.2 | 1.1 | 1.2 | 1.2 ∗ | 1.2 ∗ |
| Linear probe | 13.5 | 7.0 | 3.2 | 2.0 | 1.7 | 1.4 |
| Attention probe | 4.5 | 3.2 | 2.3 | 1.5 | 1.3 |
| FPR | 0.1% | 0.5% | 1% | 2% | 5% | 10% |
|---|---|---|---|---|---|---|
| (SD) | 3.44 | 2.45 | 2.11 | 1.80 | 1.39 | 1.09 |
| Refused above | .41 | .58 | .66 | .72 | .81 | .86 |
| Answered above | .08 | .19 | .24 | .29 | .42 | .50 |
| Frozen | Matched 1% FPR | |||||
|---|---|---|---|---|---|---|
| Pair | Leaks | Stack FPR | OR leaks [95% CI] | Indep. | Mean-rank | |
| WG + GOS | 108 | 1.8% | 163 [145, 180] | 149 | 133 | |
| GOS + LatG | 118 | 1.8% | 190 [170, 208] | 192 | 211 | |
| WG + LatG | 152 | 1.9% | 194 [175, 214] | 177 | 228 | |
| GOS + Attn | 116 | 1.7% | 201 [182, 217] | 203 | 192 | |
| WG + Q3G | 166 | 1.5% | 203 [185, 222] | 180 | 181 | |
| Level | Group | Harm | Refusal | |
| L1 | refused | 307 | [2.84, 3.27] | [6.08, 6.44] |
| L1 | answered | 340 | [1.07, 1.37] | [2.36, 2.66] |
| L2 | refused | 576 | [3.69, 4.02] | [7.36, 7.60] |
| L2 | answered | 71 | [1.21, 1.86] | [2.72, 3.46] |
| L3 | refused | 635 | [4.74, 5.06] | [8.17, 8.34] |
| L3 | answered | 12 | [0.70, 2.39] | [1.76, 4.17] |
| Level | Group | Llama Guard 3 | Qwen3Guard | |
|---|---|---|---|---|
| L1 | refused | 307 | [1.89, 2.14] | [2.81, 3.05] |
| L1 | answered | 340 | [0.69, 0.91] | [2.09, 2.34] |
| L2 | refused | 576 | [2.93, 3.08] | [3.95, 4.04] |
| L2 | answered | 71 | [1.48, 1.95] | [3.16, 3.53] |
| L3 | refused | 635 | [3.75, 3.88] | [4.37, 4.42] |
| L3 | answered | 12 | [1.77, 3.08] | [3.29, 4.07] |
| Arm | Harmful (unsafe) | Harmless (safe) | Rows/side |
|---|---|---|---|
| B | 450 originals ∗ | 1,350 prompts | 1,350 |
| B + OR | 450 originals ∗ | 1,350 prompts + 1,350 OR-Bench | 2,700 |
| C | 450 L1, L2, L3 | 1,350 prompts | 1,350 |
| D | 450 L1, L2, L3 | 450 L1, L2, L3 | 1,350 |
| D + OR | 450 L1, L2, L3 ∗ | D’s ladder + 1,350 OR-Bench | 2,700 |
| E | 450 L1, L2, L3 | 450 harm-word L1, L2, L3 | 1,350 |
| Arm | Recall L1 | Leak | Direct | Indirect | XSTest | External |
|---|---|---|---|---|---|---|
| A | .238 [.156, .327] | .784 [.694, .864] | 1.4 [0.9, 1.9] | 0.6 [0.3, 0.9] | 0.4 [0.0, 1.3] | .442 [.424, .459] |
| B | .422 [.336, .512] | .572 [.471, .671] | 0.6 [0.3, 0.9] | 0.2 [0.0, 0.4] | 13.7 [9.7, 17.7] | .461 [.443, .478] |
| B+OR | .035 [.006, .072] | .985 [.957, 1.00] | 0.4 [0.2, 0.7] | 0.0 [0.0, 0.0] | 9.6 [6.5, 13.0] | .334 [.319, .350] |
| C | .978 [.947, 1.00] | .000 [.000, .000] | 0.7 [0.4, 1.0] | 16.1 [14.7, 17.7] | 11.2 [7.7, 15.0] | .640 [.620, .658] |
| D | .889 [.827, .941] | .057 [.020, .102] | 0.8 [0.4, 1.2] | 1.1 [0.7, 1.6] | 9.4 [6.3, 13.1] | .553 [.532, .573] |
| D+OR | .914 [.859, .964] | .045 [.008, .091] | 0.7 [0.4, 1.1] | 1.6 [1.1, 2.1] | 8.3 [5.5, 11.6] | .441 [.425, .457] |
| FPR | 0.1% | 0.5% | 1% | 2% | 5% | 10% | |
|---|---|---|---|---|---|---|---|
| A | Recall | .00 | .08 | .24 | .45 | .62 | .68 |
| Leak | 1.00 | .92 | .78 | .55 | .35 | .31 | |
| XSTest | 0.0 | 0.0 | 0.4 | 1.6 | 8.0 | 11.2 | |
| External | .05 | .34 | .44 | .53 | .62 | .65 | |
| D | Recall | .47 | .81 | .89 | .95 | .96 | .98 |
| Leak | .48 | .14 | .06 | .03 | .02 | .01 |
| Set | Raw | Kept |
| SORRY-Bench, 21 classes | 9,236 | 9,155 |
| WildJailbreak adversarial harmful | 2,000 | 1,687 |
| WildGuardTest adversarial harmful | 341 | 298 |
| WildGuardTest adversarial benign | 455 | 435 |
| ToxicChat harmful | 362 | 333 |
| ToxicChat benign | 4,721 | 4,478 |
| Guard | XSTest FPR (%) |
|---|---|
| WildGuard | 0.0 |
| Qwen3Guard, untouched | 0.4 |
| Llama Guard 3 | 1.6 |
| gpt-oss-safeguard | 4.0 |
| Qwen3Guard, arm H | 4.8 |
| Qwen3Guard, arm D | 9.4 |
| Arm | Gap | Above | Verdict | L1 refused | L1 answered | Benign shift |
|---|---|---|---|---|---|---|
| A untouched | .72 | 13% | 26% | 0 | ||
| B | 1.20 | 42% | 51% | |||
| B + OR-Bench | 1.20 | 2% | 3% | |||
| C | .10 | 99% | 99% | |||
| D | .12 | 95% | 95% | |||
| D + OR-Bench | .15 | 94% | 96% |
| Held-out | Harmless FPR (%) | External | ||||
|---|---|---|---|---|---|---|
| Arm | Recall | Leak | Direct | Indirect | XSTest | Recall |
| A untouched | .048 | .943 | 1.9 | 0.5 | 1.6 | .302 |
| B originals | .406 | .617 | 0.7 | 0.1 | 37.3 | .363 |
| C harmful ladders | .990 | .000 | 1.0 | 38.3 | 36.3 | .588 |
| D + harmless ladder | .905 | .045 | 0.7 | 1.3 | 29.6 | .440 |
| H + harm-word ladder, distillation | .902 | .045 | 1.0 | 0.8 | 13.3 | .475 |