A Safe Prototype Is Not a Safety Direction: Reference Dependence and Prompt Confounds in Response-Safety Embeddings
Organizations: New York University
Abstract
Can response safety be scored by cosine similarity to the mean embedding of known-safe responses? A recent sleeper-agent detector proposes exactly this score, yet the raw positive-centroid rule is not identified: positive observations locate the safe class relative to an encoder origin, but do not determine which direction separates safe from unsafe responses. We audit the rule on two prompt-controlled, human-labeled corpora and one auxiliary jury-labeled source control, using four frozen encoders and prompt-grouped splits. On the human-labeled corpora the safe prototype reaches ROC-AUC 0.457-0.545, with two cells significantly below chance and one above, while an explicit safe-minus-unsafe reference reaches 0.588-0.738 on the same embeddings; on the jury control the prototype is inverted (0.358-0.405) and the reference reaches 0.754-0.793. At validation-calibrated 5% false-safe thresholds, the reference accepts more safe responses on PKU-SafeRLHF (0.153-0.263 versus 0.039-0.061 across encoders) and Aegis (0.189-0.291 versus 0.004-0.045), but not reliably on BeaverTails. A fully unlabeled held-out reference recovers part to most of the referenced ranking, much less when only 5% of the pool is unsafe, whereas 80-634 labeled unsafe responses recover most of it. Prompt-only ablations show that prompt-label composition can inflate uncontrolled evaluations. This is a bounded result about a raw positive centroid, not all one-class methods or safety-specialized guards. A class mean is a location, not necessarily a safety direction; a declared reference with enough unsafe mass identifies orientation.
Figures & tables
| Readout | Train information | BeaverTails | PKU-SafeRLHF | Aegis jury |
|---|---|---|---|---|
| Safe prototype | Safe labels; encoder origin | .509–.545 | .457–.501 | .358–.405 |
| Log response length | Label sign only | .529 | .595 | .554 |
| Refusal-template indicator | Label sign only | .523 | .513 | .506 |
| Stylistic features (12) | Response-level labels | .561 | .640 | .615 |
| Difference of means | Safe and unsafe labels | .588–.622 | .668–.738 | .754–.793 |
| Mixture-centered safe mean | Safe labels; pooled train mean | .588–.622 | .668–.738 | .754–.793 |
| Corpus | Target | Prototype test FSR | Prototype safe accept. | Diff-means test FSR | Diff-means safe accept. |
|---|---|---|---|---|---|
| BeaverTails | 5% | .050 [.021,.074] | .082 [.038,.109] | .045 [.025,.085] | .084 [.056,.130] |
| 10% | .091 [.057,.130] | .124 [.084,.171] | .104 [.068,.159] | .163 [.112,.229] | |
| PKU-SafeRLHF | 5% | .064 [.047,.078] | .061 [.043,.077] | .052 [.039,.066] | .170 [.133,.198] |
| 10% | .119 [.090,.147] | .117 [.089,.139] | .090 [.072,.115] | .250 [.222,.299] | |
| Aegis jury | 5% | .042 [.019,.066] | .019 [.004,.037] | .053 [.029,.085] | .245 [.160,.318] |
| 10% | .081 [.055,.108] | .034 [.012,.063] | .117 [.081,.150] | .358 [.288,.476] |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.