False Floors: LLM Safety Routing Evaluations Break Under Distribution Shift
Organizations: Quantiphi Analytics
Abstract
Safety routers send each request to one of several models and are judged against the best single model. A major routing benchmark picks that comparator on the evaluation data. In the benchmark's own setting this is harmless, but under distribution shift it is not. On HELM Safety the selection cost is 0.003-0.030 of harm under random splits and 0.045-0.113 under held-out categories, comparable to the whole deficit attributed to routing, with its direction holding under either published judge alone. It rises seven- to ninefold on AgentDojo when suites are held out. Across seven safety corpora chosen by rules fixed in advance, three meet a registered interval test and four beat a later permutation null, and three of the four interval misses are corpora where some models have zero observed harm. Prior work proves the direction of this bias. We size it on harm and accuracy, show that it is larger under the held-out splits we measure, and bound it by optimism plus a shift-dependent regret. Scored honestly under shift, routing buys little on these benchmarks. In most pool cells the nested router serves the honest baseline's model, and on the nearly saturated AgentDojo corpus a perfect pre-dispatch router is worth at most two points of harm. We also find a model's expressed recognition of a late injection steerable. On held-out reruns an attacker who knows which model it faces lowers GPT-5.4's judged recognition by 19.6 points, confirmed by an independent label. In an offline counterfactual composition into a controller, the same attack raises or lowers estimated harm depending on the fallback model. Safety routing should be evaluated under shift, against a baseline chosen without the test labels, and recognition-based defences should be scored on harm against an attacker who chooses what the model sees.
Figures & tables
| corpus | raw release | analysis set |
|---|---|---|
| HELM Safety [ 32 ] | models behaviours, judges | behaviours, categories |
| XSTest [ 46 ] | prompts | safe, unsafe |
| HarmBench [ 40 ] | cells, | complete subgrid |
| AgentDojo [ 10 ] | runs, configs, attacks | scenarios; on the complete |
| skill injection [ 47 ] | episodes | undefended; defended per defence |
| 2 | 3 | 5 | 44 | |
| pin (oracle) | 0.1873 | 0.1278 | 0.0654 | 0.0000 |
| pin (honest) | 0.2299 | 0.1951 | 0.1622 | 0.1127 |
| winner’s curse | ||||
| router pin (oracle) | ||||
| router pin (honest) |
| corpus | models items | groups | held out | random | difference, | null | headroom | |
|---|---|---|---|---|---|---|---|---|
| HELM harm_bench | 7 | 0.03290.0491 | ||||||
| SORRY-Bench | 4 | 0.00070.0098 | ||||||
| AIR-Bench 2024 (E61) | 16 | 0.00480.0055 | ||||||
| HarmBench | 16 | -0.00020.0023 | ||||||
| HELM xstest | 8 | -0.00130.0035 | ||||||
| HELM simple_safety_tests | 5 | -0.00410.0047 |
| router | pin | random | oracle | r. pin | beats pin | |
|---|---|---|---|---|---|---|
| 2 | 0.2241 | 0.2157 | 0.3317 | 0.1572 | 4.0% | |
| 3 | 0.1904 | 0.1662 | 0.3437 | 0.0906 | 3.5% | |
| 5 | 0.1545 | 0.1062 | 0.3268 | 0.0293 | 2.5% | |
| 44 | 0.0840 | 0.0331 | 0.3363 | 0.0000 | 0 of 1 |
| setting | conditional edge vs. | median, deferring | ||
|---|---|---|---|---|
| in-sample pin | honest pin | router pin | ||
| chat (HELM) | 44 | |||
| agentic | 28 | n/a | n/a | |
| attack family | cfg | scen. | fixed | oracle | headroom |
|---|---|---|---|---|---|
| important_instr. | 5 | 949 | |||
| direct | 4 | 949 | |||
| ignore_previous | 4 | 949 | |||
| harm | utility | ||||
| marginal pin ( ) | |||||
| honest pin | |||||
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| corpus | grid | observed interaction | no-interaction null (90%) |
|---|---|---|---|
| HarmBench | 0.1854 | [0.0580, 0.0715] | |
| SORRY-Bench | 0.2025 | [0.1624, 0.1759] |
| AUROC | spread | vs. a - AUROC move | |||
|---|---|---|---|---|---|
| 0.70 | 0.0833 | 3.13 | |||
| 0.80 | 0.0643 | 2.41 | |||
| 0.85 | 0.0497 | 1.87 |
| in-sample pin | honest pin | |||||
|---|---|---|---|---|---|---|
| base-rate band | base rate | tail edge | beats pin | tail edge | beats pin | |
| 0.00–0.15 | 108 | 0.082 | 0.3% | 11.0% | ||
| 0.15–0.30 | 67 | 0.221 | 1.3% | 10.7% | ||
| 0.30–0.50 | 113 | 0.397 | 1.3% | 4.7% | ||
| 0.50–0.70 | 90 | 0.574 | 1.0% | 3.7% | ||
| decision point | AUROC |
|---|---|
| pre-execution, prompt only (28-config grid) | 0.647 |
| mid-trajectory, before the injection is visible | 0.722 |
| the step injected content enters context | 0.808 |
| one step later | 0.804 |
| tasks and attacks held out, before | 0.705 |
| tasks and attacks held out, at injection | 0.703 |
| featuriser | AUROC | [p05, p95] ∗ |
|---|---|---|
| length only (non-semantic) | 0.5264 | [0.327, 0.777] |
| MiniLM ( M) | 0.6060 | [0.500, 0.710] |
| TF-IDF (1–2 gram) | 0.6509 | [0.499, 0.797] |
| Qwen3-Embedding ( M) | 0.6565 | [0.520, 0.772] |
| Qwen3-Embedding ( B) | 0.6831 | [0.508, 0.835] |
| Qwen3-Embedding ( B) | 0.6901 | [0.529, 0.794] |
| safety floor | ||||
|---|---|---|---|---|
| no cascade | no cascade | |||
| no cascade | ||||
| attacker | what it knows | attack success | flag rate |
|---|---|---|---|
| static | nothing, fixed template | ||
| adaptive | one template for the pool, chosen on train | ||
| targeted | which model it faces , chosen per model on train |
| defended cell | defended | undefended success ( ) |
|---|---|---|
| sonnet-4-6 v32 | 30 | — (0) |
| sonnet-4-6 v35 | 30 | 0.988 (80) |
| gpt-5.4 v39 | 10 | 0.633 (30) |
| gpt-5.4-mini v32 | 30 | 0.000 (30) |
| gpt-5.4-mini v35 | 30 | 0.400 (70) |
| attack template | flag rate | attack success | |
|---|---|---|---|
| claude_v41 | 70 | 1.000 | 0.000 |
| claude_v53 | 20 | 1.000 | 0.000 |
| claude_v39 | 100 | 0.990 | 0.000 |
| claude_v54 | 20 | 0.450 | 0.000 |
| claude_v35 | 90 | 0.067 | 0.878 |
| defence | attack success | |
|---|---|---|
| normal (all models) | 1342 | 0.5052 |
| normal (matched to defended runs) | 360 | 0.4722 |
| ask_user | 130 | 0.0000 |
| no_network | 130 | 0.0000 |
| script_audit | 130 | 0.0000 |
| two_pass | 130 | 0.0000 |
| id | question | headline | status |
| E0 | headroom at full pool | vacuous at | superseded |
| E0b | deployable- vs 1-D null | 6/6 cells no excess | kept |
| E1 | grader swap | safest model changes on 2/4 scenarios | kept |
| E2 | Gate 0: predictability + routing | AUROC 0.6509; in-sample-pin comparison | superseded |
| E3 | interaction vs additive null | real on one corpus, 84% noise on another | kept |
| E4b | threshold , raw arm | 0.843/0.845/0.843/0.821 (calibrated 0.946–0.867) | kept † |