Safety routers send each request to one of several models and are judged against the best single model. A major routing benchmark picks that comparator on the evaluation data. In the benchmark's own setting this is harmless, but under distribution shift it is not. On HELM Safety the selection cost is 0.003-0.030 of harm under random splits and 0.045-0.113 under held-out categories, comparable to the whole deficit attributed to routing, with its direction holding under either published judge alone. It rises seven- to ninefold on AgentDojo when suites are held out. Across seven safety corpora chosen by rules fixed in advance, three meet a registered interval test and four beat a later permutation null, and three of the four interval misses are corpora where some models have zero observed harm. Prior work proves the direction of this bias. We size it on harm and accuracy, show that it is larger under the held-out splits we measure, and bound it by optimism plus a shift-dependent regret. Scored honestly under shift, routing buys little on these benchmarks. In most pool cells the nested router serves the honest baseline's model, and on the nearly saturated AgentDojo corpus a perfect pre-dispatch router is worth at most two points of harm. We also find a model's expressed recognition of a late injection steerable. On held-out reruns an attacker who knows which model it faces lowers GPT-5.4's judged recognition by 19.6 points, confirmed by an independent label. In an offline counterfactual composition into a controller, the same attack raises or lowers estimated harm depending on the fallback model. Safety routing should be evaluated under shift, against a baseline chosen without the test labels, and recognition-based defences should be scored on harm against an attacker who chooses what the model sees.
Figures & tables
Fig. 1: The baseline, not the router, decides the verdict. Phase 1, evaluation. One fitted router is scored against two fixed-model comparators on identical folds (E24, union label). The fold protocol shows where each comparator is chosen, the honest pin on the four training folds, the in-sample pin on the held-out fold it is then scored on. Against the in-sample pin the router never beats it. Against the honest pin it ties in the median and is ahead on average. The gap is the selection cost, 0.0430.113 , whose sign is guaranteed (Observation 1) and whose size depends on the judge (E24j). Phase 2, deployment. The router sees only the request text and decides once, before any tool output exists. What routing can buy, each on its own surface, is small. The fitted router serves the honest pin’s own model on every held-out request in 91true% of two-model pool cells, a perfect pre-dispatch router on AgentDojo stays within two points of harm (E43), and the best honest cascade is one cheap model (E34). After the router commits, an injection arrives in a tool result and defence rests on the model’s recognition. On held-out reruns a targeted template lowers gpt-5.4 ’s judged recognition by 19.6 points, confirmed by an independent label (E54b, E55). Four action-level policy settings record zero judged successes on one shared set of episodes.
corpus
raw release
analysis set
HELM Safety [ 32 ]
44 models ×393 behaviours, 2 judges
393 behaviours, 7 categories
XSTest [ 46 ]
450 prompts
250 safe, 200 unsafe
HarmBench [ 40 ]
54581 cells, 16×29
complete 19×16 subgrid
AgentDojo [ 10 ]
33119 runs, 28 configs, 15 attacks
478 scenarios; 949 on the 5 complete
skill injection [ 47 ]
1862 episodes
1342 undefended; 130 defended per defence
TABLE I: The four corpora and the analysis set every number is computed on. XSTest enters as HELM Safety’s over-refusal scenario, and its human-labelled release also supplies one row of Table III . Analysis sets are smaller than the raw releases because a pool comparison is defined only where every model is scored on every item; that complete-case filter is what removes the difference.
Fig. 2: Baseline choice changes the verdict. All panels use within-fold in-sample pins, panel A from E24 (chat) and E27 (agentic, the lower line), panels B and C from E26 on its own grid. (A) The harm advantage a pin gains purely from being selected on the evaluation data rather than on training folds. (B) The conditional tail edge changes sign once the pin is selected honestly, which dissolves an apparent chat-versus-agentic reversal. (C) The AUROC a router must reach at full strength ( λ=1 ), by pin protocol. The shrunken arm is reported in the text rather than drawn.
k
2
3
5
44
pin (oracle)
0.1873
0.1278
0.0654
0.0000
pin (honest)
0.2299
0.1951
0.1622
0.1127
winner’s curse
+0.0426
+0.0674
+0.0968
+0.1127
router − pin (oracle)
+0.0355
+0.0536
+0.0732
+0.0819
router − pin (honest)
−0.0071
−0.0138
−0.0236
−0.0308
TABLE II: HELM harm on identical category-held-out folds (seed 24 ). Means over 200 pools ×5 folds for k≤5 ; k=44 is one pool ×5 folds. The in-sample pin is selected on each evaluation fold, the honest pin on training folds; λ is tuned on training data. Table IV instead uses a full-sample pin and λ=1 .
Fig. 3: The comparator’s cost appears under shift, on both safety corpora. Selection cost in harm under a random split and with whole HELM categories or AgentDojo suites held out (E25, E25a).
corpus
models × items
groups
held out
random
difference, 95true%
null p
headroom
R
HELM harm_bench
44×393
7
0.0454
0.0026
[0.0329,0.0491]
0.005
0.0797
0.57
SORRY-Bench
43×44
4
0.0076
0.0036
[0.0007,0.0098]
0.005
0.0192
0.40
AIR-Bench 2024 (E61)
87×5694
16
0.0053
0.0002
[0.0048,0.0055]
0.005
0.0431
0.12
HarmBench
19×2210
16
0.0019
0.0006
[-0.0002,0.0023]
0.005
0.0617
0.03
HELM xstest
44×200
8
0.0044
0.0022
[-0.0013,0.0035]
0.010
0.0084
0.52
HELM simple_safety_tests
44×99
5
0.0064
0.0074
[-0.0041,0.0047]
0.343
0.0106
0.60
TABLE III: Selection cost at k=2 on every eligible corpus (E58). The registered criterion, an interval above zero, is met on the first three rows. The group-permutation null, added after verification, is beaten at p=0.005$$ on the first four, at 0.010 on HELM xstest , where it depends on the pool draw, and not on the last two. AIR-Bench was registered separately and its row enumerates every pair of models (E62), except its permutation p , which comes from E61’s sampled pairs, while the other rows average 200 sampled pairs per fold. The null shuffles group labels at the same group and fold sizes. Headroom is the held-out harm of the honest pin minus that of a perfect per-item router over the pool, and R is the selection cost over headroom, both with groups held out (E59).
Fig. 4: Four axes of routing fragility, none of them accuracy. (A) Policy class. Shrinking a per-request router toward the pool marginal lowers the AUROC it needs by 0.147 under an in-sample pin, which is not itself a deployable threshold. (B) Operating point. At one fixed AUROC, changing only the shape of the score distribution flips the sign of excess harm. (C) Difficulty. Across bands of empirical difficulty the conditional tail edge falls, under either pin, so deferral is worst on the requests most models fail. (D) Attack template. Pooled flagging varies by an order of magnitude across templates. Templates are not scenario matched, so template and domain are confounded here. (E) Protocol sensitivity. The upper scorecard compares protocols. Each lower waffle cell is one percentage point of the gap to the in-sample pin. STAT is the stationary protocol and SHIFT the category-held-out one, and labels give the ratio. The STAT/SHIFT comparison is a protocol contrast rather than an additive decomposition.
k
router
pin
random
oracle
r. − pin
beats pin
2
0.2241
0.2157
0.3317
0.1572
+0.0084
4.0%
3
0.1904
0.1662
0.3437
0.0906
+0.0243
3.5%
5
0.1545
0.1062
0.3268
0.0293
+0.0483
2.5%
44
0.0840
0.0331
0.3363
0.0000
+0.0509
0 of 1
TABLE IV: HELM harm under out-of-fold, full-strength routing ( λ=1 ). The pin is selected on all 393 evaluation behaviours; the oracle is the per-request minimum. “Beats pin” covers 200 pools for k≤5 and one full pool at k=44 . Table II uses a different pin protocol.
setting
k
conditional edge vs.
median, deferring
in-sample pin
honest pin
router − pin
chat (HELM)
44
−0.1023
+0.0770
−0.0770
agentic
28
−0.0149
n/a
n/a
TABLE V: Conditional tail edge, measured only when the router defers. The chat comparison retains 2 / 5 folds ( 146 / 393 behaviours). The agentic comparison retains 0 / 5 : at the full pool the router issues the honest pin’s decision on every held-out scenario, so nothing is deferred and the edge is undefined rather than zero. Unconditionally the agentic figure is 0.0000 , the router being the pin by construction in all five folds.
attack family
cfg
scen.
fixed
oracle
headroom
important_instr.
5
949
0.0211
0.0053
0.0158
direct
4
949
0.0126
0.0074
0.0053
ignore_previous
4
949
0.0137
0.0095
0.0042
harm
utility
marginal pin ( λ=0 )
0.0238
0.8115
honest pin
0.0232
0.8093
TABLE VI: AgentDojo by attack family, without cross-family complete-case filtering. Headroom is fixed-pin harm minus per-request oracle harm. In the lower block nested tuning chooses λ=0 in all five folds: the marginal pin is a fixed model selected from mean predicted scores, not a request-level router; its per-request AUROC 0.8839 is not the performance of the scored fixed policy. Fitted without sight of the evaluation scenarios and read in canonical order it selects the honest pin’s own configuration in every fold, and the figures average permutations, one fold tying (E48).
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Fig. 5: Swapping only the judge changes which model is safest. Two independently published annotators on identical instances. Kendall τ=0.859 on harm_bench , where the safest model still changes and the judges share 3 of the top 5, and 0.498 on the over-refusal axis, where they share 1 of 5.
corpus
grid
observed interaction
no-interaction null (90%)
HarmBench
19×16
0.1854
[0.0580, 0.0715]
SORRY-Bench
43×44
0.2025
[0.1624, 0.1759]
Appendix
TABLE VII: Interaction variance against a purely additive (zero-interaction) null. The interaction is real on one corpus and largely sampling noise on the other.
Fig. 6: Accuracy alone does not decide whether routing helps. Excess harm of a full-strength router over an in-sample pin, against signal accuracy and score asymmetry κ=σ+/σ− , at pool size k=3 . Dots are measured cells and the surface between them is interpolated. Zero is break-even, and above it the router is worse than pinning one model. The key reading is at a fixed AUROC of 0.80 , where moving asymmetry from 0.25 to 4 carries the router from winning to losing while its accuracy never changes.
AUROC
κ=0.25
κ=1
κ=4
spread
vs. a 0.05 - AUROC move
0.70
+0.0319
+0.0738
+0.1152
0.0833
3.13 ×
0.80
−0.0095
+0.0233
+0.0548
0.0643
2.41 ×
0.85
−0.0276
−0.0034
+0.0221
0.0497
1.87 ×
Appendix
TABLE VIII: At fixed AUROC, varying only the asymmetry κ=σ+/σ− of the class-conditional score distributions over a 16× range, at k=3 . κ changes the shape of the ROC curve at fixed area; the router is an argmin over models and applies no threshold, so it does not select an operating point and no row describes a decision rule. The last column compares two analyst-chosen ranges, a 16× change in κ against a 0.05 change in AUROC at κ=1 , and is not scale-free. Deficits are means over 150 pools ×30 repetitions against a full-sample in-sample pin; per-cell dispersion is large relative to the cells themselves (SD 0.0560.074 ) and no interval is reported. The spread across κ is invariant to the baseline choice, since the pin does not depend on κ ; the sign of individual cells is not.
Fig. 7: For raw scores the break-even contour is nearly vertical. Deficit over signal accuracy × pool size, with marginals. In E4b’s raw-score arm against a full-sample in-sample pin, the zero contour runs almost parallel to the pool-size axis, so here the accuracy a router needs barely depends on pool size. E4b’s rank-calibrated arm needs more, falling with k (Table XVII ). Over the swept accuracy range the mean deficit does grow with pool size ( 0.0246 at k=2 , 0.0360 at k=3 , 0.0482 at k=5 , 0.0485 at k=44 ), which is the right-hand marginal. At fixed accuracy it is not monotone in k . k=44 sits below k=5 from 0.60 to 0.85 , and above break-even every cell changes sign. The right-hand marginal thus averages a deficit below the contour against an advantage above it.
in-sample pin
honest pin
base-rate band
n
base rate
tail edge
beats pin
tail edge
beats pin
0.00–0.15
108
0.082
−0.0159
0.3%
−0.0071
11.0%
0.15–0.30
67
0.221
−0.0689
1.3%
−0.0425
10.7%
0.30–0.50
113
0.397
−0.2007
1.3%
−0.1726
4.7%
0.50–0.70
90
0.574
−0.3325
1.0%
−0.3325
3.7%
Appendix
TABLE IX: Empirical difficulty is the third axis and it runs the wrong way, under either baseline convention. Bands are cut on each request’s mean harm across the pool, its base rate, so they compare requests of different observed difficulty rather than one deployment at different prevalence. Pool size k=3 , 300 random pool draws per band; n is the number of scenarios in the band, and each tail edge is a mean over the draws in which the router defers at least once. The correlation between band base rate and tail edge is r=−0.99 (in-sample pin) and $-$0.98 (honest pin), but it is a Pearson coefficient over the 4 band means, with no null and no interval, so the monotonicity is the finding and the coefficient is not. The honest pin on this axis is selected on a random half of each band’s scenarios, not on a category-grouped split, which makes it the weaker of our two out-of-sample protocols; in the top band the two conventions agree to four decimals, and the stored aggregates do not record whether the two pins ever differed there. The “beats pin” columns, fractions of the 300 draws, are where the conventions part company: the claim that no band has the router beating the pin holds only against the in-sample pin.
Fig. 8: At request time, the router scores x and commits to model m∗ before seeing tool output. In Phase 2, the agent plans, calls tools, reads lower-trust results, interprets them, and acts across N model and tool steps without rerouting. Red arrows show attacker-controlled instructions entering a tool result, and teal shading marks where that result enters model context. If source text is treated as a command, the next action can change. A1’s pre-routing steering (Section II-A ) is outside this picture. The lower panels show three content-dependent candidate defences (one evaluated, one priced, one composed into controllers, E56b) and four policy settings with 0/130 judged successes each under one automated judge. Those results motivate the proposed external action gate (Section VII ).
decision point
AUROC
pre-execution, prompt only (28-config grid)
0.647
mid-trajectory, before the injection is visible
0.722
the step injected content enters context
0.808
one step later
0.804
tasks and attacks held out, before
0.705
tasks and attacks held out, at injection
0.703
Appendix
TABLE X: What the decision point can see. AUROC within model. The three in-loop rows hold the injection step fixed ( n=872 runs, 5 models); the pre-execution row is a different population, the 28 -configuration complete-case grid over 478 scenarios, so the within-population comparison is 0.722 against 0.808 . These rows measure whether the signal is available at each decision point. No agent was re-executed with a different model mid-run, which this corpus cannot support. With user tasks and attack families held out together (E18d) the two steps score 0.705 and 0.703 , so the signal is available but its rise is not robust.
featuriser
AUROC
[p05, p95] ∗
length only (non-semantic)
0.5264
[0.327, 0.777]
MiniLM ( 22 M)
0.6060
[0.500, 0.710]
TF-IDF (1–2 gram)
0.6509
[0.499, 0.797]
Qwen3-Embedding ( 596 M)
0.6565
[0.520, 0.772]
Qwen3-Embedding ( 4 B)
0.6831
[0.508, 0.835]
Qwen3-Embedding ( 8 B)
0.6901
[0.529, 0.794]
Appendix
TABLE XI: Encoder ablation and scale sweep on identical category-grouped splits (E23, E41). Sentence embeddings at small scale are worse than TF-IDF, and the honest-pin bar of 0.6842 at full strength is cleared by only the 8 B arm, and thinly. Precision is not matched across the sweep: the local arms are fp32, the 4 B arm’s quantisation is not declared by its provider and the 8 B arm is fp8, and the single crossing is thin. ∗ The bracketed range is the 5 th– 95 th percentile of the per-model AUROC across the 44 models: a dispersion across models, not a confidence interval on the mean.
safety floor
α=0
α=0.25
α=0.5
α=1
≤0.0021
+20.3%
+43.3%
no cascade
no cascade
≤0.0042
+35.2%
+43.3%
+58.7%
no cascade
≤0.0105
+0.0%
+0.0%
+0.0%
+0.0%
≤0.0200
+0.0%
+0.0%
+0.0%
+0.0%
≤0.0500
+0.0%
+0.0%
+0.0%
+0.0%
Appendix
TABLE XII: The full adversary sweep. Cost saving against the cheapest single model that meets the floor, for every safety floor and every adversary strength α we ran. “no cascade” means no cascade meets the floor at any price. The α∈{0.25,0.5} columns are reported here rather than in the body because they are not readable without the artefact explanation in the surrounding text: the saving they show is a saving conditional on the floor still being met, and the adversary they model forces routing on a random rather than a chosen subset.
attacker
what it knows
attack success
flag rate
static
nothing, fixed template
0.7027
0.0624
adaptive
one template for the pool, chosen on train
0.7018
0.0631
targeted
which model it faces , chosen per model on train
0.8062
0.0225
Appendix
TABLE XIII: An attacker that adapts on held-out reruns. Templates are selected on half of the reruns and scored on the rest, pooled over 200 splits of 130 rerun indices across 5 models. The corpus has no task identifier, so held-out means held-out reruns of the same configurations, not new tasks. The two adaptive arms differ only in what the attacker knows. Choosing a single template for the whole pool gains nothing, though the static control is claude_v35 , the corpus-modal template, which is also the corpus’s most successful, so that arm is scored against an already-optimal control. Choosing per model lowers the flag on gpt-5.4 alone ( claude-sonnet-4-6 flags at the same rate in both arms, three models do not flag in either compared arm) and raises success by 10.35 points unmatched. Matched on run configuration the flag fall holds under every weighting and the success gain does not (Appendix C-G ), and the pooled rates are unweighted means over the five models.
Fig. 9: Through a flag-gated controller, the fallback decides. (A) Change in system harm against the static template under escalate-on-silence, for an attacker that knows only the target and one that also knows the fallback, in four matching constructions, with rerun-bootstrap 95true% intervals (E56b). (B) The target-only change with each fallback left out, one marker per construction. Only leaving out claude-sonnet-4-6 flips the sign. (C) Attack success over all undefended episodes by template, for the three templates every model has. claude-sonnet-4-6 is vulnerable to the static template alone.
Fig. 10: Four settings at zero, consistent with recognition. Each of four action-level settings records zero judged successes on its observed support, on the run cells and models they were actually run on, under one automated judge. An external gate can apply trusted policy to untrusted arguments while injected instructions cannot change its authorisation, but mechanical enforcement of these four could not be verified from the released harness, and the episodes are consistent with recognition, so that property is not shown for them.
defended cell
defended n
undefended success ( n )
sonnet-4-6 × v32
30
— (0)
sonnet-4-6 × v35
30
0.988 (80)
gpt-5.4 × v39
10
0.633 (30)
gpt-5.4-mini × v32
30
0.000 (30)
gpt-5.4-mini × v35
30
0.400 (70)
Appendix
TABLE XIV: The matched comparator of Table XVI , resolved per (model, template) cell, identically for all four defences. 30 defended episodes have no undefended comparator and 30 sit in a cell whose undefended success is exactly 0.000 , leaving 70 of 130 informative about a reduction. Per-cell model × template breakdown from e38_defence_matched.json:defences.*.per_cell .
attack template
n
flag rate
attack success
claude_v41
70
1.000
0.000
claude_v53
20
1.000
0.000
claude_v39
100
0.990
0.000
claude_v54
20
0.450
0.000
claude_v35
90
0.067
0.878
Appendix
TABLE XV: Flagging and attack success for claude-sonnet-4-6 across five templates (no v32 runs). Templates are not scenario matched, so template and domain are confounded; v35 has both the lowest flag rate and the highest success.
defence
n
attack success
normal (all models)
1342
0.5052
normal (matched to defended runs)
360
0.4722
ask_user
130
0.0000
no_network
130
0.0000
script_audit
130
0.0000
two_pass
130
0.0000
Appendix
TABLE XVI: Skill-injection attack success by policy setting. Each defended setting has 130 episodes on three of five models. The 360 -run undefended comparator matches scenario and model coverage, but not execution mode. Matching execution mode as well leaves 10 informative episodes and a Clopper–Pearson upper bound of 0.308 on residual judged success. Zeros depend on one automated judge, and mechanical enforcement was not verified.
id
question
headline
status
E0
headroom at full pool
vacuous at M=44
superseded
E0b
deployable- k vs 1-D null
6/6 cells no excess
kept
E1
grader swap
safest model changes on 2/4 scenarios
kept
E2
Gate 0: predictability + routing
AUROC 0.6509; in-sample-pin comparison
superseded
E3
interaction vs additive null
real on one corpus, 84% noise on another
kept
E4b
threshold a∗ , raw arm
0.843/0.845/0.843/0.821 (calibrated 0.946–0.867)
kept †
Appendix
TABLE XVII: Experiment register. Superseded and void runs are listed rather than dropped.