cs.CROct 1, 2026

False Floors: LLM Safety Routing Evaluations Break Under Distribution Shift

Authors: Amit Singh Bhatti, Vishal Vaddina

Organizations: Quantiphi Analytics

Abstract

Safety routers send each request to one of several models and are judged against the best single model. A major routing benchmark picks that comparator on the evaluation data. In the benchmark's own setting this is harmless, but under distribution shift it is not. On HELM Safety the selection cost is 0.003-0.030 of harm under random splits and 0.045-0.113 under held-out categories, comparable to the whole deficit attributed to routing, with its direction holding under either published judge alone. It rises seven- to ninefold on AgentDojo when suites are held out. Across seven safety corpora chosen by rules fixed in advance, three meet a registered interval test and four beat a later permutation null, and three of the four interval misses are corpora where some models have zero observed harm. Prior work proves the direction of this bias. We size it on harm and accuracy, show that it is larger under the held-out splits we measure, and bound it by optimism plus a shift-dependent regret. Scored honestly under shift, routing buys little on these benchmarks. In most pool cells the nested router serves the honest baseline's model, and on the nearly saturated AgentDojo corpus a perfect pre-dispatch router is worth at most two points of harm. We also find a model's expressed recognition of a late injection steerable. On held-out reruns an attacker who knows which model it faces lowers GPT-5.4's judged recognition by 19.6 points, confirmed by an independent label. In an offline counterfactual composition into a controller, the same attack raises or lowers estimated harm depending on the fallback model. Safety routing should be evaluated under shift, against a baseline chosen without the test labels, and recognition-based defences should be scored on harm against an attacker who chooses what the model sees.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Safety-Oriented Routing Analysis of Mixtral MoE Under Benign and Harmful Prompts

    May 22, 2026Md Nurul Absar SiddikyMixture-Of-Experts Large Language ModelsLarge Language Model Routing

  2. When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning

    Sep 1, 2026Yitong Guo, Xiaoyi Chen, Siyuan Zhang +2FragilityModel Fine-Tuning

  3. RouteHijack: Routing-Aware Attack on Mixture-of-Experts LLMs

    May 1, 2026Zhiyuan Xu, Joseph Gardiner, Sana Belguith +1Jailbreak Success RatesMixture-Of-Experts