Fig. 5: Per-request gaps are cheap to close when labels are accurate; user-level inference is not. Each marker is RouteLLM on WildChat (judge-labeled medical vs. other requests, 50% operating point) with no defense or one post-processing defense; per-category parity on its own is orange and all other defenses blue. Lines join no defense and random mixing (
η=0.25,0.5 ; solid), the per-user budget bands (
±4→±2→±1→ exact rate; dashed), and parity (cheap classifier
→ perfect labels; dash-dot). Marker shape gives what deployment needs, as in Table VIII ; gray dashed lines in (b) mark chance and the undefended A1. The
x axis, skill kept, is the router’s own expected benefit relative to random routing at the same budget (0 = random, 1 = undefended). It is the router’s self-assessment, not measured quality: on RouterBench, undefended RouteLLM is 1.30 accuracy points worse than random routing at the same budget (Section VII-D ). (a) Per-category parity with perfect labels removes the per-request gap (
∣Δadj∣=0.001 ) at a skill of 0.99, but it is not deployable; with the cheap classifier the gap stays at 0.151. Whiskers are 95% user-cluster bootstrap CIs (1,000 resamples) shown as absolute values; the perfect-label interval spans zero, so its whisker runs from 0 to 0.028. (b) User-level AUC after
T=20 requests for the bill-only attacker A1 (filled; whiskers are 95% percentile bootstrap CIs over test users, 1,000 resamples) and the original position-agnostic itemized-log attacker A2 (hollow; CI not drawn); a solid segment joins each defense’s A1 and A2. The defense-study split has 185 positive and 2,891 negative test users; there the undefended A1 is 0.718, against 0.707 on the E-5 split. Only the exact per-user rate, with or without parity, brings A1 to chance (0.500, degenerate CI), at a skill of 0.23, and only because
T is even (at
T=5 : 0.63 and 0.59). Itemized logs still accumulate under it: A2 is 0.53–0.54 at
T=20 and rises to 0.60–0.62 at
T=50 (Table VIII ). A post hoc odd-position attacker, not plotted here, reaches 0.730 under the exact rate and 0.667 with cheap-classifier parity first at
T=20 (Table IX ). Status: the defense circled in red in (b), cheap-classifier parity plus a
±2 band, is the pre-registered primary endpoint and failed all four of its criteria; the undefended baseline and the other nine defenses are pre-registered secondary analyses; the exact-rate variants ( ∗ ) and the skill-kept measure are post hoc. As for every medical result, the labels come from an LLM judge that failed its pre-registered precision gate, so all results shown are exploratory (Section X ). Data: E-8 results.json and posthoc_e8.json .