cs.CRJun 1, 2026

Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittleness Under Paraphrasing

Authors: Alexandre Cristovão Maiorano

Abstract

Production LLM applications stack several defense families -- refusal-phrase filters, token-budget controls, model allowlists, rate limits, tool-registry authentication -- yet existing breach-and-attack-simulation (BAS) benchmarks report a single aggregate coverage number, hiding which family closes which threat. We measure attribution. We add four OWASP-LLM-Top-10-aware agents to a 21-agent baseline scanner and target a lattice of four synthetic LLM endpoints: L0L_0 (no defenses), L1L_1 (refusal-only), L2L_2 (budget-only), and L3L_3 (full stack). L1L_1 and L2L_2 are sibling single-axis ablations, not subsets of each other; L3L_3 is their union plus tool-registry authentication and credential scrubbing. Across N=10N=10 replications, the per-OWASP finding count is clean: refusal alone removes all LLM01 (jailbreak) and LLM07 (system-prompt leakage) findings; budget alone removes all LLM02 (sensitive-info disclosure) and LLM10 (unbounded consumption) findings by terminating multi-step sequences; LLM06 (excessive agency) requires the full stack. We probe brittleness under paraphrasing: with 300 Gemini-generated paraphrases (K=5K=5 over a 60-template brittleness corpus), L1L_1 refusal block rate falls 15 pp on LLM01 and 25 pp on LLM07. A fifth target, L4L_4-real, swaps the stub backend for Gemini-2.5-flash behind the same L3L_3 regex and matches L1L_1 exactly, indicating no measurable alignment contribution beyond the regex (not a general claim about alignment). Budget controls show no drop (0 pp once the rate-limit floor is factored out). A refusal whitelist that clears a static benchmark can be defeated by an LLM-driven paraphraser without changing attack intent; a budget control resists the same mutation.

Explore similar work

May 14, 2026cs.CR

Talk is (Not) Cheap: A Taxonomy and Benchmark Coverage Audit for LLM Attacks

We introduce a reusable framework for auditing whether LLM attack benchmarks collectively cover the threat surface: a 4×\times6 Target ×\times Technique matrix grounded in STRIDE, constructed from a 507-leaf taxonomy -- 401 data-populated and 106 threat-model-derived leaves -- of inference-time attacks extracted from 932 arXiv security studies (2023--2026). The matrix enables benchmark-external validation -- auditing collective coverage rather than individual benchmark consistency. Applying it to six public benchmarks reveals that the three primary frameworks (HarmBench, InjecAgent, AgentDojo) occupy non-overlapping cells covering at most 25% of the matrix, while entire STRIDE threat categories (Service Disruption, Model Internals) lack any standardized evaluation, despite published attacks in these categories achieving 46×\times token amplification and 96% attack success rates through mechanisms which no benchmark tests. The corpus of 2,521 unique attack groups further reveals pervasive naming fragmentation (up to 29 surface forms for a single attack) and heavy concentration in Safety & Alignment Bypass, structural properties invisible at smaller scale. The taxonomy, attack records, and coverage mappings are released as extensible artifacts; as new benchmarks emerge, they can be mapped onto the same matrix, enabling the community to track whether evaluation gaps are closing.
Karthik Raghu Iyer, Yazdan Jamshidi, Nicholas Bray +1
Jul 20, 2026cs.CR

Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security

LLM-based agents process external content, exposing them to prompt injection and multi-turn manipulation. Most safety benchmarks evaluate defenders against fixed attack pools collected before evaluation, single-turn or multi-turn. We present a 21-scenario benchmark for \emph{adaptive multi-round attacks against memoryless LLM defenders}: an autonomous LLM attacker observes prior defender responses and pivots across rounds, while each defender response is evaluated as a fresh interaction. Holding the 21 scenarios, attackers, defenders, and structured-output scoring fixed, restricting scoring to the first attacker turn yields 00-1%1\% attack success rate (ASR); allowing 15 rounds of adaptive attack yields 5.45.4-14.0%14.0\%. Pooling three frontier attacker LLMs uncovers 1.41.4-2.2×2.2\times as many unique successful attacks as the best single attacker, and the generated attacks have low cosine similarity (0.020.02-0.140.14) to attacks in existing benchmarks. Claude Opus 4.6 and GPT-5.4 are tied in aggregate (5.4%5.4\% each; overlapping 95%95\% CIs), but their weaknesses differ sharply: on one scenario Opus reaches 60%60\% ASR (95%95\% CI 3636--80%80\%) while GPT-5.4 and Gemini each stay at 7%7\% (CI 11-30%30\%; the gap is preserved in a higher-NN replication). 1313 of 2121 scenarios distinguish at least one defender pair, yet rankings disagree across scenarios (Kendall's W=0.19W = 0.19). We release the benchmark -- 21 evaluation scenarios, 10 public development scenarios, the orchestrator, baseline harnesses, and a multi-attacker CLI -- plus 945 transcripts from the 3×\times3 frontier matrix, an attack-replay dataset, and 18{,}422 gpt-oss-20b battles from an open competition's final scoring rounds.
Devina Jain, David Hartmann, Chuan Li
Jul 27, 2026cs.CR

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs

Jailbreak defenses are essential for protecting large language models (LLMs), but they can also introduce secondary costs that weaken model utility. We present a systematic study of these defense trade-offs along three dimensions: performance impact, over-refusal on benign inputs, and inference cost. Rather than treating defenses as a single class, we organize them by operational strategy and examine how different strategies correlate with different side-effect profiles. Across state-of-the-art defense methods, widely used benchmark datasets, and representative open-source LLMs, we find that defenses rarely improve downstream capability, but instead vary in how they trade safety gains against usability and efficiency. In particular, rule-based defenses best preserve task performance, highly conservative self-reflective defenses often increase over-refusal, and multi-round defenses incur the largest runtime overhead. These results provide both a benchmark for evaluating defense side effects and practical guidance for selecting defenses under deployment constraints.
Tong Zhang, Zexin Li, Simin Chen +1