cs.CRSep 18, 2026

CASCADE Against Jailbreaks: Combination Across Stages with Controlled Attack-Defense Evaluation

Authors: Jiale Luo, Eric Han

Organizations: School of Computing National University of Singapore

Abstract

Defenses against jailbreak attacks on Large Language Models (LLMs) operate at different pipeline stages, such as input modification or output guard, but it remains unclear which defenses to deploy at each stage and how to combine them. Prior empirical studies, fragmented by inconsistent attack-success-rate definitions and experimental settings, have evaluated defenses largely in isolation. Here we present the first systematic study, to our knowledge, of defense combinations both within and across pipeline stages, under a consistent threat model of direct, black-box, single-turn attacks. Our decision framework standardizes evaluation through a principled attack-success-rate formulation with controlled query budgets, together with explicit fairness rules. Across 19 attacks and 15 defenses, we find that no single defense is universally best, but well-chosen combinations achieve substantial safety with minimal utility degradation, yielding practical recommendations for layered defense pipelines.

Explore similar work

CardsList
  1. SoK: Robustness in Large Language Models against Jailbreak Attacks

    May 6, 2026Feiyue Xu, Hongsheng Hu, Chaoxiang He +9Language Model Safety EvaluationLLM Security

  2. When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs

    Jul 27, 2026Tong Zhang, Zexin Li, Simin Chen +1LLM Inference EfficiencyLanguage Model Safety Evaluation

  3. Breaking the Assumptions: Auditing Input-Side Jailbreak Defenses Against Semantic Attacks

    Aug 22, 2026Aaditya Pratap, Harsh Kasyap, Somanath TripathyJailbreak DefenseLLM Safety Evaluation