cs.CRAug 22, 2026

Breaking the Assumptions: Auditing Input-Side Jailbreak Defenses Against Semantic Attacks

Authors: Aaditya Pratap, Harsh Kasyap, Somanath Tripathy

Organizations: NCE Chandi Nalanda, India · IIT (BHU) Varanasi, India · IIT Patna Patna, India

Abstract

Locally deployed Large Language Models (LLMs) via inference engines such as Ollama run without the moderation and abuse detection present in API-served models. Therefore, the safety of LLMs depends on the defense mechanisms used, and their effectiveness depends on the assumptions on which they were designed. This paper does an audit of defense mechanisms under jailbreak attacks on locally deployed models. Some defenses provide formal guarantees (SmoothLLM, Erase-and-Check, Sequential Monitors), while others rely on empirical detection results (Semantic Smoothing, Self-Denoised Smoothing, Perplexity Filtering). Instead of merely observing that defenses fail, we trace each failure back to the specific assumption: for every defense, we extract the condition it relies on, derive the empirical pattern a violation should produce, and test that prediction on six open-weight models (14B to 35B parameters) with a corpus of 100 jailbreak prompts taken from more than 40 public sources, totalling 13,800 evaluation records.

Explore similar work

CardsList
  1. SoK: Robustness in Large Language Models against Jailbreak Attacks

    May 6, 2026Feiyue Xu, Hongsheng Hu, Chaoxiang He +9Language Model Safety EvaluationLLM Security

  2. SAFEGuard: Detect Optimization-Based Jailbreak Attacks Through Harmful Semantic Analysis and Fluency Measurement

    Sep 5, 2026Quoc Viet Vo, Trung Le, Damith C. Ranasinghe +1LLM SecurityJailbreak Attacks