cs.LGSep 5, 2026

SAFEGuard: Detect Optimization-Based Jailbreak Attacks Through Harmful Semantic Analysis and Fluency Measurement

Authors: Quoc Viet Vo, Trung Le, Damith C. Ranasinghe, Ehsan Abbasnejad

Organizations: Australian Institute for Machine Learning · Adelaide University · Monash University

Abstract

Despite the significant efforts devoted to aligning large language models (LLMs) with human values and ensuring safe deployment, recent work has revealed that LLMs remain vulnerable to adversarial jailbreak attacks that can bypass safety guardrails and elicit harmful responses. Many defense methods are proposed to detect jailbreaks but they are limited in their effectiveness to counter wide-range optimization-based jailbreak mechanisms that can yield highly fluency-optimized or harmful semantic obfuscated prompts. To tackle this challenge, we propose a unified detection framework SAFEGuard which incorporates a hybrid fluency measurement based on cross-layer distribution distance and perplexity, and the analysis of harmful semantics through gradient matching. Our method is grounded in a paramount observation: high fluency prompts maintain their malicious intention close to harmful prompts while harmful semantic obfuscated prompts often inject gibberish token sequences. Our evaluation demonstrates that SAFEGuard consistently outperforms state-of-the-art baselines and achieves significant improvement in accuracy across different optimization-based jailbreaks. This underscores the effectiveness of SAFEGuard against evolving jailbreak attacks.

Explore similar work

CardsList
  1. SafeDream: Safety World Model for Proactive Early Jailbreak Detection

    Apr 18, 2026Bo Yan, Weikai Lin, Yada Zhu +1Language Model Safety EvaluationWorld Models

  2. Re-Triggering Safeguards within LLMs for Jailbreak Detection

    May 11, 2026Zheng Lin, Zhenxing Niu, Haoxuan Ji +2LLM Jailbreak AttacksJailbreak Detection

  3. The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring

    May 9, 2026Ismail Hossain, Tanzim Ahad, Md Jahangir Alam +3Adversarial Prompt GenerationLanguage Model Safety Evaluation