cs.CRSep 6, 2026

SRD-GUARD: A Defense Framework of LLMs via Semantic Rewriting and Joint Multi-Model Scoring for Latent Intent Exposure

Authors: Qi Wang, Chengcheng Wan, Jiangtao Wang

Organizations: East China Normal University Shanghai, China · East China Normal University, Shanghai Innovation Institute Shanghai, China · Software Engineering Institute, East China Normal University Shanghai, China

Abstract

Large language models (LLMs) are increasingly deployed in safety-critical applications, yet jailbreak attacks can conceal harmful intent through role-playing, fictional scenarios, or seemingly benign motivations. Existing inference-time defenses may miss disguised attacks or excessively refuse legitimate requests. We propose SRD-GUARD, a parameter-free, black-box defense framework that exposes concealed intent through semantic rewriting and consensus-based risk assessment. Given an input prompt, SRD-GUARD generates five semantically related rewrites that preserve the underlying objective while removing unnecessary contextual packaging. The original prompt and rewrites are jointly evaluated by multiple independent LLM-based safety scorers on a continuous risk scale. A decision module combines absolute risk thresholds with relative risk changes between the original and rewritten prompts to adaptively intercept, preserve, or warn on requests. We evaluate SRD-GUARD against UNIATTACK, CIPHER, and DeepInception on Llama-3-8B-Uncensored and DeepSeek-V4-Flash using AdvBench and OR-Bench-Hard. SRD-GUARD achieves average DSRs of 91.44% and 100%, with ORRs of 8.00% and 12.00%, respectively. Compared with evaluated baselines, it provides a more favorable DSR--ORR trade-off. Ablation studies show that rewriting exposes concealed harmful intent, joint scoring improves robustness to individual evaluator behavior, and risk-adaptive decision making enables selective handling of ambiguous inputs. These results demonstrate that semantic intent exposure, consensus-based risk assessment, and relative-risk-aware routing provide an effective and selective approach to black-box jailbreak defense. The artifact is available at https://anonymous.4open.science/status/CICD-Guard-D648.

Explore similar work

CardsList
  1. Reactivating Alignment: Defending LLMs from Jailbreaks via Intention-Aware Input-Output Matching

    Oct 3, 2026Luoyu Chen, Weiqi Wang, Chenhan Zhang +2Jailbreak DefenseLLM Safety Alignment

  2. SAFEGuard: Detect Optimization-Based Jailbreak Attacks Through Harmful Semantic Analysis and Fluency Measurement

    Sep 5, 2026Quoc Viet Vo, Trung Le, Damith C. Ranasinghe +1LLM SecurityJailbreak Attacks

  3. Breaking the Assumptions: Auditing Input-Side Jailbreak Defenses Against Semantic Attacks

    Aug 22, 2026Aaditya Pratap, Harsh Kasyap, Somanath TripathyJailbreak DefenseLLM Safety Evaluation