cs.CRAug 18, 2026

Reflex-Guard: A Low-Latency Guardrail for LLM Prompt Safety Using Dense Semantic Embeddings

Authors: Istiaque Ahmed, Afia Anjum Borsha, Ranat Das Prangon, Abu-fuad Ahmad, Thi Hong Tran

Organizations: Graduate School of Informatics, Osaka Metropolitan University, Osaka 558-8585, Japan · Dept. of Computer Science and Engineering, BRAC University, Dhaka 1212, Bangladesh · Dept. of Chemical Engineering, Bangladesh University of Engineering and Technology (BUET), Dhaka 1000, Bangladesh · New Mexico State University, Las Cruces NM 88001, USA

Abstract

Large Language Models (LLMs) in real-world applications often face the risks of specially crafted prompts designed to bypass the safety controls. Existing guardrail methods, such as LLM-as-a-judge and cloud-based safety APIs are able to detect unsafe content. However, they often add a delay of about 250-900 ms to each request. This delay is too high for real-time applications, when the system usually needs to respond in less than 100 ms. Furthermore, routing user prompts through external moderation endpoints raises significant data privacy concerns. This paper introduces Reflex-Guard, a lightweight guardrail that runs locally. It uses jailbreak-aware preprocessing, compact sentence-transformer embeddings, and seven fast binary classifiers. Together, these components enable high-accuracy prompt safety filtering with much lower latency than existing solutions. Through systematic evaluation on a strategically balanced dataset of 30,568 samples drawn from five complementary sources, we demonstrate that Reflex-Guard achieves 95.9% recall on harmful prompts at 37.6 ms end-to-end latency. It is faster than existing baselines, including Llama Guard 2 at 255 ms and SafeDecoding at 723 ms. It can detect 100% of GCG suffix attacks and Base64-encoded prompts using the default threshold. However, DrAttack structured prompts required lowering the threshold to 0.03 for optimal detection, as they produced a distinct probability distribution. Reflex-Guard achieves Reflex Efficiency Score (RES) scores up to 16.79, significantly outperforming Llama Guard 2 (11.90) and SafeDecoding (9.80). This analysis offers practical deployment advice and shows that different attack types occupy distinct regions in the embedding probability space.

Figures & tables

Explore similar work

CardsList
  1. kNNGuard: Turning LLM Hidden Activations into a Training-Free Configurable Guardrail

    Jul 2, 2026Mahmoud Abdelfattah, Hamid Nasiri, Peter GarraghanLarge Language Model SafetyAdversarial Prompts

  2. Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection

    May 24, 2026Lixing Lin, Juli You, Yue Li +4Large Language Model SafetyAdversarial Prompts

  3. Robust and Efficient Guardrails with Latent Reasoning

    May 27, 2026Siddharth Sai, Xiaofei Wen, Muhao ChenLarge Language Model SafetyGuardrail