MOMAT: Mixture of Multiple Atlases for Low-Power Jailbreak Defense of Quantized LLMs
Authors: Boyang Li, Bingyu Shen, Weihao Hong, Zhiyuan Jiang, Xinlei Guan, Yan Ma, Miles Q. Li, Yi Sheng, +1 more
Organizations: Department of Computer Science, Kean University, 1000 Morris Ave, Union, 07083, NJ, USA · Department of Computer Science and Engineering, University of Notre Dame, 257 Fitzpatrick Hall of Engineering, Notre Dame, 46556, IN, USA · McGill University, 845 Sherbrooke Street, Montréal, H3A 0G4, Québec, Canada · Department of Computer Science and Engineering, University of South Florida, 4202 E Fowler Ave, Tampa, 33620, FL, USA · Department of Computing Sciences, Villanova University, 800 E Lancaster Ave, Villanova, 19085, PA, USA
Quantized large language models are increasingly deployed on edge devices for their low latency and energy efficiency. However, model quantization weakens alignment safeguards, leaving qLLMs (quantized large language models) highly vulnerable to jailbreak attacks. To address this challenge, we present MOMAT (Mixture of Multiple Atlases), a hardware-enhanced safety framework that combines structured knowledge retrieval with low-power defense acceleration. Each atlas represents a semantic cluster of harmful or benign sample sets and policy templates, enabling domain-localized Retrieval-Augmented Generation guarding that mitigates the curse of dimensionality and the resulting semantic sparsity problem in large, heterogeneous safety databases. MOMAT retrieves top-k similarity features from all atlases for each prompt and evaluates them using a lightweight MoE (Mixture of Experts) detector, while a CiM (Compute-in-Memory)-accelerated similarity engine performs fast, low-power atlas-local retrieval. MOMAT's CiM-based retrieval accelerates a 100-query batch from 15,052.44 ms to 3,207.21 ns (a 4.69×106× speedup) and reduces energy from 8.1×107μJ to 3.32 μJ, yielding an approximately 2.5×105× energy reduction over DRAM-based (Raspberry Pi) baselines. Red-team evaluations across standard benchmarks show that MOMAT matches the defense performance of state-of-the-art methods while avoiding benign overkill and providing substantial efficiency gains, demonstrating that CiM-based modular defenses can make edge-deployed qLLMs both safer and more energy-efficient. We will release the full 223.2k-sample dataset to foster future research.
Figures & tables
Figure 1: Overview of MOMAT. The pipeline consists of two phases: an offline phase comprising (1) atlas construction from labeled examples and (2) training a lightweight MoE (1.03M parameters) with frozen embeddings, and an online inference phase where (3) MOMAT performs real-time multi-atlas CiM retrieval and MoE safety scoring on edge devices.
Size (GB)
ASR (%)
Model
Dataset
FP16
W4A8
FP16
W4A8
Δ
Llama2-13B
AdvBench
26.03
7.00
9.0
8.3
-0.7
Llama2-7B
AdvBench
13.38
3.76
28.1
32.5
+4.4
Llama2-13B
Malicious
26.03
7.00
21.0
19.0
-2.0
Llama2-7B
Malicious
13.38
3.76
25.0
28.0
+3.0
Table 1: Impact of Quantization on Model Size and ASR
Figure 2: Attack Success Rate (ASR) of different defense methods on the Llama2-7B model (W4A8). Green indicates near-zero ASR; red indicates high vulnerability without defense.
AdvBench
Malicious
Model
Defense
ASR
FRR
ASR
FRR
Llama2-7B
-
32.5
13.3
28.0
14.0
Llama2-7B
MOMAT
0.00
13.3 (+0.0)
0.00
14.0 (+0.0)
Mistral-7B
-
68.2
12.7
67.5
17.0
Mistral-7B
MOMAT
0.00
12.7 (+0.0)
0.0
17.0 (+0.0)
Table 2: ASR and FRR Results on W4A8 Quantized Models
DRAM window
DRAM size
Chunks
Avg time
(vectors)
(MB)
(#)
(ms / 100 queries)
500
1.95
325
15,052.44
1,000
3.91
163
15,763.70
2,000
7.81
82
15,788.72
4,000
15.62
41
15,827.99
10,000
39.06
17
15,911.68
Table 3: DRAM window sweep on Raspberry Pi 5 for 100-query batches (162,084 vectors, 1024-D, top- k=10 ). Latency is averaged over 10 runs.
Setting
Latency
Throughput
Energy / query
RPi 5 (100Q)
15,052.44 ms
6.6 QPS
≥8.1×107μJ
CiM (100Q)
3,207.21 ns
3.1×107 QPS
3.32 μ J
CiM (500Q)
3,213.15 ns
1.6×108 QPS
3.32 μ J
Table 4: Latency and energy for Raspberry Pi 5 DRAM retrieval vs. SRAM-based CiM retrieval (1024-D, top- k=10 ). DRAM energy is not measured; CiM energy is reported from the circuit-level simulator.
Model quantization is essential for the efficient deployment of Large Language Models (LLMs), but introduces a critical vulnerability: Quantization-Conditioned Backdoor (QCB) attacks. In these attacks, malicious behaviors remain dormant in full-precision models and activate only after specific quantization distortions, bypassing standard security audits. To mitigate this, we introduce FlipGuard, a proactive defense framework that selectively perturbs model weights prior to quantization. By breaking the adversary's precise alignment between weight patterns and quantization boundaries, FlipGuard suppresses backdoor activation without requiring access to training data or trigger samples. We further propose the Defense Effectiveness Ratio (DER), a unified metric to jointly evaluate security gains, utility preservation, and computational cost. Extensive experiments across seven LLMs (including StarCoder and LLaMA-family models) and three quantization schemes (INT8, FP4, NF4) demonstrate that FlipGuard effectively neutralizes QCBs across three scenarios, i.e., vulnerable code generation, content injection, and over-refusal, achieving high security with negligible performance degradation.
Aoying Zheng, Anqi Du, Zizhuang Deng +1
School of Cyber Science and Technology Shandong University · Suzhou Research Institute of Shandong University
Key-value (KV) cache quantization is widely used to reduce Large Language Model (LLM) inference memory, yet existing evaluations solely focus on measuring perplexity and accuracy without assessing the safety impact. In this study, we explore alignment preservation under KV cache quantization. Across eleven instruction-tuned models (3.8B-72B) and five benchmarks (1,894 prompts), we find that low-bit quantization can silently destroy safety alignment: Mistral-7B loses 15.2% of its refusals at only 1.03x perplexity, and no universal safe bit-width exists, with sharp model-specific phase transitions invisible to standard metrics. We identify that the root cause is geometric: safety features occupy a low-dimensional activation subspace 10^2-10^3x more vulnerable to quantization noise than the full representation space perplexity averages over. Inspired by this observation, we propose Per-Channel Reduction (PCR), a diagnostic that classifies each model into one of three mechanistic failure modes: outlier-crushes-safety, where safety lives in non-outlier channels collaterally damaged by outlier-driven scale factors; outlier-as-safety, where safety overlaps outlier channels and finer granularity cannot rescue it; and multi-layer dilution, where safety is distributed across many layers and per-layer fixes fail. PCR predicts the correct mitigation direction on all nine primary models and one held-out model from an independent family using 20 calibration prompts. PCR generalizes across unseen prompts, models, and production quantizers, including KIVI with up to 97.2% recovery, succeeding where attention-based allocation methods fail. The resulting training-free protocol, requiring approximately 35 GPU-minutes, recovers up to 97% of lost alignment at minimal memory overhead, addressing vulnerabilities confirmed in production vLLM serving with FP8 KV cache on NVIDIA GPUs.
Bruce Changlong Xu, Adarsh Kumarappan, Mu Zhou
†Stanford University · ‡California Institute of Technology
LLM quantization has become essential for memory-efficient deployment. Recent work has shown that quantization schemes can pose critical security risks: an adversary may release a model that appears benign in full precision but exhibits malicious behavior once quantized by users. However, existing quantization-conditioned attacks have been limited to relatively simple quantization methods, where the attacker can estimate weight regions that remain invariant under the target quantization. Notably, prior attacks have consistently failed to compromise more popular and sophisticated schemes, limiting their practical impact. In this work, we introduce the first quantization-conditioned attack that consistently induces malicious behavior that can be triggered by a broad range of advanced quantization techniques, including AWQ, GPTQ, and GGUF I-quants. Our attack exploits a simple property shared by many modern quantization methods: large outliers can cause other weights to be rounded to zero. Consequently, by injecting outliers into specific weight blocks, an adversary can induce a targeted, predictable weight collapse in the model. This effect can be used to craft seemingly benign full-precision models that exhibit a wide range of malicious behaviors after quantization. Through extensive evaluation across three attack scenarios and LLMs, we show that our attack achieves high success rates against a broad range of quantization methods on which prior attacks fail. Our results demonstrate, for the first time, that the security risks of quantization are not restricted to simpler schemes but are broadly relevant across complex, widely-used quantization methods.