MOMAT: Mixture of Multiple Atlases for Low-Power Jailbreak Defense of Quantized LLMs
Authors: Boyang Li, Bingyu Shen, Weihao Hong, Zhiyuan Jiang, Xinlei Guan, Yan Ma, Miles Q. Li, Yi Sheng, +1 more
Organizations: Department of Computer Science, Kean University, 1000 Morris Ave, Union, 07083, NJ, USA · Department of Computer Science and Engineering, University of Notre Dame, 257 Fitzpatrick Hall of Engineering, Notre Dame, 46556, IN, USA · McGill University, 845 Sherbrooke Street, Montréal, H3A 0G4, Québec, Canada · Department of Computer Science and Engineering, University of South Florida, 4202 E Fowler Ave, Tampa, 33620, FL, USA · Department of Computing Sciences, Villanova University, 800 E Lancaster Ave, Villanova, 19085, PA, USA
Quantized large language models are increasingly deployed on edge devices for their low latency and energy efficiency. However, model quantization weakens alignment safeguards, leaving qLLMs (quantized large language models) highly vulnerable to jailbreak attacks. To address this challenge, we present MOMAT (Mixture of Multiple Atlases), a hardware-enhanced safety framework that combines structured knowledge retrieval with low-power defense acceleration. Each atlas represents a semantic cluster of harmful or benign sample sets and policy templates, enabling domain-localized Retrieval-Augmented Generation guarding that mitigates the curse of dimensionality and the resulting semantic sparsity problem in large, heterogeneous safety databases. MOMAT retrieves top-k similarity features from all atlases for each prompt and evaluates them using a lightweight MoE (Mixture of Experts) detector, while a CiM (Compute-in-Memory)-accelerated similarity engine performs fast, low-power atlas-local retrieval. MOMAT's CiM-based retrieval accelerates a 100-query batch from 15,052.44 ms to 3,207.21 ns (a 4.69×106× speedup) and reduces energy from 8.1×107μJ to 3.32 μJ, yielding an approximately 2.5×105× energy reduction over DRAM-based (Raspberry Pi) baselines. Red-team evaluations across standard benchmarks show that MOMAT matches the defense performance of state-of-the-art methods while avoiding benign overkill and providing substantial efficiency gains, demonstrating that CiM-based modular defenses can make edge-deployed qLLMs both safer and more energy-efficient. We will release the full 223.2k-sample dataset to foster future research.
Figures & tables
Figure 1: Overview of MOMAT. The pipeline consists of two phases: an offline phase comprising (1) atlas construction from labeled examples and (2) training a lightweight MoE (1.03M parameters) with frozen embeddings, and an online inference phase where (3) MOMAT performs real-time multi-atlas CiM retrieval and MoE safety scoring on edge devices.
Size (GB)
ASR (%)
Model
Dataset
FP16
W4A8
FP16
W4A8
Δ
Llama2-13B
AdvBench
26.03
7.00
9.0
8.3
-0.7
Llama2-7B
AdvBench
13.38
3.76
28.1
32.5
+4.4
Llama2-13B
Malicious
26.03
7.00
21.0
19.0
-2.0
Llama2-7B
Malicious
13.38
3.76
25.0
28.0
+3.0
Table 1: Impact of Quantization on Model Size and ASR
Figure 2: Attack Success Rate (ASR) of different defense methods on the Llama2-7B model (W4A8). Green indicates near-zero ASR; red indicates high vulnerability without defense.
AdvBench
Malicious
Model
Defense
ASR
FRR
ASR
FRR
Llama2-7B
-
32.5
13.3
28.0
14.0
Llama2-7B
MOMAT
0.00
13.3 (+0.0)
0.00
14.0 (+0.0)
Mistral-7B
-
68.2
12.7
67.5
17.0
Mistral-7B
MOMAT
0.00
12.7 (+0.0)
0.0
17.0 (+0.0)
Table 2: ASR and FRR Results on W4A8 Quantized Models
DRAM window
DRAM size
Chunks
Avg time
(vectors)
(MB)
(#)
(ms / 100 queries)
500
1.95
325
15,052.44
1,000
3.91
163
15,763.70
2,000
7.81
82
15,788.72
4,000
15.62
41
15,827.99
10,000
39.06
17
15,911.68
Table 3: DRAM window sweep on Raspberry Pi 5 for 100-query batches (162,084 vectors, 1024-D, top- k=10 ). Latency is averaged over 10 runs.
Setting
Latency
Throughput
Energy / query
RPi 5 (100Q)
15,052.44 ms
6.6 QPS
≥8.1×107μJ
CiM (100Q)
3,207.21 ns
3.1×107 QPS
3.32 μ J
CiM (500Q)
3,213.15 ns
1.6×108 QPS
3.32 μ J
Table 4: Latency and energy for Raspberry Pi 5 DRAM retrieval vs. SRAM-based CiM retrieval (1024-D, top- k=10 ). DRAM energy is not measured; CiM energy is reported from the circuit-level simulator.