Security-Enhanced Seed-Based Weight Quantization for Large Language Models
Organizations: Department of Electrical and Computer Engineering (ECE) University of Florida Gainesville, FL 32611, USA
Abstract
Large language models (LLMs) incur substantial storage, memory-bandwidth and energy costs, motivating compact weight representations. Existing seed-based compression methods reconstruct weights from compact pseudo-random representations but do not explicitly account for the non-uniform sensitivity of model weights. We introduce Seed-Q, a security-enhanced sensitivity-aware seed-based weight compression framework that uses lightweight Linear Feedback Shift Register (LFSR)-based weight generation with non-uniform bit allocation. Our approach assigns larger representation budgets to sensitive weights while aggressively compressing less sensitive regions. Importantly, this non-uniform allocation requires no side-information: the decoder deterministically reconstructs the bit-allocation schedule, with no rung depending on the decoded weights, eliminating the need to store per-block metadata or use calibration data while preserving the baseline coding rate. Experiments across diverse LLMs show that Seed-Q matches 4-bit perplexity of SeedLM with fewer bits, while at the same 4 bits/weight it reduces both perplexity degradation and zero-shot accuracy loss relative to SeedLM. We also show that Seed-Q simultaneously achieves high security against bit-flip attacks on model parameters, as bit corruption affects multiple reconstructed weights, greatly amplifying its impact and making it easier to detect. We further implement Seed-Q in an ASIC-based accelerator and demonstrate modest hardware overhead compared to prior seed-based approaches.
Figures & tables
| Method | Model Perplexity (Bits/Weight) | Mean | |||
|---|---|---|---|---|---|
| LLaMA-2-7B | LLaMA-2-13B | LLaMA-3-8B | Mistral-7B | Bits | |
| Baseline | 5.5 (16) | 4.9 (16) | 6.1 (16) | 5.25 (16) | 16 |
| Seed-Q (Mode 1)* | 5.85 (3.77) | 5.18 (3.75) | 7.17 (3.88) | 6.32 (3.78) | 3.79 |
| S-Quant | 5.7 (3.8) | 5.0 (3.8) | 6.8 (3.8) | NR | 3.8 |
| Seed-Q (Mode 2) | 5.68 (4) | 5.06 (4) | 6.92 (4) | 5.96 (4) | 4 |
| SeedLM | 5.85 (4) | 5.18 (4) | 7.17 (4) | 6.32 (4) | 4 |
| Model | Method | Zero-Shot Task Accuracy (%) | ||||||
|---|---|---|---|---|---|---|---|---|
| Bits | ARC-Easy | ARC-Challenge | HellaSwag | WinoGrande | BoolQ | Mean | ||
| LLaMA-2-7B | Baseline | 16 | 74.58 | 46.33 | 75.98 | 69.06 | 77.74 | 68.74 |
| Seed-Q (Mode 1)* | 3.77 | 71.97 | 43.43 | 73.60 | 68.03 | 74.98 | 66.40 | |
| S-Quant | 3.8 | 73.36 | 44.55 | 74.51 | 68.47 | 77.34 | 67.65 | |
| Seed-Q (Mode 2) | 4 | 72.31 | 44.45 | 74.69 | 69.53 | 75.90 | 67.38 | |
| SeedLM ‡ | 4 | 72.39 | 43.60 | 73.79 | 67.72 | 76.85 | 66.87 | |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Global | ||||
|---|---|---|---|---|
| part of the ranking | Seed-Q mass | Fisher mass | Seed-Q exposure | random |
| top | ||||
| top | ||||
| top | ||||
| bottom | ||||
| bottom | ||||
| model | grid | moved | objective | measured |
|---|---|---|---|---|
| Llama-2-7B | full ( ) | — | 0.0412 | |
| Llama-2-7B | reduced ( ), dry run | 0.0663 | — | |
| Llama-2-13B | reduced | 0.0579 | 0.0543 | |
| Llama-3-8B | 0.1526 | 0.1317 | ||
| Llama-3-8B, | 0.3800 | 0.2886 | ||
| Llama-3-8B, | 0.2368 | 0.1847 |
| rung | rate | Llama-2-7B | Llama-2-13B | Llama-3-8B | Mistral-7B | Qwen2.5-7B |
|---|---|---|---|---|---|---|
| 3.000 | 1.8746 | — | 2.2518 ∗ | — | — | |
| 3.250 | 0.5852 | † | 0.7014 | 2.5967 | 0.2975 | |
| 3.500 | 0.2249 | † | 0.3892 | 0.9774 | 0.1754 | |
| 3.750 | 0.1182 | 0.0874 | 0.2445 | 0.3544 | 0.1090 | |
| 4.000 | 0.0669 | 0.0583 | 0.1556 | 0.1854 | 0.0744 | |
| 3.750 | 0.1445 | — | — | — | — |
| rung | median | sd of ratio | ||
|---|---|---|---|---|
| 3.43 | 1.75 | 6.62 | 0.54 | |
| 2.21 | 1.18 | 4.21 | 0.52 | |
| 1.49 | 0.80 | 2.77 | 0.51 | |
| 0.79 | 0.41 | 1.49 | 0.52 | |
| 0.29 | 0.15 | 0.59 | 0.55 |