How Fragile Is On-Device Language Model Safety? Localizing Safety-Critical Parameters for Sparse Fault Analysis
Organizations: Virginia Tech
Abstract
As small language models (SLMs) are increasingly deployed on resource-constrained and on-device platforms, including as components of agentic systems, the integrity of locally stored model parameters becomes an important safety concern. We investigate whether safety-sensitive behavior in LLaMA-2-7B-Chat is concentrated within a sparse subset of parameters, creating a reduced fault surface for targeted analysis. We study two complementary localization methods: low-rank safety-associated subspace analysis and parameter-level safety--utility importance filtering. Both approaches reveal highly non-uniform safety sensitivity across the network, with the MLP down_proj consistently emerging as a prominent safety-sensitive component and o_proj providing a smaller contribution. Using parameter-level localization, modifying only 0.19% of model weights in down_proj yields 53% Basic ASR and 56% GCG ASR, while tinyBenchmarks accuracy remains at 51.6% compared with a 52.2% unmodified baseline. These results motivate targeted fault analysis and selective integrity protection for language models deployed in resource-constrained, on-device, and agentic settings.
Figures & tables
| Modules | Top fraction | Basic ASR | GCG ASR | Utility |
|---|---|---|---|---|
| All modules | 1.0 | 87% | 98% | 49.3% |
| All modules | 0.1 | 71% | 85% | 50.1% |
| All modules | 0.01 | 36% | 25% | 50.7% |
| All modules | 0.001 | 20% | 18% | 51.8% |
| MLP only | 1.0 | 77% | 89% | 49.8% |
| MLP only | 0.1 | 54% | 64% | 51.1% |
| Target component | Weights (%) | Basic ASR | GCG ASR | Utility |
|---|---|---|---|---|
| All modules | 2.95 | 91% | 97% | 47.1% |
| MLP only | 1.87 | 86% | 94% | 49.7% |
| down_proj + o_proj | 0.23 | 58% | 60% | 51.3% |
| down_proj | 0.19 | 53% | 56% | 51.6% |
| o_proj | 0.04 | 20% | 14% | 52.0% |