cs.AIOct 6, 2026

How Fragile Is On-Device Language Model Safety? Localizing Safety-Critical Parameters for Sparse Fault Analysis

Authors: Muhammad Zeeshan Karamat, Christiana Chamon Garcia

Organizations: Virginia Tech

Abstract

As small language models (SLMs) are increasingly deployed on resource-constrained and on-device platforms, including as components of agentic systems, the integrity of locally stored model parameters becomes an important safety concern. We investigate whether safety-sensitive behavior in LLaMA-2-7B-Chat is concentrated within a sparse subset of parameters, creating a reduced fault surface for targeted analysis. We study two complementary localization methods: low-rank safety-associated subspace analysis and parameter-level safety--utility importance filtering. Both approaches reveal highly non-uniform safety sensitivity across the network, with the MLP down_proj consistently emerging as a prominent safety-sensitive component and o_proj providing a smaller contribution. Using parameter-level localization, modifying only 0.19% of model weights in down_proj yields 53% Basic ASR and 56% GCG ASR, while tinyBenchmarks accuracy remains at 51.6% compared with a 52.2% unmodified baseline. These results motivate targeted fault analysis and selective integrity protection for language models deployed in resource-constrained, on-device, and agentic settings.

Figures & tables

Explore similar work

CardsList
  1. Local Sparsity Enables Unsupervised LLM Safety Detection

    Sep 17, 2026Xin Chen, Gil Kur, Alexander Shevchenko +1Large Language Model SafetySparsity

  2. Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models

    Sep 16, 2026Alizishaan Khatri, Chiquita Prabhu, Omkar NeogiLarge Language Model SafetyGuardrail

  3. Distilling Safe LLM Systems via Soft Prompts for On Device Settings

    Jun 8, 2026Motasem Alfarra, Cristina Pinneri, Dana Kianfar +2Large Language Model SafetySafety Alignment