cs.LGDec 3, 2025

Data-Free Pruning of Self-Attention Layers in LLMs

Authors: Dhananjay Saikumar, Blesson Varghese

Organizations: School of Computer Science University of St Andrews United Kingdom

Abstract

Many self-attention sublayers in large language models (LLMs) can be removed with little to no loss. We attribute this to the Attention Suppression Hypothesis: during pre-training, some deep attention layers learn to mute their own contribution, leaving the residual stream and the MLP to carry the representation. We propose Gate-Norm, a one-shot, weight-only criterion that ranks attention sublayers by query-key coupling and removes the least coupled ones, requiring no calibration data, no forward passes, no fine-tuning, and no specialized kernels. On 40-layer, 13B-parameter LLaMA models, Gate-Norm prunes the model in under a second. Pruning 8-16 attention sublayers yields up to 1.30×1.30\times higher inference throughput while keeping average zero-shot accuracy within 1.5 percentage points of the unpruned baseline across BoolQ, RTE, HellaSwag, WinoGrande, ARC-Easy/Challenge, and OpenBookQA. Across these settings, Gate-Norm matches data-driven pruning methods in accuracy while being ∼1000×\sim 1000\times faster to score layers, enabling practical, data-free compression of LLMs.

Figures & tables

Explore similar work

CardsList
  1. Garbage Attention in Large Language Models: BOS Sink Heads and Sink-aware Pruning

    Jan 11, 2026Jaewon Sok, Jewon Yeom, Seonghyeon Park +2Large Language Model CompressionUnstructured Pruning

  2. Prune, Update and Trim: Robust Structured Pruning for Large Language Models

    May 18, 2026Diego Coello de Portugal Mecke, Tom Hanika, Lars Schmidt-ThiemeUnstructured PruningFeed-Forward

  3. GRASPrune: Global Gating for Budgeted Structured Pruning of Large Language Models

    Apr 21, 2026Ziyang Wang, Jiangfeng Xiao, Chuan Xiao +3Large Language Model CompressionUnstructured Pruning