Many self-attention sublayers in large language models (LLMs) can be removed with little to no loss. We attribute this to the Attention Suppression Hypothesis: during pre-training, some deep attention layers learn to mute their own contribution, leaving the residual stream and the MLP to carry the representation. We propose Gate-Norm, a one-shot, weight-only criterion that ranks attention sublayers by query-key coupling and removes the least coupled ones, requiring no calibration data, no forward passes, no fine-tuning, and no specialized kernels. On 40-layer, 13B-parameter LLaMA models, Gate-Norm prunes the model in under a second. Pruning 8-16 attention sublayers yields up to 1.30× higher inference throughput while keeping average zero-shot accuracy within 1.5 percentage points of the unpruned baseline across BoolQ, RTE, HellaSwag, WinoGrande, ARC-Easy/Challenge, and OpenBookQA. Across these settings, Gate-Norm matches data-driven pruning methods in accuracy while being ∼1000× faster to score layers, enabling practical, data-free compression of LLMs.
Figures & tables
Method
Vicuna-7B
Vicuna-13B
LLaMA-3.1-8B
# pruned attention layers
# pruned attention layers
# pruned attention layers
1
4
7
10
13
1
4
7
10
13
1
4
7
10
13
Random
6.82
7.84
13.84
355.87
473.40
6.00
10.63
67.09
587.90
594.83
6.51
8.11
14.59
54.51
98.55
Data-driven
7.23
7.62
8.78
12.55
20.46
5.99
6.11
6.24
6.39
7.54
6.35
6.83
7.39
9.79
16.48
Gate-Norm
7.23
7.42
8.57
10.02
16.64
5.97
6.05
6.24
6.47
7.05
6.30
6.67
7.74
11.55
15.65
Table 1: WikiText-2 perplexity on additional checkpoints under attention pruning. Lower is better. “Random” denotes random attention-layer removal; “Data-driven” is a cosine-similarity baseline; “Gate-Norm” is our data-free score.
LLaMA-13B v1
Method
#Pruned
Type
BoolQ
RTE
HellaSwag
WinoG
ARC-E
ARC-C
OpenBookQA
Avg Acc
SpeedUp
Data Required
Baseline
0
–
79.85
68.95
79.32
73.40
77.36
48.38
47.40
67.81
1.00×
–
Random-Block
4
Block
68.62
54.15
71.30
64.80
72.94
41.13
40.20
59.02
1.10
✗
Random-Block
8
Block
47.22
53.43
26.36
49.72
28.07
26.28
27.80
36.98
1.25
✗
ShortGPT-Block
4
Block
63.49
68.23
76.60
72.77
74.03
45.99
46.20
63.90
1.10
✓
ShortGPT-Block
8
Block
38.01
64.26
70.92
71.19
68.35
42.32
41.60
56.67
1.25
✓
Table 2: Zero-shot accuracies (%) for LLaMA-13B v1 and v2 under different pruning schemes and intensities. Columns: Method, #Pruned (layers/blocks removed), Type (block vs. attention), per-task accuracy, average accuracy, inference speed-up relative to baseline, and whether calibration data are required. Pruning schemes include random removal, structured block removal, data-driven attention pruning, and data-free Gate-Norm attention pruning.
Model
Unpruned
Gate-Norm
Gate-Norm + LoRA
Data-driven
Data-driven + LoRA
Vicuna-7B
6.78
321.49
11.96
225.13
11.50
Vicuna-13B
5.95
14.13
6.57
25.88
6.65
LLaMA-2-13B (v2)
4.88
14.11
6.35
14.73
6.27
Table 3: WikiText-2 perplexity with 20 attention layers pruned, with and without brief LoRA fine-tuning.
Large Language Models (LLMs) are known to contain significant redundancy, yet a systematic explanation for why certain components, particularly in higher layers, are more redundant has remained elusive. In this work, we identify the BOS sink phenomenon as a key mechanism driving this layer-wise sensitivity. We show that attention heads with high BOS sink scores are strongly associated with functional redundancy: such heads, especially in deeper layers, contribute little to predictive performance and effectively serve as dumping grounds for superfluous attention weights. Leveraging this insight, we introduce a simple pruning strategy that removes high-BOS sink heads. Experiments on Gemma-3, Llama-3.1, and Qwen3 demonstrate that this approach identifies redundant transformer components more reliably than weight- and activation-based criteria in terms of downstream task retention, remaining close to dense baselines at low-to-moderate pruning ratios. We further find that high-scoring sink heads sustain their focus on BOS as context length grows. Overall, our results suggest that structural properties of attention offer a more direct basis for model compression than magnitude-based methods.
Jaewon Sok, Jewon Yeom, Seonghyeon Park +2
Department of Rural Systems Engineering, Seoul National University · Graduate School of Data Science, Seoul National University · Department of Aerospace Engineering, Seoul National University
Large Language Models (LLMs) have experienced significant growth and development in recent years. However, performing inference on LLMs remains costly, especially for long-context inference or in resource-constrained devices. This motivates the development of new post-training pruning (PTP) methods. These methods reduce LLMs' requirements by removing a substantial part of the model's parameters. The discarded weights are selected depending on their impact on the models performance. Current PTP methods prune the models by removing the less informative hidden nodes from the FFN layers, and the least important attention layers. We propose Putri, a PTP method that introduces three changes to the State-of-the-art. First, we update the un-pruned weights of the FFN to compensate for the introduced pruning error. Second, the FFN layers are pruned sequentially, taking into account the updates done to the previous layers. Third, instead of removing full attention layers, we remove individual attention-heads. We extend this method such that it can also address Grouped-Query Attention. In summary, Putri is a structure pruning method which remains simple while showing SOTA performance. Pruning experiments on multiple models with a wide variety of sparsity ranges and on different datasets, validate the generality of Putri. Notably, we demonstrate that, unlike previous methods, Putri can prune LLMs on extreme sparsity ratios. The code is available at: https://github.com/Coello-dev/Putri.
Diego Coello de Portugal Mecke, Tom Hanika, Lars Schmidt-Thieme
ISMLL & DARC VWFS · University of Hildesheim · Hildesheim, Germany +1
Large language models (LLMs) are expensive to serve because model parameters, attention computation, and KV caches impose substantial memory and latency costs. We present GRASPrune, a structured pruning framework applied after pretraining that jointly prunes FFN channels and KV head groups under a single global budget. Instead of learning importance scores without constraints and applying the budget only after training, GRASPrune learns lightweight gate scores with a projected straight-through estimator that enforces a hard mask satisfying the budget at every step while keeping the backbone weights frozen. After the mask is fixed, we calibrate scaling factors on the retained units to mitigate scale mismatch caused by pruning, and fold these factors into the pruned weights to obtain a smaller dense checkpoint with no extra parameters at inference. On LLaMA-2-7B, GRASPrune removes 50% of parameters and achieves 12.18 perplexity on WikiText-2 while maintaining competitive average zero-shot accuracy on five benchmarks, using four epochs on 512 unlabeled calibration sequences on a single NVIDIA A100 80GB GPU without any full model fine-tuning.
Ziyang Wang, Jiangfeng Xiao, Chuan Xiao +3
Beijing Institute of Technology, Zhuhai · Shenzhen University · Osaka University