Semi-structured pruning compresses large language models (LLMs) while keeping a regular sparse structure, but the prevailing N:M pattern fixes the same local sparsity ratio in every layer. Layer-adaptive sparsity allocation improves unstructured pruning, yet it has been reported to be less effective under N:M sparsity, leaving open whether adaptive allocation is of limited value for semi-structured pruning in general or only under the fine-grained N:M pattern. We examine this question with group-level sparsity, which partitions each weight matrix into regular groups, retains or prunes each group as a whole, and allows each layer's sparsity ratio to vary under a global budget. We propose GroupMask, which generates the group selectors of all layers with a lightweight hypernetwork, relaxes them with a Gumbel-Sigmoid parameterization and a straight-through estimator, and learns them through sparsity-budget regularization and self-distillation while keeping the pretrained weights frozen. On LLaMA-2-7B at 50% sparsity with the same 1×256 group size, learned layer-adaptive allocation reduces WikiText-2 perplexity from 10.02 to 8.30 and raises the average zero-shot accuracy from 0.455 to 0.496 relative to a uniform per-layer ratio. GroupMask obtains the lowest WikiText-2 perplexity on LLaMA-2-7B and the highest average zero-shot accuracy with Alpaca calibration among the evaluated baselines on five LLaMA and Qwen models. Our code is available at https://github.com/ZhengaoLi/GroupMask.
Figures & tables
Figure 1: Comparison of pruning granularities: (a) unstructured pruning, (b) structured pruning, (c) fixed N:M sparsity, and (d) group-level sparsity with different group shapes.
Figure 2: Structural configuration of our group-sparse pruning framework within a single Transformer block. For each linear projection, the mask generation module yields a structural group selector matrix B , which is expanded into a full-resolution binary mask M to determine the pruned weight W^l via element-wise multiplication.
Table 3: Layer-adaptive versus uniform sparsity allocation on LLaMA-2-7B at 50% sparsity. Both settings use the same 1×256 group structure, WikiText-2 calibration, and training schedule; the uniform setting constrains every linear layer to sl=50% .
Figure 3: Ablation analysis of GroupMask. (a) Effect of the hypernetwork on mask optimization. (b–c) Perplexity and zero-shot performance under increasing sparsity.
Setting
Cal.
ARC-C
ARC-E
BoolQ
Hella.
PIQA
Wino.
Avg.
Dense
–
0.462
0.745
0.778
0.760
0.788
0.692
0.704
Hypernetwork ablation
w/o HN
Wiki
0.240
0.322
0.478
0.314
0.552
0.511
0.403
Alp.
0.242
0.283
0.528
0.269
0.522
0.479
0.387
Group granularity and calibration
1×128
Wiki
0.269
0.441
0.631
0.469
0.642
0.560
0.502
Table 4: Zero-shot ablation on LLaMA-2-7B at 50% sparsity ( ↑ ).
Figure 4: Learned preserved rates of LLaMA-2-7B at 50% sparsity. (a) Preserved rate of each projection across layers. (b) Average preserved rates of attention and MLP projections; the dashed line marks the global 50% budget.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Value
Optimizer
AdamW / AdamW-8bit
Learning rate
1×10−3
Weight decay
0.05
Adam β
(0.9,0.999)
Batch size
1
Sequence length
2048
Appendix
Table 5: Main hyperparameters used for hypernetwork-based mask learning.
Model
Sparsity
ARC-C
ARC-E
BoolQ
Hella.
PIQA
Wino.
Avg ↑
LLaMA-2-7B
30%
0.3481
0.6124
0.6606
0.6203
0.7258
0.6101
0.596
40%
0.302
0.5109
0.5829
0.5394
0.6817
0.5738
0.532
LLaMA-3-8B
30%
0.3396
0.5934
0.6602
0.5913
0.7171
0.5809
0.580
40%
0.2969
0.5004
0.6434
0.5092
0.6741
0.5509
0.529
Appendix
Table 6: Zero-shot downstream performance under different sparsity ratios. We report results on six benchmarks and the average score.
Sparsity
LLaMA-2-7B ↓
LLaMA-2-13B ↓
30%
7.32
5.98
40%
8.25
6.79
50%
9.37
7.79
Appendix
Table 7: WikiText-2 perplexity under different sparsity ratios using the 32×32 group setting. Lower values indicate better language modeling performance.
Group
Cal.
ARC-C
ARC-E
BoolQ
Hella.
PIQA
Wino.
Avg ↑
LLaMA-2-13B
32×32
Wiki
0.3131
0.5640
0.6336
0.5621
0.6986
0.5691
0.557
32×32
Alpaca
0.3311
0.5800
0.6309
0.5486
0.7029
0.5801
0.562
64×64
Wiki
0.3131
0.5505
0.5728
0.5600
0.6986
0.5762
0.545
64×64
Alpaca
0.3217
0.5703
0.6330
0.5505
0.6855
0.5833
0.557
128×128
Wiki
0.3080
0.5122
0.5982
0.5654
0.6806
0.5706
0.539
Appendix
Table 8: Additional zero-shot downstream results under square group configurations.
Figure 5: Additional analysis of GroupMask. (a) Training dynamics of the sparsity regularization loss with and without the hypernetwork. (b) Kernel speedup under different group sizes at 50% sparsity. (c) Kernel speedup under group size 64 at different sparsity ratios.
Unstructured sparsity is now natively accelerated by recent GPU kernels and dataflow hardware, shifting the bottleneck from inference execution to the pruning algorithm. State-of-the-art methods for unstructured LLM pruning are layer-wise surrogates derived from the Optimal Brain Surgeon principle, and they sacrifice end-to-end accuracy, especially under aggressive sparsity. End-to-end alternatives such as MaskLLM and PATCH show that learnable masks can close this gap, but their categorical-over-patterns parameterization scales with the number of valid masks per row and does not port to the unstructured setting. We introduce LEAP, which replaces this intractable parameterization with a per-weight Bernoulli-via-Gumbel-sigmoid relaxation that makes end-to-end unstructured mask learning tractable. Across five LLM families from 0.5B to 8B parameters at 50% and 60% sparsity, LEAP improves six-task average zero-shot accuracy by +2.59 points on average over ADMM, the best layer-wise baseline in our sweep.
Mohammad Mozaffari, Younes Hourri, Mohammad Rastegari +1
Elastix AI · Department of Computer Science, University of Toronto, Toronto, Ontario, Canada.
One-shot pruning methods like Wanda and SparseGPT apply the same sparsity ratio to every layer of a transformer, ignoring known variation in layer importance. We propose PALS (Percentile-Aware Layerwise Sparsity), which adjusts per-layer sparsity based on the 99th percentile of activation magnitudes, bounded to ±5% around the target ratio. On LLaMA-2-7B at 50% sparsity, PALS achieves 10.96 WikiText-2 perplexity versus 12.92 for uniform Wanda (mean over 9 runs, p<0.001). The benefit is architecture-dependent: LLaMA-3-8B shows marginal gains and Mistral-7B shows none. We also find that gradient-based allocation -- the seemingly more principled approach -- produces results worse than random, suggesting that gradient magnitude does not predict the impact of discrete weight removal. PALS adds negligible cost to the pruning pipeline and requires no fine-tuning.
Semi-structured sparsity provides a practical path to accelerate large language models (LLMs) with native hardware support, but post-training semi-structured pruning often suffers from substantial quality degradation due to strong structural coupling. Existing methods rely on large-scale sparse retraining to recover accuracy, resulting in high computational cost. We propose SparseForge, a post-training framework that improves recovery efficiency by directly optimizing the sparsity mask rather than scaling up retraining tokens. SparseForge combines Hessian-aware importance estimation with progressive annealing of soft masks into hardware-executable structured sparsity, enabling stable and efficient sparse recovery. On LLaMA-2-7B under 2:4 sparsity, SparseForge achieves 57.27% average zero-shot accuracy with only 5B retraining tokens, surpassing the dense model's 56.43% accuracy and approaching the 57.52% result of a state-of-the-art method using 40B tokens. Such improvements on the accuracy-efficiency trade-off from SparseForge are shown to be consistent across model families.