Semi-structured pruning compresses large language models (LLMs) while keeping a regular sparse structure, but the prevailing N:M pattern fixes the same local sparsity ratio in every layer. Layer-adaptive sparsity allocation improves unstructured pruning, yet it has been reported to be less effective under N:M sparsity, leaving open whether adaptive allocation is of limited value for semi-structured pruning in general or only under the fine-grained N:M pattern. We examine this question with group-level sparsity, which partitions each weight matrix into regular groups, retains or prunes each group as a whole, and allows each layer's sparsity ratio to vary under a global budget. We propose GroupMask, which generates the group selectors of all layers with a lightweight hypernetwork, relaxes them with a Gumbel-Sigmoid parameterization and a straight-through estimator, and learns them through sparsity-budget regularization and self-distillation while keeping the pretrained weights frozen. On LLaMA-2-7B at 50% sparsity with the same 1×256 group size, learned layer-adaptive allocation reduces WikiText-2 perplexity from 10.02 to 8.30 and raises the average zero-shot accuracy from 0.455 to 0.496 relative to a uniform per-layer ratio. GroupMask obtains the lowest WikiText-2 perplexity on LLaMA-2-7B and the highest average zero-shot accuracy with Alpaca calibration among the evaluated baselines on five LLaMA and Qwen models. Our code is available at https://github.com/ZhengaoLi/GroupMask.
Figures & tables
Figure 1: Comparison of pruning granularities: (a) unstructured pruning, (b) structured pruning, (c) fixed N:M sparsity, and (d) group-level sparsity with different group shapes.
Figure 2: Structural configuration of our group-sparse pruning framework within a single Transformer block. For each linear projection, the mask generation module yields a structural group selector matrix B , which is expanded into a full-resolution binary mask M to determine the pruned weight W^l via element-wise multiplication.
Table 3: Layer-adaptive versus uniform sparsity allocation on LLaMA-2-7B at 50% sparsity. Both settings use the same 1×256 group structure, WikiText-2 calibration, and training schedule; the uniform setting constrains every linear layer to sl=50% .
Figure 3: Ablation analysis of GroupMask. (a) Effect of the hypernetwork on mask optimization. (b–c) Perplexity and zero-shot performance under increasing sparsity.
Setting
Cal.
ARC-C
ARC-E
BoolQ
Hella.
PIQA
Wino.
Avg.
Dense
–
0.462
0.745
0.778
0.760
0.788
0.692
0.704
Hypernetwork ablation
w/o HN
Wiki
0.240
0.322
0.478
0.314
0.552
0.511
0.403
Alp.
0.242
0.283
0.528
0.269
0.522
0.479
0.387
Group granularity and calibration
1×128
Wiki
0.269
0.441
0.631
0.469
0.642
0.560
0.502
Table 4: Zero-shot ablation on LLaMA-2-7B at 50% sparsity ( ↑ ).
Figure 4: Learned preserved rates of LLaMA-2-7B at 50% sparsity. (a) Preserved rate of each projection across layers. (b) Average preserved rates of attention and MLP projections; the dashed line marks the global 50% budget.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Value
Optimizer
AdamW / AdamW-8bit
Learning rate
1×10−3
Weight decay
0.05
Adam β
(0.9,0.999)
Batch size
1
Sequence length
2048
Appendix
Table 5: Main hyperparameters used for hypernetwork-based mask learning.
Model
Sparsity
ARC-C
ARC-E
BoolQ
Hella.
PIQA
Wino.
Avg ↑
LLaMA-2-7B
30%
0.3481
0.6124
0.6606
0.6203
0.7258
0.6101
0.596
40%
0.302
0.5109
0.5829
0.5394
0.6817
0.5738
0.532
LLaMA-3-8B
30%
0.3396
0.5934
0.6602
0.5913
0.7171
0.5809
0.580
40%
0.2969
0.5004
0.6434
0.5092
0.6741
0.5509
0.529
Appendix
Table 6: Zero-shot downstream performance under different sparsity ratios. We report results on six benchmarks and the average score.
Sparsity
LLaMA-2-7B ↓
LLaMA-2-13B ↓
30%
7.32
5.98
40%
8.25
6.79
50%
9.37
7.79
Appendix
Table 7: WikiText-2 perplexity under different sparsity ratios using the 32×32 group setting. Lower values indicate better language modeling performance.
Group
Cal.
ARC-C
ARC-E
BoolQ
Hella.
PIQA
Wino.
Avg ↑
LLaMA-2-13B
32×32
Wiki
0.3131
0.5640
0.6336
0.5621
0.6986
0.5691
0.557
32×32
Alpaca
0.3311
0.5800
0.6309
0.5486
0.7029
0.5801
0.562
64×64
Wiki
0.3131
0.5505
0.5728
0.5600
0.6986
0.5762
0.545
64×64
Alpaca
0.3217
0.5703
0.6330
0.5505
0.6855
0.5833
0.557
128×128
Wiki
0.3080
0.5122
0.5982
0.5654
0.6806
0.5706
0.539
Appendix
Table 8: Additional zero-shot downstream results under square group configurations.
Figure 5: Additional analysis of GroupMask. (a) Training dynamics of the sparsity regularization loss with and without the hypernetwork. (b) Kernel speedup under different group sizes at 50% sparsity. (c) Kernel speedup under group size 64 at different sparsity ratios.