Large language models (LLMs) deliver impressive performance but incur prohibitive memory and compute costs at deployment. Model pruning is an effective way to reduce these overheads, yet existing approaches face challenges: unstructured sparsity, where nonzeros can appear anywhere, preserves accuracy but yields irregular access patterns that prevent GPU acceleration, while semi-structured 2:4 sparsity is hardware-friendly but enforces a rigid 50% pattern that degrades model quality. To bridge this gap, we introduce PATCH, a hybrid sparsity framework that enables a continuous sparsity ratio between 0% and 50%. PATCH partitions weight matrices into tiles, assigning each tile to be either dense or 2:4 sparse via a learnable mask selection mechanism. This design provides fine-grained control over accuracy-acceleration tradeoffs and supports non-uniform sparsity across layers, leading to superior overall quality. Across models from 0.5B to 13B parameters, PATCH consistently narrows the gap to dense accuracy while delivering practical speedups. For instance, on LLaMA-2 7B with an A6000 GPU, PATCH achieves 1.18x-1.38x end-to-end speedup over dense baselines while improving accuracy by 0.37%-2.96% compared to the state-of-the-art 2:4 pruning method, MaskLLM.
Figures & tables
Figure 1 : Illustration of the PATCH learning process for generating tile-level hybrid masks. Each tile is parameterized by a learnable distribution and sampled with Gumbel Softmax to produce M~tile . The dense probability is expanded and merged with a 2:4 mask M~2:4 , which can be fixed or jointly learned during training, yielding M~ . The final mask assigns each tile to remain dense or follow the 2:4 pattern, enabling flexible sparsity across the weight matrix.
Sparsity
Method
Pattern
Qwen-2.5 0.5B
LLaMA-3.2 1B
Gemma-3 1B
Acc (% ↑ )
PPL ( ↓ )
Acc (% ↑ )
PPL ( ↓ )
Acc (% ↑ )
PPL ( ↓ )
0%
Dense
-
46.00
12.08
47.70
9.06
47.01
11.67
50%
Magnitude
2:4
30.16
6734.97
29.66
563.44
31.66
5005.56
Wanda
2:4
32.97
72.48
31.61
78.18
34.16
69.41
SparseGPT
2:4
34.81
36.59
35.55
32.73
35.58
44.59
Thanos
2:4
31.31
37.32
35.71
33.03
35.09
62.63
Table 1: Model quality (average accuracy across eight zero-shot tasks and perplexity on WikiText2 dataset) for different pruning methods. By jointly optimizing the location of dense tiles and the sparsity pattern within the sparse tiles, PATCH Joint allows for a continuous sparsity ratio for the models, providing a flexible tradeoff between sparsity and model quality.
Sparsity
Method
Pattern
LLaMA-2 7B
LLaMA-2 13B
LLaMA-3.1 8B
Acc (% ↑ )
PPL ( ↓ )
Acc (% ↑ )
PPL ( ↓ )
Acc (% ↑ )
PPL ( ↓ )
0%
Dense
-
54.61
5.12
58.38
4.89
60.31
5.84
50%
Magnitude
2:4
43.44
54.39
45.94
8.89
35.93
765.92
Wanda
2:4
44.30
11.15
47.95
8.91
41.77
21.29
SparseGPT
2:4
45.09
10.12
49.67
8.86
45.53
15.11
Thanos
2:4
44.80
11.19
49.33
8.80
45.72
16.09
Table 2: Model quality (average accuracy across eight zero-shot tasks and perplexity on WikiText2 dataset) for different pruning methods. By only optimizing the location of dense tiles while keeping sparsity pattern within the sparse tiles frozen, PATCH Tile provides a memory efficient variant for PATCH Joint , allowing for a continuous sparsity ratio for the models and providing a flexible tradeoff between sparsity and model quality.
Table 4
Sparsity (0.5B)
MaskLLM
SparseGPT (no update)
Wanda
Mag.
PATCH Joint
45%
15.06
21.84
21.83
21.33
14.57
35%
14.55
17.29
17.96
19.90
13.84
25%
14.17
14.89
15.09
16.05
13.47
Table 5 : Impact of fixed 2:4 mask selection for PATCH Tile , compared with joint optimization ( ↓ is better). PATCH Joint achieves the lowest perplexity overall, while for PATCH Tile , MaskLLM provides the best frozen mask. Mag . refers to Magnitude pruning.
Figure 2 : Layer-wise sparsity allocation under different global sparsity budgets for various models. PATCH achieves the target global sparsity while flexibly distributing pruning across transformer layers.
Sparsity
Method
Bit
LoRA
LLaMA-2-7B
LLaMA-3.1-8B
Comp.
Acc (% ↑ )
PPL ( ↓ )
Acc (% ↑ )
PPL ( ↓ )
Ratio
0%
Dense
-
-
54.61
5.12
60.31
5.84
1x
50%
MaskLLM
4
-
47.98
7.64
51.12
9.92
5.33x
45%
PATCH Tile
4
-
48.19
7.34
52.47
9.68
5.16x
45%
PATCH Tile
4
SLiM -LoRA
50.71
6.83
54.04
9.12
4.10x
35%
PATCH Tile
4
-
49.38
6.92
53.81
9.26
4.85x
Table 6 : Average accuracy ( ↑ indicates better) across eight zero-shot downstream tasks and WikiText2 perplexity ( ↓ indicates better) of compressed models with 4-bit weight-only quantization . Please note that using LoRA adds additional parameters to the model. Comp. Ratio refers to the theoretical weight memory compression factor relative to the dense model; LoRA adapters are quantized to 4-bit integers.
Figure 3 : Accuracy–efficiency trade-off on LLaMA-2 7B (A6000, batch size 16, prefill and generation length 128). Eight-task average accuracy versus end-to-end speedup over dense (left) and versus weight memory relative to dense (right). Dense and full 2:4 (MaskLLM, speedup measured with STOICC) are single operating points; PATCH at 25%, 35%, and 45% sparsity fills the region between them, and every PATCH point lies above the full-2:4 accuracy (dashed).
Appendix figures & tables30 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Sparsity
Avg Acc (% ↑ )
Speedup ( ↑ )
Weight memory ( ↓ )
Dense
0%
54.61
1.00 ×
1.00 ×
MaskLLM (full 2:4)
50%
48.62
1.47 × / 1.40 ×
0.56 ×
PATCH
45%
48.99
1.38 ×
0.59 ×
PATCH
35%
50.08
1.27 ×
0.68 ×
PATCH
25%
51.58
1.18 ×
0.76 ×
Appendix
Table 7 : Accuracy, speedup, and weight memory of LLaMA-2 7B on an A6000 GPU (batch size 16, prefill length 128, generation length 128). Throughput of a full 2:4 model is determined by the sparsity format, so the full-2:4 throughput of Table 23 applies to MaskLLM; its speedup is given as STOICC / cuSPARSELt.
Sparsity
Prefill length
Tokens generated
Throughput (tok/s)
Speedup vs. dense
0%
128
128
1023.80
1.00 ×
25%
128
128
1212.79
1.18 ×
35%
128
128
1304.46
1.27 ×
45%
128
128
1410.20
1.38 ×
0%
128
1024
435.42
1.00 ×
25%
128
1024
493.33
1.13 ×
Appendix
Table 8 : Throughput of LLaMA-2 7B with mixed sparsity compared to the dense model. Measurements taken on an A6000 GPU with batch size 16. Throughput is reported in tokens processed/sec.
Model
Sparsity
Prefill
Generated
Throughput (tok/s)
Speedup vs. dense
LLaMA-2 7B
0%
128
128
1876.24
1.00 ×
25%
128
128
2002.02
1.07 ×
35%
128
128
2088.98
1.11 ×
45%
128
128
2180.88
1.16 ×
0%
128
1024
812.55
1.00 ×
25%
128
1024
864.66
1.06 ×
Appendix
Table 9 : Throughput of LLaMA-2 7B and LLaMA-2 13B with mixed sparsity compared to the dense model. Measurements taken on an A100 GPU with batch size 16. Throughput is reported in tokens processed/sec.
Component
Dense
PATCH 25%
PATCH 35%
PATCH 45%
Weights
13.5 GB (1.00 × )
≈ 10.3 GB (0.76 × )
≈ 9.2 GB (0.68 × )
≈ 8.0 GB (0.59 × )
KV cache (128 generated)
≈ 2 GB
≈ 2 GB
≈ 2 GB
≈ 2 GB
KV cache (1024 generated)
≈ 9 GB
≈ 9 GB
≈ 9 GB
≈ 9 GB
Transient activations
< 0.5 GB
< 0.5 GB
< 0.5 GB
< 0.5 GB
Analytical peak (128 generated)
≈ 16.0 GB
≈ 12.8 GB
≈ 11.7 GB
≈ 10.5 GB
Empirical peak (128 generated)
16.721 GB
14.189 GB
12.530 GB
11.769 GB
Appendix
Table 10 : Peak GPU memory breakdown of LLaMA-2 7B in FP16 at inference (A6000, batch size 16, prefill length 128). PATCH reduces only the weights; the KV cache and transient activations are identical to the dense model. Analytical peaks sum the components; empirical peaks are measured.
Sparsity
Method
Pattern
MMLU
PIQA
ARC-E
ARC-C
WinoG.
OBQA
RACE
HellaS.
Avg
0%
Dense
-
47.71
70.24
64.48
29.52
56.20
24.20
35.02
40.63
46.00
50%
Magnitude
2:4
23.00
54.24
31.23
19.20
49.96
13.60
23.44
26.59
30.16
Wanda
2:4
24.43
58.71
43.18
17.75
51.62
12.20
26.32
29.58
32.97
SparseGPT
2:4
22.93
60.77
46.60
20.82
52.88
14.00
29.57
30.93
34.81
Thanos
2:4
22.97
60.17
45.37
19.20
53.59
15.20
31.00
31.31
34.85
ProxSparse
2:4
23.00
57.34
40.53
18.26
48.62
14.00
25.65
29.02
32.05
Appendix
Table 11: Model quality (task accuracy across eight zero-shot tasks, reported in %) for Qwen-2.5 0.5B with different pruning methods. PATCH Joint optimizes dense tile locations and sparsity patterns, enabling a flexible sparsity-quality tradeoff.
Sparsity
Method
Pattern
MMLU
PIQA
ARC-E
ARC-C
WinoG.
OBQA
RACE
HellaS.
Avg
0%
Dense
-
41.82
78.07
76.35
43.52
69.06
31.40
39.52
57.13
54.61
50%
Magnitude
2:4
25.82
70.02
61.78
30.12
61.01
21.80
31.48
45.45
43.44
Wanda
2:4
25.80
71.00
63.80
30.29
61.09
25.20
35.50
41.75
44.30
SparseGPT
2:4
26.17
70.73
63.80
30.63
65.04
24.00
37.13
43.18
45.09
Thanos
2:4
25.27
70.78
63.43
30.97
64.56
23.80
36.46
43.11
44.80
ProxSparse
2:4
26.77
71.60
65.70
33.02
62.90
24.20
35.31
47.84
45.92
Appendix
Table 12: Model quality (task accuracy across eight zero-shot tasks, reported in %) for LLaMA-2 7B with different pruning methods. PATCH Tile optimizes tile-based sparsity, enabling a flexible sparsity-quality tradeoff.
Sparsity
Method
Pattern
MMLU
PIQA
ARC-E
ARC-C
WinoG.
OBQA
RACE
HellaS.
Avg
0%
Dense
-
52.07
79.16
79.42
48.46
71.98
35.40
40.48
60.08
58.38
50%
Magnitude
2:4
27.53
72.03
62.46
32.17
62.35
24.20
36.65
50.10
45.94
Wanda
2:4
29.51
73.01
68.90
35.07
66.93
24.80
38.76
46.61
47.95
SparseGPT
2:4
33.42
73.56
68.60
36.86
69.61
28.00
39.52
47.78
49.67
Thanos
2:4
33.51
73.50
68.90
36.92
66.90
28.00
39.03
47.86
49.33
ProxSparse
2:4
34.86
75.68
71.46
38.31
66.85
28.60
37.51
53.09
50.80
Appendix
Table 13: Model quality (task accuracy across eight zero-shot tasks, reported in %) for LLaMA-2 13B with different pruning methods. PATCH Tile optimizes tile-based sparsity, enabling a flexible sparsity-quality tradeoff. The Avg column reports the eight-task average. MaskLLM is omitted from this table because no public LLaMA-2 13B checkpoint is available and reproducing it would require approximately 2304 A100 GPU-hours, beyond our compute budget.
Sparsity
Method
Pattern
MMLU
PIQA
ARC-E
ARC-C
WinoG.
OBQA
RACE
HellaS.
Avg
0%
Dense
-
63.57
80.09
81.44
51.37
73.48
33.40
39.14
60.02
60.31
50%
Magnitude
2:4
23.06
63.82
45.33
25.94
53.91
15.20
26.70
33.49
35.93
Wanda
2:4
27.85
68.88
58.33
26.71
60.93
19.00
33.78
38.70
41.77
SparseGPT
2:4
31.82
70.46
63.85
31.74
64.56
21.60
37.22
42.99
45.53
Thanos
2:4
34.23
70.40
63.13
31.40
63.61
23.20
37.03
42.75
45.72
ProxSparse
2:4
29.89
71.71
62.63
33.28
58.56
23.80
35.22
46.03
45.14
Appendix
Table 14: Model quality (task accuracy across eight zero-shot tasks, reported in %) for LLaMA-3.1 8B with different pruning methods. PATCH Tile optimizes tile-based sparsity, enabling a flexible sparsity-quality tradeoff.
Sparsity
Method
Pattern
MMLU
PIQA
ARC-E
ARC-C
WinoG.
OBQA
RACE
HellaS.
Avg
0%
Dense
-
37.57
74.54
65.53
31.32
60.62
26.40
37.89
47.76
47.70
50%
Magnitude
2:4
23.31
53.81
27.74
18.94
51.38
11.80
24.02
26.26
29.66
Wanda
2:4
22.90
58.11
37.08
19.20
49.09
13.20
25.17
28.11
31.61
SparseGPT
2:4
22.93
61.43
45.03
22.35
54.93
15.80
29.86
32.08
35.55
Thanos
2:4
23.12
62.40
44.91
21.76
54.30
16.00
31.10
32.09
35.71
ProxSparse
2:4
22.96
60.83
39.44
20.31
51.54
16.80
25.17
31.37
33.55
Appendix
Table 15: Model quality (task accuracy across eight zero-shot tasks, reported in %) for LLaMA-3.2 1B with different pruning methods. PATCH Joint optimizes dense tile locations and sparsity patterns, enabling a flexible sparsity-quality tradeoff.
Sparsity
Method
Pattern
MMLU
PIQA
ARC-E
ARC-C
WinoG.
OBQA
RACE
HellaS.
Avg
0%
Dense
-
24.95
75.03
71.84
34.90
58.64
28.60
34.83
47.26
47.01
50%
Magnitude
2:4
23.08
59.79
37.29
17.66
50.59
14.00
22.87
27.97
31.66
Wanda
2:4
23.96
59.52
48.02
18.34
51.22
14.20
27.85
30.18
34.16
SparseGPT
2:4
23.62
62.79
49.83
19.03
51.54
15.20
30.62
31.99
35.58
Thanos
2:4
23.44
62.24
48.86
18.34
50.12
15.60
30.81
31.28
35.09
ProxSparse
2:4
23.10
64.25
50.72
21.59
53.43
18.00
29.09
32.86
36.63
Appendix
Table 16: Model quality (accuracy across eight zero-shot tasks) for Gemma-3 1B with different pruning methods. PATCH Joint optimizes dense tile locations and sparsity patterns, enabling a flexible sparsity-quality tradeoff.
Sparsity (0.5B)
Nothing
SparseGPT
Wanda
Magnitude
Random
45%
14.80
14.57
14.50
14.48
14.51
35%
13.97
13.84
13.87
13.85
13.79
25%
13.47
13.47
13.37
13.44
13.33
Appendix
Table 17: Perplexity ( ↓ ) under different tile prior initializations. All priors yield nearly identical performance, suggesting that the global sparsity target allows dynamic reallocation of sparsity during training, overriding the influence of fixed initialization.
Sparsity
Method
Optimizer
Logits Init
Gumbel Scaling
Gumbel
Prior(Strength)
Sparse Reg.
Weight Reg.
25%
PATCH Joint
Adam(0.001)
N(0,0.014)
25→350
2→0.05
SparseGPT( 3 )
7
10
35%
PATCH Joint
Adam(0.001)
N(0,0.014)
25→350
2→0.05
SparseGPT( 3 )
7
10
45%
PATCH Joint
Adam(0.001)
N(0,0.014)
25→350
4→0.05
SparseGPT( 3 )
7
10
25%
PATCH Tile
Adam(0.0001)
N(0,0.014)
100→500
2→0.05
SparseGPT( 3 )
3
0.1
35%
PATCH Tile
Adam(0.0001)
N(0,0.014)
100→500
2→0.05
SparseGPT( 3 )
3
0.1
45%
PATCH Tile
Adam(0.0001)
N(0,0.014)
100→500
2→0.05
SparseGPT( 3 )
3
0.1
Appendix
Table 18 : Hyper-parameters used for PATCH Joint and PATCH Tile across sparsity ratios. All hyper parameters were tuned on Qwen-2.5-0.5B.
Figure 4 : Sparsity distribution across Attention and MLP layers under varying global sparsity budgets in Qwen-2.5 0.5B.
Figure 5 : Sparsity distribution across Attention and MLP layers under varying global sparsity budgets in Gemma-3 1B.
Figure 6 : Sparsity distribution across Attention and MLP layers under varying global sparsity budgets in LLaMA-3.2 1B.
Sparsity
Method
Pattern
Qwen-2.5 0.5B
LLaMA-3.2 1B
Gemma-3 1B
LLaMA-2 7B
LLaMA-3.1 8B
Acc (% ↑ )
PPL ( ↓ )
Acc (% ↑ )
PPL ( ↓ )
Acc (% ↑ )
PPL ( ↓ )
Acc (% ↑ )
PPL ( ↓ )
Acc (% ↑ )
PPL ( ↓ )
45%
PATCH
Dense/2:4 Tiles
40.29
14.57
42.08
12.23
42.80
11.96
48.99
6.55
53.60
8.20
45%
Wanda
Unstructured
41.45
18.81
40.76
16.56
42.87
25.38
52.72
6.36
55.67
8.24
45%
SparseGPT
Unstructured
42.31
17.65
42.66
15.01
43.52
22.26
52.77
6.46
56.70
8.21
35%
PATCH
Dense/2:4 Tiles
41.15
13.84
42.72
11.67
43.30
11.48
50.08
6.18
55.28
7.89
35%
Wanda
Unstructured
43.46
15.04
44.60
11.95
45.50
16.98
54.37
5.87
58.68
7.02
Appendix
Table 19 : Model quality (average accuracy across eight zero-shot tasks and perplexity on WikiText2 dataset) for PATCH , Wanda, and SparseGPT. For models with less than or equal to 1B parameters, PATCH Joint optimizes both dense tile locations and sparsity patterns, while for larger models PATCH Tile optimizes only dense tile locations with frozen sparsity patterns, both using Dense/2:4 Tiles pattern allowing continuous sparsity ratios and flexible tradeoffs between sparsity and model quality. Wanda and SparseGPT are unstructured pruning methods.
Figure 7 : Layer-wise sparsity distribution of OWL across models and global sparsity budgets.
Figure 8 : Layer-wise sparsity distribution of AlphaPruning across models and global sparsity budgets.
Figure 9 : AlphaPruning sparsity distribution across attention and MLP layers under varying global sparsity budgets in Gemma-3 1B.
Model
Sparsity (%)
Seed 0 (default)
Seed 25
Seed 26
Seed 42
Qwen-2.5 0.5B
25
13.47
13.41
13.38
13.36
35
13.84
13.89
13.85
13.84
45
14.57
14.59
14.61
14.49
Llama-3.2 1B
25
11.00
11.09
11.03
11.21
35
11.67
11.72
11.56
11.86
45
12.23
12.32
12.26
12.55
Appendix
Table 20 : Perplexity across seeds for Qwen-2.5 0.5B and Llama-3.2 1B.
Figure 10 : Layer-wise sparsity distribution of PATCH across seeds.
Figure 11 : Sparsity distribution across attention and MLP layers under varying global sparsity budgets in Llama-3.2 1B across seeds.
Figure 12 : Sparsity distribution across attention and MLP layers under varying global sparsity budgets in Qwen-2.5 0.5B across seeds.
Model
Sparsity
Method
Wiki PPL ( ↓ )
Avg Acc (% ↑ )
Qwen-2.5 0.5B
45%
PATCH Joint
14.57
40.29
PATCH Joint + FT
14.96
40.87
35%
PATCH Joint
13.84
41.15
PATCH Joint + FT
14.32
41.59
25%
PATCH Joint
13.47
42.39
PATCH Joint + FT
13.85
42.55
Appendix
Table 21 : Model quality (average accuracy across eight zero-shot tasks and perplexity on WikiText2 dataset) for PATCH after a short fine-tuning.
Model
Method
Sparsity (%)
PPL ( ↓ )
Avg Acc (% ↑ )
Qwen-2.5 0.5B
PATCH
45
14.57
40.29
SparseGPT (FFN only)
44
24.79
36.84
MaskLLM (FFN only)
44
14.54
39.34
PATCH
35
13.84
41.15
SparseGPT (FFN except last 5)
35
20.46
38.33
MaskLLM (FFN except last 5)
35
13.92
40.59
Appendix
Table 22 : Comparison of PATCH against an FFN-only 2:4 pruning baseline at matched global sparsity. The baseline applies SparseGPT 2:4 pruning to FFN modules only, optionally keeping the last few decoder blocks dense (denoted “except last k ”) to reach the target global sparsity. Rows labeled MaskLLM apply MaskLLM’s learned 2:4 masks to the FFNs with the same heuristic allocation. PATCH achieves higher average accuracy than every heuristic baseline at comparable sparsity, and lower perplexity in all but three cases, where a MaskLLM-based heuristic is marginally lower. Best value per model and sparsity group in bold .
Model
Method
Sparsity (%)
Backend
Throughput (tok/s)
Speedup
LLaMA-2 7B
Dense
0
cuBLAS
1023.80
1.00 ×
PATCH
25
STOICC
1212.79
1.18 ×
PATCH
35
STOICC
1304.46
1.27 ×
PATCH
45
STOICC
1410.20
1.38 ×
SparseGPT (FFN only)
33
cuSPARSELt
1276.67
1.25 ×
SparseGPT (FFN only)
33
STOICC
1302.27
1.27 ×
Appendix
Table 23 : Throughput comparison of PATCH against FFN-only 2:4 pruning under the same STOICC backend on an A6000 GPU (batch size 16, prefill length 128, generation length 128). At matched global sparsity, the FFN-only baseline matches PATCH ’s throughput, but PATCH achieves significantly higher accuracy (Table 22 ). For reference, we also report fully 2:4 sparse models, which represent the upper bound of attainable acceleration.
Model
Method
Sparsity (%)
PPL ( ↓ )
Avg Acc (% ↑ )
Qwen-2.5 0.5B
SparseGPT (3:4)
25
13.99
43.34
PATCH
25
13.47
42.39
LLaMA-3.2 1B
SparseGPT (3:4)
25
10.88
44.66
PATCH
25
11.00
43.81
Gemma-3 1B
SparseGPT (3:4)
25
13.96
45.27
PATCH
25
11.17
44.07
Appendix
Table 24 : Sparsity-matched comparison between PATCH at 25% sparsity and SparseGPT pruned to a 3:4 pattern (75% density). While SparseGPT 3:4 achieves slightly lower perplexity due to its higher per-group flexibility, the 3:4 pattern is incompatible with NVIDIA Sparse Tensor Cores (which require exactly two non-zeros per group of four) and therefore cannot be accelerated, running at dense speeds. PATCH at 25% sparsity, in contrast, mixes dense and 2:4 tiles and unlocks ∼ 1.18 × end-to-end speedup on standard GPUs.
Model
Method
C4 PPL ( ↓ )
C4 Avg Acc (% ↑ )
SlimPajama PPL ( ↓ )
SlimPajama Avg Acc (% ↑ )
Qwen-2.5 0.5B
MaskLLM (50%)
16.27
38.78
15.22
39.33
PATCH (45%)
15.30
40.02
14.57
40.29
PATCH (35%)
14.34
40.94
13.84
41.15
PATCH (25%)
13.81
41.91
13.47
42.39
LLaMA-3.2 1B
MaskLLM (50%)
14.23
40.97
12.93
41.04
PATCH (45%)
13.10
41.73
12.23
42.08
Appendix
Table 25 : Robustness of PATCH mask training to the choice of calibration corpus. We train identical PATCH configurations on the C4 dataset and on SlimPajama and report WikiText2 perplexity and average zero-shot accuracy. Despite the domain shift, downstream accuracy remains within ∼ 0.6% across datasets, and PATCH outperforms MaskLLM under both calibration corpora.
Sparsity (%)
Max. Execution Tile Size
Throughput (tok/s)
Speedup vs. Dense
0
—
1023.80
1.00 ×
25
64 × 64
1207.08
1.18 ×
32 × 64
1075.98
1.05 ×
35
64 × 64
1288.54
1.26 ×
32 × 64
1142.65
1.12 ×
45
64 × 64
1382.13
1.35 ×
Appendix
Table 26 : Effect of restricting STOICC’s autotuner to specific maximum execution tile sizes on PATCH ’s end-to-end throughput, measured on LLaMA-2 7B with an A6000 GPU (batch size 16, prefill length 128, generation length 128). Limiting the autotuner to 64 × 64 has a small effect on throughput (the unrestricted autotuner already prefers this size in most layers), while restricting it to 32 × 64 underutilizes Tensor Cores and yields slower kernels. Tiles smaller than 32 × 64 are not supported by STOICC’s metadata layout and compression constraints. As reported in Table 4 , model accuracy is largely insensitive to tile size, so the chosen 128 × 128 mask-tile granularity (which the compiler subdivides at execution) provides the best balance between accuracy and hardware efficiency.
Tokens generated
Sparsity (%)
2K
2K (FA2)
8K
8K (FA2)
16K
16K (FA2)
128
25
1.18 ×
1.24 ×
1.09 ×
1.11 ×
1.02 ×
1.03 ×
35
1.27 ×
1.33 ×
1.14 ×
1.17 ×
1.07 ×
1.08 ×
45
1.38 ×
1.45 ×
1.22 ×
1.26 ×
1.14 ×
1.17 ×
1024
25
1.13 ×
1.21 ×
1.13 ×
1.16 ×
1.12 ×
1.14 ×
35
1.18 ×
1.26 ×
1.16 ×
1.22 ×
1.15 ×
1.18 ×
45
1.25 ×
1.38 ×
1.19 ×
1.29 ×
1.19 ×
1.23 ×
Appendix
Table 27 : End-to-end throughput speedup of PATCH over a dense PyTorch baseline for LLaMA-2 7B on an A6000 GPU (batch size 16) at prefill lengths of 2K, 8K, and 16K tokens, with and without the FlashAttention-2 (FA2) attention backend; FA2 speedups are relative to the dense model using the same backend. Speedups decrease as the prefill grows, since attention computation begins to dominate runtime and reduces the relative contribution of accelerating the FFN GEMMs. FA2 reduces the attention overhead and increases the speedup in every setting.
Semi-structured 2:4 sparsity is widely supported by modern accelerators, providing up to a 2x theoretical speedup. However, its strict 50% sparsity constraint often causes non-negligible accuracy degradation under post-training pruning. Meanwhile, existing relaxed sparsity formats either require specialized compiler support or introduce runtime overheads that limit end-to-end speedup. We propose Spense, a practical hybrid sparse-dense format that splits each weight matrix into a 2:4 sparse region and a dense region. This design relaxes the effective sparsity constraint while remaining compatible with existing high-performance sparse and dense GEMM libraries, avoiding both custom compiler support and input activation expansion. Building on this format, we introduce SpenseGPT, a one-shot post-training pruning method that produces sparse and dense regions. Notably, we show that selecting the right dense regions is important, and we devise two different strategies to choose them. Experiments on Qwen3-32B and Seed-OSS-36B demonstrate that our method achieves up to 1.2x end-to-end decoding speedup on B200 GPUs with FP8 precision, while preserving accuracy. To the best of our knowledge, this is the first one-shot pruning demonstration of real-world end-to-end LLM decoding speedup from semi-structured sparse tensor cores on recent GPUs such as B200s, while maintaining model quality.
Jaeseong Lee, Seung-won Hwang, Samyam Rajbhandari
Snowflake AI Research* · Seoul National University†
Semi-structured sparsity provides a practical path to accelerate large language models (LLMs) with native hardware support, but post-training semi-structured pruning often suffers from substantial quality degradation due to strong structural coupling. Existing methods rely on large-scale sparse retraining to recover accuracy, resulting in high computational cost. We propose SparseForge, a post-training framework that improves recovery efficiency by directly optimizing the sparsity mask rather than scaling up retraining tokens. SparseForge combines Hessian-aware importance estimation with progressive annealing of soft masks into hardware-executable structured sparsity, enabling stable and efficient sparse recovery. On LLaMA-2-7B under 2:4 sparsity, SparseForge achieves 57.27% average zero-shot accuracy with only 5B retraining tokens, surpassing the dense model's 56.43% accuracy and approaching the 57.52% result of a state-of-the-art method using 40B tokens. Such improvements on the accuracy-efficiency trade-off from SparseForge are shown to be consistent across model families.
One-shot pruning methods like Wanda and SparseGPT apply the same sparsity ratio to every layer of a transformer, ignoring known variation in layer importance. We propose PALS (Percentile-Aware Layerwise Sparsity), which adjusts per-layer sparsity based on the 99th percentile of activation magnitudes, bounded to ±5% around the target ratio. On LLaMA-2-7B at 50% sparsity, PALS achieves 10.96 WikiText-2 perplexity versus 12.92 for uniform Wanda (mean over 9 runs, p<0.001). The benefit is architecture-dependent: LLaMA-3-8B shows marginal gains and Mistral-7B shows none. We also find that gradient-based allocation -- the seemingly more principled approach -- produces results worse than random, suggesting that gradient magnitude does not predict the impact of discrete weight removal. PALS adds negligible cost to the pruning pipeline and requires no fine-tuning.