Residual stream pruning methods reduce inference cost by shrinking the model's hidden dimension, but existing approaches typically choose these dimensions by minimizing activation reconstruction error. This criterion implicitly treats all perturbation directions as equally important, ignoring the sensitivity of downstream layers. We introduce a sensitivity-aware approach to residual-stream pruning that directly accounts for this direction-dependent sensitivity. Using a second-order approximation to the output KL divergence, we characterize the effect of a residual-stream perturbation through both its activation covariance and the local sensitivity of the model output. The resulting subspace selection objective couples these two quantities, but is difficult to optimize directly. We derive a tractable spectral upper bound that reduces subspace selection to an eigendecomposition of a sensitivity-weighted covariance matrix, retaining the efficiency and structural simplicity of rotation-based pruning methods. Across several instruction-tuned language model families, our method consistently reduces calibration KL divergence relative to activation-only pruning and improves perplexity and downstream task performance over a range of compression levels. Our results show that preserving activation energy alone is insufficient for residual-stream pruning, and that explicitly accounting for how perturbations propagate to the model output provides a more effective criterion for selecting dimensions to remove.
Figures & tables
Figure 1: Overview of our structured residual-stream pruning method. (a) We remove residual-stream dimensions by rotating unimportant directions into removable coordinates, allowing corresponding columns and rows of the attention and MLP weight matrices to be deleted. (b) Directions are selected using both activation variance and output sensitivity. (c) Our objective balances activation reconstruction and output sensitivity through τ , interpolating between the two extremes.
Figure 2: Validation of the proposed approach on Llama 3.2 3B. (a) measures the correlation between our sensitivity approximation with KL. We fix activation reconstruction error to the error induced at 25% pruning rate and vary perturbation direction. (b) shows the value of Equation 3 attained for bases which minimize Equation 5 for various values of τ . Results are shown across different layers at a pruning rate of 25% . We compare with the loss achieved by the SliceGPT solution.
Llama-3.x Instruct
Mistral Instruct
Phi-3 Instruct
Sparsity
Method
3.2 3B
3.1 8B
7B v0.3
Nemo
Medium
DKL
PPL
Acc.
DKL
PPL
Acc.
DKL
PPL
Acc.
DKL
PPL
Acc.
DKL
PPL
Acc.
0%
N/A
0.0
11.76
60.55
0.0
7.21
68.53
0.0
5.49
69.72
0.0
6.09
70.05
0.0
4.30
72.95
10%
SliceGPT
0.5059
19.90
58.11
0.7507
15.63
65.78
0.1944
6.58
67.91
0.2791
8.11
67.06
0.5192
6.38
72.84
Ours
0.2017
14.30
58.53
0.4040
11.04
66.27
0.1737
6.47
68.02
0.2170
7.62
67.76
0.4373
6.05
72.56
20%
SliceGPT
0.9729
31.40
53.49
1.3165
27.87
61.06
0.4883
8.75
64.25
0.5991
11.23
62.28
1.0421
10.68
68.63
Table 1: Comparison across model families and compression rates. We report KL divergence, perplexity (PPL), and average downstream accuracy (Acc.).
Llama-3.2 3B Instruct
Llama-3.1 8B Instruct
Mistral Nemo Instruct
Sparsity
Method
IFEval
GSM8K
IFEval
GSM8K
IFEval
GSM8K
Prompt
Inst.
Strict
Prompt
Inst.
Strict
Prompt
Inst.
Flexible
0%
N/A
71.4
79.5
76.6
74.5
81.8
85.3
56.0
67.4
80.0
10%
SliceGPT
58.8
69.3
66.0
63.4
72.4
80.3
52.7
63.1
77.2
Ours
63.4
73.9
69.7
65.8
75.7
81.4
54.7
65.4
78.5
20%
SliceGPT
50.5
62.4
61.6
54.2
66.0
74.0
47.3
58.4
67.1
Table 2: Comparison of SliceGPT and our method on mathematical reasoning and instruction following tasks. For IFEval, we report strict prompt-level (Prompt) and instruction-level (Inst.) accuracy. For GSM8K, we report strict extraction accuracy for Llama models and flexible extraction for Mistral. All results are reported as percentages.
Figure 3: Average common-sense reasoning performance of the proposed method and SliceGPT using varying numbers of calibration samples on Llama-3.1 8B Instruct.
Sparsity
Add. Opt.
DKLcal
DKLwiki
PPL
Avg. Acc.
10%
×
0.1369
0.4040
11.04
66.27
✓
0.1371
0.4038
11.11
66.37
20%
×
0.3289
0.8820
18.04
62.13
✓
0.3252
0.9178
18.75
62.73
25%
×
0.4611
1.1569
23.79
59.02
✓
0.4520
1.1856
24.57
59.44
Table 3: Effect of additional optimization of L(U) on Llama-3.1-8B-Instruct. DKLcal and DKLwiki denote KL divergence on the calibration data and WikiText-2 test set, respectively.
Model
Method
Sensitivity Estimation
Direction Selection
Total
Llama-3.1 8B
SliceGPT
–
01:03
01:03
Ours
00:09
01:05
01:14
Phi-3 Medium
SliceGPT
–
01:57
01:57
Ours
00:18
02:01
02:19
Table 4: Calibration runtime (HH:MM) comparison on a single NVIDIA H100 GPU. Direction selection includes layerwise covariance accumulation and pruning-basis construction.
Method
10%
20%
25%
30%
LLM-Pruner
62.91
53.97
45.08
42.98
OSSCAR
64.51
58.74
55.16
46.27
Wanda-sp
65.96
59.58
52.91
49.78
FLAP
63.59
58.14
53.72
51.21
SliceGPT
65.78
61.06
57.97
53.99
Ours
66.27
62.13
59.02
55.72
Table 5: Average common-sense reasoning accuracy of additional structured pruning methods on Llama-3.1-8B-Instruct across pruning rates.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Sparsity
Method
WG
ARC-E
ARC-C
HS
OBQA
PIQA
MMLU
Avg.
0%
N/A
67.32
67.85
46.16
70.46
36.00
75.52
60.53
60.55
10%
SliceGPT
65.59
68.90
42.06
65.49
37.00
73.67
54.09
58.11
Ours
67.17
67.98
42.58
65.54
37.40
73.61
55.42
58.53
20%
SliceGPT
64.40
63.89
37.54
56.64
35.40
70.13
46.45
53.49
Ours
65.27
64.39
40.10
58.61
35.60
70.89
49.55
54.92
25%
SliceGPT
62.43
59.01
34.56
51.95
34.40
68.17
43.46
50.57
Appendix
Table 6: Per-task common-sense reasoning accuracy on Llama-3.2 3B Instruct.
Sparsity
Method
WG
ARC-E
ARC-C
HS
OBQA
PIQA
MMLU
Avg.
0%
N/A
73.88
79.59
54.95
79.25
43.00
80.96
68.10
68.53
10%
SliceGPT
73.16
77.90
51.96
74.66
42.60
78.67
61.52
65.78
Ours
72.77
78.32
52.82
75.27
42.00
79.00
63.70
66.27
20%
SliceGPT
68.75
73.48
46.16
66.71
41.80
74.37
56.16
61.06
Ours
70.56
75.51
46.93
67.15
41.80
74.97
58.00
62.13
25%
SliceGPT
67.01
69.44
42.49
61.60
39.60
71.93
53.75
57.97
Appendix
Table 7: Per-task common-sense reasoning accuracy on Llama-3.1 8B Instruct.
Table 9: Per-task common-sense reasoning accuracy on Mistral Nemo Instruct.
Sparsity
Method
WG
ARC-E
ARC-C
HS
OBQA
PIQA
MMLU
Avg.
0%
N/A
76.56
81.36
61.60
82.76
50.60
81.66
76.14
72.95
10%
SliceGPT
76.56
85.98
63.74
80.37
49.60
81.83
71.79
72.84
Ours
77.43
84.34
61.77
80.38
48.60
81.18
74.21
72.56
20%
SliceGPT
73.24
82.66
58.36
74.26
46.00
79.05
66.81
68.63
Ours
74.98
82.62
58.02
75.46
46.20
79.71
69.80
69.54
25%
SliceGPT
71.74
78.83
54.69
70.33
44.00
77.04
59.58
65.17
Appendix
Table 10: Per-task common-sense reasoning accuracy on Phi-3 Medium Instruct.
Llama-3.2 3B Instruct
Llama-3.1 8B Instruct
Sparsity
Method
DKL
PPL
Acc.
DKL
PPL
Acc.
0%
N/A
0.0
11.76
60.55
0.0
7.21
68.53
10%
SliceGPT
0.2985
16.43
56.81
0.5505
12.60
64.60
Ours
0.1540
13.95
57.25
0.4458
11.49
65.18
20%
SliceGPT
0.6387
23.04
50.55
1.0121
20.31
57.25
Ours
0.4209
18.31
51.44
0.8330
17.09
57.93
Appendix
Table 11: Results using RedPajama as the calibration set. We report KL divergence, perplexity (PPL), and average downstream accuracy (Acc.).
Llama-2 7B
Llama-3.2 3B
Llama-3.1 8B
Sparsity
Method
DKL
PPL
Acc.
DKL
PPL
Acc.
DKL
PPL
Acc.
0%
N/A
0.0
5.47
61.56
0.0
7.82
62.19
0.0
6.24
68.03
10%
SliceGPT
0.3736
8.16
60.01
0.1770
15.59
57.44
0.9215
15.92
64.75
Ours
0.2007
6.86
60.10
0.1340
12.56
58.31
0.5836
11.42
65.53
20%
SliceGPT
0.9157
14.68
56.57
0.4287
30.46
51.78
1.4205
26.42
58.99
Ours
0.5178
9.62
56.82
0.3377
19.59
53.25
1.1735
20.82
60.42
Appendix
Table 12: Results on non-instruction-tuned models. We report KL divergence, perplexity (PPL), and average downstream accuracy (Acc.).
Figure 4: Comparison of dataset-averaged and token-wise sensitivity across candidate pruning bases. For each basis, we rescale its induced perturbations by a single scalar such that all bases have equal total reconstruction error, isolating differences in output sensitivity from differences in perturbation magnitude. Each panel reports the correlation between the objective computed using the dataset-averaged curvature matrix H and the corresponding objective computed using token-specific matrices Hi(s) . The strong agreement indicates that averaging sensitivity across the calibration distribution preserves the relative sensitivity of pruning-relevant directions.
Pruning reduces the inference cost of large language models, but existing criteria primarily preserve large activations or reconstruct layer outputs. We argue that this overlooks a key computation performed by particularly sparsity-sensitive neurons in the MLP up and gate projections: separating similar inputs into dissimilar outputs. This suggests that effective pruning should preserve not only activations, but also the differences between outputs more broadly. We introduce a family of difference-informed pruning methods built upon this principle. Wisp is a first-order, update-free method that scores weights using input-difference norms, and Wisp+ refines this score neuronwise using the input pairs each neuron separates most strongly. Finally, Whisper is a second-order method that uses a lightly regularized difference Hessian as its reconstruction objective. Across Llama 2 and 3.1 models from 7B to 405B parameters, our second-order variant consistently improves over strong reconstruction-based baselines, while our update-free variants improve over activation-aware baselines, especially in constrained settings. The improvements over Wanda and SparseGPT extend to structured sparsity, downstream evaluations, and other model families. Augmenting stronger techniques such as RIA and ALPS with our difference-informed criteria yields further improvements, shifting the overall accuracy-runtime frontier outward at negligible additional cost. These results suggest that preserving output differences is a broadly useful and composable signal for post-training LLM sparsification.
Structured pruning of large language models (LLMs) offers hardware-efficient compression, yet existing methods require calibration data, gradient computation, or large auxiliary policy networks at pruning time. LILA (\emph{Latent-Informed Layer Analysis}) scores neuron importance via the Kolmogorov--Smirnov (KS) distance between empirical singular value distributions of the full and neuron-ablated feed-forward network (FFN) weight matrix, providing a closed-form spectral rule requiring no training, calibration data, or auxiliary network. Without any fine-tuning, LILA surpasses PruneNet (45M-parameter RL policy) by 1.57pp in zero-shot accuracy on LLaMA-2-7B at 25% sparsity, and outperforms WikiText-2-calibrated SliceGPT by up to 6.0pp across all sparsity levels, while preserving the original architecture. After one epoch of LoRA recovery fine-tuning, LILA achieves highly competitive performance, matching the heavily calibrated SliceGPT baseline to within a 0.48~pp margin across LLaMA-2-7B and Phi-2, despite using zero calibration data. A Neural Tangent Kernel analysis confirms a 22× reduction in functional distortion versus random pruning, providing theoretical grounding for the spectral importance criterion. Finally, extending LILA to dynamically allocate sparsity budgets via KS-scores yields state-of-the-art generative preservation at moderate compression, while uncovering fundamental single-layer architectural bottlenecks at higher compression regimes.
Depth pruning reduces large language model (LLM) inference cost by removing complete Transformer blocks. Activation-based methods collect hidden states through forward passes on calibration data, while existing forward-free methods score each Transformer block separately without measuring similarity between blocks. We propose Weight-Redundancy Pruning (WRP), a forward-free depth-pruning method that estimates inter-layer redundancy from checkpoint weights to select blocks without calibration data or model forward passes. WRP compares attention output and MLP down-projection weights across layers and combines their pairwise similarities with relative projection-scale information. The resulting all-pairs similarity matrix guides layer grouping and block selection. Across multiple pruning settings, model families, and downstream tasks, WRP consistently outperforms existing forward-free magnitude pruning and approaches the performance of activation-based methods.
Vincent-Daniel Yun, Woosang Lim
University of Southern California, United States · 2Seoul National University, Republic of Korea