Residual stream pruning methods reduce inference cost by shrinking the model's hidden dimension, but existing approaches typically choose these dimensions by minimizing activation reconstruction error. This criterion implicitly treats all perturbation directions as equally important, ignoring the sensitivity of downstream layers. We introduce a sensitivity-aware approach to residual-stream pruning that directly accounts for this direction-dependent sensitivity. Using a second-order approximation to the output KL divergence, we characterize the effect of a residual-stream perturbation through both its activation covariance and the local sensitivity of the model output. The resulting subspace selection objective couples these two quantities, but is difficult to optimize directly. We derive a tractable spectral upper bound that reduces subspace selection to an eigendecomposition of a sensitivity-weighted covariance matrix, retaining the efficiency and structural simplicity of rotation-based pruning methods. Across several instruction-tuned language model families, our method consistently reduces calibration KL divergence relative to activation-only pruning and improves perplexity and downstream task performance over a range of compression levels. Our results show that preserving activation energy alone is insufficient for residual-stream pruning, and that explicitly accounting for how perturbations propagate to the model output provides a more effective criterion for selecting dimensions to remove.
Figures & tables
Figure 1: Overview of our structured residual-stream pruning method. (a) We remove residual-stream dimensions by rotating unimportant directions into removable coordinates, allowing corresponding columns and rows of the attention and MLP weight matrices to be deleted. (b) Directions are selected using both activation variance and output sensitivity. (c) Our objective balances activation reconstruction and output sensitivity through τ , interpolating between the two extremes.
Figure 2: Validation of the proposed approach on Llama 3.2 3B. (a) measures the correlation between our sensitivity approximation with KL. We fix activation reconstruction error to the error induced at 25% pruning rate and vary perturbation direction. (b) shows the value of Equation 3 attained for bases which minimize Equation 5 for various values of τ . Results are shown across different layers at a pruning rate of 25% . We compare with the loss achieved by the SliceGPT solution.
Llama-3.x Instruct
Mistral Instruct
Phi-3 Instruct
Sparsity
Method
3.2 3B
3.1 8B
7B v0.3
Nemo
Medium
DKL
PPL
Acc.
DKL
PPL
Acc.
DKL
PPL
Acc.
DKL
PPL
Acc.
DKL
PPL
Acc.
0%
N/A
0.0
11.76
60.55
0.0
7.21
68.53
0.0
5.49
69.72
0.0
6.09
70.05
0.0
4.30
72.95
10%
SliceGPT
0.5059
19.90
58.11
0.7507
15.63
65.78
0.1944
6.58
67.91
0.2791
8.11
67.06
0.5192
6.38
72.84
Ours
0.2017
14.30
58.53
0.4040
11.04
66.27
0.1737
6.47
68.02
0.2170
7.62
67.76
0.4373
6.05
72.56
20%
SliceGPT
0.9729
31.40
53.49
1.3165
27.87
61.06
0.4883
8.75
64.25
0.5991
11.23
62.28
1.0421
10.68
68.63
Table 1: Comparison across model families and compression rates. We report KL divergence, perplexity (PPL), and average downstream accuracy (Acc.).
Llama-3.2 3B Instruct
Llama-3.1 8B Instruct
Mistral Nemo Instruct
Sparsity
Method
IFEval
GSM8K
IFEval
GSM8K
IFEval
GSM8K
Prompt
Inst.
Strict
Prompt
Inst.
Strict
Prompt
Inst.
Flexible
0%
N/A
71.4
79.5
76.6
74.5
81.8
85.3
56.0
67.4
80.0
10%
SliceGPT
58.8
69.3
66.0
63.4
72.4
80.3
52.7
63.1
77.2
Ours
63.4
73.9
69.7
65.8
75.7
81.4
54.7
65.4
78.5
20%
SliceGPT
50.5
62.4
61.6
54.2
66.0
74.0
47.3
58.4
67.1
Table 2: Comparison of SliceGPT and our method on mathematical reasoning and instruction following tasks. For IFEval, we report strict prompt-level (Prompt) and instruction-level (Inst.) accuracy. For GSM8K, we report strict extraction accuracy for Llama models and flexible extraction for Mistral. All results are reported as percentages.
Figure 3: Average common-sense reasoning performance of the proposed method and SliceGPT using varying numbers of calibration samples on Llama-3.1 8B Instruct.
Sparsity
Add. Opt.
DKLcal
DKLwiki
PPL
Avg. Acc.
10%
×
0.1369
0.4040
11.04
66.27
✓
0.1371
0.4038
11.11
66.37
20%
×
0.3289
0.8820
18.04
62.13
✓
0.3252
0.9178
18.75
62.73
25%
×
0.4611
1.1569
23.79
59.02
✓
0.4520
1.1856
24.57
59.44
Table 3: Effect of additional optimization of L(U) on Llama-3.1-8B-Instruct. DKLcal and DKLwiki denote KL divergence on the calibration data and WikiText-2 test set, respectively.
Model
Method
Sensitivity Estimation
Direction Selection
Total
Llama-3.1 8B
SliceGPT
–
01:03
01:03
Ours
00:09
01:05
01:14
Phi-3 Medium
SliceGPT
–
01:57
01:57
Ours
00:18
02:01
02:19
Table 4: Calibration runtime (HH:MM) comparison on a single NVIDIA H100 GPU. Direction selection includes layerwise covariance accumulation and pruning-basis construction.
Method
10%
20%
25%
30%
LLM-Pruner
62.91
53.97
45.08
42.98
OSSCAR
64.51
58.74
55.16
46.27
Wanda-sp
65.96
59.58
52.91
49.78
FLAP
63.59
58.14
53.72
51.21
SliceGPT
65.78
61.06
57.97
53.99
Ours
66.27
62.13
59.02
55.72
Table 5: Average common-sense reasoning accuracy of additional structured pruning methods on Llama-3.1-8B-Instruct across pruning rates.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Sparsity
Method
WG
ARC-E
ARC-C
HS
OBQA
PIQA
MMLU
Avg.
0%
N/A
67.32
67.85
46.16
70.46
36.00
75.52
60.53
60.55
10%
SliceGPT
65.59
68.90
42.06
65.49
37.00
73.67
54.09
58.11
Ours
67.17
67.98
42.58
65.54
37.40
73.61
55.42
58.53
20%
SliceGPT
64.40
63.89
37.54
56.64
35.40
70.13
46.45
53.49
Ours
65.27
64.39
40.10
58.61
35.60
70.89
49.55
54.92
25%
SliceGPT
62.43
59.01
34.56
51.95
34.40
68.17
43.46
50.57
Appendix
Table 6: Per-task common-sense reasoning accuracy on Llama-3.2 3B Instruct.
Sparsity
Method
WG
ARC-E
ARC-C
HS
OBQA
PIQA
MMLU
Avg.
0%
N/A
73.88
79.59
54.95
79.25
43.00
80.96
68.10
68.53
10%
SliceGPT
73.16
77.90
51.96
74.66
42.60
78.67
61.52
65.78
Ours
72.77
78.32
52.82
75.27
42.00
79.00
63.70
66.27
20%
SliceGPT
68.75
73.48
46.16
66.71
41.80
74.37
56.16
61.06
Ours
70.56
75.51
46.93
67.15
41.80
74.97
58.00
62.13
25%
SliceGPT
67.01
69.44
42.49
61.60
39.60
71.93
53.75
57.97
Appendix
Table 7: Per-task common-sense reasoning accuracy on Llama-3.1 8B Instruct.
Table 9: Per-task common-sense reasoning accuracy on Mistral Nemo Instruct.
Sparsity
Method
WG
ARC-E
ARC-C
HS
OBQA
PIQA
MMLU
Avg.
0%
N/A
76.56
81.36
61.60
82.76
50.60
81.66
76.14
72.95
10%
SliceGPT
76.56
85.98
63.74
80.37
49.60
81.83
71.79
72.84
Ours
77.43
84.34
61.77
80.38
48.60
81.18
74.21
72.56
20%
SliceGPT
73.24
82.66
58.36
74.26
46.00
79.05
66.81
68.63
Ours
74.98
82.62
58.02
75.46
46.20
79.71
69.80
69.54
25%
SliceGPT
71.74
78.83
54.69
70.33
44.00
77.04
59.58
65.17
Appendix
Table 10: Per-task common-sense reasoning accuracy on Phi-3 Medium Instruct.
Llama-3.2 3B Instruct
Llama-3.1 8B Instruct
Sparsity
Method
DKL
PPL
Acc.
DKL
PPL
Acc.
0%
N/A
0.0
11.76
60.55
0.0
7.21
68.53
10%
SliceGPT
0.2985
16.43
56.81
0.5505
12.60
64.60
Ours
0.1540
13.95
57.25
0.4458
11.49
65.18
20%
SliceGPT
0.6387
23.04
50.55
1.0121
20.31
57.25
Ours
0.4209
18.31
51.44
0.8330
17.09
57.93
Appendix
Table 11: Results using RedPajama as the calibration set. We report KL divergence, perplexity (PPL), and average downstream accuracy (Acc.).
Llama-2 7B
Llama-3.2 3B
Llama-3.1 8B
Sparsity
Method
DKL
PPL
Acc.
DKL
PPL
Acc.
DKL
PPL
Acc.
0%
N/A
0.0
5.47
61.56
0.0
7.82
62.19
0.0
6.24
68.03
10%
SliceGPT
0.3736
8.16
60.01
0.1770
15.59
57.44
0.9215
15.92
64.75
Ours
0.2007
6.86
60.10
0.1340
12.56
58.31
0.5836
11.42
65.53
20%
SliceGPT
0.9157
14.68
56.57
0.4287
30.46
51.78
1.4205
26.42
58.99
Ours
0.5178
9.62
56.82
0.3377
19.59
53.25
1.1735
20.82
60.42
Appendix
Table 12: Results on non-instruction-tuned models. We report KL divergence, perplexity (PPL), and average downstream accuracy (Acc.).
Figure 4: Comparison of dataset-averaged and token-wise sensitivity across candidate pruning bases. For each basis, we rescale its induced perturbations by a single scalar such that all bases have equal total reconstruction error, isolating differences in output sensitivity from differences in perturbation magnitude. Each panel reports the correlation between the objective computed using the dataset-averaged curvature matrix H and the corresponding objective computed using token-specific matrices Hi(s) . The strong agreement indicates that averaging sensitivity across the calibration distribution preserves the relative sensitivity of pruning-relevant directions.