SVD-based pruning and quantization have recently emerged as a promising strategy for the ultra-efficient compression of large language models. In these methods, compression is performed in two stages: components are first truncated, and the remaining ones are subsequently quantized. Although this decoupled pipeline benefits from both pruning and quantization, it requires separate optimization for each stage and fails to fully exploit their balance, which can lead to suboptimal performance under aggressive compression. To address this limitation, we propose a new LLM compression method that co-optimizes pruning and quantization in a unified framework. Our key idea is a differentiable method for learning component-wise bit-widths, allowing less important components to be assigned 0-bit precision and pruned away. Notably, our method performs favorably against two-stage baselines, even when subjected to extreme quantization settings (1.61 bits) designed for ultra-efficiency. Code: https://github.com/MMAI-Laboratory/DBW.
Figures & tables
Figure 1 : Our compression concept illustration. We co-optimize pruning and quantization in a single step by updating bit-widths in our differentiable method.
Figure 2 : Our differentiable method illustration. (a) Non-differentiable method uses a standard quantizer [ 3 ] for uniform precision. (b) DBW learns bit-width using the bit-width quantizer (BQ) and learned bit-width quantizer (LBQ).
Figure 3 : Our mixed-precision GEMM illustration. Ours (a) compress the weight memory by bit-group-wise packing and (b) use parallel execution of bit-group GEMMs to maximize GPU.
LLaMA
Method
Group
Bits
PIQA
ARC-e
HellaS
Wing
Race
ARC-c
LAMB-o
LAMB-s
Avg.
1-7B
FP16
-
16
78.67
75.29
56.99
70.01
40.29
41.81
73.57
67.82
63.06
GPTQ
128
2
59.41
36.11
32.59
50.43
27.56
20.31
12.79
9.02
31.03
SliM-LLM +
128
2
65.72
52.74
39.49
55.33
34.35
23.81
38.15
22.96
41.57
OmniQuant
128
2
67.74
59.51
41.24
56.75
33.30
28.41
40.89
25.75
44.20
PB-LLM
128
1.7(+1)
55.71
29.12
28.31
48.86
26.41
19.80
10.42
10.09
28.59
BiLLM
128
1(+1.1)
61.10
40.99
31.80
53.67
30.14
20.64
23.15
16.48
36.00
Table 1 : Commonsense reasoning accuracy evaluation. We report the commonsense reasoning accuracy using the LLaMA model families. Following the recent practice in PTQ1.61, we report the bit configuration as weight bits(+mask bits). Specifically, PB-LLM uses 1.7-bit weights with an additional 1-bit mask overhead, while BiLLM employs 1-bit weights with a mask cost of 1.1 bits.
Dataset
Method
Group
Bits
LLaMA
LLaMA-2
LLaMA-3
7B
13B
30B
65B
*Avg
7B
13B
70B
*Avg
8B
WikiText2
FP16
-
16
5.68
5.09
4.10
3.53
4.60
5.47
4.88
3.31
4.55
6.14
AWQ
128
2
2.6e5
2.8e5
2.4e5
7.4e4
2.1e5
2.2e5
1.2e5
-
1.7e5
1.7e6
GPTQ
128
2
152.31
20.44
13.01
9.51
48.82
60.45
28.14
8.78
32.46
210.00
QuIP
128
2
29.74
12.48
11.57
7.83
15.41
39.73
13.48
6.64
19.95
84.97
OmniQuant
128
2
9.72
7.93
7.12
5.95
7.68
11.06
8.26
6.55
8.62
5.3e3
Table 2 : Perplexity evaluation. We report perplexity on WikiText2 and C4 across LLaMA models. *Avg is the arithmetic mean over available model sizes, excluding missing (-) or NaN entries.
Table 6
Figure 4 : Memory-accuracy trade-off evaluation. We compare our DBW with representative ultra-low-bit baselines on LLaMA-2-7B, LLaMA-3-8B, and Mistral-7B. Each point shows the model memory footprint (GB) and the mean accuracy on the eight commonsense reasoning benchmarks. The shaded regions highlight the extreme memory-compression regimes ( e.g. , 7.5× , 5.5× , and 7.5× ).
Figure 5 : End-to-end latency evaluation.
Type
Method (Bit)
LLaMA-7B
LLaMA-2-7B
LLaMA-3-8B
Ratio
Mem
PPL ↓
Acc. ↑
Ratio
Mem
PPL ↓
Acc. ↑
Ratio
Mem
PPL ↓
Acc. ↑
-
Baseline (FP16)
0.00
12.31
6.38
63.06
0.00
12.31
6.22
63.19
0.00
13.98
7.51
65.98
sequential optimization
ASVD (FP16)
46.33
6.72
3.0e4
25.02
46.33
6.72
4.5e4
23.94
46.33
7.96
2.6e4
24.43
ASVD+GPTQ (1.61)
46.33
1.56
3.0e4
25.00
46.33
1.56
4.5e4
24.63
46.33
2.40
2.6e4
24.26
ASVD+OmniQ (1.61)
46.33
1.56
3.1e4
25.02
46.33
1.56
8.9e4
24.89
46.33
2.40
5.3e4
24.97
SVD-LLM (FP16)
46.33
6.72
14.08
47.21
46.33
6.72
59.28
31.59
46.33
7.96
53.76
36.55
Table 5 : Sequential vs. co-optimization of pruning and quantization. We evaluate compression methods using SVD with PTQ approaches. ASVD and SVD-LLM apply the sequential optimization of pruning and quantization, while ours co-optimizes them concurrently. Note that methods in parentheses with FP16 denote the pruning-only results. We report the average perplexity ( PPL ↓ ) of the WikiText-2 and C4 datasets and the mean accuracy ( Acc. ↑ ) on the eight commonsense benchmarks. Mem is the memory footprint (GB) of the compressed model. For a controlled comparison, we match the pruning Ratio and bit-width of the baselines to those of DBW , providing more results in Section D.3 .
Method (Bit)
LLaMA-7B
LLaMA-2-7B
LLaMA-3-8B
Ratio
Mem
PPL ↓
Acc. ↑
Ratio
Mem
PPL ↓
Acc. ↑
Ratio
Mem
PPL ↓
Acc. ↑
Baseline (FP16)
0.00
12.31
6.38
63.06
0.00
12.31
6.22
63.19
0.00
13.98
7.51
65.98
LLM-Pruner (FP16)
44.40
6.84
96.46
39.70
44.40
6.84
108.33
38.43
43.50
7.90
93.37
41.45
StreamLine (FP16)
45.10
6.76
2.0e3
29.98
45.10
6.76
126.99
35.85
43.50
7.90
63.66
36.85
Our DBW (1.61)
47.21
1.56
10.65
48.98
46.95
1.56
10.54
49.06
44.71
2.40
19.38
44.38
Table 6 : Structured pruning method evaluation under matched pruning ratio. We compare DBW with structured pruning baselines by matching their pruning Ratios . The structured pruning baselines keep the remaining weights in FP16, whereas DBW combines SVD-rank pruning with low-bit quantization; therefore, Mem is reported to make the storage difference explicit.
Figure 6 : Learned bit-width analysis.
Table 7 : Controlled factor analysis of SVD-space bit allocation. We disentangle two factors of DBW: (a) the bit-allocation method and (b) the set of learnable targets. All variants are evaluated under the same training protocol; only the factor specified in each sub-table is changed. In (a), Uniform denotes a non-adaptive allocation that assigns a fixed precision pattern to SVD components, while Heuristic denotes a non-differentiable SVD-aware mixed-bit allocation based on singular-value importance under the same budget. Learnable replaces this static/discrete allocation with our differentiable bit-width learning. In (b), we progressively enable learning of component-level bit-widths β , matrix-level bit budgets α , and singular values σ . PPL is the average perplexity on WikiText2 and C4, and Acc. is the average accuracy on eight commonsense reasoning benchmarks.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Type
Value
dataset
RedPajama-1T
#samples
4096
sequence length
2048
epoch
2
optimizer
adam
lr scheduler
cosine
Appendix
Table 8 : Hyperparameter specification. We provide an in-depth hyperparameter specification.
Method
Dataset
#SeenSamples
SVD-LLM
Alpaca
50,000
OmniQuant
WikiText-2
5,120
SliM-LLM +
WikiText-2
6,400
PTQ1.61
RedPajama-1T
20,000
DBW (Ours)
RedPajama-1T
8,192
Appendix
Table 9 : Calibration setting comparison. We compare our calibration setting against the existing SVD pruning and post-training quantization method. For the methods using multiple datasets, we only report the main dataset that requires the most time and GPU memory. The #SeenSamples denotes the total number of samples seen during training, which is computed as the number of epochs multiplied by the number of samples.
Metric
Mean
Std.
95% CI
PPL ↓
10.58
± 0.04
± 0.09
Acc. ↑
49.32
± 0.23
± 0.57
Appendix
Table 10 : Statistical variation analysis. We report statistical values of LLaMA-2-7B models over three calibration seeds such as 44, 45, and 46.
Figure 7 : Optimal trade-off concept illustration. (a) We depict the reconstruction error of truncation and quantization as a function of singular-value entropy. (b) For the low SVD entropy, the quantization leads to a high error rate due to assigning uniform precision across components. (c) When the SVD entropy is high, the truncation error becomes severe due to removing a considerable information that is evenly distributed among components.
Figure 8 : Adaptive mixed-bit allocation illustration. While the singular value distribution dictates the optimal bit-width, conventional static methods rely on non-differentiable thresholds. In contrast, our adaptive approach assigns differentiable thresholds based on the singular value distribution, naturally facilitating both pruning (bit=0) and quantization (bit>0) in a unified manner.
Method
Leftover bit allocation
PPL ↓
Acc. ↑
Floor only
hi=0
11.83
46.13
Random
hi={1,0,i∈γrand,i∈/γrand,
11.86
46.18
Hamilton
hi={1,0,i∈γham,i∈/γHam.
10.54
49.06
Appendix
Table 11 : Hamilton correction analysis. We compare different leftover bit allocation methods for assigning the leftover bit indicator hi after flooring the continuous bit allocation. γrand is a uniformly sampled subset of feasible components with ∣γrand∣=R . γham denotes the top- R components with the largest remainders. Floor only ignores the leftover bits, Random assigns them to random feasible components, and Hamilton uses the largest-remainder rule. Hamilton gives a deterministic budget-preserving projection and yields stable final performance. This analysis uses the LLaMA-2-7B model.
Figure 9 : Our learned bit quantizer illustration. We use the learned bit quantizer to allow the backpropagation for the bit-width parameters.
Figure 10 : Gradient analysis. We analyze the gradient of bit-width parameters as a function of the weight values. Compared to the commonly used LSQ method, the proposed method can provide the gradient of rounding error to the bit-width parameters.
Figure 11 : Trajectory of learned bit budget. We visualize the learned bit budgets during calibration on the first MLP layer of LLaMA-2-7B under two matrix-level bit learning rates. Across both settings, the learned bit-budgets remain far from zero and gradually converge to a stable range, indicating that the matrix-level budget does not collapse during optimization. This supports our observation that the reconstruction loss acts as an implicit lower-bound regularizer for the bit-budget parameters.
Figure 12 : Our on-the-fly dequantization kernel. Our on-the-fly dequantization kernel converts a pair of separated integer values ( e.g. , INT2 and INT1 in this case) into FP16 format by efficiently fusing them via (a) shift and (b) logical & operator.
Bit
Memory
Compression
Runtime
Speed up
FP16
1,544MB
-
3.08ms
-
INT5
483MB
3.2×
2.87ms
1.1×
INT4
386MB
4.0×
2.24ms
1.4×
INT3
290MB
5.3×
2.42ms
1.3×
INT2
193MB
8.0×
1.83ms
1.7×
INT1
97MB
16.0×
1.72ms
1.8×
Appendix
Table 12 : Our bit-wise GEMM analysis. We report the total weight memory footprint (MB) of one LLaMA-65B block, which consists of seven matrices ( e.g. , 8192×8192 or 8192×22016 ) and measure the runtime of GEMM operation (ms) during generation tasks, across sequence lengths ranging from 1 to 128. For stable results, the runtime is the average values of 30 iterations with a 10-warm-up. We demonstrate that the bit-width non-evenly packable to 32-bit word ( e.g. , 5, 3-bit) can achieve the theoretical compression ratio as well as achieve fast GEMM latency.
Group
Method
Bit
PPL ↓
Acc. ↑
–
FP16
16
6.22
63.19
128
BiLLM
1(+1.1)
33.10
33.28
128
DBW(Ours)
1.61
10.54
49.06
256
BiLLM
1(+1.1)
39.55
29.65
256
DBW(Ours)
1.61
11.02
47.72
Appendix
Table 13 : Group size analysis. We evaluate the DBW and BiLLM with different group sizes in terms of perplexity and accuracy. We apply the quantization on the LLaMA-2-7B model.
Model
Method
Bit
Group
PPL ↓
Acc. ↑
LLaMA-3.1-8B-Instruct
FP16
16
–
9.30
66.37
OmniQuant
2
128
112.38
27.19
SliM-LLM +
2
128
301.27
26.08
BiLLM
1(+1.1)
128
51.55
34.97
DBW(Ours)
1.61
128
18.01
45.95
DeepSeek-R1-Distill-LLaMA-8B
FP16
16
–
17.21
58.09
Appendix
Table 14 : Instruction-tuned model evaluation. This analysis assesses the perplexity and accuracy of the sub-2-bit quantization methods on the instruction-tuned LLMs. Our DBW works well on the post-tuned LLMs.
Method
Group
Bits
LLaMA-7B
LLaMA-2-7B
MMLU
GSM8K
MMLU
GSM8K
PB-LLM
128
1.7(+1)
23.00
0.23
23.00
0.00
BiLLM
128
1(+1.1)
22.90
0.00
22.90
0.00
PTQ1.61
-
1.61
23.00
0.15
22.90
0.61
DBW (Ours)
128
1.61
24.47
1.82
23.89
2.12
Appendix
Table 15 : PTQ knowledge-mathematical reasoning accuracy evaluation. We evaluate zero-shot reasoning accuracy for the ultra-efficient method on the MMLU and GSM8K datasets. The numbers of PB-LLM and BiLLM come from PTQ1.61.
Type
Method
Trade-off
Perplexity
Ratio
R-bit
Bit
WikiText2
C4
Avg.
sequential optim.
SVD-LLM+AWQ
20.00
4.00
3.20
7.97
11.71
9.84
SVD-LLM+GPTQ
20.00
4.00
3.20
8.14
11.45
9.79
SVD-LLM+AWQ
20.00
3.00
2.40
12.47
18.57
15.52
SVD-LLM+GPTQ
20.00
3.00
2.40
14.47
15.42
14.94
co-optimization (our DBW )
35.54
3.10
2.00
8.27
10.85
9.56
Appendix
Table 16 : Extended sequential vs. co-optimization evaluation. We compare sequential vs. co-optimization on different bit-widths. We use the LLaMA-2-7B model. The sequential optimization method employs SVD-LLM for pruning and AWQ and GPTQ for quantizing the remaining components. For a comprehensive evaluation, we attempt to vary pruning ratio , bit-width of remaining components ( R-bit ), and quantization method of the sequential optimization method. The proposed co-optimization method works well against the sequential optimization baseline on different bit-widths.
LLaMA
Trade-off
Weight memory (GB)
Ratio
R-bit
Bit
FP16
Ours
Size ↓
1-7B
47.21
3.05
1.61
12.31
1.56
7.89×
1-13B
47.99
3.10
1.61
23.94
2.88
8.31×
1-30B
48.90
3.15
1.61
60.19
6.89
8.74×
1-65B
42.91
2.82
1.61
121.11
13.66
8.87×
2-7B
46.95
3.03
1.61
12.31
1.56
7.89×
Appendix
Table 17 : Model-level pruning-quantization trade-off analysis. We report the model-level pruning ratio and quantization bit-width for remaining components ( R-bit ) trade-off as well as an bits-per-weight ( Bit ). For weight memory footprint (GB), we include an embedding layer and a normalization layer, which leads to a total size reduction ( Size ↓ ) of less than 10× against FP16, even though we compress the weight matrix by the bits-per-weight of 1.61 .
Figure 13 : Block-level truncation-quantization trade-off analysis. We report the block-level truncation and quantization trade-off. We use the LLaMA-2-7B model.
Module
SVD entropy
Trade-off
Weight memory (MB)
Ratio
R-bit
Bit
FP16
Ours
Size ↓
Q
6.64
91.80
1.89
0.08
32.00
0.32
103.37×
K
6.53
86.62
1.78
0.12
32.00
0.48
67.15×
V
7.84
45.61
3.34
0.91
32.00
3.63
8.81×
O
7.85
62.26
3.52
0.66
32.00
2.66
12.04×
Up
8.17
45.53
3.06
1.21
86.00
8.96
9.61×
Appendix
Table 18 : Matrix-level pruning-quantization trade-off analysis. We evaluate the matrix-level trade-off of pruning (ratio) and quantization for remaining components (r-bit) as well as an bits-per-weight (Bit). We use the LLaMA-2-7B model.
Module
FP16 (MB)
SVD entropy
Single-level ( Acc.: 48.77 )
Bi-level ( Acc.: 49.06 )
Bit
Mem
Bit
Mem
Q
32.00
6.64
1.61
3.22
0.08
0.32
K
32.00
6.53
1.61
3.22
0.12
0.48
V
32.00
7.84
1.61
3.22
0.91
3.63
O
32.00
7.85
1.61
3.22
0.66
2.66
Up
86.00
8.17
1.61
8.65
1.21
8.96
Appendix
Table 19 : Bi-level differentiable bit-width analysis. We compare matrix-level bits-per-weight ( Bit ) and memory footprint (MB) between single-level (component-level) and the bi-level (component- and matrix-level) bit-width learning method. We employ the LLaMA-2-7B model. Our bi-level method optimizes the matrix-level bits-per-weight to reduce the errors. 1.61 † denotes the weighted average bit-width to account for the differences between the number of elements in the weight matrices.
Figure 14 : Low SVD entropy due to our loss. We show that our entropy minimization loss effectively reduces the SVD entropy compared to the baseline using LLaMA-2-7B model.