Weight compression helps large neural networks fit deployment memory budgets, but common fixed-width formats offer only coarse storage choices. Entropy coding supports finer rates, yet the achieved size depends on the quantized weight distribution and coding overhead. Exploiting this flexibility requires accurate rate selection and efficient weight reconstruction for inference. We present EntroPack, an entropy-coded weight compressor that supports arbitrary target bitrates without activation calibration or fine-tuning. It combines row-normalized E8 lattice quantization with a conditional probability model of lattice coordinates. Sampled storage estimates select the quantization resolution without repeated full-stream encoding. The final coordinates are entropy-coded in independently decodable tiles, enabling fast, fused symbol decoding and numerical weight reconstruction on the GPU. EntroPack supports floating-point and integer weight containers, such as BF16, FP16, FP8, and INT8, with storage bitrate controlled independently of numerical precision. Online decoding adds latency that grows with weight count, making the method well suited to compute-intensive workloads such as diffusion denoising and Transformer prefill. Experiments demonstrate fast encoding and modest inference overhead in these settings. When compressing the linear-layer weights of the image generator Z-Image-Turbo, EntroPack achieves substantially lower weight and denoiser output errors than fixed-width formats at comparable storage rates, with modest denoising-step overhead. Targeting 4 bits per parameter, it achieves lower weight and denoiser output errors than NF4, including about 24% lower relative L2 weight error, with less storage. Source code is available at https://github.com/modelscope/entropack.
Figures & tables
Figure 1: Storage and reconstruction error of compressed linear weights in Z-Image-Turbo.
Figure 2: EntroPack encoding and decoding. (a) Sampled scale selection precedes full-tensor quantization, row-scale fitting, optional RDO, and tiled rANS encoding. (b) Symbol decoding, lattice reconstruction, row rescaling, and dtype conversion are fused on the GPU.
method
bpp
weights
network out
images (3 seeds, 2 prompts)
inference
conversion
rel. L2 (%) ↓
rel. L2 (%) ↓
PSNR ↑
SSIM ↑
step (ms) ↓
(s) ↓
BF16 reference
16.00
0.00
0.00
99.00
1.00
505.7
0.0
NF4 (bitsandbytes)
4.50
9.41
25.76
18.97 ± 3.08
0.77 ± 0.09
518.2
0.1
FP4 (bitsandbytes)
4.50
12.73
31.55
17.74 ± 2.08
0.74 ± 0.08
520.0
0.1
NVFP4 (torchao)
4.50
9.47
22.44
19.85 ± 3.45
0.79 ± 0.09
802.6
0.6
MXFP4 (torchao)
4.25
12.02
25.74
18.00 ± 1.86
0.73 ± 0.07
798.8
0.6
Table 1: Weight compression, image fidelity, and runtime for Z-Image-Turbo. Network output error compares first-step denoising outputs with those of the BF16 model.
method
network bpp
rel. L2 (%) ↓
WikiText-2 PPL ↓
prefill
decode
conversion
(ms / 32k tokens) ↓
(ms / 8 tokens) ↓
(s) ↓
BF16 reference
16.00
0.00
5.98
18639
155.1
0.0
NF4 (bitsandbytes)
4.50
9.28
6.08
18634
210.8
0.3
FP4 (bitsandbytes)
4.50
12.41
6.23
18638
210.7
0.3
NVFP4 (torchao)
4.50
9.49
6.09
19727
1456.7
2.0
MXFP4 (torchao)
4.31
11.80
6.16
19693
1436.4
2.0
Table 2: Weight reconstruction, WikiText-2 perplexity, and runtime on Qwen3.8-27B.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
method
bpp
weights
video
audio
inference
conversion
rel. L2 (%) ↓
PSNR ↑
SSIM ↑
log-mel rel. L2 (%) ↓
step (ms) ↓
(s) ↓
BF16 reference
16.00
0.00
99.00
1.00
0.00
26209.4
0.0
NF4 (bitsandbytes)
4.50
9.14
15.86 ± 2.34
0.60 ± 0.10
34.14 ± 3.26
26269.1
0.4
FP4 (bitsandbytes)
4.50
12.03
14.05 ± 2.45
0.52 ± 0.10
29.94 ± 5.10
26274.0
0.5
NVFP4 (torchao)
4.50
9.77
16.79 ± 2.04
0.67 ± 0.07
28.66 ± 2.24
27785.9
2.7
MXFP4 (torchao)
4.25
11.51
15.19 ± 2.27
0.59 ± 0.08
27.42 ± 4.40
27739.9
2.7
Appendix
Table 3: Weight compression, video and audio fidelity, and runtime on MiniMax-H3. Fidelity metrics are reported as mean ± standard deviation over three seeds.
source model
target 2 bpp
target 4 bpp
target 8 bpp
9 pooled tables (no coset conditioning)
2.07
4.03
8.01
direct z8 (no parity reduction)
2.12
4.15
8.13
EntroPack (17 tables + parity reduction)
2.02
4.02
8.00
Appendix
Table 4: Measured stored rates for coding ablations on Z-Image-Turbo’s reference linear layer. Each column uses the same quantized points and row reconstruction scales.
Figure 3: RDO sweep count on Z-Image-Turbo at a target of 3 bpp with five candidate scales. Left: network weight error. Right: total model encoding time.
Figure 4: RDO candidate count on Z-Image-Turbo at a target of 3 bpp with one allocation sweep. Left: network weight error. Right: total encoding time.
Figure 5: Reconstruction cost on the reference layer at a 4 bpp target. Left: decode throughput and stored rate versus tile size. Right: linear-layer latency overhead relative to resident BF16 weights, versus the number of input tokens processed together.
target bpp
tile (symbols)
tile (vectors)
tiles
b
tile metadata (bytes)
tile metadata (%)
2
32,768
3,640
1,351
11
178,336
1.80
4
16,384
1,820
2,701
11
356,536
1.80
7
8,192
910
5,402
11
713,068
2.06
8
4,096
455
10,803
12
1,426,000
3.60
Appendix
Table 5: Tile configurations and metadata storage on the reference layer. Tile metadata comprises rANS states and payload offsets, with percentages relative to total stored bytes.
container
target bpp
stored bpp
rel. L2 (%) ↓
encode (ms) ↓
decode (ms) ↓
FP32
4.00
4.02
7.19
11.7
0.31
FP32
8.00
8.05
0.51
35.5
0.41
FP16
4.00
4.02
7.19
11.7
0.31
FP16
8.00
8.05
0.51
36.3
0.40
BF16
4.00
4.02
7.19
11.5
0.23
BF16
8.00
8.05
0.53
36.2
0.34
Appendix
Table 6: Stored rate, reconstruction error, and codec time for seven dtypes on the reference layer. Errors use each dtype’s input as the reference, with the UINT8 offset removed.
route
stored bpp
rel. L2 (%) ↓
PSNR ↑
step (ms) ↓
BF16 reference
16.00
0.00
99.00
505.7
FP8 e4m3 container, uncoded
8.01
2.65
22.21
343.1
INT8 container, uncoded
8.01
1.07
20.10
342.5
EntroPack (4 bpp), FP8 container
4.02
7.98
19.52
391.7
EntroPack (4 bpp), INT8 container
4.02
7.34
18.88
392.5
EntroPack (4 bpp), direct BF16
4.02
7.18
19.62
544.7
Appendix
Table 7: Storage, weight error, image fidelity, and runtime with compressed FP8 and INT8 weights on Z-Image-Turbo. Fidelity uses the uncompressed BF16 model as the reference.
allocation
network bpp
rel. L2 (%) ↓
PSNR ↑
SSIM ↑
uniform 2.5 bpp
2.50
20.64
16.82
0.64
sensitive@3 + rest@2.47
2.50
20.76
16.42
0.67
Appendix
Table 8: Uniform and mixed per-layer rate allocation on Z-Image-Turbo at approximately 2.5 stored bpp. Both allocations are evaluated against the same BF16 reference.
Figure 6: Z-Image-Turbo generations with uniform and mixed rate allocation at 2.50 stored bpp, using the same prompt, random seed, and denoising settings.