Bits Under ZK-LLM: Evaluating Zero-Knowledge-Friendly Quantization for Verifiable Private LLM Inference
Organizations: Siebel School of Computing and Data Science University of Illinois Urbana-Champaign
Abstract
Zero-knowledge proofs are emerging as a promising approach for enabling private, verifiable LLM governance and auditing, where regulators, users, and auditors need to verify claims about training-data usage or LLM inference-time behavior, while model providers must protect proprietary model parameters. However, despite the growing interest in ZK-LLMs, the understanding of ZK-friendly quantization remains limited. This gap matters because in the ZK setting, quantization directly shapes the arithmetic structure, constraint complexity, and proving cost of ZK inference. ZK protocols operate over finite fields and incur costs that depend heavily on the number and type of arithmetic operations, nonlinearities, and lookup constraints. Understanding ZK-friendly quantization is therefore essential for making ZK-LLMs practical. In this work, we present the first systematic study of ZK-friendly quantization for LLMs. We first formalize the definition of ZK-friendly quantization, capturing the properties required for ZK proof generation. We then evaluate nine language models, including Qwen2.5-14B and the mixture-of-experts model Qwen3-30B-A3B, across a broad design space of weight, activation, and nonlinear lookup table precision. Our results show that activation precision is substantially more sensitive than weight precision, while nonlinear lookup approximations can become the dominant source of utility degradation. Also, we identify RMSNorm inverse-square-root lookups as a recurring bottleneck in several large models and recover near-baseline utility by selectively increasing precision only at the bottleneck. Finally, we show that reducing bit-width or lookup-table size does not necessarily yield proportional end-to-end proving savings, showing that conventional low-bit quantization heuristics do not directly translate to ZK proving efficiency and motivating operator-aware precision selection.
Figures & tables
| WikiText-2 C4 Model B8 B12 B16 B8 B12 B16 GPT-2 S 1866.2 1.011 1.000 1439.4 1.018 1.002 GPT-2 M 3213.6 1.002 1.001 1465.4 1.000 0.999 Llama 8B 1247.9 1126.8 422.2 380.5 Qwen 14B 32.5 18.3 Qwen3 30B 148.2 131.3 78.7 120.7 | LUT configuration Model Full B16 RMS only RMS exact RMS B24 Llama 8B 1126.8453 1126.7972 0.9993 1.0002 Qwen 7B 1496.9087 1528.4988 0.9999 0.9987 Qwen 14B 32.4609 32.6139 1.0002 0.9997 Qwen3 30B 131.3352 1.0398 1.0034 1.0014 |
| (a) LUT precision sensitivity | (b) Targeted precision adjustment |
| Projection | Floating-point | Integer |
|---|---|---|
| (each) | (estimated) | (measured) |
| Q, O | 12,158.78 | 11.877 |
| K, V | 2,431.76 | 4.603 |
| FFN gate, up | 32,828.71 | 20.985 |
| FFN down | 27,105.54 | 20.526 |
| Sum | 121,944.05 | 95.457 |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Simulation error | ||||
|---|---|---|---|---|
| Scale | Precision | Quantization error | FP32 | BF16 |
| Standard | W8A8 | |||
| W12A12 | ||||
| W16A16 | ||||
| PoT | W8A8 | |||
| W12A12 | ||||
| Dataset | GPT-2 | TinyLlama | Llama-3.1 | Qwen2.5 | Qwen3 | |||||
|---|---|---|---|---|---|---|---|---|---|---|
| Small | Medium | XL | 1.1B | 8B | 3B | 7B | 14B | 30B-A3B | ||
| WikiText-2 | 8 | 1866.2 | 3213.6 | 704.3 | ||||||
| 12 | 1.011 | 1.002 | 1.002 | 2356.3 | 1247.9 | 125.5 | 5254.5 | 148.2 | ||
| 16 | 1.000 | 1.001 | 1.000 | 3.073 | 1126.8 | 1.000 | 1496.9 | 32.5 | 131.3 | |
| C4 | 8 | 1439.4 | 1465.4 | 316.7 | ||||||
| 12 | 1.018 | 1.000 | 0.989 | 2258.8 | 422.2 | 78.8 | 3472.9 | 78.7 | ||
| LUT configuration | Llama-3.1-8B | Qwen2.5-7B | Qwen2.5-14B | Qwen3-30B-A3B |
|---|---|---|---|---|
| Full B16 | 1126.8453 | 1496.9087 | 32.4609 | 131.3352 |
| RMS B16 only; others exact | 1126.7972 | 1528.4988 | 32.6139 | 1.0398 |
| Full B16; RMS exact | 0.9993 | 0.9999 | 1.0002 | 1.0034 |
| Full B16; RMS B24 | 1.0002 | 0.9987 | 0.9997 | 1.0014 |
| GPT-2 Small | GPT-2 Medium | GPT-2 Large | |||||||
| Configuration | Mean | Min | Max | Mean | Min | Max | Mean | Min | Max |
| W16A16 | 104.474 | 103.257 | 105.087 | 222.355 | 218.421 | 227.839 | 525.555 | 517.006 | 531.283 |
| W8A16 | 102.718 | 102.015 | 103.790 | 221.179 | 218.017 | 224.541 | 524.634 | 523.111 | 525.836 |
| W10A16 | 103.205 | 102.106 | 103.921 | 221.231 | 216.977 | 225.548 | 526.105 | 519.276 | 535.732 |
| W12A16 | 104.079 | 101.009 | 106.103 | 217.248 | 212.212 | 220.218 | 527.309 | 524.410 | 530.007 |
| W14A16 | 104.683 | 102.837 | 106.115 | 215.175 | 212.812 | 217.813 | 535.667 | 529.128 | 539.949 |
| GPT-2 Small | GPT-2 Medium | |||
|---|---|---|---|---|
| Context length | W16A16 | W12A16 | W16A16 | W12A16 |
| 4 | 73.116 | 74.093 | 158.405 | 150.433 |
| 8 | 87.156 | 85.499 | 183.162 | 179.619 |
| 16 | 103.666 | 100.756 | 225.444 | 218.262 |
| 32 | 125.902 | 130.727 | 274.920 | 274.536 |
| WikiText-2 | C4 | ||||||
|---|---|---|---|---|---|---|---|
| Model | Configuration | KL | Top-5 | Bottom-20% | KL | Top-5 | Bottom-20% |
| Qwen2.5-3B | W16A16, linear | 0.001847 | 0.9726 | 0.9872 | 0.001570 | 0.9769 | 0.9884 |
| W16A16, full | 0.003710 | 0.9612 | 0.9835 | 0.002999 | 0.9695 | 0.9853 | |
| W12A12, full | 0.040696 | 0.8842 | 0.9505 | 0.030019 | 0.9032 | 0.9536 | |
| Qwen2.5-7B | W16A16, linear | 0.015128 | 0.9354 | 0.9719 | 0.019180 | 0.9307 | 0.9670 |
| W16A16, full | 0.014987 | 0.9363 | 0.9701 | 0.013916 | 0.9395 | 0.9744 | |
| W12A12 | W16A16 | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Model | Dataset | Base | PoT | Conv. | Asym | Dyn | Base | PoT | Conv. | Asym | Dyn | ||
| GPT-2 Small | WikiText-2 | 27.05 | 28.53 | 27.18 | 26.99 | 26.60 | 26.45 | 26.35 | 26.32 | 26.41 | 26.28 | 26.33 | 26.37 |
| C4 | 31.99 | 33.43 | 32.00 | 31.99 | 31.31 | 31.42 | 31.20 | 31.27 | 31.30 | 31.26 | 31.14 | 31.31 | |
| GPT-2 Medium | WikiText-2 | 19.77 | 19.93 | 19.80 | 19.79 | 19.57 | 19.35 | 19.31 | 19.34 | 19.34 | 19.32 | 19.36 | 19.29 |
| C4 | 24.99 | 25.17 | 24.92 | 25.06 | 25.02 | 24.85 | 24.82 | 24.80 | 24.86 | 24.88 | 24.84 | 24.75 | |
| GPT-2 XL | WikiText-2 | 15.33 | 15.27 | 15.33 | 15.32 | 15.30 | 15.30 | 15.31 | 15.31 | 15.31 | 15.31 | 15.31 | 15.30 |
| Ours (Plonky3) | SP1, fixed point | SP1, FP32 | |||
|---|---|---|---|---|---|
| Prover (s) | Prover (s) | Cycles | Prover (s) | Cycles | |
| 64 | 3.15 | 25.61 | 12,794 | 25.29 | 107,677 |
| 128 | 3.59 | 26.00 | 17,863 | 27.11 | 207,150 |
| 256 | 3.28 | 25.97 | 28,001 | 30.50 | 405,742 |
| 512 | 2.98 | 26.48 | 48,268 | 35.58 | 806,221 |
| 1024 | 3.58 | 27.33 | 88,802 | 44.61 | 1,622,618 |
| Llama-3.1-8B | Qwen2.5-7B | Qwen2.5-14B | Qwen3-30B-A3B | ||||||
| Operator / configuration | Mode | Wiki | C4 | Wiki | C4 | Wiki | C4 | Wiki | C4 |
| Full B16 | – | 1126.8453 | 380.5457 | 1496.9087 | 769.6439 | 32.4609 | 18.2997 | 131.3352 | 120.6984 |
| RMS InvSqrt | O | 1126.7972 | 380.6613 | 1528.4988 | 759.8001 | 32.6139 | 18.1795 | 1.0398 | 1.0562 |
| E | 0.9993 | 0.9997 | 0.9999 | 0.9978 | 1.0002 | 0.9997 | 1.0034 | 0.9995 | |
| SiLU | O | 0.9999 | 1.0005 | 0.9998 | 0.9993 | 0.9998 | 1.0003 | 1.0038 | 0.9992 |
| E | 1117.7300 | 382.1309 | 1415.1683 | 757.1557 | 31.7131 | 18.5148 | 132.4172 | 116.0887 | |
| Proving time (seconds) | Analysis: GeLU , Exp | ||||||
|---|---|---|---|---|---|---|---|
| GeLU Exp 16 | GeLU 16 Exp | GeLU Exp | LUT payload (MiB) | Direct queries (M) | All queries (M) | Joint table proof (s) | |
| 8 | 103.58 (0.77) | 104.70 (0.53) | 104.75 (0.37) | 0.102 | 0.823 | 18.238 | 0.483 |
| 12 | 103.37 (0.34) | 104.47 (1.25) | 103.28 (0.73) | 1.625 | 0.823 | 17.906 | 0.592 |
| 16 | 102.65 (3.04) | 102.65 (3.04) | 102.65 (3.04) | 26.000 | 0.823 | 17.104 | 1.325 |
| 24 | 248.44 (3.51) | 120.49 (1.13) | 265.54 (2.93) | 6,656.000 | 0.823 | 18.218 | 103.761 |
| Proving time (seconds) | Peak RAM usage (GiB) | ||||
| Projection | Weight shape | Floating-point | Integer | Floating-point | Integer |
| Direct measurements | |||||
| Q/K/V/O, gate/up | 18.998 | 1.115 | 14.8816 | 0.2048 | |
| FFN down | 42.352 | 1.359 | 40.7634 | 0.2943 | |
| Full-width costs: floating-point estimated, integer measured | |||||
| Q | 12,158.78 | 11.877 | – | – | |