Quantization Error Is Spectrally Flat: A Single Random Probe Is a Calibrated, Data-Free Sensitivity Estimator, with Application to Budget-Targeted Mixed-Precision Quantization
Organizations: baa.ai Auckland, New Zealand
Abstract
A single random Gaussian probe gives an unbiased estimate of the squared Frobenius norm of a layer's quantization error. The estimator is well-behaved because round-to-nearest error is spectrally flat. Across 1,683 tensors from a 35B MoE and a 9B dense model, effective dimensionality is 0.93 to 0.96 times the i.i.d. noise value of the same shape, and on the MoE the median is unchanged from 2-bit to 8-bit. The probe coefficient of variation is predictable from tensor shape. One probe measures per-tensor sensitivity to within 4 to 7%; twenty probes reach 1.3 to 1.4%.RAM applies the propagated form of this estimator to budget-targeted mixed-precision quantization with no calibration data. Gaussian probes carrying the network's own input statistics score every tensor at six bit-widths. A knapsack solver allocates bits under an exact byte budget, with guardrails against catastrophic 2-bit assignments. One probe pass serves any budget. Isolated and propagated scores rank tensors independently on Qwen3.5-35B-A3B (Spearman -0.01), yet the propagated probe rank-correlates 0.81 to 0.83 with the GPTQ layer objective from real activations, while the isolated estimator is uncorrelated with it. That objective is the wrong allocation target: at matched bytes on Qwen3.8-27B, a block-output probe beats a vendor IQ3_M mix and an oracle that allocates from the real-activation objective. On Qwen3-8B the propagated probe ties HAWQ-V2 at matched bytes. Across seven architectures from 8B to 122B, with probe timing up to a 400B model in nine minutes on one workstation, RAM reaches 3.5 to 13.6% lower median WikiText-2 perplexity than size-comparable uniform 4-bit builds on the tested MoE models. (Black Sheep Ai baa.ai)
Figures & tables
| Model | Arch | BF16 (GB) | RAM (GB) | RAM Med. PPL | Unif. 4b (GB) | Unif. 4b Med. PPL | Time | |
|---|---|---|---|---|---|---|---|---|
| Qwen3.5-35B-A3B | MoE 256exp | 69.3 | 19.5 | 6.585 | 17.2 | 6.928 | 4.9% | 188 s |
| Qwen3-30B-A3B a | MoE 128exp | 61.1 | 16.8 | 8.877 | — | 9.158 | 3.1% | 53 s |
| MiniMax-M2.5 | MoE 256exp FP8 | 230.1 | 112.3 | 9.070 | 119.8 | 9.399 b | 3.5% | 301 s |
| Llama-4-Scout | MoE 16exp | 217.3 | 56.0 | 7.806 | 56.9 | 8.219 b | 5.0% | 214 s |
| Qwen3.5-122B-A10B c | MoE 128exp | 250.2 | 107.6 | 5.327 | 60.4 | 5.601 | 4.9% | 345 s |
| GLM-4.7-Flash | MoE + MLA | 60.0 | 16.2 | 8.700 | — | 10.075 | 13.6% | 99 s |
| Method | Size (GB) | Median PPL | vs BF16 |
|---|---|---|---|
| BF16 | 69.3 | 6.494 | — |
| RAM (probe allocation) | 19.5 | 6.585 | 1.4% |
| GPTQ-style (protect attention) | 20.5 | 6.697 | 3.1% |
| Uniform 4-bit g64 | 17.4 | 6.764 | 4.2% |
| Uniform 4-bit g128 | 17.2 | 6.928 | 6.7% |
| Qwen3.5-35B-A3B (1,317 tensors) | Qwen3.5-9B (366 tensors) | |||
| Probes ( ) | Spearman | Rel. MAE | Spearman | Rel. MAE |
| 1 | 0.671 | 6.2% | 0.627 | 5.5% |
| 5 | 0.855 | 2.8% | 0.864 | 2.5% |
| 10 | 0.903 | 2.0% | 0.921 | 1.8% |
| 20 | 0.938 | 1.4% | 0.957 | 1.3% |
| 50 | 0.968 | 0.90% | 0.982 | 0.75% |
| Model | Signal | Tensors | CV [IQR] | Spread (dex) | MAE 50 | ||||
|---|---|---|---|---|---|---|---|---|---|
| Qwen3.5-9B | isolated | 366 | 0.037 | 1,476 | — | 0.049 | 0.627 | 0.982 | 0.75% |
| Qwen3.5-9B | propagated | 312 | 0.16 [0.09, 0.22] | 79 | 0.04 | 1.06 | 0.994 | 1.000 | 1.3% |
| Qwen3.5-35B-A3B | isolated | 1,317 | 0.072 | 386 | — | 0.065 | 0.671 | 0.968 | 0.90% |
| Qwen3.5-35B-A3B | propagated | 590 | 0.24 [0.13, 0.59] | 34 | 0.03 | 0.86 | 0.954 | 0.998 | 2.6% |
| Qwen3-8B | propagated | 324 | 2.45 [1.60, 3.16] | 0.3 | 1.67 | 0.64 | 0.742 | 0.967 | 19% |
| Model | Tensors | HAWQ-V2 | Block cosine | Domain ceiling | Ratio | ||
|---|---|---|---|---|---|---|---|
| Qwen3-8B | 252 | 0.83 | 0.25 | 0.50 | 0.99 | 1.29 [1.05, 1.48] | |
| attention | 144 | 0.90 | 0.36 | — | — | 1.35 | |
| MLP | 108 | 0.71 | 0.08 | 0.07 | — | — | 1.13 |
| Qwen3.5-9B | 248 | 0.81 | — | 0.11 | 0.99 | 1.12 [0.92, 1.43] | |
| attention | 32 | 0.88 | — | — | — | 0.99 | |
| linear attention | 120 | 0.72 | — | — | — | 1.25 |
| Signal | Data | File (GB) | Perplexity | vs. IQ3_M | vs. best |
|---|---|---|---|---|---|
| IQ3_M (llama.cpp mix) | none | 12.58 | 6.336 0.073 | — | 3 |
| Q3_K_M (llama.cpp mix) | none | 13.30 | 6.207 0.075 | 31 | 3 |
| Propagated block cosine | none | 12.72 | 6.041 0.071 | 36 | 6 |
| Block cosine isolated CKA | none | 12.79 | 5.991 0.071 | 37 | — |
| Quadratic form, adaptation rule on | none | 12.61 | 6.204 0.075 | 34 | 0 |
| Quadratic form, adaptation rule off | none | 12.66 | 6.238 0.076 | 30 | 1 |
| Build | Size (GB) | 2-bit | Mean PPL | Median PPL |
|---|---|---|---|---|
| BF16 | 15.3 | 0% | 9.624 | 9.648 |
| Uniform 4-bit g64 | 4.3 | 0% | 9.965 | 9.982 |
| RAM 4 GB target, + veto | 5.8 | 7.4% | 10.231 | 10.299 |
| RAM 4 GB target, six widths + veto | 5.7 | 1.4% | 10.473 | 10.494 |
| RAM 6 GB target, six widths | 7.3 | 0% | 9.642 | 9.628 |
| Attention | Routers | Shared experts | Routed experts | Other | |
| Propagated signal (avg bits) | 8.5 | 16.0 | 8.0 | 3.8 | 11.8 |
| Isolated estimator (avg bits) | 5.7 | 8.0 | 8.0 | 3.9 | 10.6 |
| Build | Size (GB) | Median PPL | MMLU | ||
| Propagated (Table 1 build) | 19.5 | 6.585 | 0.657 | ||
| Propagated (rebuilt, same runtime) | 19.5 | 6.544 | 0.668 | ||
| Isolated estimator | 19.5 | 6.550 | 0.679 | ||
| Signal | gate | up | down | Agreement | |||||
|---|---|---|---|---|---|---|---|---|---|
| Propagated probe (data-free) | 5.3 | 6.4 | 7.8 | 6.3 | 5.4 | 5.5 | 6.0 | — | — |
| Isolated estimator (data-free) | 6.0 | 8.0 | 8.0 | 6.0 | 5.5 | 5.4 | 5.7 | 50% | 0.04 |
| HAWQ-V2 (calibration) | 5.6 | 6.6 | 7.6 | 5.4 | 5.6 | 5.8 | 5.6 | 41% | 0.19 |
| Strategy | Size (GB) | Median PPL | vs BF16 |
|---|---|---|---|
| BF16 | 58.2 | 8.470 | — |
| Probe, MLA-aware (latent stage) | 16.2 | 8.700 | 2.7% |
| Probe, adaptive | 14.6 | 9.354 | 10.4% |
| Probe, propagated | 14.6 | 9.486 | 12.0% |
| Uniform 4-bit | — | 10.075 | 18.9% |
| Model | BF16 Size | Probe Time | RD Time | Speedup |
|---|---|---|---|---|
| Qwen3-8B | 15 GB | 36 s | 195 s | 5.4 |
| Qwen3-30B-A3B | 61 GB | 53 s* | 2,634 s | 50 |
| Qwen3.5-35B-A3B | 69 GB | 188 s | — | — |
| GLM-4.7-Flash | 60 GB | 99 s | 1,847 s | 19 |
| Llama-4-Scout | 217 GB | 214 s* | — | — |
| MiniMax-M2.5 | 230 GB | 301 s* | — | — |
| Build | Size (GB) | Med. PPL | MMLU | Note |
| Reference points | ||||
| BF16 | 69.3 | 6.494 | 0.721 | |
| Uniform 8-bit g64 | 34.3 | 6.517 | 0.720 | |
| Uniform 4-bit g64 | 17.4 | 6.764 | 0.704 | |
| RAM budget sweep (one probe pass) | ||||
| RAM, 34 GB target | 33.6 | 6.504 | 0.731 | beats BF16 by 1.0 pp |
| Tensor set | Spearman | |
|---|---|---|
| All matched tensors | 390 | |
| Attention projections | 190 | |
| Shared-expert projections | 120 | |
| Routed expert stacks | 40 | |
| Routers | 40 | |
| Overlap of the two top-30 lists | 7 of 30 |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Class | Tensors | Params | 2b | 3b | 4b | 5b | 6b | 8b/16b | Avg bits |
| Qwen3.5-35B-A3B | Attention | 260 | 1.28B | 22 | 0 | 0 | 3 | 13 | 129/93 | 8.5 |
| Router | 40 | 0.02B | 0 | 0 | 0 | 0 | 0 | 0/40 | 16.0 | |
| Shared expert | 160 | 0.13B | 0 | 0 | 0 | 0 | 3 | 139/18 | 8.0 | |
| Routed experts | 120 | 32.2B | 3 | 31 | 74 | 12 | 0 | 0/0 | 3.8 | |
| Other | 40 | 0.01B | 0 | 0 | 0 | 0 | 0 | 21/19 | 11.8 | |
| GLM-4.7-Flash | Attention | 329 | 1.02B | 0 | 5 | 12 | 17 | 26 | 131/138 | 9.0 |
| Probe calibration (median) | |||||
|---|---|---|---|---|---|
| Model | median | IQR | med. [IQR] | CV /CV | |
| Qwen3.5-35B-A3B | 386 | [354, 405] | 0.93 [0.86, 0.97] | 0.999 | 0.9998 |
| Qwen3.5-9B | 1,476 | [722, 2,709] | 0.96 [0.90, 0.98] | 1.000 | 1.0000 |