Post-training quantization (PTQ) reduces the cost of on-device text-to-speech (TTS), but published evaluations cover one system or method. We evaluate PTQ across TTS architectures under one protocol with three core models, weight and activation ablations of eight more, and two held-out models quantized blind. Four-bit per-channel weights reduce UTMOS, a predicted mean opinion score, by 2.8 on Supertonic and 0.07 on Kokoro, and per-tensor scaling can cause severe degradation even at 8 bits. The same bit width yields different outcomes, because the sensitive component is model-specific and not reliably predicted from the model class. A staged ablation procedure identifies it, and per-layer GPTQ can restore it to within 0.1 UTMOS. Real int8 and int4 kernels reproduce the simulated ordering at hardware-dependent cost. On a Mac mini, a 4-bit weight kernel runs Supertonic at 0.60x the fp32 latency while int8 is slower, so each configuration requires validation on the target runtime.
Figures & tables
Figure 1: Sensitivity map. Magnitude of the paired UTMOS loss at 4-bit weights against each model’s own fp baseline for the whole model under three scale granularities, for the most sensitive component alone at W4 per-channel (component, named in the row label), and for the rest of the model at W4 per-channel with that component at fp (rest). OmniVoice uses its CUDA fp16 baseline, and the last two rows are the held-out models of Sec. 4.7 .
Component
Scales
RTN
Scaling
GPTQ
F5-TTS vocoder
per-channel
-0.94
-0.61
-0.15
group:128
-0.48
-0.22
-0.07
Kyutai depth transformer
per-channel
-2.99 a
-1.82 c
-0.29
group:128
-2.98 b
-0.20
-0.08
VoxCPM local DiT
per-channel
-2.73 d
–
-0.71
group:128
-1.26
–
-0.41
Table 1: Calibrated PTQ on the most sensitive component of each model. Paired Δ UTMOS against the model’s own fp baseline under RTN, activation-aware weight scaling ( α=0.5 ), and per-layer GPTQ at per-channel and group:128 scales. WER stays at its fp level (0.035 for F5-TTS, 0.034 for Kyutai, 0.043 for VoxCPM) in every cell except the four marked ones, whose WER is a 1.14, b 1.01, c 0.13, and d 0.14.
Condition
RTF
× fp
RSS (MB)
P (W)
E
Supertonic fp32 (NFE 8)
0.148
–
607
19.1
3.25
dyn8 full (real int8)
0.310
2.09 ×
329
9.7
3.21
dyn8, vocoder excl.
0.218
1.47 ×
436
19.1
4.64
int4 W, vocoder excl.
0.089
0.60 ×
360
–
–
Kokoro fp32
0.069
–
2712
7.5
0.99
dyn8 (real int8)
0.077
1.12 ×
2722
8.9
1.25
Table 2: System metrics on the quiet Mac mini M4 Pro over 5 repeats that agree within 2%. RTF and its multiple of the fp row of the same model, peak RSS, power P, and the marginal energy E of the whole 20-sentence process in J per audio-second (Sec. 3 , not measured for the int4 row); dyn8 is dynamic int8.
Configuration
CI
WER +
Kokoro real int8 (PyTorch)
[ -0.000 , 0.003]
0.001
OmniVoice W8A8, compiled, NFE 8
[ -0.038 , 0.048]
0.002
Kokoro W4 group:128
[ -0.039 , -0.035 ]
0.002
Supertonic int8, no vocoder
[ -0.056 , -0.018 ] ∗
0.005
Chatterbox T3 W4, S3Gen W8
[ -0.032 , 0.001]
0.005
Chatterbox real int4 on T3
[ -0.012 , 0.020]
0.003
Table 3: The quality-preserving criterion applied. The paired 95% interval of Δ UTMOS (CI) must lie within [−0.05,0.05] , and the upper endpoint of the paired Δ WER interval (WER + ) must be at most 0.01 , unrounded, against the fp baseline of each configuration. An asterisk marks a violated condition.
Weight-only post-training quantization is the cheapest way to shrink a retrieval embedder, and the received advice for applying it -- protect the embedding table, allocate bits by module sensitivity, prefer a ranking-aware objective over weight reconstruction -- was carried into LLM quantization largely intact. We test that advice on retrieval embedders directly, quantizing five checkpoints from four architecture families across a grid of bit widths and group sizes, and isolating the embedding, attention and feed-forward blocks at each width. Every heuristic fails to transfer as stated. The embedding table never emerges as the dominant isolated protection priority in any family, despite being the largest tensor in several of them. Module sensitivity does not survive as a transferable ordering: at INT4/g16 the spread between modules is too small to allocate against, at INT3 the ordering becomes family-dependent and joint damage stops being the sum of its parts, and at INT2 comparable reconstruction error accompanies retention ranging from 1.3 to 65.9 percent of full precision. A cheap reconstruction proxy is useful for screening uniform bit widths but substantially less reliable for choosing which tensors to protect; its apparent strength across the whole grid is a range-extension artifact. A distilled 109M student at INT3 holds 78.04 NDCG@10 in 68.4 MB and dominates the extreme-PTQ arm of its own 0.6B teacher, 297.9 MB at 64.46, on both size and quality -- but only inside the task it was distilled for. Sizes are byte counts of files that exist rather than arithmetic estimates, and the measurement repository carries the byte provenance for every one of them.
Weight-only post-training quantization (PTQ) can alleviate the computational burden of serving large language models (LLMs) at scale. However, existing PTQ methods often fail to generalize across models and suffer severe accuracy loss below 2 bits. Many leverage unstructured sparsity to mitigate this loss, but at the cost of regularity and GPU-friendly execution. We present QTEA, a sub-2-bit PTQ framework that quantizes weights into ternary values and uses salient weights as residual error compensators. To maintain hardware efficiency, residuals are assigned to selected columns with semi-structured 1:4 sparsity within the salient columns. We further add column-wise rescale refinement to GPTQ-style column-by-column quantization, alternately updating per-column scales and ternary assignments to reduce reconstruction error. We also identify order-dependent error propagation in GPTQ and introduce error decay to attenuate late-stage error accumulation. On Qwen3-14B, QTEA compresses all weights to an effective 1.7 bits per weight while improving average accuracy over the strongest ternary PTQ baseline by 16.7%. It also achieves 1.40× and 2.61× lower perplexity on WikiText and C4 respectively. This trend holds on Llama3-8B, where QTEA obtains a 6.6% accuracy gain and 1.34× / 1.95× lower perplexity on the same datasets. Finally, we develop a lookup-table based kernel that achieves 7.2× faster per-token generation over an FP16 baseline. Code is available at https://github.com/Intelligent-Microsystems-Lab/QTEA.
Post-training quantization (PTQ) compresses large language models by mapping weights to low-bit representations. The scaling factor that defines the quantization grid is typically chosen using simple, data-free heuristics. In this work, we present PiSO (Piecewise Scale Optimization), an algorithm that leverages calibration data to compute the optimal channel-wise weight scales exactly and efficiently under round-to-nearest quantization. PiSO partitions the scale search space into finitely many intervals on which the objective admits a closed-form minimizer. We extend PiSO to group-wise quantization via principled heuristics and propose effective strategies for interleaving scale optimization with error correction. Experiments on Llama and Qwen models across multiple model sizes and target weight bit-widths demonstrate consistent improvements in perplexity and downstream zero-shot accuracy, both standalone and combined with error correction. In particular, we observe increased benefits as the target bit-width narrows and quantization becomes more challenging.
Juan Amboage, Pablo Monteagudo-Lago, Ian Colbert +2