Post-training quantization (PTQ) reduces the cost of on-device text-to-speech (TTS), but published evaluations cover one system or method. We evaluate PTQ across TTS architectures under one protocol with three core models, weight and activation ablations of eight more, and two held-out models quantized blind. Four-bit per-channel weights reduce UTMOS, a predicted mean opinion score, by 2.8 on Supertonic and 0.07 on Kokoro, and per-tensor scaling can cause severe degradation even at 8 bits. The same bit width yields different outcomes, because the sensitive component is model-specific and not reliably predicted from the model class. A staged ablation procedure identifies it, and per-layer GPTQ can restore it to within 0.1 UTMOS. Real int8 and int4 kernels reproduce the simulated ordering at hardware-dependent cost. On a Mac mini, a 4-bit weight kernel runs Supertonic at 0.60x the fp32 latency while int8 is slower, so each configuration requires validation on the target runtime.
Figures & tables
Figure 1: Sensitivity map. Magnitude of the paired UTMOS loss at 4-bit weights against each model’s own fp baseline for the whole model under three scale granularities, for the most sensitive component alone at W4 per-channel (component, named in the row label), and for the rest of the model at W4 per-channel with that component at fp (rest). OmniVoice uses its CUDA fp16 baseline, and the last two rows are the held-out models of Sec. 4.7 .
Component
Scales
RTN
Scaling
GPTQ
F5-TTS vocoder
per-channel
-0.94
-0.61
-0.15
group:128
-0.48
-0.22
-0.07
Kyutai depth transformer
per-channel
-2.99 a
-1.82 c
-0.29
group:128
-2.98 b
-0.20
-0.08
VoxCPM local DiT
per-channel
-2.73 d
–
-0.71
group:128
-1.26
–
-0.41
Table 1: Calibrated PTQ on the most sensitive component of each model. Paired Δ UTMOS against the model’s own fp baseline under RTN, activation-aware weight scaling ( α=0.5 ), and per-layer GPTQ at per-channel and group:128 scales. WER stays at its fp level (0.035 for F5-TTS, 0.034 for Kyutai, 0.043 for VoxCPM) in every cell except the four marked ones, whose WER is a 1.14, b 1.01, c 0.13, and d 0.14.
Condition
RTF
× fp
RSS (MB)
P (W)
E
Supertonic fp32 (NFE 8)
0.148
–
607
19.1
3.25
dyn8 full (real int8)
0.310
2.09 ×
329
9.7
3.21
dyn8, vocoder excl.
0.218
1.47 ×
436
19.1
4.64
int4 W, vocoder excl.
0.089
0.60 ×
360
–
–
Kokoro fp32
0.069
–
2712
7.5
0.99
dyn8 (real int8)
0.077
1.12 ×
2722
8.9
1.25
Table 2: System metrics on the quiet Mac mini M4 Pro over 5 repeats that agree within 2%. RTF and its multiple of the fp row of the same model, peak RSS, power P, and the marginal energy E of the whole 20-sentence process in J per audio-second (Sec. 3 , not measured for the int4 row); dyn8 is dynamic int8.
Configuration
CI
WER +
Kokoro real int8 (PyTorch)
[ -0.000 , 0.003]
0.001
OmniVoice W8A8, compiled, NFE 8
[ -0.038 , 0.048]
0.002
Kokoro W4 group:128
[ -0.039 , -0.035 ]
0.002
Supertonic int8, no vocoder
[ -0.056 , -0.018 ] ∗
0.005
Chatterbox T3 W4, S3Gen W8
[ -0.032 , 0.001]
0.005
Chatterbox real int4 on T3
[ -0.012 , 0.020]
0.003
Table 3: The quality-preserving criterion applied. The paired 95% interval of Δ UTMOS (CI) must lie within [−0.05,0.05] , and the upper endpoint of the paired Δ WER interval (WER + ) must be at most 0.01 , unrounded, against the fp baseline of each configuration. An asterisk marks a violated condition.