Text-output scores alone do not show whether quantization preserves performance on speech tasks whose target labels cannot be recovered from the transcript. We evaluate fixed mixed 4/8-bit Qwen2-Audio-7B-Instruct allocations averaging 6 and 7 bits per parameter on 508 English-to-German FLEURS utterances and on 512 RAVDESS emotion clips from 16 speakers. The BLEU and chrF differences from half precision (FP16) have intervals that include zero for both allocations. On RAVDESS, the same two sentences occur equally often with every emotion label. The absolute accuracy differences from FP16 are -3.71% for 6 bit and -1.17% for 7 bit. The 6-bit speaker interval excludes zero and an exact two-sided sign-flip test gives p=0.0148; the 7-bit interval includes zero. Same-budget controls do not identify either selected allocation as best. This case study shows why translation scores and performance on tasks beyond the transcript need separate evaluation.
Figures & tables
Figure 1: Why quantized speech models need more than a transcript score. A: Schematic waveforms illustrate two recordings with the same words but different emotion labels. B: Translation scores do not measure performance on labels that the transcript cannot determine. C: The selected 6-bit allocation has lower RAVDESS accuracy than FP16, with an absolute difference of −3.71 % (95% speaker interval [−6.05,−1.37] %; exact two-sided sign-flip p=0.0148 ). The 7-bit interval includes zero.
Allocation
BLEU ↑
chrF ↑
Accuracy (%) ↑
FP16
20.36
49.73
70.12
Translation-selected, 6 bit
20.22
49.60
66.41
Translation-selected, 7 bit
20.33
49.68
68.95
Table 1: FLEURS translation scores and RAVDESS emotion accuracy for the same precision allocations. BLEU and chrF use SacreBLEU 2.5.1; emotion accuracy (%) uses the label-normalization rule in Section 3 . Differences are calculated before rounding. Shading identifies the translation-selected allocations.
Figure 2: Selected allocation minus comparator with 95% intervals. Panels A and B show FLEURS translation-score differences from FP16, using paired utterance resampling (508 examples; 2,000 resamples). Panels C and D show RAVDESS accuracy differences from FP16 and same-budget controls, using speaker resampling (16 speakers; 10,000 resamples). A negative value favors the comparator.
Allocation
Δ BLEU ↑ (95% CI)
Δ chrF ↑ (95% CI)
6 bit
−0.15[−0.51,+0.14]
−0.13[−0.41,+0.10]
7 bit
−0.04[−0.38,+0.26]
−0.05[−0.33,+0.19]
Table 2: FLEURS translation-score differences, selected allocation minus FP16. Paired 95% bootstrap intervals resample the 508 utterances 2,000 times with seed 1729. All intervals include zero. CI denotes confidence interval.
Actor
FP16 (%) ↑
Selected 6 (%) ↑
Selected 7 (%) ↑
6 − FP16 (%) ↑
7 − FP16 (%) ↑
01
56.25
53.12
53.12
−3.12
−3.12
02
75.00
65.62
75.00
−9.38
0.00
03
84.38
78.12
84.38
−6.25
0.00
04
68.75
62.50
68.75
−6.25
0.00
05
71.88
68.75
71.88
−3.12
0.00
06
65.62
62.50
65.62
−3.12
0.00
Table 3: RAVDESS emotion accuracy by actor. Each actor contributes 32 clips. The final two columns are absolute accuracy differences, selected allocation minus FP16, for that actor, computed before rounding. The uncertainty analyses use these actor-level means.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Allocation
Accuracy (%) ↑
Selected − comparator (%) ↑ (95% CI)
Sign-flip p
6-bit budget
FP16
70.12
−3.71[−6.05,−1.37]
0.0148
Selected
66.41
Reference
—
Front-block
68.16
−1.76[−4.30,+0.78]
0.2754
Stratified-uniform
68.36
−1.95[−4.30,+0.39]
0.1836
7-bit budget
Appendix
Table 4: RAVDESS emotion accuracy and absolute accuracy differences, selected allocation minus comparator. Each interval is a 95% speaker bootstrap interval over 16 actors with 10,000 resamples. The two-sided exact sign-flip test uses the paired actor means; its p values are not adjusted for multiple comparisons. Differences are calculated before rounding. Shading identifies the selected allocations.
Post-training quantization (PTQ) reduces the cost of on-device text-to-speech (TTS), but published evaluations cover one system or method. We evaluate PTQ across TTS architectures under one protocol with three core models, weight and activation ablations of eight more, and two held-out models quantized blind. Four-bit per-channel weights reduce UTMOS, a predicted mean opinion score, by 2.8 on Supertonic and 0.07 on Kokoro, and per-tensor scaling can cause severe degradation even at 8 bits. The same bit width yields different outcomes, because the sensitive component is model-specific and not reliably predicted from the model class. A staged ablation procedure identifies it, and per-layer GPTQ can restore it to within 0.1 UTMOS. Real int8 and int4 kernels reproduce the simulated ordering at hardware-dependent cost. On a Mac mini, a 4-bit weight kernel runs Supertonic at 0.60x the fp32 latency while int8 is slower, so each configuration requires validation on the target runtime.
Although low-bit quantization provides practical means to deploy speaker verification on resource-constrained devices, its effects on speaker verification performance remain poorly understood. In this paper, we study uniform K-means quantization-aware training of ResNet-36 and ResNet-200 through joint layer-wise and score-level analyses. Our layer-wise analysis highlights fragile components and shows that score degradation is not fully explained by weight distortion alone. We identify a clear knee point at 2 bits, with larger score drift and harmful decision flips concentrated near the FP32 threshold. Our score-level analysis reveals where and how score errors emerge under extreme quantization. Building on these findings, we propose a calibrated multi-precision cascade that resolves most trials at 2 bits and escalates only ambiguous cases, achieving performance close to FP32 while preserving the efficiency benefits of low-bit inference with substantially lower compute and memory costs.
Hugo Leguillier, Driss Matrouf, Guillaume Lechien +1
Avignon University, LIA, UPR 4128, France · Aday, France
We present HydraQE, our contribution to the IWSLT 2026 Speech Translation Metrics shared task. HydraQE is an end-to-end, reference-free quality estimation (QE) system for speech translation built on a Qwen3-ASR backbone, which accepts source audio and a translation hypothesis as joint input. Hidden states from all backbone layers are combined via a learnable sparsemax scalar mix, then re-encoded by a lightweight bidirectional Transformer to enable full cross-modal interaction prior to pooling into a shared embedding. Three independent prediction heads are trained on complementary supervision signals: human direct assessment (DA) annotations, MetricX-24 pseudo-labels, and xCOMET pseudo-labels. To address the scarcity of human-annotated data, we train on a combination of synthetically corrupted examples and silver pseudo-labeled machine translation outputs, using a curriculum that begins on synthetic and silver data and gradually shifts toward human-annotated examples. HydraQE outperforms cascaded text-based baselines and prior direct speech QE systems, demonstrating that end-to-end speech translation QE is competitive with cascaded approaches.
Kevin Krahn, Eric Fosler-Lussier
Dept. of Computer Science and Engineering The Ohio State University