Text-output scores alone do not show whether quantization preserves performance on speech tasks whose target labels cannot be recovered from the transcript. We evaluate fixed mixed 4/8-bit Qwen2-Audio-7B-Instruct allocations averaging 6 and 7 bits per parameter on 508 English-to-German FLEURS utterances and on 512 RAVDESS emotion clips from 16 speakers. The BLEU and chrF differences from half precision (FP16) have intervals that include zero for both allocations. On RAVDESS, the same two sentences occur equally often with every emotion label. The absolute accuracy differences from FP16 are -3.71% for 6 bit and -1.17% for 7 bit. The 6-bit speaker interval excludes zero and an exact two-sided sign-flip test gives p=0.0148; the 7-bit interval includes zero. Same-budget controls do not identify either selected allocation as best. This case study shows why translation scores and performance on tasks beyond the transcript need separate evaluation.
Figures & tables
Figure 1: Why quantized speech models need more than a transcript score. A: Schematic waveforms illustrate two recordings with the same words but different emotion labels. B: Translation scores do not measure performance on labels that the transcript cannot determine. C: The selected 6-bit allocation has lower RAVDESS accuracy than FP16, with an absolute difference of −3.71 % (95% speaker interval [−6.05,−1.37] %; exact two-sided sign-flip p=0.0148 ). The 7-bit interval includes zero.
Allocation
BLEU ↑
chrF ↑
Accuracy (%) ↑
FP16
20.36
49.73
70.12
Translation-selected, 6 bit
20.22
49.60
66.41
Translation-selected, 7 bit
20.33
49.68
68.95
Table 1: FLEURS translation scores and RAVDESS emotion accuracy for the same precision allocations. BLEU and chrF use SacreBLEU 2.5.1; emotion accuracy (%) uses the label-normalization rule in Section 3 . Differences are calculated before rounding. Shading identifies the translation-selected allocations.
Figure 2: Selected allocation minus comparator with 95% intervals. Panels A and B show FLEURS translation-score differences from FP16, using paired utterance resampling (508 examples; 2,000 resamples). Panels C and D show RAVDESS accuracy differences from FP16 and same-budget controls, using speaker resampling (16 speakers; 10,000 resamples). A negative value favors the comparator.
Allocation
Δ BLEU ↑ (95% CI)
Δ chrF ↑ (95% CI)
6 bit
−0.15[−0.51,+0.14]
−0.13[−0.41,+0.10]
7 bit
−0.04[−0.38,+0.26]
−0.05[−0.33,+0.19]
Table 2: FLEURS translation-score differences, selected allocation minus FP16. Paired 95% bootstrap intervals resample the 508 utterances 2,000 times with seed 1729. All intervals include zero. CI denotes confidence interval.
Actor
FP16 (%) ↑
Selected 6 (%) ↑
Selected 7 (%) ↑
6 − FP16 (%) ↑
7 − FP16 (%) ↑
01
56.25
53.12
53.12
−3.12
−3.12
02
75.00
65.62
75.00
−9.38
0.00
03
84.38
78.12
84.38
−6.25
0.00
04
68.75
62.50
68.75
−6.25
0.00
05
71.88
68.75
71.88
−3.12
0.00
06
65.62
62.50
65.62
−3.12
0.00
Table 3: RAVDESS emotion accuracy by actor. Each actor contributes 32 clips. The final two columns are absolute accuracy differences, selected allocation minus FP16, for that actor, computed before rounding. The uncertainty analyses use these actor-level means.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Allocation
Accuracy (%) ↑
Selected − comparator (%) ↑ (95% CI)
Sign-flip p
6-bit budget
FP16
70.12
−3.71[−6.05,−1.37]
0.0148
Selected
66.41
Reference
—
Front-block
68.16
−1.76[−4.30,+0.78]
0.2754
Stratified-uniform
68.36
−1.95[−4.30,+0.39]
0.1836
7-bit budget
Appendix
Table 4: RAVDESS emotion accuracy and absolute accuracy differences, selected allocation minus comparator. Each interval is a 95% speaker bootstrap interval over 16 actors with 10,000 resamples. The two-sided exact sign-flip test uses the paired actor means; its p values are not adjusted for multiple comparisons. Differences are calculated before rounding. Shading identifies the selected allocations.