cs.SDSep 20, 2026

Text Scores Do Not Establish Performance on Lexically Non-Diagnostic Speech Tasks: A Qwen2-Audio Quantization Case Study

Authors: Mengzhe Geng, Jinxi Ji, Junhao Xu

Organizations: National Research Council Canada · The Hong Kong Polytechnic University · The Chinese University of Hong Kong

Abstract

Text-output scores alone do not show whether quantization preserves performance on speech tasks whose target labels cannot be recovered from the transcript. We evaluate fixed mixed 4/8-bit Qwen2-Audio-7B-Instruct allocations averaging 6 and 7 bits per parameter on 508 English-to-German FLEURS utterances and on 512 RAVDESS emotion clips from 16 speakers. The BLEU and chrF differences from half precision (FP16) have intervals that include zero for both allocations. On RAVDESS, the same two sentences occur equally often with every emotion label. The absolute accuracy differences from FP16 are -3.71% for 6 bit and -1.17% for 7 bit. The 6-bit speaker interval excludes zero and an exact two-sided sign-flip test gives p=0.0148; the 7-bit interval includes zero. Same-budget controls do not identify either selected allocation as best. This case study shows why translation scores and performance on tasks beyond the transcript need separate evaluation.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. On Low-Bit Quantization Errors in Speaker Verification: Diagnostic and Mitigation

    Jun 6, 2026Hugo Leguillier, Driss Matrouf, Guillaume Lechien +1Automatic Speaker VerificationQuantization-Aware Training

  2. HydraQE: OSU's Submission for the IWSLT 2026 Speech Translation Metrics Shared Task

    Jun 7, 2026Kevin Krahn, Eric Fosler-LussierSpeech TranslationImage Quality Assessment