Automatic speech recognition models are audited for demographic fairness at full precision, yet the models that ship to production have been quantized, pruned, and distilled. We ask whether post-training weight compression, which alters model weights rather than the audio signal or its feature representation, redistributes error burden across demographic groups. Across the Whisper family on Fair-Speech, Common Voice 25, and AfriSpeech-200, 50% Wanda pruning of Whisper-large-v3 sharply widens the Black/AA-vs-Asian temporal-taxation differential on Fair-Speech: the absolute word-error-rate gap between the worst- and best-served groups more than doubles; at an assumed cost of five seconds of correction effort per transcription error this is a rise from 30 to 64 seconds of correction time per minute of speech. This +111% relative increase is invariant to the assumed per-error cost, survives an audio-quality control, and is only partly mitigated by beam-search decoding, which still leaves an +86% increase. At edge model size, INT4 HQQ quantization compounds catastrophic transcript loops on West African accents by factors of five to seven. Distillation, by contrast, narrows demographic gaps in 21 of 27 evaluated settings (teacher-student pair, precision, and dataset), with the exceptions concentrated on a single model pair. We cast the temporal-taxation construct of Choi and Choi (2025) as a quantitative metric, and show that single-snapshot fairness audits on full-precision models do not capture the deployment-time burden that compression places on already-marginalized speakers.
Figures & tables
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Precision
Kendall τ
AfriSpeech
INT8
1.000
AfriSpeech
INT4 NF4
1.000
AfriSpeech
INT4 HQQ
1.000
Common Voice 25
INT8
1.000
Common Voice 25
INT4 NF4
1.000
Common Voice 25
INT4 HQQ
1.000
Appendix
Table 1: Kendall’s τ between FP16 model rankings by MMR and rankings under each compressed precision. The single adjacent-pair swap at Fair-Speech + INT4 HQQ is the only deviation from a perfect ordering.
Model
Cell
C=2
C=5
C=8
rel.
whisper-large-v3
Wanda, FS Black/AA vs Asian
12.06 → 25.49
30.14 → 63.73
48.23 → 101.97
+111%
whisper-medium
Wanda, FS Black/AA vs Asian
14.99 → 28.21
37.46 → 70.53
59.94 → 112.85
+88%
whisper-base
Wanda, CV25 indian vs canada
29.86 → 50.79
74.64 → 126.97
119.43 → 203.16
+70%
whisper-large-v3
Wanda, CV25 indian vs canada
14.08 → 23.79
35.20 → 59.47
56.32 → 95.15
+69%
whisper-small
Wanda, FS Black/AA vs Asian
23.48 → 35.67
58.69 → 89.17
93.91 → 142.68
+52%
whisper-tiny
INT4 HQQ, CV25 african vs canada
38.82 → 53.99
97.04 → 134.97
155.27 → 215.95
+39%
Appendix
Table 2: Worst-vs-best temporal-taxation differential in s/min, FP16 → compressed, at three cost-per-edit values. The relative change column is invariant in C .
(rep5, len-ratio)
FP16 → HQQ
BH-sig
CI
Loose (4,2.5)
2.07% → 6.79%
6/62
+0.73
Baseline (5,3.0)
1.50% → 5.78%
7/62
+0.68
Strict (6,3.5)
1.42% → 5.46%
8/62
+0.56
Appendix
Table 3: H2 cell ( whisper-tiny + INT4 HQQ on AfriSpeech) under three loop-flag thresholds. BH-sig is the count of accents (out of 62 with n≥30 ) reaching q<0.05 in two-proportion z-tests with Benjamini-Hochberg correction. CI is the smoothed loop-rate MMR compounding index. Aggregate loop rate decreases as the threshold tightens, but the FP16-to-HQQ contrast and the per-accent ranking are preserved.
Accent ( n )
Loose
Baseline
Strict
Kanuri (66)
4.00×⋆
6.67×⋆
6.33×⋆
Yoruba (648)
3.50×⋆
5.12×⋆
5.00×⋆
Hausa (196)
4.75×⋆
6.00×⋆
5.00×⋆
Swahili (521)
2.71×
3.40×
4.25×⋆
Igbo (355)
2.20×
2.83×
2.83×
Appendix
Table 4: INT4 HQQ over FP16 loop-rate ratio under whisper-tiny on AfriSpeech, for the five headline West African accents at each loop-flag setting. ⋆ marks q<0.05 after BH-FDR correction within the 62-accent slate. The three accents driving the H2 claim (Kanuri, Yoruba, Hausa) are significant at every setting.