Automatic Speech Recognition (ASR) systems often show uneven performance across demographic groups, and errors can be especially difficult to address for speakers belonging to multiple demographic groups. This work studies demographic-aware model merging for fair Speech-LLM-based ASR. Starting from a SLAM-ASR-based model, we fine-tune only the connector on demographic-specific subsets and merge the resulting subgroup-adapted connectors into a global model. We then identify critical cross-axis demographic pairs using subgroup WER and task-vector conflict, and apply intersection-specific correction vectors to the global merged model. Experiments on Fair-Speech show that global demographic merging improves overall WER over the base model, while intersection correction provides additional gains for several merging strategies. In particular, TIES with WER-based correction achieves the best overall WER, reducing it from 7.38% to 5.13%. Subgroup and disparity analyses further show that the proposed approach improves performance across demographic axes, while highlighting that lower average WER does not always imply reduced subgroup disparity.
Figures & tables
Merging Method
Global
Rand
M1
M2
M3
Base
7.38
Linear
6.4312.9
6.4113.1
6.3114.5
6.3414.1
6.3713.7
TIES
6.4213.0
6.4812.2
5.1330.5
6.3813.6
5.5524.8
Task Arithmetic
6.3813.6
6.4013.3
6.3214.4
6.3314.2
6.3214.4
Breadcrumbs
6.778.30
6.4412.7
5.7522.1
6.4812.2
5.6323.7
DARE-TIES
6.4312.9
6.4512.6
6.3414.1
6.3314.2
6.3414.1
Table 1: Overall WER (%). Relative WER improvements over Base are shown as subscripts.
Merging method
Variant
Age
Gender
Ethnicity
SES
First Lang.
Base
–
7.07
7.65
6.84
7.08
6.65
Linear
Global
6.31
6.63
6.04
6.23
5.87
Linear
M1
6.20
6.52
5.91
6.13
5.76
Linear
M2
6.28
6.58
5.89
6.19
5.82
Linear
M3
6.25
6.56
5.89
6.12
5.78
TIES
Global
6.27
6.62
6.08
6.27
5.88
Table 2: Demographic-level WER (%).
Figure 1: Relative reduction in max–min WER disparity compared with the base model.
Figure 2: Average WER relative to global merging. Lower than 100% is better.
Speech-LLMs are expensive to run, making compression important for real-world deployment. However, compressed models are usually selected using aggregate word error rate (WER), which can hide how pruning affects different demographic groups. In this work, we systematically study the effect of audio encoder pruning on SLAM-ASR for different demographic groups. Using the Fair-Speech and Common Voice datasets, we found that the pruning does not affect all demographic groups equally; the gap between best- and worst-performing groups increases in fold. These disparities appear across all three encoder scales, but only the largest model initially hides them behind aggregate WER. LoRA adaptation improves WER for every group, but benefits groups already performing well more strongly and widens for certain groups. On Common Voice English, Danish, and Dutch, accent gaps persist but do not clearly widen, showing that the fairness effects of pruning vary across datasets and must be measured directly. Our findings suggest that for pruned models, deployment decisions should include per-group WER, with the worst-performing group's error rate as an explicit criterion.
Ganesh Pavan Kartikeya Bharadwaj Kolluri, Michael Kampouridis, Ravi Shekhar
School of Computer Science and Electronic Engineering, University of Essex, UK
Speech technology penalizes some voices: recognition errs nearly twice as often for Black speakers, and accuracy declines for second-language accents and older speakers. We introduce TRIAD, an audit grid crossing 120 texts, 24 rendered demographic voice profiles (gender, age band, accent), and ten expressive styles via controllable text-to-speech, isolating perceived demographic attributes from content and affect. For ten open-weights encoders we define axis-fidelity functionals, principal-angle leakage between axis subspaces, and group-conditional gaps; a proposition proves that average probe disparity grows with the same aggregate voice-semantic leakage Λ we measure, and a corollary shows that peak leakage forces worst-case disparity inside an active region. The measured mean-square probe disparity tracks Λ (Pearson r = 0.93), and a black-box protocol exposes the same signature in two closed-source models. ORCA, an adapter combining axis-specific contrastive heads, an orthogonality penalty, and group-balanced sampling, cuts leakage 72% and roughly halves the gaps.
Automatic speech recognition (ASR) systems, trained on paired speech-text data, have been improved by leveraging language models (LMs) trained on text-only data. LM fusion methods such as shallow fusion and density ratio are well-established methods that incorporate external LMs during ASR decoding. However, they incur additional computational costs due to LM inference, which is particularly problematic for recent larger LMs. In this study, we propose incorporating external LMs via model merging. This method integrates the LMs directly into the parameters of an LLM-based ASR model, requiring no additional computational cost at inference. We formulate domain extension and transfer via arithmetic operations on LoRA parameters. Experimental evaluations were conducted for the domain adaptation of LLM-based ASR trained on CSJ and LibriSpeech. We show that our LM merging consistently improved the ASR performance in the target domains, without degrading inference speed or memory footprint.
Hayato Futami, Tatsuya Kawahara
Graduate School of Informatics, Kyoto University, Japan