Speech Language Models (SLMs) inherit strong instruction-following capabilities from pretrained language models, yet ASR specialization can substantially degrade them. To address this ASR--QA trade-off, we propose Task-Specific On-Policy Distillation (TS-OPD), which leverages models before and after ASR specialization as complementary QA and ASR teachers. The student generates separate task-conditioned trajectories for ASR and QA, each supervised only by its corresponding teacher, thereby reducing direct competition between the two supervision signals. Experiments on basic ASR, contextual ASR, and QA demonstrate that TS-OPD improves recognition while preserving QA capability. Moreover, TS-OPD remains robust across different balancing coefficients and continues to benefit from increased distillation data.
Figures & tables
System
Training Strategy
Basic ASR ↓
ContextASR-Bench
QA (TELEVAL) ↑
LS- test-clean
LS- test-other
GigaSpeech
Overall
WER ↓
NE-FNR ↓
LlamaQA-en
TriviaQA-en
WebQ-en
Overall
A1
QA Adapter
5.26
10.15
15.91
14.19
9.38
12.28
71.67
34.77
41.43
42.57
A2
A1 ASR Full-FT
2.59
5.50
10.52
9.16
5.65
20.31
1.67
4.78
2.22
2.86
A3
ASR Adapter
4.02
8.03
13.04
11.57
9.29
26.02
0.33
2.63
1.81
1.89
A4
ASR+QA Adapter
5.09
9.52
15.08
13.44
8.70
10.54
73.33
36.68
43.09
44.29
B1
TS-OPD
3.55
7.15
11.28
10.03
6.19
16.82
74.00
35.48
44.74
45.07
Table 1 : Main results on basic ASR (WER), contextual ASR (WER and NE-FNR), and QA (ACC). A1, A3, and A4 start from randomly initialized adapters. A1 is trained with QA supervision, A3 with ASR only, and A4 jointly with ASR and QA, using each utterance for both tasks. A2 is obtained by full-parameter ASR fine-tuning from A1. B1 is initialized from A1 and further trained with TS-OPD using 500 hours of distillation data with λ=0.5 . The best and second-best results are shown in bold and underlined text, respectively.
System
Training Strategy
ASR
ContextASR-Bench
QA
Teacher
Teacher
Routing
Overall ↓
WER ↓
NE-FNR ↓
Overall ↑
ASR
QA
A1
-
-
-
14.19
9.38
12.28
42.57
B3
√
-
-
9.96
6.84
22.37
1.89
B2
√
√
ST
10.92
8.96
17.38
42.73
B1
√
√
TS
10.03
6.19
16.82
45.07
Table 2 : Ablation of teacher supervision and routing strategies. B3 uses only the ASR teacher, while B2 and B1 use both teachers. TS and ST denote TS-OPD and ST-OPD, defined in Secs. 2.3 and 3.3 , respectively. The best and second-best results are shown in bold and underlined text, respectively.
Building competitive automatic speech recognition (ASR) models usually requires large-scale au- dio supervision, which makes reproduction and specialization expensive. We study Ark-ASR, a 0.6B- parameter audio-conditioned language model trained with 100k hours of speech, and examine whether a strong Qwen-ASR teacher can transfer additional recognition capability through on-policy distillation. Across Mandarin and English ASR benchmarks, the proposed training recipe consistently improves over supervised fine-tuning alone and outperforms the same-scale Qwen3-ASR-0.6B baseline on four of five evaluation sets. This is achieved with only 100k hours of speech, compared with the 20M hours of super- vised audio reported for the Qwen3-Omni AuT encoder. The larger Qwen3-ASR-1.7B remains stronger, but the results show that teacher-guided on-policy training can substantially close the gap for compact ASR models under a much smaller audio budget. A support-overlap diagnostic further suggests that the teacher-data stage improves local student-teacher compatibility, matching recent analyses of when on-policy distillation is effective.
Modern LLM-based ASR systems have established multilingual capability as a standard feature, leveraging large-scale multilingual corpora and LLMs' cross-lingual knowledge to achieve competitive performance across multilingual benchmarks. However, jointly modeling languages with heterogeneous acoustic, phonological, and lexical characteristics inevitably introduces optimization conflicts, undermining language-wise specialization. To address this challenge, we propose Language-Specialized Multi-Teacher On-Policy Distillation (LS-MOPD), which decouples language-specific knowledge acquisition from multilingual capability integration: language-specialized teachers are independently optimized via reinforcement learning (RL), with their expertise then integrated into a generalist multilingual student through language routing and token-level multi-teacher distillation, thereby reducing direct cross-lingual optimization conflicts. We further explore static and dynamic acoustic-prefix configurations to examine how teacher-student prefix consistency influences the efficacy of on-policy distillation. Experiments on benchmarks covering Mandarin, Mandarin subdialects, Cantonese, and English demonstrate that LS-MOPD substantially outperforms RL baselines and surpasses the empirical performance envelope defined by the best-performing RL teachers on nearly all benchmarks, revealing its potential to generalize beyond all teachers in multilingual ASR.
Yuan Xie, Jiaqi Song, Xianliang Wang +3
Advanced Intelligent Systems Group, NIO, Beijing, China
Spoken language models (SLMs) unify speech perception and reasoning, but adapting them to sensitive domains is underexplored, especially when the original training data is inaccessible and the use case demands multilingual, spoken-query interaction. We adapt an open-source SLM to the Singaporean Home Team context across five speech tasks in Singapore's four official languages, combining LoRA fine-tuning, a surrogate text-QA dataset that guards against catastrophic forgetting, and a multi-task objective that adapts the CoBa reweighting scheme to speech. We also build HTD-multilingual-QA, a 504,853 sample multilingual QA dataset in text and spoken form. The resulting HT-Moonstone (5B) matches or outperforms SLMs up to 7x its size on most tasks, attains the best accent and gender recognition among all models evaluated, and loses under 2% of its original speech QA ability.
Ng Jia Sheng Jason
Language AI R&D, xData Home Team Science & Technology Agency (HTX), Singapore