Speech-LLMs often exhibit prompt overfitting, where models solely trained on automatic speech recognition (ASR) instruction fail to generalize to new instructions such as speech translation and continue to behave primarily as ASR system. We propose DirectSpeech2LLM, a simple end-to-end framework that preserves the instruction-following ability of the LLM on unseen tasks when conditioned on speech. It computes distance-based CTC loss over the frozen LLM embedding matrix and uses greedy CTC labels to derive geometrically and temporally aligned speech embeddings respectively as an input to the LLM. Trained solely on 960 hours of LibriSpeech ASR data, DirectSpeech2LLM outperforms the cascaded system on ASR (seen task) and generalizes zero-shot to speech translation and emotion recognition (two unseen tasks), closely matching the cascaded system upper bound on these two new instructions despite seeing neither during training. We also find that geometric alignment strength plays a smaller role than previously assumed, as our modified CTC loss is shown to provide sufficient implicit geometric grounding without requiring an explicit regression loss. Results are consistent across two LLM families and scale with both more training data and model capacity.
Figures & tables
Figure 1 : Overview of DirectSpeech2LLM . Given an input speech signal, the SFM produces frame representations st∈Rd , where d matches the dimensionality of the LLM input embeddings Q . Logits are computed as negative squared distances ( Zt,v ) followed by log-softmax and CTC loss ( Lctc ) is applied. Next, the resulting greedy CTC predicted labels are then used to downsample SFM output embeddings via mean pooling. The LLM is fed with the resulting updated sequence embeddings from SFM and is trained with cross-entropy loss ( Lllm ). The reader should keep in mind that the LLM-IES serves only as a geometric anchor for the SFM-OES, without being replaced or discretized with the LLM input embeddings. Commitment loss is not shown in the figure for clarity.
Model
LLM (Learnable)
Train data duration (hrs)
ASR (Librispeech)
ST (CoVoST2)
SQA (IEMOCAP)
WER ↓
BLEU ↑
ACC ↑
t-clean / t-other
en-de
en-zh
Emotion
Cascaded (CTC + LLM) (50k steps)
Whisper + LLM [ 11 ]
LLama-7B
-
2.7 / 5.2
18.2
-
-
Ours ( Backbone )
Phi3.1-3.8B
1k
3.13 / 5.12
19.37
17.52
44.76 (IFR = 0.99)
Qwen3-4B
2.35 / 4.57
19.88
34.03
42.98 (IFR = 1.00)
Table 1 : Models marked with SFT uses supervised finetuning for that task. DirectSpeech2LLM shows near perfect instruction following on two unseen tasks, ST and ER as performance approaches to the upper bound of cascade system showing LLM’s generalization ability to new instructions is preserved. IFR, as explained in [ 13 ] is shown for SQA task. [( → ) means continued training after 38k hours.]
Hparams
ASR (WER ↓ )
ST (BLEU ↑ )
Cascaded
end-to-end
Cascaded
CTC → +LLM
CTC / W (ith) B (lank) / N (o) B (lank)
WB / NB
CTC + LLM
Libri-other (in-domain)
VP (out-of-domain))
CoVoST2 (en-de)
Stage 1: Modified -CTC (20k steps)
→Lctc
4.88 → 4.86
- / - / 100.00
16.14 / - / 100.00
- / 0.00
19.89
→+Lcommit
4.94 → 4.85
- / - / 16.05
16.18 / - / 26.43
- / 10.44
19.92
Table 2 : Ablation study of DirectSpeech2LLM on Stage 1 and Stage 2. Assume Qwen3-4B model and LibriSpeech training data (1k) unless mentioned otherwise. Results compare a cascaded interface (CTC + LLM) with an end-to-end interface. WB and NB corresponds to With blank and No blank setup in training and inference both. Different rows show training setup and columns show inference setup. Hparams (rows) denote the effect of different losses, extended training steps, additional speech data (CV) or different LLM backbone. For stage 2 results, a fair comparison to Linear-CTC loss is Modified-CTC loss −Lcommit . Gray rows show the best results for each stage. Row marked with ( ∗ ) is the closest equivalent to AlignFormer minus the input ordering trick.
en-es
en-de
Wav2Prompt (CIF)
13.8
-
− MSE
5.4
-
DirectSpeech2LLM (Modified CTC)
-
19.97
−Lcommit
-
19.69
Table 3 : The BLEU scores show the sensitivity to the MSE loss in Wav2Prompt or commitment loss in DirectSpeech2LLM. Both are regression loss for geometric alignment.
Figure 2 : Validation commitment loss of DirectSpeech2LLM with and without commitment loss.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Tasks
WB
WB decoding_E
NB
CTC + LLM
Libri-clean
2.4
2.95
2.8
2.39
Libri-other
4.39
5.73
5.51
4.61
en-de
21.45
21.61
21.68
20.29
en-zh
34.20
33.91
35.23
34.98
Emotion
40.89
39.35
41.53
43.55
Appendix
Table 4 : Multitask training setup with Librispeech and en-de training data without commitment loss.
Duration
WB
NB
CTC
1 min
.62
0.62
0.62
5 min
2.97
2.6
3.71
10 min
14.26
4.22
5.08
20 min
38.48
33.12
12.37
Appendix
Table 5 : Longform speech inference results.
Figure 3 : Left: Stage 2 LLM WER vs LLM backbone shows training longer would improve performance further. Right: Commitment loss vs LLM backbone.
Figure 4 : TBP method overview. The authors of Byte Latent Transformer [ 47 ] use dynamic boundaries at step 2, while TBP uses static tokenizer dependent boundaries.
-
-
-
t-clean / t-other
en-de
en-zh
Emotion
end-to-end
Ours
Qwen-3-4B
1k
2.21 (2.05) / 4.28 (4.13)
19.97
33.74
40.16 (IFR = 1.00)
Ours −Lcommit
2.13 (1.96) / 4.09 (3.93)
19.69
32.86
39.52 (IFR = 1.00)
Ours + TBP
1.96 (1.81) / 4.06 (3.90)
19.79
32.74
40.97 (IFR = 1.00)
Appendix
Table 6 : Extension to Table 1 . Numbers with underline shows the best result for each test set. The improvements except on the ASR task are marginal. Overall, the results suggest that input tokenization has no substantial effect on the LLM’s task-solving ability. Values in parentheses report ASR results computed using the Whisper EnglishTextNormalizer compared to mms normalization.
Hparams
ASR (WER ↓ )
ST (BLEU ↑ )
Cascaded
end-to-end
Cascaded
CTC → +LLM
CTC / W (ith) B (lank) / N (o) B (lank)
WB / NB
CTC + LLM
Libri-other (in-domain)
VP (out-of-domain))
CoVoST2 (en-de)
Stage 2
Ours
4.60 → 4.57
- / 4.28 / 5.83
15.95 / 13.86 / 15.09
5.80 / 19.97 (49.4/26.1/15.1/9.1)
19.88 (48.8/25.7/14.8/8.9)
Ours_TBP
4.38 → 4.38
- / 4.06 / 4.43
15.74 / 13.89 / 14.42
11.44 / 19.79 (49.0/25.8/15.0/9.0)
20.10 (49.3/26.2/15.2/9.1)
Appendix
Table 7 : Extension to Table 2 . Numbers with underline shows the best result for each test set.
Figure 5 : Updated DirectSpeech2LLM_TBP method.
Zt,v=−(∥st∥22−2st⊤qv+∥qv∥22)
Appendix
Algorithm 1 DirectSpeech2LLM method
Figure 6 : Prompt template used for Stage 2 training and inference with Qwen3, including task-specific prompts for ASR, speech translation, and emotion recognition.
Stage 1 – Failure: Formatting
Real: i’d recommend him to you instead of blackstone thanks laughed kenneth
CTC: i’d recommend him to you instead of blackstone thanks laupped kenneth
LLM (NB): i d e r c o m e n d h i m t o y o u i n s t e a d o f b a c k s t o n e t h a n s l a l p e d k e n n e t h
Stage 2 – Failure: Formatting / Translation
Real: if they tried to run they were hit from behind if they stood still they were clubbed carefully
CTC: if they tried to run they were hit from behind if they stood still they were clubbed carefully
Appendix
Table 8 : Dominant failure cases of DirectSpeech2LLM on the ASR task when using the NB decoding strategy. Stage 1 formatting failures are expected as we ask the LLM to repeat. Even after stage 2 LLM finetuning, the formatting failures of limited vocabulary of speech tokenizer are still visible, very rare. We also found degradation in instruction following as sometimes, instead of repeating, the llm starts translating.
Model
Vocab
WER (%)
Base comparison
Phi-3.1-3.8B
Vllm
0.91
Qwen-3-4B
0.02
Phi-3.1-3.8B
Vctc
6.91
Qwen-3-4B
4.40
Qwen scaling under Vctc
Appendix
Table 9 : Effect of using a ∼ 1k subset of the LLM vocabulary for tokenizing SFM-OES. When evaluated on LibriSpeech test-other using ground-truth transcripts as input to the ASR prompt. Performance degrades substantially under Vctc , with Phi-3.1-3.8B showing higher sensitivity than Qwen-3-4B. No clear scaling trend is observed across Qwen models. The reader should keep in mind that inflated WER is because of formatting errors as shown in Table 8 in the Appendix. Furthermore, these errors to a large extent are resolved after finetuning.
School of Data Science, The Chinese University of Hong Kong, Shenzhen, China. · Microsoft Research, Redmond, WA, USA. · Microsoft Research Asia, Hong Kong, China.