Contextual biasing improves rare-word recognition in speech large language models (SpeechLLMs), but efficiently exploiting large bias lists remains challenging. We propose PTC-Bias, a two-stage framework based on phoneme-level temporal competition. At the prefill stage, PTC Retrieval performs frame-synchronous phoneme decoding and temporal competition among candidate pronunciations, producing a compact bias-word shortlist and corresponding speech intervals. After SpeechLLM decoding, PTC Correction conducts a second local competition between the retrieved candidates and mismatched transcript spans within these intervals. Selective correction reduces near-homophone and word-segmentation errors while preserving correct transcriptions. Both stages share the same phoneme posteriors and require no additional SpeechLLM forward pass. Experiments on LibriSpeech show consistent gains across two SpeechLLMs and bias lists of up to 2000 words. With Prompt-SLAM-ASR-7B and 2000 bias words, PTC-Bias reduces B-WER by 23.4%/23.9% relative to CTC-Filter on test-clean/test-other, while keeping U-WER nearly unchanged.
Figures & tables
Figure 1: Overview of PTC-Bias. A lightweight phoneme-CTC branch extracts frame-level phoneme posteriors from the frozen audio encoder. PTC Retrieval performs temporal competition to select acoustically supported bias words and locate their speech intervals before SpeechLLM decoding. PTC Correction then reuses the intervals and cached posteriors to correct mismatched transcript spans through local competition.
Contextual ASR Model
N=100
N=500
N=1000
N=2000
WER
B-WER
U-WER
WER
B-WER
U-WER
WER
B-WER
U-WER
WER
B-WER
U-WER
DB-NNLM [ 12 ]
1.98 / 5.86
5.70 / 14.10
1.50 / 4.90
2.09 / 6.09
6.20 / 15.10
1.60 / 5.10
2.14 / 6.35
6.70 / 17.20
1.60 / 5.10
2.27 / 6.58
7.30 / 18.90
1.60 / 5.20
USTR-CT [ 17 ]
2.06 / 5.38
2.00 / 4.40
2.10 / 5.50
2.09 / 5.62
2.20 / 5.60
2.10 / 5.60
2.16 / 5.75
2.50 / 6.30
2.10 / 5.70
2.17 / 5.84
3.00 / 7.60
2.10 / 5.60
CB-QwenAudio [ 7 ]
1.60 / 3.80
5.50 / 13.50
1.30 / 2.60
1.90 / 3.90
6.00 / 14.20
1.40 / 2.70
–
–
–
–
–
–
Prompt-SLAM-ASR-7B [ 20 ]
7.40 / 17.90
24.80 / 44.10
–
–
–
–
–
–
–
–
–
–
+ CTC-Filter [ 20 ]
1.27 / 2.72
3.67 / 8.02
1.00 / 2.16
1.33 / 3.04
3.92 / 9.04
1.03 / 2.40
1.33 / 2.99
4.16 / 9.33
1.00 / 2.31
1.38 / 3.20
4.41 / 10.02
1.03 / 2.47
Table 1: Contextual ASR results (%) on LibriSpeech with different numbers of distractors. Each entry reports test-clean / test-other. In the Qwen3-ASR block, † denotes WavLM-based phoneme posteriors; otherwise, the AuT front-end is used. Lower is better.
Front-end
LS-460
LS-960
test-clean
test-other
test-clean
test-other
DS-KWS [ 3 , 2 ]
4.44
13.39
4.45
11.80
AuT [ 18 ]
2.39
5.88
2.06
5.13
WavLM [ 5 ]
1.29
2.44
1.13
2.27
Table 2: Phoneme error rates (PER, %) of different phoneme front-ends on LibriSpeech. Lower is better.
Method
RecallB@99↓
RecallB#50 (%) ↑
RecallH#50 (%) ↓
BR-ASR (Acoustic) [ 8 ]
42.2
99.7
69.3
BR-ASR (Textual) [ 8 ]
47.3
99.1
58.4
PTC Retrieval
16.9
99.3
22.0
Table 3: Bias retrieval on LibriSpeech test-other with N=2000 distractors. Acoustic and Textual are BR-ASR [ 8 ] variants .
Figure 2: Temporal competition among phoneme paths in PTC Retrieval. Solid paths are selected; dashed paths are competing alternatives.
Figure 3: CPU search latency of PTC Retrieval for a 5.09-s utterance across bias-list sizes and worker counts. Phoneme-posterior extraction and trie construction are excluded.