Contextual biasing supplies an ASR system with a list of expected words at inference time, but existing methods rely on word boundaries that Japanese and Chinese do not provide. We present a boundary-free biasing decoder for frozen public CTC models, built on a character-level Aho-Corasick automaton, with no training and no second pass. Two evidence-based mechanisms replace the boundary: a depth-adaptive gate that sets how hard to push from match depth, and reading-space matching for when the audio is right but the characters are wrong. On Aishell-1 NE's hard R1 subset we reach 66.5% recall, above the trained CLAS baseline (64%), transferring to WenetSpeech and to a second architecture without retuning. We release the first open Japanese contextual-biasing benchmark, where biasing lifts rare-word recall by 25 points at precision above 97%, and still by 19 and 22 points against 1,000-word lists.
Figures & tables
Figure 1: (a) Character-level Aho-Corasick automaton for 東京 and 京都; solid arrows are trie edges, the dashed one a failure link. Circled numbers trace the decoded stream 東京都 along the thick path, the failure link moving from the end of 東京 into the middle of 京都 without consuming a character, so both are matched. (b) The precomputed token-level table and per-frame decode flow; cells give the state reached and the bias ψ(d′)−ψ(d) at w=1 , plus a lump c per hotword completed.
corpus
utts
w/ target
hotwords
occurrences
JSUT
5,000
422 (8.44%)
477
489
CV ja
300,077
25,876 (8.62%)
4,357
27,108
Table 1: JA biasing benchmark statistics. JSUT is fully human-verified; CV is partly human-reviewed, the unreviewed residual noise rate is measured by a seeded blind sample and reported in the dataset card.
Aishell-1 NE
WenetSpeech
JSUT
Common Voice ja
method
config
CER
Rec. all/R1
P
CER
Rec.
P
config
CER
Rec.
P
CER
Rec.
P
greedy
–
7.54
52.4/8.9
99.6
5.62
88.6
99.7
–
8.42
38.0
100
18.50
29.1
99.9
+ masked
w=2.5
5.70
77.4/53.1
98.8
5.51
93.9
98.8
w=2
8.52
59.7
99.3
17.77
53.9
99.7
+ two-stage
w=4
5.91
83.5/66.5
96.9
6.43
94.5
93.5
w=2
8.47
54.6
99.6
17.62
50.5
99.8
+ adaptive
w=3 , d⋆=2
5.48
81.0/60.6
98.7
5.69
94.1
98.3
w=2 , d⋆=3
8.52
60.3
99.3
17.69
53.5
99.7
+ reading
w=2
5.20
79.0/58.1
98.9
5.41
93.5
98.6
w=1.5
8.68
57.5
98.0
18.41
42.2
99.5
Table 2: Decode-only biasing on frozen SenseVoice-small: Aishell-1 NE test, recall over the full list and over the hard R1 subset (trained references, R1 recall [ 15 ] : CLAS 64, SeACo 80, +ASF 87); its untuned transfer to WenetSpeech [ 19 ] ; the JA benchmark (§3; N=100, configs frozen on internal benchmark). K=10 ; P=biased-word precision. Reading and combined (adaptive+reading) rows are frequency-damped (§2.3).
Figure 2: Gate-family regimes. (a) Threshold sweep at matched weight ( w=2 , N=1000 ); each endpoint degrades in the band the other protects. (b) Effect of distractor list-size scaling on recall.
Transcribing domain-specific entities and rare proper nouns remains a major challenge in automatic speech recognition (ASR). In this paper, we propose BaLEEN (Biasing with Latent Encoded Entities), a lightweight, hypernetwork-based framework for dynamic contextual adaptation without fine-tuning the underlying ASR model. BaLEEN encodes variable-length contextual keywords using a pretrained language model, compresses them into a fixed sequence of latent vectors via a Perceiver bottleneck, and injects context-dependent bias vectors directly into the intermediate encoder representations of the ASR model. Because both the language model and the backbone ASR model remain entirely frozen during training, BaLEEN operates as a plug-and-play adapter that incurs zero computational overhead at inference time when context biases are precomputed. We evaluate our method on a CTC-based ASR model using a Wikipedia-derived corpus with annotated named entities and synthetic speech. Experimental results demonstrate that BaLEEN reduces keyword miss rate by 8.7% on the test set relative to the unbiased baseline while simultaneously improving overall word error rate by 21% and character error rate by 28%.
Chihiro Taguchi, Yotaro Kubo, Rujikorn Charakorn
University of Notre Dame, Department of Computer Science and Engineering, IN, USA · Sakana AI, Tokyo, Japan
Recognizing new and rare words - named entities, acronyms, domain specific special words, and other items scarce in training data - remains a key challenge for automatic speech recognition (ASR). We compare two strategies for this: context biasing methods, where an ASR model is extended such that during inference a word list can be supplied, and speech large language models (LLMs) prompted with context directly. We evaluate two context biasing methods based on Whisper against three speech LLMs across read and non-read speech, reporting biased, unbiased, and overall word error rate (WER). The context biasing methods cut biased WER by up to 88% relative while leaving other words largely unaffected. Speech LLMs excel on read speech but generalize less well to non-read speech, and prove sensitive to distractor count and prompt word order. We characterize the resulting trade-offs to guide method selection.
Christian Huber, Alexander Waibel
Interactive Systems Lab · Karlsruhe Institute of Technology · Karlsruhe, Germany +2
Contextual biasing improves rare-word recognition in speech large language models (SpeechLLMs), but efficiently exploiting large bias lists remains challenging. We propose PTC-Bias, a two-stage framework based on phoneme-level temporal competition. At the prefill stage, PTC Retrieval performs frame-synchronous phoneme decoding and temporal competition among candidate pronunciations, producing a compact bias-word shortlist and corresponding speech intervals. After SpeechLLM decoding, PTC Correction conducts a second local competition between the retrieved candidates and mismatched transcript spans within these intervals. Selective correction reduces near-homophone and word-segmentation errors while preserving correct transcriptions. Both stages share the same phoneme posteriors and require no additional SpeechLLM forward pass. Experiments on LibriSpeech show consistent gains across two SpeechLLMs and bias lists of up to 2000 words. With Prompt-SLAM-ASR-7B and 2000 bias words, PTC-Bias reduces B-WER by 23.4%/23.9% relative to CTC-Filter on test-clean/test-other, while keeping U-WER nearly unchanged.
Zhiqi Ai, Han Cheng, Shiyi Mu +2
Shanghai University, Shanghai, China · Xi’an Jiaotong-Liverpool University, Suzhou, China