Boundary-Free Contextual Biasing: Depth-Adaptive Gating and Reading-Space Matching for Unsegmented Languages
Organizations: R&D Team, Recho Inc., Tokyo, Japan
Abstract
Contextual biasing supplies an ASR system with a list of expected words at inference time, but existing methods rely on word boundaries that Japanese and Chinese do not provide. We present a boundary-free biasing decoder for frozen public CTC models, built on a character-level Aho-Corasick automaton, with no training and no second pass. Two evidence-based mechanisms replace the boundary: a depth-adaptive gate that sets how hard to push from match depth, and reading-space matching for when the audio is right but the characters are wrong. On Aishell-1 NE's hard R1 subset we reach 66.5% recall, above the trained CLAS baseline (64%), transferring to WenetSpeech and to a second architecture without retuning. We release the first open Japanese contextual-biasing benchmark, where biasing lifts rare-word recall by 25 points at precision above 97%, and still by 19 and 22 points against 1,000-word lists.
Figures & tables
| corpus | utts | w/ target | hotwords | occurrences |
|---|---|---|---|---|
| JSUT | 5,000 | 422 (8.44%) | 477 | 489 |
| CV ja | 300,077 | 25,876 (8.62%) | 4,357 | 27,108 |
| Aishell-1 NE | WenetSpeech | JSUT | Common Voice ja | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| method | config | CER | Rec. all/R1 | P | CER | Rec. | P | config | CER | Rec. | P | CER | Rec. | P |
| greedy | – | 7.54 | 52.4/8.9 | 99.6 | 5.62 | 88.6 | 99.7 | – | 8.42 | 38.0 | 100 | 18.50 | 29.1 | 99.9 |
| + masked | 5.70 | 77.4/53.1 | 98.8 | 5.51 | 93.9 | 98.8 | 8.52 | 59.7 | 99.3 | 17.77 | 53.9 | 99.7 | ||
| + two-stage | 5.91 | 83.5/66.5 | 96.9 | 6.43 | 94.5 | 93.5 | 8.47 | 54.6 | 99.6 | 17.62 | 50.5 | 99.8 | ||
| + adaptive | , | 5.48 | 81.0/60.6 | 98.7 | 5.69 | 94.1 | 98.3 | , | 8.52 | 60.3 | 99.3 | 17.69 | 53.5 | 99.7 |
| + reading | 5.20 | 79.0/58.1 | 98.9 | 5.41 | 93.5 | 98.6 | 8.68 | 57.5 | 98.0 | 18.41 | 42.2 | 99.5 | ||