cs.CLSep 29, 2026

BaLEEN: Biasing with Latent Encoded Entities for Context-Aware ASR

Authors: Chihiro Taguchi, Yotaro Kubo, Rujikorn Charakorn

Organizations: University of Notre Dame, Department of Computer Science and Engineering, IN, USA · Sakana AI, Tokyo, Japan

Abstract

Transcribing domain-specific entities and rare proper nouns remains a major challenge in automatic speech recognition (ASR). In this paper, we propose BaLEEN (Biasing with Latent Encoded Entities), a lightweight, hypernetwork-based framework for dynamic contextual adaptation without fine-tuning the underlying ASR model. BaLEEN encodes variable-length contextual keywords using a pretrained language model, compresses them into a fixed sequence of latent vectors via a Perceiver bottleneck, and injects context-dependent bias vectors directly into the intermediate encoder representations of the ASR model. Because both the language model and the backbone ASR model remain entirely frozen during training, BaLEEN operates as a plug-and-play adapter that incurs zero computational overhead at inference time when context biases are precomputed. We evaluate our method on a CTC-based ASR model using a Wikipedia-derived corpus with annotated named entities and synthetic speech. Experimental results demonstrate that BaLEEN reduces keyword miss rate by 8.7% on the test set relative to the unbiased baseline while simultaneously improving overall word error rate by 21% and character error rate by 28%.

Figures & tables

Explore similar work

CardsList
  1. How to Recognize New Words: A Comparison Between Context Biasing Methods and Speech LLMs

    Aug 6, 2026Christian Huber, Alexander WaibelAutomatic Speech RecognitionBiases