Abstract
Open-vocabulary keyword spotting (KWS) must detect arbitrary keywords without retraining. Cross-attention-based models achieve state-of-the-art performance but require pairwise interaction between the audio and each keyword, recomputing the fused representation for every audio--keyword pair. Embedding-based models avoid this computational cost through independent encoding and similarity scoring, yet remain inferior on challenging benchmarks such as LibriPhrase-hard. We show that this gap is primarily a training issue rather than a limitation of model capacity. While contrastive learning encourages correct keywords to rank above negatives, it does not explicitly calibrate absolute similarity scores, which are critical for applying a fixed threshold across keyword detection. We propose Calibrated, Order-aware Detection KWS (CORD-KWS), a training framework that introduces two complementary objectives while retaining the same encoders, cosine scoring, and O(1) keyword enrollment: a calibrated detection head that supervises absolute scores, and a frame-level CTC objective that preserves phonetic order information. CORD-KWS achieves EERs of 0.43% and 9.64% on LibriPhrase-easy and LibriPhrase-hard, respectively, outperforming existing embedding- and cross-attention-based models.
Explore similar work
Jun 9, 2026cs.SD
User-defined keyword spotting (KWS) enables personalized voice interaction by detecting user-specified keywords. A key challenge in this task is distinguishing target keywords from phonetically confusable alternatives. To address this challenge, we propose KFC-KWS, a multimodal framework that leverages connectionist temporal classification (CTC)-guided keyframe selection. Specifically, we exploit the peaky posterior distributions of CTC to identify high-confidence phoneme frames, enabling precise alignment across audio, phoneme, and text modalities. These keyframes are then fused with full-utterance representations through cross-attention to capture both local discriminative cues and global contextual information. On LibriPhrase, KFC-KWS achieves the best-balanced performance (98.73% AUC) and substantially outperforms advanced baselines on the challenging hard subset (97.65% AUC and 7.75% EER), demonstrating its effectiveness in discriminating between highly confusable keywords.
Jin Li, Wenbin Jiang, Ji Hu
School of Electronics and Information Engineering, Hangzhou Dianzi University, Hangzhou, China · School of Communication Engineering, Hangzhou Dianzi University, Hangzhou, China
Jun 18, 2026eess.AS
User-defined keyword spotting (UD-KWS) enables zero-shot wake-word detection from text, but existing systems learn speaker-invariant representations that cannot reject impostors uttering the correct keyword. We address this dual zero-shot setting -- unseen keywords and unseen speakers -- with ZP-KWS, a lightweight framework combining a phoneme-supervised audio encoder with a GE2E-pretrained compact speaker encoder (about 0.9M parameters). Multiplicative late fusion at inference grants each branch independent veto power, supporting modes from conventional detection to strict speaker-gated activation without retraining. On LibriPhrase, Google Speech Commands, and Qualcomm datasets, ZP-KWS reduces target-only FRR at 1% FAR by up to 60% relative to the strongest baseline while maintaining competitive keyword detection, all within a 1.55M parameter budget for edge deployment.
Ming-Hsiang Hu, Kuan-Tang Huang, Chien-Chun Wang +2
Dept. Computer Science and Information Engineering, National Taiwan Normal University, Taiwan · E.SUN Financial Holding Co., Ltd., Taiwan · United Link Co., Ltd., Taiwan
Aug 11, 2026cs.CL
In this paper, we present PromptKWS, a novel Prompt-guided keyword spotting (KWS) framework to improve the accuracy of open vocabulary KWS systems. In specific terms, we introduce the Prompt Phrases Prediction Network (PPN), an encoder-decoder architecture designed to effectively extract keyword prompts embeddings. we employ the PPN encoder to encode the keyword prompts and infuse the prompt embedding into the Prompt-guided KWS encoder by utilizing a Prompt-acoustic Multi-head Cross-attention (MHCA). Experiments show that PromptKWS improves the wakeup rate by over 10% compared to baseline system. Notably, another strength of PromptKWS is its ability to effectively leverage keyword prompts for adapting to complex real-world environments involving noise and pronunciation variations. In comparison to purely acoustic models, which often struggle in such situations, PromptKWS demonstrates remarkable performance, with an average accuracy improvement of over 15% in test sets.
Gaopeng Xu, Chengfei Li, Xianliang Wang +5