cs.SDSep 25, 2026

CORD-KWS: Calibrated, Order-Aware Detection for Open-Vocabulary Keyword Spotting

Authors: Ramesh Gundluru, Adarsh Arigala, Sri Rama Murty Kodukula

Abstract

Open-vocabulary keyword spotting (KWS) must detect arbitrary keywords without retraining. Cross-attention-based models achieve state-of-the-art performance but require pairwise interaction between the audio and each keyword, recomputing the fused representation for every audio--keyword pair. Embedding-based models avoid this computational cost through independent encoding and similarity scoring, yet remain inferior on challenging benchmarks such as LibriPhrase-hard. We show that this gap is primarily a training issue rather than a limitation of model capacity. While contrastive learning encourages correct keywords to rank above negatives, it does not explicitly calibrate absolute similarity scores, which are critical for applying a fixed threshold across keyword detection. We propose Calibrated, Order-aware Detection KWS (CORD-KWS), a training framework that introduces two complementary objectives while retaining the same encoders, cosine scoring, and O(1) keyword enrollment: a calibrated detection head that supervises absolute scores, and a frame-level CTC objective that preserves phonetic order information. CORD-KWS achieves EERs of 0.43% and 9.64% on LibriPhrase-easy and LibriPhrase-hard, respectively, outperforming existing embedding- and cross-attention-based models.

Explore similar work

Jun 9, 2026cs.SD

KFC-KWS: Keyframe Fusion with CTC for User-Defined Keyword Spotting

User-defined keyword spotting (KWS) enables personalized voice interaction by detecting user-specified keywords. A key challenge in this task is distinguishing target keywords from phonetically confusable alternatives. To address this challenge, we propose KFC-KWS, a multimodal framework that leverages connectionist temporal classification (CTC)-guided keyframe selection. Specifically, we exploit the peaky posterior distributions of CTC to identify high-confidence phoneme frames, enabling precise alignment across audio, phoneme, and text modalities. These keyframes are then fused with full-utterance representations through cross-attention to capture both local discriminative cues and global contextual information. On LibriPhrase, KFC-KWS achieves the best-balanced performance (98.73% AUC) and substantially outperforms advanced baselines on the challenging hard subset (97.65% AUC and 7.75% EER), demonstrating its effectiveness in discriminating between highly confusable keywords.
Jun 18, 2026eess.AS

Personalized Keyword Spotting for User-Defined Keywords Leveraging Text-Independent Speaker Verification

User-defined keyword spotting (UD-KWS) enables zero-shot wake-word detection from text, but existing systems learn speaker-invariant representations that cannot reject impostors uttering the correct keyword. We address this dual zero-shot setting -- unseen keywords and unseen speakers -- with ZP-KWS, a lightweight framework combining a phoneme-supervised audio encoder with a GE2E-pretrained compact speaker encoder (about 0.9M parameters). Multiplicative late fusion at inference grants each branch independent veto power, supporting modes from conventional detection to strict speaker-gated activation without retraining. On LibriPhrase, Google Speech Commands, and Qualcomm datasets, ZP-KWS reduces target-only FRR at 1% FAR by up to 60% relative to the strongest baseline while maintaining competitive keyword detection, all within a 1.55M parameter budget for edge deployment.
Aug 11, 2026cs.CL

PromptKWS: A Novel Prompt-Guided Open-Vocabulary Keyword Spotting Framework

In this paper, we present PromptKWS, a novel Prompt-guided keyword spotting (KWS) framework to improve the accuracy of open vocabulary KWS systems. In specific terms, we introduce the Prompt Phrases Prediction Network (PPN), an encoder-decoder architecture designed to effectively extract keyword prompts embeddings. we employ the PPN encoder to encode the keyword prompts and infuse the prompt embedding into the Prompt-guided KWS encoder by utilizing a Prompt-acoustic Multi-head Cross-attention (MHCA). Experiments show that PromptKWS improves the wakeup rate by over 10% compared to baseline system. Notably, another strength of PromptKWS is its ability to effectively leverage keyword prompts for adapting to complex real-world environments involving noise and pronunciation variations. In comparison to purely acoustic models, which often struggle in such situations, PromptKWS demonstrates remarkable performance, with an average accuracy improvement of over 15% in test sets.