eess.ASOct 5, 2026

GIVE-KWS: Gated Injection of Visual Evidence for Noise-Robust Query-by-Example Keyword Spotting

Authors: Ming-Hsiang Hu, Kuan-Tang Huang, Hung-Shin Lee, Berlin Chen

Abstract

Visual speech promises noise-robust keyword spotting, yet a visual stream is not necessarily used. On a tri-modal query-by-example keyword spotting (QbyE-KWS) benchmark, we find that a system with a task-trained visual encoder comes within 2 percentage points of a text-and-audio system in equal error rate (EER) at -10 dB, and link this gap to the encoder's lack of phonemic information. We present GIVE-KWS, whose fusion stage, GIVE (Gated Injection of Visual Evidence), conditions query audio on lip motion through gated cross-attention. We show that visual robustness depends on two interacting conditions: a phoneme-bearing visual representation, and fusion that injects visual evidence rather than rescaling audio features. Under a phoneme-bearing encoder, injection yields an effective SNR gain of 4.0-9.3 dB over masking at -10 dB, whereas under a phoneme-poor one it nearly vanishes. Relative to the benchmark system, GIVE-KWS reduces unseen-keyword EER by 72.9% at -10 dB and 62.8% on average.

Explore similar work

CardsList
  1. Personalized Keyword Spotting for User-Defined Keywords Leveraging Text-Independent Speaker Verification

    Jun 18, 2026Ming-Hsiang Hu, Kuan-Tang Huang, Chien-Chun Wang +2Zero-Shot LearningKeyword Spotting

  2. CORD-KWS: Calibrated, Order-Aware Detection for Open-Vocabulary Keyword Spotting

    Sep 25, 2026Ramesh Gundluru, Adarsh Arigala, Sri Rama Murty Kodukula

  3. Massive Open-Vocabulary Keyword Spotting

    Jun 9, 2026Leonor Barreiros, Raul Monteiro, Afonso Mendes +1Keyword SpottingAutomatic Speech Recognition