cs.CVSep 24, 2026

Pose Adaptive Dynamic FiLM Modulation for Visual Speech Recognition

Authors: Matthew Kit Khinn Teng, Haibo Zhang, Takeshi Saitoh

Organizations: Kyushu Institute of Technology, Japan

Abstract

Head-pose variation introduces substantial appearance transformations in visual speech recognition (VSR), making pose-aware feature modulation desirable. However, performance degradation and unwanted feature interactions may result from using numerous Feature-wise Linear Modulation (FiLM) circuits with fixed modulation intensity. We propose a Pose Adaptive Dynamic FiLM framework with a Dynamic Residual FiLM (DR-FiLM) modulator that predicts input-dependent weights to adaptively control the strength of pose-conditioned modulation. Experiments on LRS2 and LRS3 demonstrate that unweighted multi-pathway modulation substantially degrades phoneme recognition, increasing PER to 20.33% and 29.42%, respectively, compared with 16.20% and 20.96% for the single ResFiLM configuration. In contrast, the proposed DR-FiLM with dynamic Deep-Res weighting reduces PER to 15.74% on LRS2 and 23.91% on LRS3, substantially mitigating the adverse effects of unweighted modulation. The analysis of the learned weights further reveals a consistent tendency to assign greater weight to the deeper FiLM pathway as head-pose variation increases. These results show that merging pose-conditioned FiLM circuits is more efficient when the modulation strength is dynamically controlled.

Figures & tables

Explore similar work

CardsList
  1. Head-Pose-Aware Visual Speech Recognition with FiLM Modulation

    May 30, 2026Matthew Kit Khinn Teng, Haibo Zhang, Takeshi SaitohGrapheme-To-Phoneme

  2. FiLM-Based Speaker Conditioning of a SpeechLLM for Pathological Speech Recognition

    Jun 4, 2026Fernando López, Santosh Kesiraju, Jordi LuqueSpeech EncoderDysarthric Speech

  3. Diffusion Large Language Models for Visual Speech Recognition

    May 27, 2026Jeong Hun Yeo, Chae Won Kim, Hyeongseop Rha +1Diffusion Language ModelsConstrained Decoding