Head-pose variation introduces substantial appearance transformations in visual speech recognition (VSR), making pose-aware feature modulation desirable. However, performance degradation and unwanted feature interactions may result from using numerous Feature-wise Linear Modulation (FiLM) circuits with fixed modulation intensity. We propose a Pose Adaptive Dynamic FiLM framework with a Dynamic Residual FiLM (DR-FiLM) modulator that predicts input-dependent weights to adaptively control the strength of pose-conditioned modulation. Experiments on LRS2 and LRS3 demonstrate that unweighted multi-pathway modulation substantially degrades phoneme recognition, increasing PER to 20.33% and 29.42%, respectively, compared with 16.20% and 20.96% for the single ResFiLM configuration. In contrast, the proposed DR-FiLM with dynamic Deep-Res weighting reduces PER to 15.74% on LRS2 and 23.91% on LRS3, substantially mitigating the adverse effects of unweighted modulation. The analysis of the learned weights further reveals a consistent tendency to assign greater weight to the deeper FiLM pathway as head-pose variation increases. These results show that merging pose-conditioned FiLM circuits is more efficient when the modulation strength is dynamically controlled.
Figures & tables
Figure 1: Overview of the proposed architecture. Head-pose features dynamically determine the contribution of the Deep FiLM pathways at ResNet 18 Layers 3–4 and the ResFiLM pathway for pose-adaptive feature modulation.
Method
Modality
FiLM Layers
LRS2
LRS3
Total Hours
WER ↓
Total Hours
WER ↓
Hyb.-Conf. [ 6 ]
Video
–
223
39.1
438
46.9
Hyb.-Conf. [ 6 ]
Video
–
381
37.9
590
43.3
VTP [ 1 ]
Video
–
2676
22.6
2676
30.7
Auto-AVSR [ 2 ]
Video
–
818
27.9
818
33.0
Auto-AVSR [ 2 ]
Video
–
3448
14.6
3448
19.1
Table 1: Performance comparison (PER % and WER %) of representative VSR methods and our proposed framework on the LRS2 and LRS3 datasets. The total hours indicate the combined duration or size of the pretraining and training datasets.
Method
Dynamic
LRS2
LRS3
Weights
PER ↓
WER ↓
PER ↓
WER ↓
Baseline
✗
20.33
28.84
29.42
38.98
DR-FiLM (Layer-wise)
✓
20.14
29.01
25.88
35.54
DR-FiLM (Deep + Res)
✓
15.74
23.24
23.91
34.43
DR-FiLM (L4 + Res)
✓
16.79
24.98
27.08
37.58
Abs FiLM (R → D)
✓
18.79
27.55
23.67
33.90
Table 2: Ablation study of DR-FiLM configurations and dynamic weighting on LRS2 and LRS3. D denotes the Deep FiLM pathway applied at ResNet 18 Layers 3–4, while R denotes the ResFiLM pathway applied after the 2D CNN frontend. R → D shifts the weighting from ResFiLM toward Deep FiLM as the absolute yaw angle increases, whereas D → R applies the reverse weighting.
Dataset
Yaw Pose
Avg. wdeep
Avg. wres
Std.
Avg. PER
No. of Samples
LRS2
<15∘
0.7948
0.2052
0.0043
14.03
693
15 – 30∘
0.8027
0.1973
0.0068
17.50
341
≥30∘
0.8234
0.1766
0.0162
19.36
209
LRS3
<15∘
0.7586
0.2414
0.0031
23.22
430
15 – 30∘
0.7631
0.2369
0.0050
23.35
488
≥30∘
0.7794
0.2206
0.0168
25.42
403
Table 3: Average DR-FiLM (Deep + Res) weights and PER across different head-pose ranges on LRS2 and LRS3. wdeep and wres denote the weights assigned to the deeper L3–L4 FiLM and ResFiLM branches, respectively. The standard deviation (Std.) is identical for both weights because wdeep+wres=1 .
Visual Speech Recognition (VSR) aims to recognize speech from visual cues such as lip movements, but its performance is fundamentally limited by viseme ambiguity and pose-induced variations that introduce geometric distortions and occlusions. Existing approaches mainly rely on linguistic context or implicit invariance, leaving visual representations insufficiently robust under non-frontal views. In this work, we propose a pose-aware phoneme-level framework, termed HP-VSR-ResFiLM, that explicitly incorporates head-pose information into visual feature extraction. The proposed framework adopts a two-stage pipeline consisting of a pose-conditioned visual encoder in Stage 1 and a pretrained NLLB language model in Stage 2 for phoneme-to-text reconstruction. Specifically, Stage 1 incorporates a pose-conditioned residual Feature-wise Linear Modulation (FiLM) block after the 2D CNN frontend to adaptively refine visual representations using head-pose information. Experiments on LRS2 and LRS3 demonstrate that HP-VSR-ResFiLM achieves competitive performance under comparable training conditions, attaining word error rates (WER) of 25.0% and 33.2%, respectively, without relying on additional training data. Ablation studies further show that a single residual FiLM block consistently improves overall WER, while deeper modulation at Layers 3 and 4 provides larger gains for samples with yaw angles greater than 30° without degrading performance for smaller pose variations. These findings demonstrate that explicit pose-aware feature modulation offers an effective and computationally efficient solution for improving VSR robustness in unconstrained settings.
Matthew Kit Khinn Teng, Haibo Zhang, Takeshi Saitoh
Department of Artificial Intelligence, Kyushu Institute of Technology, Iizuka, 820-8502, Fukuoka, Japan.
Automatic speech recognition (ASR) has advanced remarkably for standard speech; however, pathological speech from neurological conditions remains a significant challenge. We investigate speaker conditioning via Feature-wise Linear Modulation (FiLM), injecting x-vector-derived information into each transformer layer of a frozen ASR encoder to adapt internal representations to individual pathological speakers without modifying base model weights. We benchmark this for the ASR task against standard and parameter-efficient fine-tuning baselines, complemented by post-processing, on Spanish and English pathological speech. Additionally, we evaluate if the adapted model preserves the ability to answer speech-related questions. Results show that speaker-conditioned ASR is competitive with established adaptation strategies while retaining performance on non-conditioned speech.
Fernando López, Santosh Kesiraju, Jordi Luque
Scientific Research, Telefónica Innovación Digital, Spain · AUDIAS, Universidad Autónoma de Madrid, Spain
Existing Visual Speech Recognition (VSR) systems commonly rely on left-to-right autoregressive decoding, which can force premature decisions on visually ambiguous tokens before sufficient context is available. We propose DLLM-VSR, to the best of our knowledge, the first Diffusion Large Language Model (DLLM)-based VSR framework, formulating transcription as iterative masked denoising with flexible-order decoding. With confidence-based unmasking, DLLM-VSR commits high-confidence positions early and uses the committed tokens as bidirectional context to refine ambiguous ones. To adapt DLLMs to VSR, we introduce a two-stage masked-denoising training strategy that separates visual-to-text content alignment from length modeling. We further observe a performance gap compared with an upper-bound setting where the ground-truth transcript length is provided at inference, allowing the model to focus on transcript content decoding. To reduce this gap, we develop length-guided candidate decoding, which uses video duration to construct plausible transcript-length hypotheses and reranks the decoded candidates using length plausibility and decoding confidence. The proposed method achieves a 19.4% word error rate on LRS3, establishing state-of-the-art performance among methods using only LRS3 as labeled training data.
Jeong Hun Yeo, Chae Won Kim, Hyeongseop Rha +1
Integrated Vision Language Lab, KAIST, South Korea