cs.SDNov 1, 2025

Physics-Informed Neural Networks for Speech Production

Authors: Kazuya YokotaRyosuke HarakawaMasaaki BabaMasahiro Iwahashi

Organizations: Department of Mechanical Engineering, Nagaoka University of Technology, 1603-1, Kamitomioka, Nagaoka, Niigata, Japan · Department of Electrical, Electronics and Information Engineering, Nagaoka University of Technology

Abstract

The analysis of speech production based on physical models of the vocal folds and vocal tract is essential for studies on vocal-fold behavior and linguistic research. This paper proposes a speech production analysis method using physics-informed neural networks (PINNs). The networks are trained directly on the governing equations of vocal-fold vibration and vocal-tract acoustics. Vocal-fold collisions introduce nondifferentiability and vanishing gradients, challenging phenomena for PINNs. We demonstrate, however, that introducing a differentiable approximation function enables the analysis of vocal-fold vibrations within the PINN framework. The period of self-excited vocal-fold vibration is generally unknown. We show that by treating the period as a learnable network parameter, a periodic solution can be obtained. Furthermore, by implementing the coupling between glottal flow and vocal-tract acoustics as a hard constraint, glottis-tract interaction is achieved without additional loss terms. We confirmed the method's validity through forward and inverse analyses, demonstrating that the glottal flow rate, vocal-fold vibratory state, and subglottal pressure can be simultaneously estimated from speech signals. Notably, the same network architecture can be applied to both forward and inverse analyses, highlighting the versatility of this approach. The proposed method inherits the advantages of PINNs, including mesh-free computation and the natural incorporation of nonlinearities, and thus holds promise for a wide range of applications.

Explore similar work

CardsList
  1. Physics-Informed Neural Operator for Speech Production Analysis

    Jun 21, 2026Kazuya Yokota, Xinmeng Luan, Debasish Ray Mohapatra +2VocalizationsConsonants

  2. An Optimal Contact-Mechanically Consistent and Flow-Separation Adapted Modeling of Vocal Fold Dynamics

    Jun 27, 2026Sardar Nafis Bin Ali, Maryam Naghibolhosseini, Mohsen ZayernouriVocalizations