cs.SDSep 29, 2026

Reconstructing the Vocal Tract with Differentiable Acoustic Simulation

Authors: Eric Ming Chen, Jin Woo Lee, Vincent Sitzmann

Organizations: MIT CSAIL

Abstract

The vocal tract is the region of the human body responsible for filtering one's voice to create speech. In this paper, we present a differentiable and GPU accelerated acoustic simulator for the vocal tract. The differentiable simulator synthesizes speech by propagating sound along an acoustic tube model of the vocal tract, and via its gradients, can solve the inverse problem: reconstructing the shape of the vocal tract solely from the sound it produces. Although the inverse mapping between geometry and sound is notoriously non-convex, we discover that gradient descent succeeds with three technical contributions: (1) we design a frequency domain formulation of the vocal tract's fluid dynamics that is 70x more GPU parallelizable than finite differences in time, (2) we integrate a differentiable model for turbulence to synthesize consonants, and (3) similar to prior work in implicit neural representations (INRs) and neural fields, we find that parameterizing the geometry with a neural network accelerates convergence and escapes local minima that trap discrete representations. Because the simulator is differentiable, it is readily integrated with other deep learning pipelines to enable novel linguistics and medical imaging applications. (1) We demonstrate self-supervised autoencoding of vocal tract shapes across 11 languages, and (2) we couple our simulator with a generative model of MRI (magnetic resonance imaging) images to reconstruct one's moving vocal tract from only their speech without paired data.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Speech-Guided Multimodal Learning for Vocal Tract Segmentation in Real-Time MRI

    May 18, 2026Daiqi Liu, Lukas Mulzer, Md Hasan +11Vocal Tract ShapeMagnetic Resonance Imaging

  2. SIREM: Speech-Informed MRI Reconstruction with Learned Sampling

    May 18, 2026Md Hasan, Nyvenn Castro, Daiqi Liu +6Magnetic Resonance Imaging ReconstructionSpeaker

  3. Anatomy-aware cross-speaker adaptation of complete vocal-tract acoustic-to-articulatory inversion

    Sep 24, 2026Nhat-Nam Nguyen, Pierre-Andre Vuissoz, Yves LaprieArticulatorySpeaker