eess.ASSep 24, 2026

Anatomy-aware cross-speaker adaptation of complete vocal-tract acoustic-to-articulatory inversion

Authors: Nhat-Nam Nguyen, Pierre-Andre Vuissoz, Yves Laprie

Organizations: Université de Lorraine, CNRS, Inria, F-54000 Nancy, France · Université de Lorraine, Inserm, IADI U1254, F-54000 Nancy, France

Abstract

Cross-speaker acoustic-to-articulatory inversion requires accounting for anatomical differences between speakers. We propose a geometric adaptation framework that uses anatomical landmarks, primarily on vertebrae and dental structures,to transfer predictions from a fixed inversion model to unseen speakers. An affine transformation followed by thin-plate spline (TPS) deformation maps the predicted contours of 10 vocal-tract structures into each target speaker's geometry without retraining. Landmarks are identified in one selected /u/ frame per speaker as a common phonetic reference without assuming identical articulatory configurations across speakers, and the resulting mapping is reused across recordings. We train the model on a single-speaker rt-MRI database and evaluate adaptation on eight speakers from a separate multi-speaker rt-MRI database. We compare affine and TPS configurations using 12 or 14 landmarks. Affine12+TPS14 achieves the lowest mean point-to-closest-point error of 3.19mm. These results support the combined value of anatomical landmark information and nonrigid alignment.

Figures & tables

Explore similar work

CardsList
  1. Arti-JEPA: Adapting Video World Model to Real-Time MRI of the Vocal Tract for Speech-Production Analysis

    Sep 9, 2026Hong Nguyen, Sean Foley, Christina Hagedorn +4Vocal Tract ShapeSpeaker

  2. ArtBoost: Synthetic Articulatory Data Augmentation for Acoustic-to-Articulatory Inversion

    Jun 15, 2026Hyung Kyu Kim, Byungchan Hwang, Hak Gu KimArticulatoryAugmentation