Cross-speaker acoustic-to-articulatory inversion requires accounting for anatomical differences between speakers. We propose a geometric adaptation framework that uses anatomical landmarks, primarily on vertebrae and dental structures,to transfer predictions from a fixed inversion model to unseen speakers. An affine transformation followed by thin-plate spline (TPS) deformation maps the predicted contours of 10 vocal-tract structures into each target speaker's geometry without retraining. Landmarks are identified in one selected /u/ frame per speaker as a common phonetic reference without assuming identical articulatory configurations across speakers, and the resulting mapping is reused across recordings. We train the model on a single-speaker rt-MRI database and evaluate adaptation on eight speakers from a separate multi-speaker rt-MRI database. We compare affine and TPS configurations using 12 or 14 landmarks. Affine12+TPS14 achieves the lowest mean point-to-closest-point error of 3.19mm. These results support the combined value of anatomical landmark information and nonrigid alignment.
Figures & tables
Figure 1: Anatomical landmark configurations used for geometric adaptation. The 12-landmark set (in blue) is shown together with the two additional lower-incisor landmarks, M1 and L6 (in orange), used in the 14-landmark configuration.
Articulator
Mean ± SD
Median
Arytenoid cartilage
1.53 ± 0.76
1.38
Epiglottis
1.71 ± 1.18
1.39
Lower lip
1.29 ± 0.64
1.17
Pharyngeal wall
0.96 ± 0.44
0.85
Soft-palate midline
1.01 ± 0.55
0.88
Tongue
1.99 ± 0.74
1.84
Table 1: ASD2 model trained on ten articulators with global normalization. P2CP mean in mm.
Articulator
A12
A14
A12+T12
A12+T14
Arytenoid
5.09 ± 2.33
4.34 ± 2.16
4.52 ± 2.27
3.64 * ± 1.89
Epiglottis
4.94 ± 2.74
4.29 ± 2.40
5.00 ± 2.76
4.39 ± 2.43
Lower lip
3.26 ± 1.50
2.67 ± 1.06
2.91 ± 1.20
2.53 ± 1.11
Pharynx
2.87 ± 1.79
3.02 ± 1.91
2.29 ± 1.41
2.19 ± 1.27
Soft palate
2.81 ± 1.26
2.83 ± 1.42
2.91 ± 1.19
2.88 ± 1.15
Tongue
4.58 ± 1.61
4.06 ± 1.36
4.41 ± 1.43
3.62 ± 1.17
Table 2: Articulator-wise P2CP mean (mm). A14 is the main affine baseline; A12+T14 is the proposed configuration.
Speaker
Raw
A12
A14
A12+T12
A12+T14
P1
9.71 ± 3.63
5.85 ± 3.41
5.33 ± 2.94
5.34 ± 3.51
3.95 * ± 2.31
P3
9.69 ± 4.26
5.06 ± 3.38
4.42 ± 2.49
4.31 ± 2.89
3.46 * ± 1.76
P4
8.14 ± 4.26
3.28 ± 1.64
2.95 ± 1.36
3.06 ± 1.93
3.17 * ± 2.10
P5
4.42 ± 2.12
3.13 ± 1.36
2.93 ± 1.21
2.64 ± 1.03
2.46 * ± 1.00
P6
7.79 ± 3.37
3.41 ± 1.97
2.94 ± 1.58
3.32 ± 1.69
2.82 * ± 1.36
P7
11.23 ± 4.55
3.41 ± 2.27
3.34 ± 2.26
3.69 ± 2.15
3.17 * ± 2.14
Table 3: Speaker-wise Full10 P2CP mean (mm). A14 is the main affine baseline; A12+T14 is the proposed configuration.
Figure 2: Adaptation example for P8 containing /k/ in avec . Orange solid: predictions; cyan dashed: ground truth. Frame-level P2CP errors are shown below each panel.
Real-time MRI (rtMRI) captures the dynamics of the entire vocal tract during speech, but labeled data are scarce and the modality - single-slice, grayscale, low-resolution - differs substantially from the natural videos that video foundation models are trained on. We introduce Arti-JEPA, a joint embedding predictive architecture to model vocal tract rtMRI by continuing its self-supervised objective on about 62h of unlabelled vocal-tract videos, and evaluate the frozen representation on three tasks: cross-domain phoneme prediction (on typical speakers), fluent-vs-disfluent classification (a corpus containing stuttered speech), and characterizing pre/post-operative transfer (after partial glossectomy). Three key findings emerge. (1) A temporal video prior decisively outperforms per-frame image encoders, and latent prediction (V-JEPA) is at least as strong as pixel reconstruction (VideoMAE), with the edge on fine-grained phonemes. (2) Domain adaptation is \emph{task-dependent}: it roughly doubles cross-domain phoneme prediction κ (to 0.352) but does not help binary stuttering classification. (3) Arti-JEPA was able to recover phoneme signal from pre/post glossectomy speech --- an in-domain probe decodes patients at least as well as a typical speaker, indicating that the residual transfer gap is cross-speaker/domain misalignment, not surgical signal loss, and post-operative decoding does not fall below performance on pre-operative speech. Together, these position a frozen, domain-adapted rtMRI encoder as a reusable measurement tool for articulatory and clinical speech science.
Hong Nguyen, Sean Foley, Christina Hagedorn +4
Ming Hsieh Department of Electrical and Computer Engineering, University of Southern California, Los Angeles, CA, USA · Department of English, College of Staten Island, City University of New York, New York, NY, USA · Department of Linguistics, University of Potsdam, Germany +1
Recent acoustic-to-articulatory inversion (AAI) models rely on electromagnetic articulography (EMA) data, which are costly and limited in scale. To address this limitation, we propose \textit{ArtBoost}, a novel data augmentation strategy that leverages large-scale speech--mesh datasets originally developed for speech-driven 3D facial animation to improve AAI under limited EMA supervision. \textit{ArtBoost} extracts pseudo articulatory trajectories from visible facial anchors and uses them for pre-training before fine-tuning on real EMA data. Experiments show consistent improvements in PCC and RMSE. Trajectory analyses confirm that the pseudo articulatory signals reflect physically meaningful visible articulatory dynamics. Additional evaluations across different AAI architectures demonstrate stable performance gains, indicating that \textit{ArtBoost} can be integrated into diverse AAI models. These results suggest that speech--mesh data provide an effective and scalable source of articulatory supervision for AAI. Project page: https://cau-irislab.github.io/Interspeech26-ArtBoost/
Hyung Kyu Kim, Byungchan Hwang, Hak Gu Kim
Department of Imaging Science and Arts, Chung-Ang University, South Korea · Department of Metaverse Convergence, Chung-Ang University, South Korea
In cross-lingual zero-shot text-to-speech, the accent of the reference leaks into the target speech. We propose accent analogy guidance (AAG), a training-free sampler term that subtracts an accent direction estimated from the model's own predictions for one synthetic voice rendered in both languages, so the voice cancels and only the accent remains. By a blind LLM accent judge on real dubbing data, reweighting classifier-free guidance between reference and text, and its variants, stay near one identity-accent trade-off curve; we score a method by its speaker similarity above that curve at equal accent (ΔSIM). Across four open TTS models AAG lies above the curve: on OmniVoice ΔSIM is +0.11 to +0.27 on three test sets (accent 3.51 to 4.28 on a 1-5 scale at speaker similarity 0.29, where reweighting keeps 0.02); MaskGCT and CosyVoice 2 also lie above their curves, and on F5-TTS it is more native than any reweighting setting. An LLM-free language-ID measure and a twelve-listener panel agree. A premise test and the reach of a model's own curve indicate in advance whether and roughly how much AAG can gain, predicting the one model where it gains nothing (X-Voice).