cs.SDOct 19, 2021
SaveA mathematical model of the vowel space
Organizations: GIPSA-PCMD
Abstract
The articulatory-acoustic relationship is many-to-one and non linear and this is a great limitation for studying speech production. A simplification is proposed to set a bijection between the vowel space (f1, f2) and the parametric space of different vocal tract models. The generic area function model is based on mixtures of cosines allowing the generation of main vowels with two formulas. Then the mixture function is transformed into a coordination function able to deal with articulatory parameters. This is shown that the coordination function acts similarly with the Fant's model and with the 4-Tube DRM derived from the generic model.
Figures & tables
Figure 1: The vowel space of the generic vocal tract model and its 8 characteristic vowels (Mm. Multimedia and supplementary material ). The area functions (Mm. Multimedia and supplementary material ) are plotted for each vowel with, at the top, the pairs of cosine Fourier coefficients which are introduced in the vowel equation (Eq. 4a b). For [o,e], we have .
Figure 2: Plot of the functions expressed with the values of derived in Table 1 . This shows how are set at their respective phases and it focuses on for . The expression of is consistent with this given by Eq. 3b because and .
Figure 3: Simulation in condition together with the vowel space with obtained with the coordination function (Mm. Multimedia and supplementary material and Mm. Multimedia and supplementary material ). All (non figured) points of are inside of this. The spaces derived from the Schroeder-Ehrenfest relation and Fourier coefficients before and after rectification ("SE" from and ) and the biased relation ("est" from ) are plotted relative to the neutral reference . The bias transforms the circular shape "SE" from into the triangular shape "est".
Figure 4: (a) The design of the Fant model redrawn from ( Badin et al., 1990 , Fig. 1) and the definition of geometrical parameters used in Table 3 (b) Simulation in condition together with the vowel space with obtained with the coordination function (Mm. Multimedia and supplementary material and Mm. Multimedia and supplementary material ). All (non figured) points of are inside of this.
Figure 5: Simulations of in condition and : (a) DRM (b) Fant model. Abscissa: Fourier cosine coefficients of the area function. Ordinate: Deviations from neutral in relative values for the two first formants. The circled gray surfaces represent the configurations produced by the coordination function in and the dots represent which is a simulation of the 4 parameters of each model but uncorrelated in the same domain of variation .
Figure 6: Plot of the estimates of the deviations from neutral in relative values for DRM . For the first formant, this is close to the bisecting line thanks to the bias.
Figure 7: Overlay of DRM data with ( Boë et al., 2019 , Fig. 4 with ) . This shows that a 4-tube model driven by 7 parameters has a significantly larger space than one with only 2 parameters. The human vowel space is depicted by background ellipses and only the [i] area is not well covered by this model.
Explore similar work
This study investigates vowel-inherent spectral change (VISC) in spontaneous conversational Mandarin. Using the generalized additive model and word embeddings from distributional semantics, we show that, when controlling for variables such as vowel duration, gender, speaker identity, co-articulation, vowel identity, and utterance position, vowel formant trajectory dynamics have word-specific components that are tied to their meaning in context: The F1 and F2 trajectories of words can be predicted from their contextualized embeddings with an accuracy that substantially exceeds a permutation baseline. Challenging modular cognitive models of speech production, these results indicate that, words' semantics co-determine the fine details of their articulation.
Articulatory strategy as a source of variation in acoustic vowel dynamics
Acoustic vowel dynamics have some speaker-identifying characteristics, which have been ascribed to individual properties of articulatory strategies: formant transitions have a particular shape because speakers move their articulators, using specific and practised movements. However, there is little existing evidence that different articulatory strategies systematically affect formant dynamics. The present study corroborates the link between the two. Ultrasound tongue imaging data from 36 speakers of Northern-Anglo English are used to identify distinct articulatory strategies for the production of palatal vowel /i/. Tongue shape in /i/ is found to be a significant predictor of formant dynamics in diphthongs with a palatal offglide. The observed relationships can be explained by the characteristics of articulatory movement conditioned by vocal tract shape. Greater articulatory displacement of tongue root and/or dorsum produces greater distortion from the mean tongue shape in palatal vowels, and it also requires higher articulatory velocities, resulting in relatively earlier and steeper formant transitions. The results contribute to the conceptual understanding of individuality in speech, by illuminating the regularising and individual aspects of articulatory compensation.
Evaluating Speech Articulation Synthesis with Articulatory Phoneme Recognition
Recent advances in machine learning and the availability of articulatory datasets allow vocal tract synthesis to be conditioned on phonetic sequences, a primary task of articulatory speech synthesis. However, quality assessment needs a better definition. Generally, ranking generative models is tricky due to subjectivity. However, articulatory synthesis has the additional difficulty of requiring specialized knowledge in vocal tract anatomy and acoustics. To address this problem, this paper proposes to evaluate speech articulation synthesis using phoneme recognition as a proxy. Our hypothesis is that phoneme recognition using articulatory features better captures nuances in phoneme production, such as correct places of articulation, which traditional metrics (e.g., point-wise distance metrics) do not. We train a neural network with acoustic and articulatory features extracted from a single-speaker RT-MRI dataset. Then, we compare the recognition performance when testing the model with different synthetic articulatory features. Our results show that our articulatory feature set is phonetically rich and helps exploring additional dimensions on speech articulation synthesis.