Self-supervised speech models (S3Ms) are known to encode rich phonetic information, yet how this information is structured remains underexplored. We conduct a comprehensive study across 96 languages to analyze the underlying structure of S3M representations, with particular attention to phonological vectors. We first show that there exist linear directions within the model's representation space that correspond to phonological features. We further demonstrate that the scale of these phonological vectors correlate to the degree of acoustic realization of their corresponding phonological features in a continuous manner. For example, the difference between [d] and [t] yields a voicing vector: adding this vector to [p] produces [b], while scaling it results in a continuum of voicing. Together, these findings indicate that S3Ms encode speech using phonologically interpretable and compositional vectors, demonstrating phonological vector arithmetic. All code and interactive demos are available at https://github.com/juice500ml/phonetic-arithmetic .
Figures & tables
Figure 1: Comparing analogies for text and speech. Word representations enable semantic analogies 3 3 3 Word analogies borrowed from Ethayarajh et al. (2019) . , while speech representations enable phonological analogies. Such analogies ( Section 3 ) can be used to control speech synthesis in a phonologically grounded manner ( Section 4 ).
Figure 2: Comparing S3Ms with spectral representations on TIMIT (top) and VoxAngeles (bottom).
Figure 3: Comparing consonant-only and vowel-only quadruplets on TIMIT (top) and VoxAngeles (bottom) for WavLM. Number within the parenthesis denotes the number of quadruplets. We exclude cases where a quadruplet contains both consonants and vowels.
Phonological feature
High
Low
Back
Round
Nasal
Sonorant
Strident
Voice
Acoustic measurement
F1
F1
F2
F2
F1BW
HNR
COG
COG
Expected correlation sign
–
+
–
–
–
+
+
–
Table 1: Summary of phonological features, their associated acoustic measurements, and expected correlation signs. The phonologically expected correlation sign is denoted as + (positive) or – (negative). Five types of acoustic measurements are used: first formant (F1), second formant (F2), first-formant bandwidth (F1BW), harmonic-to-noise ratio (HNR), and center of gravity (COG). Further details are provided in § A.3 .
Figure 4: Comparison between the phonological vector scale λ and the acoustic measurements ( § A.3 ) on TIMIT. ρ denotes Spearman’s rank correlation coefficient. Blue and orange plots indicate vowels and consonants, respectively. The empirically observed correlation signs match the theoretical expectations shown in Table 1 . Further, plots demonstrate the linearity of phonological vectors, resulting in monotonic (but not necessarily linear) changes in acoustic measurements.
Figure 5: Applying round vector to front vowel [i], where there is no front rounded vowel in English. Orange and blue arrows denote F2 and F3, respectively, which are all decreasing for λ>0 .
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9: Comparing pairing consistency score (PCS) of S3Ms and spectral representations using TIMIT (upper) and VoxAngeles (lower).
Figure 10: Comparing success rates for S3Ms with spectral representations using TIMIT. We denote feature and audio slicing as feat and audio, respectively.
Figure 11: Comparing success rates for feature sliced S3Ms with spectral representations using VoxAngeles. We denote feature and audio slicing as feat and audio, respectively.
Figure 12: Comparing averaged similarities C for feature sliced S3Ms with audio sliced MFCC on TIMIT (upper) and VoxAngeles (lower). We calculate the 99% CI considering quadruplet-wise cosine averages.
Figure 13: Comparing pre-trained S3M (XLSR-53) and fine-tuned phone recognition models (Wav2vec2Phoneme and MultIPA) on TIMIT.
Figure 14: Phonological feature-wise success rates for consonants on TIMIT using WavLM representations.
Figure 16: Phonological feature-wise success rates for consonants on VoxAngeles using WavLM representations.
Figure 18: PanPhon feature distance-wise success rates on TIMIT using WavLM representations.
Figure 20: Approximating various phonological vectors on TIMIT (upper) and VoxAngeles (lower) with different sample size N={1,4,16,64,256} . Each histogram depicts the distribution of cosine similarities.
Figure 21: Approximating various phonological vectors on TIMIT (upper) and VoxAngeles (lower) using a fixed phone pair. Each histogram depicts the distribution of cosine similarities. Figure 20 is overlaid for comparison.
Figure 22: Cosine similarities between different phonological vectors drawn from TIMIT.
Figure 24: Acoustic measurements of F1, F2, and F3 for rounded (blue) and unrounded (orange) vowels on TIMIT (upper) and VoxAngeles (lower). ρ indicates Spearman’s rank correlation coefficient between the roundness and the acoustic measurements.
Figure 25: Comparing the phonological vector weight λ with acoustic measurements on TIMIT (upper) VoxAngeles (lower) using WavLM. We observe three measurements, F1, F2, and F3, for the round vector. ρ indicates Spearman’s rank correlation coefficient.
Figure 26: Comparing the phonological vector weight λ with acoustic measurements on VoxAngeles using WavLM. ρ indicates Spearman’s rank correlation coefficient. Blue and orange plots indicate vowels and consonants, respectively.
Figure 27: Comparing the acoustic measurements of original and synthesized speech ( λ=0 ) on TIMIT (upper two rows) and VoxAngeles (lower two rows). We use the same range of x-axis from y-axis in Figures 4 and 26 . We observe that the differences are highly centralized to zero, ensuring the stability of resynthesis through the vocoder.
Figure 28: Comparing the phonological vector weight λ with acoustic measurements on TIMIT (upper two rows) and VoxAngeles (lower two rows) using phonological vectors from audio-sliced MFCC. ρ indicates Spearman’s rank correlation coefficient. Blue and orange plots indicate vowels and consonants, respectively. There is little to no controllability, with the exception of weak correlation on back vowel on VoxAngeles.
Phone segmentation and recognition are inherently related tasks, yet modern approaches typically model them separately. We argue that phonetic structure is already latent in the representations of self-supervised speech models (S3Ms), and one only needs to steer them to solve both tasks. We leverage S3M-based Phonological Activation Mapping (SPAM), which maps each S3M representation frame to a vector of phonological feature activations, such as voicing and nasality. On top of SPAM, we introduce two simple but effective lightweight, gradient-descent-free prediction heads: a recognition head and a segmentation head. Our method requires less than a minute of phonetic transcriptions, and generalizes to unseen phones during training. Across a diverse range of datasets, our approach attains strong segmentation and recognition performance.
Shikhar Bharadwaj, Kwanghee Choi, Stephen McIntosh +8
Phonological features provide a language-general and linguistically grounded representation of speech. We present PhonoQ-2.0, a multilingual frame-level phonological feature recognizer built on self-supervised speech models. The system directly predicts a structured 22-dimensional feature vector per frame encoding manner, vowel quality, place, and voicing, instead of deriving features from phoneme outputs. To ensure phonologically coherent predictions, we introduce a manner-conditioned gating mechanism that activates valid feature groups. Evaluated across multiple languages and corpora, PhonoQ-2.0 achieves an average macro-F1 of 91.3% in-domain and 88.9% out-of-domain. Compared to a strong CTC phoneme baseline, it delivers consistent gains of +8.8 F1 in-domain and +8.6 out-of-domain on average. In unseen-language evaluation, PhonoQ-2.0 improves macro-F1 from 66.9% to 73.6% (+6.7 on average), with gains of up to +10.8 points.
Abner Hernandez, Tomás Arias-Vergara, Daiqi Liu +2
Pattern Recognition Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg, Germany · GITA Lab. Facultad de Ingenier´ıa. Universidad de Antioquia UdeA, Medell´ın, Colombia
Audio language models pass speech and text through a single decoder. We ask whether that decoder represents a distinctive feature in the same direction when a phoneme is heard and when it is read. For minimal pairs of phonemes differing in one feature, we take the offset between the two members' mean representations. Averaging those offsets gives a direction for each stream, and we measure the cosine between the two. Because the two streams already agree about arbitrary phoneme pairs, we compare every measure against a reference built from random pairings rather than against zero. We apply this to 6 models, 7 features and 15 languages from 11 families. Only voicing in the two Qwen2.5-Omni models exceeds that reference after correction for multiple testing, and the reference varies by a factor of seven between models. In three of the six models, voicing has one direction in audio across the 14 languages with enough minimal pairs to measure it, and every language pair agrees in two of them. The model family, not the model size, predicts which stream represents a feature.
Yuanhao Chen, Peter Chin
Thayer School of Engineering, Dartmouth College, Hanover, NH, USA