eess.ASFeb 21, 2026

[b] = [d] - [t] + [p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic

Authors: Kwanghee Choi, Eunjung Yeo, Cheol Jun Cho, David Harwath, David R. Mortensen

Organizations: UT Austin · UC Berkeley · CMU

Abstract

Self-supervised speech models (S3Ms) are known to encode rich phonetic information, yet how this information is structured remains underexplored. We conduct a comprehensive study across 96 languages to analyze the underlying structure of S3M representations, with particular attention to phonological vectors. We first show that there exist linear directions within the model's representation space that correspond to phonological features. We further demonstrate that the scale of these phonological vectors correlate to the degree of acoustic realization of their corresponding phonological features in a continuous manner. For example, the difference between [d] and [t] yields a voicing vector: adding this vector to [p] produces [b], while scaling it results in a continuum of voicing. Together, these findings indicate that S3Ms encode speech using phonologically interpretable and compositional vectors, demonstrating phonological vector arithmetic. All code and interactive demos are available at https://github.com/juice500ml/phonetic-arithmetic .

Figures & tables

Appendix figures & tables16 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Phone Segmentation and Recognition through Phonological Activation Mapping

    Jul 10, 2026Shikhar Bharadwaj, Kwanghee Choi, Stephen McIntosh +8Self-Supervised Speech ModelsGrapheme-To-Phoneme

  2. Multilingual Phonological Feature Recognition with Self-Supervised Speech Models

    May 25, 2026Abner Hernandez, Tomás Arias-Vergara, Daiqi Liu +2Self-Supervised Speech ModelsPronunciation Assessment

  3. Do Audio Language Models Hear and Read Distinctive Features Alike?

    Sep 24, 2026Yuanhao Chen, Peter ChinGrapheme-To-PhonemeAudio Understanding