cs.LGAug 19, 2026

Cluster Assignments in Soft Targets Shape Speech Representations: Evidence from S-JEPA

Authors: Wenxuan He, Yunpeng Li, Zewei Li, Yongke Yang, Yuze Li, Yin Cao, Shan Liang

Organizations: Department of Intelligent Science, School of Advanced Technology, Xi’an Jiaotong-Liverpool University, Suzhou, China

Abstract

Cluster-based prediction is widely used in self-supervised speech learning. A soft target preserves a distribution over clusters rather than a single label. This distribution specifies both the probability values and which clusters receive them. Comparisons between soft targets and hard labels do not separate the contributions of these two aspects to the learned representation. We study this in S-JEPA, a recent high-performing self-supervised speech model trained with soft Gaussian mixture model (GMM) targets. We compare its original targets with counterfactual targets that preserve the most likely cluster and all probability values but change which remaining clusters receive the other probabilities. Across three training seeds, the original soft distribution is recovered more accurately from Encoders trained with the original than counterfactual targets. Because this could reflect target matching alone, we also test low-level acoustic and phonetic information. Both are more accessible from Encoders trained with the original targets. This suggests that cluster assignments affect acoustic and phonetic properties of the learned representation, not just recovery of the training target.

Figures & tables

Explore similar work

CardsList
  1. S-JEPA : Soft Clustering Anchors for Self-Supervised Speech Representation Learning

    Jun 17, 2026Georgios Ioannides, Adrian Kieback, Judah Goldfeder +5S-JepaRepresentation Learning

  2. GLaS-JEPA: Gaussian-Regularized Speech SSL without Engineered Prediction Targets

    Sep 29, 2026Gaspard Botté, Séverin Baroudi, Samir Sadok +5Discrete Speech RepresentationsSpeech Encoder

  3. Leveraging Soft Distributions of SSL-Derived Discrete Speech Tokens for Downstream Inference

    Jun 5, 2026Kentaro Onda, Satoru Fukayama, Daisuke Saito +1Soft-Token RepresentationsAutomatic Speech Recognition