Cluster-based prediction is widely used in self-supervised speech learning. A soft target preserves a distribution over clusters rather than a single label. This distribution specifies both the probability values and which clusters receive them. Comparisons between soft targets and hard labels do not separate the contributions of these two aspects to the learned representation. We study this in S-JEPA, a recent high-performing self-supervised speech model trained with soft Gaussian mixture model (GMM) targets. We compare its original targets with counterfactual targets that preserve the most likely cluster and all probability values but change which remaining clusters receive the other probabilities. Across three training seeds, the original soft distribution is recovered more accurately from Encoders trained with the original than counterfactual targets. Because this could reflect target matching alone, we also test low-level acoustic and phonetic information. Both are more accessible from Encoders trained with the original targets. This suggests that cluster assignments affect acoustic and phonetic properties of the learned representation, not just recovery of the training target.
Figures & tables
Figure 1: Counterfactual target construction and matched training runs. REAL uses the original fixed-GMM posterior qt , whereas RANDOM-REASSIGN preserves the top-1 component and all posterior values while reassigning only the non-maximal values to form qtRR . The two runs share the masked input, architecture, and training setup, and minimize KL divergence to their corresponding targets. Encoder outputs are used for subsequent representation analysis; solid and dashed arrows denote forward and gradient flow, respectively.
Figure 2: Phase-1 differences between REAL SOFT and RANDOM-REASSIGN across three independent seeds. Points show the test-clean gap for each seed; error bars are paired 95% speaker-cluster bootstrap confidence intervals over the 40 test speakers (10,000 draws). Panel (a) reports RANDOM-REASSIGN minus REAL SOFT in DeepTailCE. Panels (b) and (c) report REAL SOFT minus RANDOM-REASSIGN in controlled dynamic R2 and phone accuracy. Positive values therefore favor REAL SOFT in every panel.
Phase 1
Phase 2
Readout
A
B
C
A (100k)
DeepTailCE reduction
0.0679
0.0135
0.0402
0.0370
Controlled dynamic R2 gain
0.0262
0.0247
0.0167
0.0052
Phone accuracy gain (pp)
1.55
0.12
1.07
1.77
Table 1: REAL SOFT gains over RANDOM-REASSIGN across three Phase 1 seeds and after 100k Phase 2 updates for the primary seed.
Figure 3: Graded exposure to the original component correspondence. Each line shows one independent Phase-1 seed across λ∈{0,0.25,0.5,0.75,1} . Panels report (a) DeepTailCE, where lower is better; (b) controlled dynamic R2 ; and (c) phone accuracy, where higher is better. At λ=0 , all frames use RANDOM-REASSIGN; at λ=1 , all frames use REAL SOFT. Intermediate levels retain the original correspondence for nested subsets of frames.