eess.ASSep 29, 2026

GLaS-JEPA: Gaussian-Regularized Speech SSL without Engineered Prediction Targets

Authors: Gaspard Botté, Séverin Baroudi, Samir Sadok, Francesco Paissan, Thomas Hueber, Xavier Alameda-Pineda, Ricard Marxer, Mirco Ravanelli

Organizations: Concordia University · Mila – Qu´ebec AI Institute · Univ Toulon, Aix Marseille Univ, CNRS, LIS · Inria, Univ. Grenoble Alpes, CNRS, LJK · Universit´e Laval · Univ. Grenoble Alpes, CNRS, Grenoble INP, GIPSA-lab, Grenoble, France · CNRS, ILLS, Univ Toulon

Abstract

Speech self-supervised learning aims to learn general-purpose representations for downstream speech tasks. However, current approaches rely on complex, carefully designed prediction targets. We challenge this necessity with GLaS-JEPA, a framework that directly predicts the current encoder's continuous representations at masked positions, without contrastive learning, discrete targets, or separate EMA target encoders. We prevent representation collapse using SIGReg representation-space regularization, eliminating the need for engineered target-generation mechanisms. Pretrained on 960 hours of LibriSpeech, our 57M-parameter model achieves a 6.89% WER on frozen-encoder SUPERB ASR and a 25.87% CER on slot filling, outperforming the best non-distilled sub-90M baselines by 43.1% and 22.0%, respectively. These results demonstrate that highly competitive speech representations can emerge from a radically simplified training recipe.

Figures & tables

Explore similar work

CardsList
  1. S-JEPA : Soft Clustering Anchors for Self-Supervised Speech Representation Learning

    Jun 17, 2026Georgios Ioannides, Adrian Kieback, Judah Goldfeder +5S-JepaRepresentation Learning

  2. OLIVE: View-Augmented Latent Prediction with Waveform Reconstruction for Speech SSL

    Jun 29, 2026Karl El Hajal, Mathew Magimai. -DossSelf-Supervised Speech ModelsSelf-Supervised Learning

  3. Cluster Assignments in Soft Targets Shape Speech Representations: Evidence from S-JEPA

    Aug 19, 2026Wenxuan He, Yunpeng Li, Zewei Li +4Self-Supervised Speech ModelsDiscrete Speech Representations