cs.SDSep 23, 2026

BiMamba2 Masked Discrete-Unit Prediction for Multilingual Speech Representation for Unsupervised Speech in the Wild Challenge

Authors: Prakriti Subedi, Howard Prioleau, Saurav K Aryal

Organizations: Howard University, Washington D.C., USA

Abstract

We describe our submission to the Unsupervised Speech in the Wild (UPS) Challenge at Interspeech 2026, a bidirectional Mamba-2 (BiMamba2) encoder trained with masked discrete-unit prediction following the HuBERT-style paradigm. The 47.88M-parameter model is trained on 250 hours of speech across 67 languages from the MLCommons Unsupervised People's Speech dataset, with no labeled data. The objective combines masked k-means pseudo-label prediction with language identification supervision and VICReg regularization. On official evaluation, the system achieves an Adjusted Rand Index of 0.735, exceeding four baselines on speaker clustering. Language identification macro-F1 (0.073) and character error rate (0.870) remain below supervised baselines. We analyze a local-official discrepancy in metric scale and checkpoint ranking, highlighting limitations of in-distribution diagnostics for predicting Dynabench probe outcomes.

Figures & tables

Explore similar work

CardsList
  1. DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech Units

    Mar 19, 2026Maxime Poli, Manel Khentout, Angelo Ortiz Tandazo +3Grapheme-To-PhonemeMultilingual Benchmark

  2. MauBERT: Universal Phonetic Inductive Biases for Few-Shot Acoustic Units Discovery

    Dec 22, 2025Angelo Ortiz Tandazo, Manel Khentout, Youssef Benchekroun +2Self-Supervised Speech ModelsArticulatory

  3. From Monolingual to Multilingual: Evaluating Mamba for ASR in South African Languages

    Jul 1, 2026Jesujoba O. Alabi, Julian Herreilers, Badr M. Abdullah +1Multilingual Automatic Speech RecognitionAutomatic Speech Recognition