BiMamba2 Masked Discrete-Unit Prediction for Multilingual Speech Representation for Unsupervised Speech in the Wild Challenge
Organizations: Howard University, Washington D.C., USA
Abstract
We describe our submission to the Unsupervised Speech in the Wild (UPS) Challenge at Interspeech 2026, a bidirectional Mamba-2 (BiMamba2) encoder trained with masked discrete-unit prediction following the HuBERT-style paradigm. The 47.88M-parameter model is trained on 250 hours of speech across 67 languages from the MLCommons Unsupervised People's Speech dataset, with no labeled data. The objective combines masked k-means pseudo-label prediction with language identification supervision and VICReg regularization. On official evaluation, the system achieves an Adjusted Rand Index of 0.735, exceeding four baselines on speaker clustering. Language identification macro-F1 (0.073) and character error rate (0.870) remain below supervised baselines. We analyze a local-official discrepancy in metric scale and checkpoint ranking, highlighting limitations of in-distribution diagnostics for predicting Dynabench probe outcomes.
Figures & tables
| en, es, de, pt, ar, fr, it, gl, ne, id, pl, zh, nl, ca, sv, am, ru, hu, hi, eu, ur, jw, ms, tr, ko, el, kn, sa, cy, ta, sw, uk, bs, ro, si, fa, te, sr, so, hr, ja, nn, la, da, tl, hy, vi, sq, pa, sk, ka, yo, be, bn, he, bg, km, br, af, sd, lt, sn, cs, mr, ps, th, ml |
| System | Macro-F1 | CER | ARI |
|---|---|---|---|
| Whisper [ 24 ] | 0.950 | 0.790 | 0.560 |
| HuBERT-large [ 2 ] | 0.690 | 0.570 | 0.600 |
| XLSR [ 4 ] | 0.260 | 0.900 | 0.100 |
| wav2vec 2.0 [ 1 ] | 0.250 | 0.980 | 0.630 |
| Ours (d=512, 8L) † | 0.052 | 0.996 | 0.291 |
| Ours (d=768, 12L, step 48k) | 0.070 | 0.998 | 0.710 |
| Metric | Step 19.5k | Step 48k |
|---|---|---|
| Macro-F1 | 0.342 | 0.395 |
| ARI | 0.225 | 0.232 |
| Intra-class sim. | 0.786 | 0.747 |
| Inter-class sim. | 0.004 | 0.002 |
| Separability ratio | 208.98 | 366.84 |