UniAE-MoE: A Unified Audio Encoder via Mixture of Experts
Organizations: Tsinghua University, China
Abstract
Large Audio Language Models (LALMs) rely on effective audio encoders for multi-task performance. We introduce UniAE-MoE, a unified audio encoder designed to model cross-domain audio representations and achieve outstanding downstream understanding performance via a Mixture-of-Experts (MoE) architecture. Specifically, we explore mainstream audio encoders and integrate those from Qwen2-Audio and Audio-Flamingo 3, which demonstrate superior downstream capabilities. To facilitate effective model fusion, we improve our encoder using SwiGLU with shared experts to decouple encoder networks, and we further introduce a two-stage instruction-tuning strategy to better adapt the model to diverse downstream tasks. Moreover, we propose the task-specific data scaling (TSDS) technique to enhance \tool's understanding capabilities. On the XARES-LLM benchmark, UniAE-MoE attains a score of 0.802, achieving state-of-the-art performance. It also delivers top-tier performance in the official Interspeech 2026 Audio Encoder Capability Challenge, further demonstrating robust generalization across diverse audio tasks. Together, these results validate the effectiveness of \tool for unified audio understanding across speech, music, and general audio domains.
Figures & tables
| Task | Dataset | Train # | New |
| Task 1 for Stage-2 Training | |||
| Spoofing Detection | ASVSpoof 2015 | 15,720 | - |
| Emotion Recognition | CREMA-D | 6,000 | - |
| IEMOCAP/TESS/RAVDESS | 3,552 | ✔ | |
| Environment Class | ESC-50 | 1,600 | - |
| Intent Classification | Fluent Speech Commands | 23,132 | - |
| Encoder | Domain | Task 1 | Task 2 | Overall |
| Audio Encoders for Specific Tasks. | ||||
| Whisper-large [ 20 ] | Speech | 0.698 | 0.441 | 0.634 |
| FunASR-nano [ 3 ] | Speech | 0.552 | 0.516 | 0.543 |
| MERT-330M [ 17 ] | Music | 0.475 | 0.236 | 0.415 |
| Dasheng-base [ 10 ] | Unified | 0.540 | 0.233 | 0.463 |
| Dasheng-1.2B [ 10 ] | Unified | 0.587 | 0.236 | 0.499 |
| (Q2A, AF3) | (Q2A, KA) | |||||
| Concat | Gating | MoE (Ours) | Concat | Add | MoE (Ours) | |
| Task 1 | 0.816 | 0.800 | 0.839 | 0.785 | 0.778 | 0.816 |
| Task 2 | 0.662 | 0.660 | 0.691 | 0.648 | 0.604 | 0.682 |
| Overall | 0.778 | 0.764 | 0.802 | 0.750 | 0.735 | 0.782 |
| Params (M) | Task 1 | Task 2 | Overall | |
| UniAE-MoE | 115.8 | 0.839 | 0.691 | 0.802 |
| w/o SwiGLU | 85.2 | 0.827 | 0.693 | 0.794 |
| UniAE-MoE-Light | 88.5 | 0.832 | 0.693 | 0.797 |
| Public Tasks | Hidden Tasks | Aggregate Results | |||
| Dataset | Score | Dataset | Score | Result | Score |
| Track A Performance | |||||
| ASVspoof 2015 | 0.9900 | Fingersnap | 0.8750 | Overall | 0.9122 |
| CREMA-D | 0.8793 | KeyScratch | 0.9990 | Public test | 0.8527 |
| ESC-50 | 0.9125 | King-ASR-457 | 0.9850 | Hidden test | 0.9717 |
| Fluent Speech Commands | 0.9939 | King-ASR-719 | 0.9810 | ||