Large Audio Language Models (LALMs) rely on effective audio encoders for multi-task performance. We introduce UniAE-MoE, a unified audio encoder designed to model cross-domain audio representations and achieve outstanding downstream understanding performance via a Mixture-of-Experts (MoE) architecture. Specifically, we explore mainstream audio encoders and integrate those from Qwen2-Audio and Audio-Flamingo 3, which demonstrate superior downstream capabilities. To facilitate effective model fusion, we improve our encoder using SwiGLU with shared experts to decouple encoder networks, and we further introduce a two-stage instruction-tuning strategy to better adapt the model to diverse downstream tasks. Moreover, we propose the task-specific data scaling (TSDS) technique to enhance \tool's understanding capabilities. On the XARES-LLM benchmark, UniAE-MoE attains a score of 0.802, achieving state-of-the-art performance. It also delivers top-tier performance in the official Interspeech 2026 Audio Encoder Capability Challenge, further demonstrating robust generalization across diverse audio tasks. Together, these results validate the effectiveness of \tool for unified audio understanding across speech, music, and general audio domains.
Figures & tables
Figure 1: The UniAE-MoE framework comprises the overall architecture (top), the SwiGLU-based expert network (bottom left), and the two-stage instruction-tuning strategy (bottom right).
Table 1: Comparison of different audio encoders on the XARES-LLM benchmark
Task
Dataset
Train #
New
Task 1 for Stage-2 Training
Spoofing Detection
ASVSpoof 2015
15,720
-
Emotion Recognition
CREMA-D
6,000
-
IEMOCAP/TESS/RAVDESS
3,552
✔
Environment Class
ESC-50
1,600
-
Intent Classification
Fluent Speech Commands
23,132
-
Table 2: Statistics of our training data
Encoder
Domain
Task 1
Task 2
Overall
Audio Encoders for Specific Tasks.
Whisper-large [ 20 ]
Speech
0.698
0.441
0.634
FunASR-nano [ 3 ]
Speech
0.552
0.516
0.543
MERT-330M [ 17 ]
Music
0.475
0.236
0.415
Dasheng-base [ 10 ]
Unified
0.540
0.233
0.463
Dasheng-1.2B [ 10 ]
Unified
0.587
0.236
0.499
Table 3: Benchmarking results of various audio encoders
(Q2A, AF3)
(Q2A, KA)
Concat
Gating
MoE (Ours)
Concat
Add
MoE (Ours)
Task 1
0.816
0.800
0.839
0.785
0.778
0.816
Task 2
0.662
0.660
0.691
0.648
0.604
0.682
Overall
0.778
0.764
0.802
0.750
0.735
0.782
Table 4: Comparison of model fusion strategies
Params (M)
Task 1
Task 2
Overall
UniAE-MoE
115.8
0.839
0.691
0.802
w/o SwiGLU
85.2
0.827
0.693
0.794
UniAE-MoE-Light
88.5
0.832
0.693
0.797
Table 5: Ablation study of the UniAE-MoE
Table 6: The table compares data-scaling performance across tasks; “w/o” denotes training without TSDS, whereas “w” denotes training with TSDS.
Public Tasks
Hidden Tasks
Aggregate Results
Dataset
Score
Dataset
Score
Result
Score
Track A Performance
ASVspoof 2015
0.9900
Fingersnap
0.8750
Overall
0.9122
CREMA-D
0.8793
KeyScratch
0.9990
Public test
0.8527
ESC-50
0.9125
King-ASR-457
0.9850
Hidden test
0.9717
Fluent Speech Commands
0.9939
King-ASR-719
0.9810
Table 7: Official leaderboard results of UniAE-MoE in the Interspeech 2026 Audio Encoder Capability Challenge, covering Track A (classification) and Track B (audio understanding)
Audio encoders are critical to modern audio applications as large language models (LLMs) increasingly rely on a single encoder for diverse inputs. While self-supervised learning (SSL) has yielded strong domain-specific encoders like speech or music experts, multi-domain approaches like USAD and SPEAR remain limited in coverage and evaluation. Recent studies also suggest supervised encoders align better with audio LLMs. We present USAD 2.0, a universal encoder integrating knowledge from both SSL and supervised foundation models. USAD 2.0 introduces domain-aware distillation to address teacher mismatch, extends coverage to the music domain, and adds second-stage supervised distillation for downstream use. We further scale the model to one billion parameters via depth scaling. Experiments show USAD 2.0 achieves strong or state-of-the-art performance across probing and LLM-based evaluations.
Heng-Jui Chang, Alexander H. Liu, Saurabhchand Bhati +4
While pre-trained models excel in specialized tasks, learning universal representations across diverse acoustic domains remains challenging. To address this, we propose WQ-Fusion, a robust dual-encoder framework for cross-domain audio representation learning. Overcoming the limitations of static concatenation, WQ-Fusion integrates whisper and qwen via an Adaptive Feature Modulation module and a novel element-wise gated attention mechanism. This design enables dynamic feature selection, allowing the model to selectively emphasize relevant acoustic and semantic dimensions. Extensive experiments on the Interspeech 2026 Audio Encoder Capability Challenge (Track A) benchmark demonstrate that by effectively routing heterogeneous information, WQ-Fusion achieves a superior overall score of 0.836, significantly outperforming the strongest single-encoder baseline.
Mingda Lin, Lei Ding, Xinyue Zhou +6
School of Electronic Information, Wuhan University, Wuhan, Hubei, China · Tencent AI Lab Seattle, Seattle, USA · CIAIC, Northwestern Polytechnical University, Xi’an, Shaanxi, China +1
Audio intelligence involves understanding, reasoning about, and generating both audio and speech. In this work, we introduce Nemotron-Labs-Audex-30B-A3B (Audex), a unified audio-text LLM built on Nemotron-Cascade-2-30B-A3B, a strong text-only MoE LLM. Audex adopts a simple unified design with a single Transformer decoder: audio inputs are encoded and projected into the text embedding space, while text tokens and quantized audio output tokens are treated uniformly during generation. This architecture enables strong audio-text fusion, seamless multimodal generation, and compatibility with standard LLM training and inference infrastructure. For training, we meticulously curate audio-text datasets comprising 157.4B audio tokens and 320.5B text tokens. We apply multi-stage supervised training on these datasets, followed by text-only Cascade RL and multi-domain on-policy distillation. Audex delivers state-of-the-art audio understanding, speech recognition and translation, text-to-speech, audio generation, and speech-to-speech generation, while preserving very compelling reasoning, alignment, knowledge, long-context, and agentic capabilities of its text-only LLM backbone with marginal or no regression. We release the model checkpoints to facilitate open research.