Audio-text retrieval is crucial for bridging acoustic signals and natural language. While contrastive dual-encoder architectures like CLAP have shown promise, they are fundamentally limited by the capacity of small-scale encoders. Specifically, the text encoders struggle to understand complex queries that require reasoning or world knowledge. In this paper, we propose AuroLA, a novel contrastive language-audio pre-trained model that re-purposes Multimodal Large Language Models (MLLMs) as a unified backbone for audio-text retrieval. Specifically, we make the following contributions: (i) we construct a scalable data pipeline that curates diverse audio from multiple sources and generates multi-granular captions, ranging from long descriptions to structured tags, via automated annotation; (ii) we adapt an MLLM for retrieval by prompting it to summarise the audio/text input and using the hidden state of a special token as audio/text embeddings. (iii) extensive experiments demonstrate that AuroLA consistently outperforms state-of-the-art dual-encoder models, including the recent PE-AV. This validates the effectiveness of MLLM as a unified backbone for audio-text retrieval.
Figures & tables
Dataset
# Source
Scale
Annotation
Avg. Audio Len.
Avg. Word Len.
# Caption / Audio
Multi- granular
AudioCaps [ 45 ]
1
46K
Human
10.0 (1.8 ∼ 10)
9.0 (2 ∼ 39)
5
✗
Clotho [ 46 ]
1
5K
Human
22.5 (15 ∼ 30)
11.0 (8 ∼ 21)
5
✗
LAION-Audio-630K [ 3 ]
8
630K
Auto
24.6
7.3
1
✗
WavCaps [ 12 ]
4
403K
Auto
67.6
7.8
1
✗
Auto-ACD [ 38 ]
2
1.5M
Auto
10.0
18.1
1
✗
FusionAudio [ 14 ]
1
1.2M
Auto
10.0
47.2
1
✗
TABLE I: Comparison of source audio–text datasets and AudioVerse . Our dataset integrates audio from multiple sources and provides multi-granular textual descriptions. Average audio length (seconds) and word length are reported with (min ∼ max).
Fig. 1: Data processing pipeline. We assemble audio from diverse platforms and datasets. Qwen3-Omni-30B-A3B [ 47 ] is used to generate multi-granular captions (including long captions, short captions and tag captions) based on raw audio, task instructions, few-shot examples and auxiliary textual clues.
Fig. 2: Architecture comparison between the classic dual encoder model (left) and our unified MLLM-based retrieval model (right). The retrieval model is trained by aligning the embedding tokens of audio and text inputs via contrastive learning.
Method
AudioCaps
Clotho
Auto-ACD
T2A
A2T
T2A
A2T
T2A
A2T
OnePeace (PT)* [ 62 ]
20.7
24.0
11.1
16.2
21.7
24.2
VAST (PT)* [ 7 ]
25.4
34.8
16.1
20.0
26.4
27.5
PE-AV (PT) [ 26 ]
33.7
48.5
17.5
26.3
32.9
33.2
AuroLA (PT)
42.1
54.2
25.1
32.5
39.8
37.5
CLAP [ 3 ]
32.7
43.9
15.6
23.7
–
–
TABLE II: Main results on AudioCaps [ 45 ] , Clotho [ 46 ] , and Auto-ACD [ 38 ] . PT stands for pre-training without downstream training sets. * refers to our reproduced results.
Method
AudioCaps
Clotho
T2A
A2T
T2A
A2T
Baseline
14.1
14.5
11.1
9.7
+ Stage-1 text-only pre-training
22.0
27.2
19.0
18.4
+ Stage-2 audio-text training
31.9
42.7
21.1
25.3
+ Multi-granular caption
40.4
49.4
25.5
32.5
TABLE III: Ablation study on AudioCaps [ 45 ] and Clotho [ 46 ] .
Dataset
Scale
AudioCaps
Clotho
Auto-ACD
T2A
A2T
T2A
A2T
T2A
A2T
CLAP as backbone
WavCaps [ 12 ]
0.4M
28.6
40.2
16.5
20.0
24.5
25.6
AudioSetCaps [ 36 ]
1.9M
32.8
44.5
10.5
14.4
29.1
30.5
AudioVerse
1.3M
40.9
52.6
16.8
23.3
32.7
34.5
AuroLA as backbone
TABLE IV: Comparison of different pre-training datasets under different backbones. For fair comparison, the training sets of AudioCaps and Clotho are excluded from the pre-training data.
Fig. 3: Visualisation of text to audio retrieval on AudioCaps . Visual information is only for reference.
Model
Latency
Peak Mem
GFLOPs
Recall@1
T2A
A2T
CLAP [ 3 ]
1.2
2.2
10.5
32.7
43.9
AuroLA (ours)
37.2
18.0
7.8k
46.0
63.1
TABLE V: Inference cost comparison on AudioCaps using a H200 GPU. We report Latency (ms / sample), Peak Mem (GB) and GFLOPs.
Audio-text retrieval systems based on Contrastive Language-Audio Pretraining (CLAP) achieve strong performance on traditional benchmarks; however, these benchmarks rely on caption-style queries that differ substantially from real-world search behavior, limiting their assessment of practical retrieval robustness. We present Omni-Embed-Audio (OEA), a retrieval-oriented encoder leveraging multimodal LLMs with native audio understanding. To systematically evaluate robustness beyond caption-style queries, we introduce User-Intent Queries (UIQs) - five formulations reflecting natural search behaviors: questions, commands, keyword tags, paraphrases, and exclusion-based negative queries. For negative queries, we develop a hard negative mining pipeline and propose discrimination metrics (HNSR, TFR) assessing models' ability to suppress acoustically similar distractors. Experiments on AudioCaps, Clotho, and MECAT show that OEA achieves comparable text-to-audio retrieval performance to state-of-the-art M2D-CLAP, while demonstrating clear advantages in two critical areas: (1) dominant text-to-text retrieval (+22% relative improvement), and (2) substantially superior hard negative discrimination (+4.3%p HNSR@10, +34.7% relative TFR@10), revealing that LLM backbones provide superior semantic understanding of complex queries.
Audio-text retrieval enables semantic alignment between audio content and natural language queries, supporting applications in multimedia search, accessibility, and surveillance. However, current state-of-the-art approaches struggle with long, noisy, and weakly labeled audio due to their reliance on contrastive learning and large-batch training. We propose a novel multimodal retrieval framework that refines audio and text embeddings using a cross-modal embedding refinement module combining transformer-based projection, linear mapping, and bidirectional attention. To further improve robustness, we introduce a hybrid loss function blending cosine similarity, L1, and contrastive objectives, enabling stable training even under small-batch constraints. Our approach efficiently handles long-form and noisy audio (SNR 5 to 15) via silence-aware chunking and attention-based pooling. Experiments on benchmark datasets demonstrate improvements over prior methods.
Recent advances in language--audio retrieval have been largely driven by contrastive dual-encoder architectures that align audio and text in a shared embedding space. While effective, existing retrieval embeddings are primarily optimized for audio--caption matching, limiting their ability to support diverse retrieval objectives and controllable retrieval behaviors. We present ALM2Vec, a universal audio embedding framework derived from pretrained large audio--language models (LALMs). By transferring the audio understanding, instruction-following, and reasoning capabilities acquired through large-scale multimodal training, ALM2Vec learns a unified embedding space for retrieval across audio domains and task types. Beyond conventional text--audio retrieval, ALM2Vec incorporates natural-language instructions into the embedding process, enabling instruction-aware retrieval for scenarios such as audio question answering and aspect-conditioned retrieval. Experimental results show that ALM2Vec achieves competitive performance on standard audio and speech retrieval benchmarks while exhibiting promising compositional and controllable retrieval capabilities, highlighting its potential as a unified audio embedding model for retrieval across domains, tasks, and user intents.