Audio-text retrieval is crucial for bridging acoustic signals and natural language. While contrastive dual-encoder architectures like CLAP have shown promise, they are fundamentally limited by the capacity of small-scale encoders. Specifically, the text encoders struggle to understand complex queries that require reasoning or world knowledge. In this paper, we propose AuroLA, a novel contrastive language-audio pre-trained model that re-purposes Multimodal Large Language Models (MLLMs) as a unified backbone for audio-text retrieval. Specifically, we make the following contributions: (i) we construct a scalable data pipeline that curates diverse audio from multiple sources and generates multi-granular captions, ranging from long descriptions to structured tags, via automated annotation; (ii) we adapt an MLLM for retrieval by prompting it to summarise the audio/text input and using the hidden state of a special token as audio/text embeddings. (iii) extensive experiments demonstrate that AuroLA consistently outperforms state-of-the-art dual-encoder models, including the recent PE-AV. This validates the effectiveness of MLLM as a unified backbone for audio-text retrieval.
Figures & tables
Dataset
# Source
Scale
Annotation
Avg. Audio Len.
Avg. Word Len.
# Caption / Audio
Multi- granular
AudioCaps [ 45 ]
1
46K
Human
10.0 (1.8 ∼ 10)
9.0 (2 ∼ 39)
5
✗
Clotho [ 46 ]
1
5K
Human
22.5 (15 ∼ 30)
11.0 (8 ∼ 21)
5
✗
LAION-Audio-630K [ 3 ]
8
630K
Auto
24.6
7.3
1
✗
WavCaps [ 12 ]
4
403K
Auto
67.6
7.8
1
✗
Auto-ACD [ 38 ]
2
1.5M
Auto
10.0
18.1
1
✗
FusionAudio [ 14 ]
1
1.2M
Auto
10.0
47.2
1
✗
TABLE I: Comparison of source audio–text datasets and AudioVerse . Our dataset integrates audio from multiple sources and provides multi-granular textual descriptions. Average audio length (seconds) and word length are reported with (min ∼ max).
Fig. 1: Data processing pipeline. We assemble audio from diverse platforms and datasets. Qwen3-Omni-30B-A3B [ 47 ] is used to generate multi-granular captions (including long captions, short captions and tag captions) based on raw audio, task instructions, few-shot examples and auxiliary textual clues.
Fig. 2: Architecture comparison between the classic dual encoder model (left) and our unified MLLM-based retrieval model (right). The retrieval model is trained by aligning the embedding tokens of audio and text inputs via contrastive learning.
Method
AudioCaps
Clotho
Auto-ACD
T2A
A2T
T2A
A2T
T2A
A2T
OnePeace (PT)* [ 62 ]
20.7
24.0
11.1
16.2
21.7
24.2
VAST (PT)* [ 7 ]
25.4
34.8
16.1
20.0
26.4
27.5
PE-AV (PT) [ 26 ]
33.7
48.5
17.5
26.3
32.9
33.2
AuroLA (PT)
42.1
54.2
25.1
32.5
39.8
37.5
CLAP [ 3 ]
32.7
43.9
15.6
23.7
–
–
TABLE II: Main results on AudioCaps [ 45 ] , Clotho [ 46 ] , and Auto-ACD [ 38 ] . PT stands for pre-training without downstream training sets. * refers to our reproduced results.
Method
AudioCaps
Clotho
T2A
A2T
T2A
A2T
Baseline
14.1
14.5
11.1
9.7
+ Stage-1 text-only pre-training
22.0
27.2
19.0
18.4
+ Stage-2 audio-text training
31.9
42.7
21.1
25.3
+ Multi-granular caption
40.4
49.4
25.5
32.5
TABLE III: Ablation study on AudioCaps [ 45 ] and Clotho [ 46 ] .
Dataset
Scale
AudioCaps
Clotho
Auto-ACD
T2A
A2T
T2A
A2T
T2A
A2T
CLAP as backbone
WavCaps [ 12 ]
0.4M
28.6
40.2
16.5
20.0
24.5
25.6
AudioSetCaps [ 36 ]
1.9M
32.8
44.5
10.5
14.4
29.1
30.5
AudioVerse
1.3M
40.9
52.6
16.8
23.3
32.7
34.5
AuroLA as backbone
TABLE IV: Comparison of different pre-training datasets under different backbones. For fair comparison, the training sets of AudioCaps and Clotho are excluded from the pre-training data.
Fig. 3: Visualisation of text to audio retrieval on AudioCaps . Visual information is only for reference.
Model
Latency
Peak Mem
GFLOPs
Recall@1
T2A
A2T
CLAP [ 3 ]
1.2
2.2
10.5
32.7
43.9
AuroLA (ours)
37.2
18.0
7.8k
46.0
63.1
TABLE V: Inference cost comparison on AudioCaps using a H200 GPU. We report Latency (ms / sample), Peak Mem (GB) and GFLOPs.