cs.SDFeb 20, 2026

Re-purposing Multimodal Large Language Models for Audio-Text Retrieval

Authors: Jilan Xu, Carl Thomé, Danijela Horak, Weidi Xie, Andrew Zisserman

Organizations: Visual Geometry Group, University of Oxford · Epidemic Sound · School of Artificial Intelligence, Shanghai Jiao Tong University

Abstract

Audio-text retrieval is crucial for bridging acoustic signals and natural language. While contrastive dual-encoder architectures like CLAP have shown promise, they are fundamentally limited by the capacity of small-scale encoders. Specifically, the text encoders struggle to understand complex queries that require reasoning or world knowledge. In this paper, we propose AuroLA, a novel contrastive language-audio pre-trained model that re-purposes Multimodal Large Language Models (MLLMs) as a unified backbone for audio-text retrieval. Specifically, we make the following contributions: (i) we construct a scalable data pipeline that curates diverse audio from multiple sources and generates multi-granular captions, ranging from long descriptions to structured tags, via automated annotation; (ii) we adapt an MLLM for retrieval by prompting it to summarise the audio/text input and using the hidden state of a special token as audio/text embeddings. (iii) extensive experiments demonstrate that AuroLA consistently outperforms state-of-the-art dual-encoder models, including the recent PE-AV. This validates the effectiveness of MLLM as a unified backbone for audio-text retrieval.

Figures & tables

Explore similar work

CardsList
  1. Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval

    Apr 20, 2026HaeJun Yoo, Yongseop Shin, Insung Lee +2Text-To-AudioEmbedder

  2. ALM2Vec: Learning Audio Embeddings for Universal Audio Retrieval with Large Audio-Language Models

    Jun 27, 2026Fengjie Lu, Chenang Jiang, Jiarui Hai +2Neural AudioLarge Audio Language Models