cs.SDOct 7, 2026

Listen-to-Reason: Listen with Experts, Retrieve over a Graph, Reason with LLMs

Authors: Pooneh Mousavi, Mirco Ravanelli, Cem Subakan

Organizations: Concordia University · Mila – Quebec AI Institute · Laval University

Abstract

Large audio-language models (LALMs) fuse an audio encoder into a large language model (LLM) through multi-stage training. This coupling means that a new domain or a stronger LLM requires retraining, and their answers cannot be traced to what the model heard: a chain-of-thought is a post-hoc account. We propose Listen-to-Reason (L2R), an interpretable-by-design pipeline that passes audio to the LLM through an explicit, human-readable tree: small heads on frozen expert encoders map each chunk of a clip to semantically meaningful nodes on the tree (for speech, music and environmental sound), and a frozen text-only LLM answers from these nodes and an ASR transcript without hearing the clip. Every answer can therefore be traced to the nodes and transcript it read, and the nodes are causal: replacing the deciding node with a distractor overturns 78% of correct answers on SAKURA. With a 7B reader, L2R outperforms all LALMs we compare against on SAKURA and trails them by 6-12 points on MMAU and MMAR, despite training about 1,400x fewer parameters on orders of magnitude less audio data. However, because any LLM can serve as the reader, we show that a stronger reader narrows this gap without retraining any audio component. A new domain is added with one small head: with five labelled clips per species, it outperforms QLoRA fine-tuning of an LALM on the same clips by 13-26 points.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Beyond Accuracy: ARIA-Rubrics for Evaluating Audio Reasoning in Large Audio Language Models

    Sep 9, 2026Yupei Li, Qiyang Sun, Mohamed Mady +4Audio-Language Model EvaluationAudio Reasoning

  2. Audio-DeepThinker: Progressive Reasoning-Aware Reinforcement Learning for High-Quality Chain-of-Thought Emergence in Audio Language Models

    Apr 20, 2026Xiang He, Chenxing Li, Jinting Wang +5RL for Language Model ReasoningAudio Reasoning

  3. EChO-Agent: Evidence Chain Orchestration Agent for Audio Reasoning

    Jun 13, 2026Siyuan Zhang, Jian Zong, Junyu Wang +7Audio QATool-Augmented Language Model Agents