Listen-to-Reason: Listen with Experts, Retrieve over a Graph, Reason with LLMs
Organizations: Concordia University · Mila – Quebec AI Institute · Laval University
Abstract
Large audio-language models (LALMs) fuse an audio encoder into a large language model (LLM) through multi-stage training. This coupling means that a new domain or a stronger LLM requires retraining, and their answers cannot be traced to what the model heard: a chain-of-thought is a post-hoc account. We propose Listen-to-Reason (L2R), an interpretable-by-design pipeline that passes audio to the LLM through an explicit, human-readable tree: small heads on frozen expert encoders map each chunk of a clip to semantically meaningful nodes on the tree (for speech, music and environmental sound), and a frozen text-only LLM answers from these nodes and an ASR transcript without hearing the clip. Every answer can therefore be traced to the nodes and transcript it read, and the nodes are causal: replacing the deciding node with a distractor overturns 78% of correct answers on SAKURA. With a 7B reader, L2R outperforms all LALMs we compare against on SAKURA and trails them by 6-12 points on MMAU and MMAR, despite training about 1,400x fewer parameters on orders of magnitude less audio data. However, because any LLM can serve as the reader, we show that a stronger reader narrows this gap without retraining any audio component. A new domain is added with one small head: with five labelled clips per species, it outperforms QLoRA fine-tuning of an LALM on the same clips by 13-26 points.
Figures & tables
| B | SAKURA | MMAU | MMAR | MMAU-Pro | ||
|---|---|---|---|---|---|---|
| Input ablation (Qwen2.5-7B) | Blind (question and options only) | 7.6 | 32.6 | 50.2 | 37.8 | 38.9 |
| Transcript only | 7.6 | 51.8 | 60.1 | 47.8 | 44.6 | |
| Tree only | 7.6 | 74.7 | 57.5 | 40.3 | 39.3 | |
| L2R (ours) (tree + transcript) by reader | Qwen2.5-7B | 7.6 | 80.1 | 68.3 | 51.6 | 45.7 |
| Qwen2.5-14B | 14.7 | 83.3 | 69.4 | 55.6 | 48.5 | |
| Qwen2.5-32B | 32.8 | 85.1 | 71.4 | 55.7 | 49.8 |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Region | Attribute | Classes | Leaves | Decided per | Example classes | Router encoder | Expert labeller | Human labels |
| all | regions present* | 3 | – | chunk | speech, music, environmental sound | CLAP | PANNs (0.30) | – |
| speech | gender | 2 | – | chunk | male voice, female voice | WavLM-SV | WavLM-SV (0.05) | – |
| age group | 4 | – | chunk | child, young adult, elderly adult | Whisper | – | – | |
| language | 7 | 50 | clip | western european ( french ), south asian | Whisper † | Whisper-LID (0.50) | – | |
| number of speakers | 3 | – | clip | one, two, three or more | Whisper | – | – | |
| emotion | 7 | 50 | chunk | angry, happy, sad, neutral | emotion2vec | emotion2vec (0.60) | – |
| Attribute | Classes | Leaves | Router encoder | Example classes |
|---|---|---|---|---|
| animal | 26 | 96 | BEATs | bird, cat, cattle, dog, insect |
| vehicle | 12 | 40 | CLAP | aircraft, boat, car, emergency vehicle, train |
| machine | 2 | 37 | CLAP | engine, mechanisms |
| tool | 7 | 12 | CLAP | hammer, jackhammer, power tool, sanding |
| household | 23 | 38 | CLAP | blender, dishes, door, tap or sink, typing |
| human non-speech | 24 | 58 | BEATs | applause, breathing, cough, laughter |
| Step | Rule |
|---|---|
| Normalize | Remove filler words ( a , some , sound ) and merge names that are nearly identical and have the same number of words. |
| Admit | Keep a leaf only if it appears in at least 2 clips from at least 2 sources, adds detail beyond its class, and contains no number. |
| Merge | The text LLM proposes groups of synonyms. A group is kept only if its names do not conflict in number, negation or polarity ( high-pitched vs low-pitched ), its chunks sound closer to each other than to the other leaves of the class, and it does not join two different sources ( goat bleating vs sheep bleating ). |
| Prune | Fold a leaf that only adds a filler word into its sibling ( alarm sound alarm ). Remove attributes whose content lies in the spoken words (topic, intent, scenario, speaker role) or that carried no usable information (recording channel, voice quality, accent, weather, number of sources). |
| Heads | Architecture | Loss | Learning rate | Batch | Training |
|---|---|---|---|---|---|
| speech, music, scene | 1–3-layer MLP body (width 256), one head per attribute | masked CE / BCE (Eq. 2 ) | – , one-cycle | 64 | 60–80 epochs |
| environmental-sound identity | LayerNorm + 1,024-unit layer, sigmoid per node | masked BCE, positive weights | 512 | 25 epochs | |
| music properties | 2-layer MLP (width 256) | class-weighted CE | 64 | 60 epochs | |
| event timeline | linear layer on PANNs frames | BCE on strong labels | 256 | 6 epochs | |
| new domain | 1-layer MLP body, classes + none | CE | 64 | 40 epochs |
| Dataset | Domain | Tree pool | Router training (labels, size) |
| WavCaps AudioSet-SL | sound | 805 | sound hierarchy: AudioSet labels, 81,325 clips |
| Clotho | sound | 645 | – |
| FSD50K | sound | 413 | sound hierarchy (CLAP heads): leaf labels, 24,238 dev clips |
| ESC-50 | sound | 280 | sound hierarchy (CLAP heads): class labels, 1,793 clips |
| WavCaps SoundBible | sound | 127 | – |
| AudioCaps | sound | 107 | – |
| removed group | SAKURA | MMAU | MMAR |
|---|---|---|---|
| every leaf (classes kept) | |||
| speech gender | |||
| sound source (event, action and material) | |||
| speech emotion | |||
| speech language | |||
| music identity (genre, instrument, mood, era, vocals) |
| Reader | B | MMAU | MMAR | SAKURA |
|---|---|---|---|---|
| Qwen2.5-0.5B | 0.5 | 40.5 | 36.5 | 47.6 |
| Qwen2.5-1.5B | 1.5 | 61.5 | 47.3 | 69.3 |
| SmolLM2-1.7B | 1.7 | 43.2 | 39.1 | 44.9 |
| Qwen2.5-3B | 3 | 63.1 | 50.1 | 72.0 |
| Phi-3.5-mini | 3.8 | 63.0 | 47.4 | 73.8 |
| Qwen3-4B | 4 | 63.9 | 50.9 | 78.2 |
| method (accuracy ) | 3 | 5 | 10 | 20 | 50 | ||
|---|---|---|---|---|---|---|---|
| Birds (48 sp.) | L2R (ours, 0.34M trainable) | 48.4 | 72.7 | 79.6 | 85.3 | 88.4 | 91.1 |
| Omni-7B + QLoRA (40.4M trainable) | – | – | 54.1 | 61.0 | 69.0 | 75.2 | |
| Omni-7B, clips in context | 43.3 | 46.1 | 47.3 | – | – | – | |
| Marine (31 sp.) | L2R (ours, 0.17M trainable) | 41.3 | 59.6 | 65.2 | 76.5 | 83.0 | – |
| Omni-7B + QLoRA (40.4M trainable) | – | – | 52.2 | 68.6 | 73.5 | 83.6 | |
| Omni-7B, clips in context | 44.1 | 57.1 | 60.5 | – | – | – |
| MMAU | MMAR | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| question type | n | ours | Omni | blind | net | n | ours | Omni | blind | net |
| counting | 100 | 47.0 | 53.0 | 43.0 | -6 | 122 | 33.6 | 45.1 | 27.0 | -14 |
| order / timing | 113 | 60.2 | 62.8 | 39.8 | -3 | 106 | 47.2 | 57.5 | 37.7 | -11 |
| scene / place / activity | 115 | 77.4 | 93.0 | 46.1 | -18 | 111 | 58.6 | 71.2 | 27.0 | -14 |
| identify source / instrument / genre | 148 | 71.6 | 83.8 | 60.1 | -18 | 30 | 43.3 | 46.7 | 40.0 | -1 |
| music theory | 102 | 57.8 | 61.8 | 52.9 | -4 | 22 | 50.0 | 31.8 | 40.9 | +4 |
| sub-category | n | ours | Omni | blind | net |
|---|---|---|---|---|---|
| Acoustic Scene Reasoning | 48 | 56.2 | 81.2 | 33.3 | -12 |
| Musical Texture Interpretation | 34 | 58.8 | 82.4 | 44.1 | -8 |
| Event-Based Sound Reasoning | 48 | 64.6 | 79.2 | 62.5 | -7 |
| Acoustic Source Inference | 48 | 75.0 | 87.5 | 72.9 | -6 |
| Temporal Event Reasoning | 48 | 58.3 | 70.8 | 25.0 | -6 |
| Melodic Structure Interpretation | 33 | 48.5 | 66.7 | 42.4 | -6 |
| sub-category | n | ours | Omni | blind | net |
|---|---|---|---|---|---|
| Environmental Perception and Reasoning | 149 | 57.7 | 73.2 | 30.9 | -23 |
| Content Analysis | 301 | 62.1 | 67.4 | 44.9 | -16 |
| Counting and Statistics | 99 | 31.3 | 46.5 | 25.3 | -15 |
| Correlation Analysis | 50 | 44.0 | 68.0 | 28.0 | -12 |
| Professional Knowledge and Reasoning | 70 | 47.1 | 62.9 | 34.3 | -11 |
| Speaker Analysis | 48 | 52.1 | 68.8 | 29.2 | -8 |
| total params | trained params | audio training data | MACs / item | s / item | |
|---|---|---|---|---|---|
| L2R (ours) | 10.50B | 7.7M | 0.2M clips | 5.3T | 0.57 |
| Qwen2.5-Omni-7B | 10.73B | 10.73B | 300B tokens | 2.7T | 0.12 |
| Audio Flamingo 3 | 8.27B | 8.27B | 9.6M QA + 13.2M captions | 3.3T | 0.11 |
| Qwen3-Omni-30B-A3B | 35.3B (3B active) | 35.3B | 20M h + 0.77T tokens | 0.76T | 0.93 |