Music understanding often requires comparing excerpts and reasoning about relationships among songs, sections, and stems. However, existing large audio-language models (LALMs) and music question-answering datasets typically operate on single recordings or compare independently sampled tracks with no known production relationship. We introduce STEMMA, a multi-audio music question-answering framework built around production provenance: whether excerpts originate from the same track or section, and which stems belong to which mixtures. Because such relations are sparse under conventional audio-first sampling, STEMMA adopts a relation-first construction strategy: it first specifies a target relation and then queries the catalog for excerpts that satisfy it and hard negatives that do not. Labels are determined directly from catalog provenance rather than generated by a language model from metadata. We build STEMMA-Bench for evaluation and a track-disjoint training set, STEMMA-Instruct. Fine-tuning two LALMs on STEMMA-Instruct improves multi-audio reasoning, with the largest gains on structural relations directly determined by the catalog, while preserving single-audio music understanding.
Figures & tables
Figure 1: Overview of the proposed relation-first generation pipeline. Hierarchical music units are indexed by structural and musical-attribute relations. A relation is stated over catalog fields, and the tuples that satisfy or violate it are retrieved by query before being rendered as multi-audio instructions.
Relation
Positive conditions
Units
Question
Structural
same track
same(track), different(section)
Sect., Stem
Which candidate audio comes from the same track as Audio A?
adjacent section
same(track), section + 1
Sect., Stem
Which candidate excerpt immediately follows Audio A in the original track?
same section
same(track), same(section), different(role)
Stem
Which candidate audio is a different isolated component from the same recorded section as Audio A?
parent–child
same(track), same(section)
Sect. → Stem
Which candidate audio is contained within Audio A?
Attribute
Table 1: The nine relations: the conditions a positive tuple must satisfy, the units each is instantiated over, and the question a candidate QA instance asks. A pair QA instance asks the same relation of two excerpts. In the Units column, Sect. denotes a section excerpt, and Stem denotes a section stem.
Multi-Audio QA
Single-Audio QA
System
MUGEN
Jamendo-MT-QA
STEMMA-Bench
MMAU-Pro
MMAU
MMAR
RUL-MuChoMusic
Music Analysis
Struct.
Attr.
Music
Mini-Music
Music
Cascaded systems
Audio Flamingo 3 [ 28 ] → Qwen3.6
0.5000
0.7169
0.4803
0.2428
0.5896
0.7006
0.4138
0.6828
Music Flamingo → Qwen3.6
0.5400
0.8716
0.4803
0.2428
0.5367
0.6677
0.4039
0.6114
MOSS-Audio [ 29 ] → Qwen3.6
0.5320
0.7741
0.4094
0.2254
0.4915
0.5689
0.4236
0.5078
Table 2: System-level accuracy across multi-audio reasoning and single-audio music understanding. Each cascade supplies independently generated audio captions to Qwen3.6-35B-A3B. Best and second-best scores are bold and underlined, respectively. Parse failures count as incorrect.
Training data
Multi-Audio Avg.
Single-Audio Avg.
Music Flamingo
Pretrained
0.3634
0.5282
Coreference only
0.4189 ( +5.55 )
0.5345 ( +0.63 )
Relation MultiQA only
0.5463 ( +18.29 )
0.5538 ( +2.56 )
Relation MultiQA + Coreference
0.6458 ( +28.24 )
0.5590 ( +3.08 )
MOSS-Music
Table 3: Ablation over the two multi-audio roles of STEMMA-Instruct; Relation MultiQA + Coreference is STEMMA-Instruct itself. Each average weights the columns of Table 2 equally within its group. Parentheses report percentage-point changes from the backbone’s pretrained model.
Large audio-language models (LALMs) fuse an audio encoder into a large language model (LLM) through multi-stage training. This coupling means that a new domain or a stronger LLM requires retraining, and their answers cannot be traced to what the model heard: a chain-of-thought is a post-hoc account. We propose Listen-to-Reason (L2R), an interpretable-by-design pipeline that passes audio to the LLM through an explicit, human-readable tree: small heads on frozen expert encoders map each chunk of a clip to semantically meaningful nodes on the tree (for speech, music and environmental sound), and a frozen text-only LLM answers from these nodes and an ASR transcript without hearing the clip. Every answer can therefore be traced to the nodes and transcript it read, and the nodes are causal: replacing the deciding node with a distractor overturns 78% of correct answers on SAKURA. With a 7B reader, L2R outperforms all LALMs we compare against on SAKURA and trails them by 6-12 points on MMAU and MMAR, despite training about 1,400x fewer parameters on orders of magnitude less audio data. However, because any LLM can serve as the reader, we show that a stronger reader narrows this gap without retraining any audio component. A new domain is added with one small head: with five labelled clips per species, it outperforms QLoRA fine-tuning of an LALM on the same clips by 13-26 points.
Pooneh Mousavi, Mirco Ravanelli, Cem Subakan
Concordia University · Mila – Quebec AI Institute · Laval University
While LALMs show promise on audio question answering, they fail to focus on question-relevant segments of audio and provide a clear, checkable reasoning process when dealing with complex audio reasoning. Reinforcement learning and tool-augmented prompting can help models better relate questions to audio but lack a reliable way to understand, integrate, and self-verify audio segments. To address this gap, we present EChO-Agent, a modular agent framework that reformulates complex audio QA as a planning, tool execution, evidence integration, and answer verification workflow. Experiments on MMAR benchmark show EChO-Agent improves both accuracy and rubric scores over baseline and ablation studies show evidence integration is the key factor.
Siyuan Zhang, Jian Zong, Junyu Wang +7
School of Artificial Intelligence, Tianjin University, Tianjin, China
Large Audio Language Models (LALMs) have shown strong performance on audio reasoning benchmarks, but accuracy alone cannot distinguish true reasoning from superficial pattern matching, often overestimating reasoning ability since high scores may result from guessing rather than genuine audio understanding. Evaluating the reasoning process itself is essential for improving LALMs' reasoning ability, yet remains challenging. Existing methods either rely on costly human annotation or opaque LLM-as-judge approaches, making them impractical, biased, and lacking transparency. Moreover, audio reasoning introduces unique challenges absent in text-based settings, perceptual hallucination and cross-modal alignment between audio understanding and textual inference, hence text-based evaluation frameworks cannot be directly applied. Therefore, we propose ARIA-Rubrics (Audio Reasoning Integrity Assessment), a lightweight, annotation-free gold reasoning chains, automatic and transparent framework comprising six complementary metrics that evaluate audio reasoning quality across perceptual grounding, reasoning coherence, and answer consistency. We use Chain-of-Thought prompting as an externalization mechanism to make the reasoning process observable. Experiments on 9 models across 2 benchmarks identify three reasoning modes of current LALMs with actionable directions for future development, with ARIA-Rubrics achieving high correlation with human judgments. The code is available at the Github Repository.
Yupei Li, Qiyang Sun, Mohamed Mady +4
Imperial College London, UK · Technische Universität München München, Germany · Mohamed bin Zayed University of Artificial Intelligence Abu Dhabi, AE +2