Music understanding often requires comparing excerpts and reasoning about relationships among songs, sections, and stems. However, existing large audio-language models (LALMs) and music question-answering datasets typically operate on single recordings or compare independently sampled tracks with no known production relationship. We introduce STEMMA, a multi-audio music question-answering framework built around production provenance: whether excerpts originate from the same track or section, and which stems belong to which mixtures. Because such relations are sparse under conventional audio-first sampling, STEMMA adopts a relation-first construction strategy: it first specifies a target relation and then queries the catalog for excerpts that satisfy it and hard negatives that do not. Labels are determined directly from catalog provenance rather than generated by a language model from metadata. We build STEMMA-Bench for evaluation and a track-disjoint training set, STEMMA-Instruct. Fine-tuning two LALMs on STEMMA-Instruct improves multi-audio reasoning, with the largest gains on structural relations directly determined by the catalog, while preserving single-audio music understanding.
Figures & tables
Figure 1: Overview of the proposed relation-first generation pipeline. Hierarchical music units are indexed by structural and musical-attribute relations. A relation is stated over catalog fields, and the tuples that satisfy or violate it are retrieved by query before being rendered as multi-audio instructions.
Relation
Positive conditions
Units
Question
Structural
same track
same(track), different(section)
Sect., Stem
Which candidate audio comes from the same track as Audio A?
adjacent section
same(track), section + 1
Sect., Stem
Which candidate excerpt immediately follows Audio A in the original track?
same section
same(track), same(section), different(role)
Stem
Which candidate audio is a different isolated component from the same recorded section as Audio A?
parent–child
same(track), same(section)
Sect. → Stem
Which candidate audio is contained within Audio A?
Attribute
Table 1: The nine relations: the conditions a positive tuple must satisfy, the units each is instantiated over, and the question a candidate QA instance asks. A pair QA instance asks the same relation of two excerpts. In the Units column, Sect. denotes a section excerpt, and Stem denotes a section stem.
Multi-Audio QA
Single-Audio QA
System
MUGEN
Jamendo-MT-QA
STEMMA-Bench
MMAU-Pro
MMAU
MMAR
RUL-MuChoMusic
Music Analysis
Struct.
Attr.
Music
Mini-Music
Music
Cascaded systems
Audio Flamingo 3 [ 28 ] → Qwen3.6
0.5000
0.7169
0.4803
0.2428
0.5896
0.7006
0.4138
0.6828
Music Flamingo → Qwen3.6
0.5400
0.8716
0.4803
0.2428
0.5367
0.6677
0.4039
0.6114
MOSS-Audio [ 29 ] → Qwen3.6
0.5320
0.7741
0.4094
0.2254
0.4915
0.5689
0.4236
0.5078
Table 2: System-level accuracy across multi-audio reasoning and single-audio music understanding. Each cascade supplies independently generated audio captions to Qwen3.6-35B-A3B. Best and second-best scores are bold and underlined, respectively. Parse failures count as incorrect.
Training data
Multi-Audio Avg.
Single-Audio Avg.
Music Flamingo
Pretrained
0.3634
0.5282
Coreference only
0.4189 ( +5.55 )
0.5345 ( +0.63 )
Relation MultiQA only
0.5463 ( +18.29 )
0.5538 ( +2.56 )
Relation MultiQA + Coreference
0.6458 ( +28.24 )
0.5590 ( +3.08 )
MOSS-Music
Table 3: Ablation over the two multi-audio roles of STEMMA-Instruct; Relation MultiQA + Coreference is STEMMA-Instruct itself. Each average weights the columns of Table 2 equally within its group. Parentheses report percentage-point changes from the backbone’s pretrained model.
Imperial College London, UK · Technische Universität München München, Germany · Mohamed bin Zayed University of Artificial Intelligence Abu Dhabi, AE +2