cs.SDOct 8, 2026

STEMMA: Song-to-Stem Multi-Audio Reasoning for Large Audio Language Models

Authors: Hoyeol Sohn, Wonil Kim, Keunhyoung Kim, Sangeun Kum, Taehyoung Kim, Dongjoo Moon, Theerasak Charoenchob, Teeratep Weerapang, +2 more

Organizations: Graduate School of Culture Technology, KAIST, Daejeon, Republic of Korea · Neutune, Seoul, Republic of Korea

Abstract

Music understanding often requires comparing excerpts and reasoning about relationships among songs, sections, and stems. However, existing large audio-language models (LALMs) and music question-answering datasets typically operate on single recordings or compare independently sampled tracks with no known production relationship. We introduce STEMMA, a multi-audio music question-answering framework built around production provenance: whether excerpts originate from the same track or section, and which stems belong to which mixtures. Because such relations are sparse under conventional audio-first sampling, STEMMA adopts a relation-first construction strategy: it first specifies a target relation and then queries the catalog for excerpts that satisfy it and hard negatives that do not. Labels are determined directly from catalog provenance rather than generated by a language model from metadata. We build STEMMA-Bench for evaluation and a track-disjoint training set, STEMMA-Instruct. Fine-tuning two LALMs on STEMMA-Instruct improves multi-audio reasoning, with the largest gains on structural relations directly determined by the catalog, while preserving single-audio music understanding.

Figures & tables

Explore similar work

CardsList
  1. Listen-to-Reason: Listen with Experts, Retrieve over a Graph, Reason with LLMs

    Oct 7, 2026Pooneh Mousavi, Mirco Ravanelli, Cem SubakanAudio-Language ModelsAudio Reasoning

  2. EChO-Agent: Evidence Chain Orchestration Agent for Audio Reasoning

    Jun 13, 2026Siyuan Zhang, Jian Zong, Junyu Wang +7Audio QATool-Augmented Language Model Agents

  3. Beyond Accuracy: ARIA-Rubrics for Evaluating Audio Reasoning in Large Audio Language Models

    Sep 9, 2026Yupei Li, Qiyang Sun, Mohamed Mady +4Audio-Language Model EvaluationAudio Reasoning