cs.SDOct 1, 2026

From Isolated Feature to Orbits: Discovering Music Concepts via Multi-SAE Alignment

Authors: Liwei Lin, Gus Xia

Organizations: Computer Science Department, NYU Shanghai · Machine Learning Department, MBZUAI

Abstract

How can we understand what a music foundation model has learned \textit{internally}? Most interpretability approaches, such as probing and Sparse Autoencoders (SAEs), focus on identifying individual features with minimal structural assumptions. We argue that many concepts are better understood as \textit{structured relations} rather than isolated features. This is especially prominent in music, where tonal structures are organized in the space of pitch and time. For example, concepts such as chords or keys are naturally expressed as structured sets (e.g., the 12 transpositions of a chord or the diatonic system within a key), rather than isolated features. In this study, \textbf{we shift from feature identification to structure-based analysis}, asking whether the learned inner representations of music foundation model emerge as organized structures over features. To this end, we introduce a framework that uses pitch transposition as an inductive bias to induce ordered orbits via multi-view SAE alignment. Concretely, we generate pitch-shifted input pairs and align their SAE representations to discover structured groups of pitch-related features. Experimental results show that this approach recovers orbit structures corresponding to chords, keys, and melodic patterns across two state-of-the-art music foundation models, while requiring only minimal grounding (e.g., a few anchor examples) to interpret entire concept families.

Figures & tables

Appendix figures & tables22 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Do Music Foundation Models Embed Pitch in Helical Structure?

    Jul 31, 2026Hayato Yagi, Shinnosuke Takamichi, Rin Sato +2Music Structure AnalysisFoundation Model

  2. MERIT: Learning Disentangled Music Representations for Audio Similarity

    May 26, 2026Abhinaba Roy, Junyi Liang, Dorien HerremansMusic Structure AnalysisAudio Understanding

  3. MIDI-RAE-JEPA: Hierarchical Representation Learning and Generation for Symbolic Music

    Jul 16, 2026Scott H. HawleySymbolic Music GenerationS-Jepa