How can we understand what a music foundation model has learned \textit{internally}? Most interpretability approaches, such as probing and Sparse Autoencoders (SAEs), focus on identifying individual features with minimal structural assumptions. We argue that many concepts are better understood as \textit{structured relations} rather than isolated features. This is especially prominent in music, where tonal structures are organized in the space of pitch and time. For example, concepts such as chords or keys are naturally expressed as structured sets (e.g., the 12 transpositions of a chord or the diatonic system within a key), rather than isolated features. In this study, \textbf{we shift from feature identification to structure-based analysis}, asking whether the learned inner representations of music foundation model emerge as organized structures over features. To this end, we introduce a framework that uses pitch transposition as an inductive bias to induce ordered orbits via multi-view SAE alignment. Concretely, we generate pitch-shifted input pairs and align their SAE representations to discover structured groups of pitch-related features. Experimental results show that this approach recovers orbit structures corresponding to chords, keys, and melodic patterns across two state-of-the-art music foundation models, while requiring only minimal grounding (e.g., a few anchor examples) to interpret entire concept families.
Figures & tables
Figure 1 : Recovered orbits over SAE features (left) and orbit recovery via two aligned SAEs (right). The i -th feature in SAE-2 is the one-semitone-up version of the i -th feature in SAE-1. We pair each feature i in SAE-2 with its semantically closest counterpart in SAE-1, measured by decoder vector similarity. This semantically matched feature in SAE-1 is defined to be the successor of the original i -th feature. Iteratively applying this successor relation recovers ordered pitch orbits.
Figure 2 : Activations of MuQ SAE features extracted from different-key songs, organized according to recovered orbits. The horizontal axis is song time with annotated chord roots. The vertical axis is feature indices. We append the the piano-roll ground truth as the bottom row. Vertical grid lines separate chord segments, and horizontal grid lines separate different orbits. Four 12-sized orbits are shown to correspond with the chord progression and one long chain mimics the melody.
Root WCR ↑
MajMin WCR ↑
Models
Representations
POP909
RWC
Slakh2100
POP909
RWC
Slakh2100
s-SAE
MusicFM (4)
71.08
67.56
70.32
65.93
64.68
68.39
(Est. Bd)
MuQ (2)
77.42
78.08
78.10
73.11
76.92
76.53
o-SAE
MusicFM (4)
74.01
69.65
75.35
70.09
67.68
73.99
(Est. Bd)
MuQ (2)
80.03
79.79
78.45
73.68
77.40
75.61
s-SAE
MusicFM (4)
80.31
78.43
79.85
75.14
75.82
79.67
Table 1: Weighted Chord Recall (WCR) across Different Models.
Figure 3 : Key detection accuracy comparing probing with standard SAE. Given labels, probing yields generally high accuracy regardless of base model and concept nature. In contrast, SAE uncovers high-contrast but consistent results that are specific to music concept.
MIREX Accuracy ↑
Accuracy ↑
Methods
Representations
FMAKv2
Gtzan
GS
FMAKv2
Gtzan
GS
Probe
MusicFM (2)
62.75
60.29
65.98
53.26
49.76
56.95
MuQ (2)
66.13
63.22
69.70
56.90
52.89
61.92
s-SAE
MusicFM (2)
68.89
63.76
63.96
60.42
55.66
54.47
MuQ (2)
68.87
66.63
69.34
61.08
58.92
62.42
o-SAE
MusicFM (2)
69.89
64.40
61.76
61.76
56.51
52.32
Table 3: Key Detection Results Derived from Chord-based Features.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4 : The Circle of Fifths. A geometric representation of the relationships among the 12 chromatic pitch classes, harmonic functions, and relative key signatures arranged by intervals of perfect fifths. Outer circle labels represent major keys; inner circle labels represent relative minor keys.
Dataset
Usage
Audio domain
Tracks
Slakh2100
Training
Synthesized, instrumental
326
Slakh2100
Validation
Synthesized, instrumental
164
FMAKv2
Key detection
Real recordings
5,434
GTZAN-Key
Key detection
Real recordings
830
GiantSteps
Key detection
Real recordings
604
POP909
Chord recognition
Real recordings
810
Appendix
Table 4 : Overview of the datasets used in our experiments. The reported numbers correspond to tracks retained after label-validity filtering.
Dataset
Major
Minor
Total
FMAKv2
2,777
2,657
5,434
GTZAN-Key
357
473
830
GiantSteps
93
511
604
Appendix
Table 5 : Major–minor distribution of the key-detection datasets.
Genre
Tracks
Genre
Tracks
Blues
98
Jazz
79
Country
96
Metal
93
Disco
97
Pop
94
Hip-hop
79
Reggae
97
Rock
97
Total
830
Appendix
Table 6 : Genre distribution of GTZAN tracks.
Figure 5 : Activations of the selected silence and chord-boundary features. Both features are identified from a few music samples; in batched experiments, we approximate the same manual procedure using a small set of labeled anchors.
Figure 6 : The recovery of three core orbits across different topK settings and different-layer representations under different τ .
Figure 7 : Directed orbit graphs recovered from MuQ at K=32 with ne=10 and varying τ . Fixed points and components with fewer than three nodes are omitted, and “forth” refers to the subdominant.
Figure 8 : Directed orbit graphs recovered from MusicFM at K=32 with ne=10 and varying τ . Fixed points and components with fewer than three nodes are omitted.
Figure 9 : t-SNE projections of normalized SAE decoder vectors across different MuQ layers at K=32 . The columns correspond to different layers; the top row shows o-SAE, and the bottom row shows s-SAE. Black, red, and green points denote features in the subdominant, major-chord, and minor-chord orbits, respectively, while blue points denote the remaining active features.
Figure 10 : t-SNE visualization of the subdominant features across MuQ layers at K=32 . The top row shows o-SAE and the bottom row shows s-SAE. Directed edges follow the orbit order, and labels denote pitch classes.
Figure 11 : Frame-level activations of recovered MuQ (2) orbits on controlled synthetic excerpts in C major and A minor, rendered with piano, strings, brass, and woodwinds, using K=32 and τ=0.5 . The comparison isolates the effects of key and mode across different instrument families.
Figure 12 : Frame-level activations of all recovered MuQ (2) orbits with more than 10 features for three POP909 excerpts, using K=32 and τ=0.5 . The horizontal axis shows time with chord annotations, while feature rows are grouped by recovered orbit. All excerpts are real multi-instrument vocal recordings with complex arrangements rather than synthetic audio.
Figure 13 : Chord recognition results.
Figure 14 : Layer-wise key detection results derived from major- and minor-chord features.
Figure 15 : Effect of Top- K sparsity on layer-wise chord-recognition performance with Oracle Bd . Results are shown for K∈{32,48,64} , with the linear probe included as a reference.
Figure 16 : Effect of Top- K sparsity on layer-wise key-detection performance derived from chord features. Results are shown for K∈{32,48,64} , with the linear probe included as a reference.
Figure 17 : Effect of Top- K sparsity on layer-wise key-signature detection from subdominant features in MuQ. Results are shown for K∈{32,48,64} , with the linear probe included as a reference.
Table 7: Performance of the model trained on POP909
Figure 18 : Downstream performance across random seeds for MuQ (2) o-SAE at K=48 .
Figure 19 : Directed orbit graphs recovered from MuQ across random seeds at K=48 , ne=10 , and τ=0.5 . Fixed points and components with fewer than three nodes are omitted. All seeds are sampled randomly one-shot.
Figure 20 : Genre-wise key-detection results on GTZAN-Key.
Figure 21 : Activations of songs across different genres.
This study analyzes the intermediate representations of music foundation models (MFMs) and reports the geometric structures used to represent pitch information. By inputting isolated musical notes into trained MFMs and analyzing their principal components, we reveal that the representations form a helical structure reflecting the octave periodicity of pitch. Furthermore, we show that the clarity and geometry of this helical structure vary not only across models but also with the acoustic properties of the input signals. Our analysis provides a novel approach for clarifying the internal mechanisms of MFMs.
Hayato Yagi, Shinnosuke Takamichi, Rin Sato +2
Keio University, Japan · The University of Tokyo, Japan · Waseda University, Japan
Current music similarity models typically compute a single, monolithic score, entangling distinct musical dimensions like melody, rhythm, and timbre. This limits user control and interpretability, making it impossible to execute nuanced queries. We introduce MERIT, a framework for learning disentangled, factor-specific music representations tailored to these three core dimensions. To overcome the lack of isolated musical variations in real-world audio, we use a novel training strategy that uses conditional audio generation and source-separated stems to strongly encourage single-factor variation in training data. Our evaluations demonstrate strong factor-wise disentanglement. Each head responds strongly to its intended perceptual dimension while remaining near chance on the others, a representational property that holds across both the synthetic training domain and independent real-world audio.
Rich internal representations of musical structure are essential for music understanding tasks such as machine-assisted music co-writing, yet self-supervised approaches for symbolic music representation remain underexplored, particularly those that encode the hierarchical multiscale nature of musical structures. We present MIDI-RAE-JEPA, combining a pitch- and time-shift equivariance objective with LeJEPA and a Swin Transformer V2 encoder to learn such hierarchical representations of symbolic music encoded as piano roll images. The time-shift equivariance objective encourages the model to internalize temporal musical relationships. The encoder is trained purely on self-supervised objectives -- including a masked embedding predictor (MEP) -- with collapse prevented via SIGReg. A separate decoder trained on the frozen encoder embeddings achieves reconstruction F1 of 0.995, and a flow matching generative model conditioned on those embeddings produces generations that closely match the pitch register and rhythmic density of the conditioning excerpt, while mismatched conditioning yields unrelated but musically plausible output. Learned representations outperform a Haar scattering transform baseline on a downstream emotion classification task, and embedding distances increase monotonically with pitch and time shift magnitude, confirming measurable equivariance. These results suggest that equivariance-based SSL objectives, combined with sufficient fine-level encoder capacity, provide a viable path toward semantically rich, generatively useful representations of symbolic music.
Scott H. Hawley
Department of Chemistry & Physics, Belmont University, Nashville, TN, USA