How can we understand what a music foundation model has learned \textit{internally}? Most interpretability approaches, such as probing and Sparse Autoencoders (SAEs), focus on identifying individual features with minimal structural assumptions. We argue that many concepts are better understood as \textit{structured relations} rather than isolated features. This is especially prominent in music, where tonal structures are organized in the space of pitch and time. For example, concepts such as chords or keys are naturally expressed as structured sets (e.g., the 12 transpositions of a chord or the diatonic system within a key), rather than isolated features. In this study, \textbf{we shift from feature identification to structure-based analysis}, asking whether the learned inner representations of music foundation model emerge as organized structures over features. To this end, we introduce a framework that uses pitch transposition as an inductive bias to induce ordered orbits via multi-view SAE alignment. Concretely, we generate pitch-shifted input pairs and align their SAE representations to discover structured groups of pitch-related features. Experimental results show that this approach recovers orbit structures corresponding to chords, keys, and melodic patterns across two state-of-the-art music foundation models, while requiring only minimal grounding (e.g., a few anchor examples) to interpret entire concept families.
Figures & tables
Figure 1 : Recovered orbits over SAE features (left) and orbit recovery via two aligned SAEs (right). The i -th feature in SAE-2 is the one-semitone-up version of the i -th feature in SAE-1. We pair each feature i in SAE-2 with its semantically closest counterpart in SAE-1, measured by decoder vector similarity. This semantically matched feature in SAE-1 is defined to be the successor of the original i -th feature. Iteratively applying this successor relation recovers ordered pitch orbits.
Figure 2 : Activations of MuQ SAE features extracted from different-key songs, organized according to recovered orbits. The horizontal axis is song time with annotated chord roots. The vertical axis is feature indices. We append the the piano-roll ground truth as the bottom row. Vertical grid lines separate chord segments, and horizontal grid lines separate different orbits. Four 12-sized orbits are shown to correspond with the chord progression and one long chain mimics the melody.
Root WCR ↑
MajMin WCR ↑
Models
Representations
POP909
RWC
Slakh2100
POP909
RWC
Slakh2100
s-SAE
MusicFM (4)
71.08
67.56
70.32
65.93
64.68
68.39
(Est. Bd)
MuQ (2)
77.42
78.08
78.10
73.11
76.92
76.53
o-SAE
MusicFM (4)
74.01
69.65
75.35
70.09
67.68
73.99
(Est. Bd)
MuQ (2)
80.03
79.79
78.45
73.68
77.40
75.61
s-SAE
MusicFM (4)
80.31
78.43
79.85
75.14
75.82
79.67
Table 1: Weighted Chord Recall (WCR) across Different Models.
Figure 3 : Key detection accuracy comparing probing with standard SAE. Given labels, probing yields generally high accuracy regardless of base model and concept nature. In contrast, SAE uncovers high-contrast but consistent results that are specific to music concept.
MIREX Accuracy ↑
Accuracy ↑
Methods
Representations
FMAKv2
Gtzan
GS
FMAKv2
Gtzan
GS
Probe
MusicFM (2)
62.75
60.29
65.98
53.26
49.76
56.95
MuQ (2)
66.13
63.22
69.70
56.90
52.89
61.92
s-SAE
MusicFM (2)
68.89
63.76
63.96
60.42
55.66
54.47
MuQ (2)
68.87
66.63
69.34
61.08
58.92
62.42
o-SAE
MusicFM (2)
69.89
64.40
61.76
61.76
56.51
52.32
Table 3: Key Detection Results Derived from Chord-based Features.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4 : The Circle of Fifths. A geometric representation of the relationships among the 12 chromatic pitch classes, harmonic functions, and relative key signatures arranged by intervals of perfect fifths. Outer circle labels represent major keys; inner circle labels represent relative minor keys.
Dataset
Usage
Audio domain
Tracks
Slakh2100
Training
Synthesized, instrumental
326
Slakh2100
Validation
Synthesized, instrumental
164
FMAKv2
Key detection
Real recordings
5,434
GTZAN-Key
Key detection
Real recordings
830
GiantSteps
Key detection
Real recordings
604
POP909
Chord recognition
Real recordings
810
Appendix
Table 4 : Overview of the datasets used in our experiments. The reported numbers correspond to tracks retained after label-validity filtering.
Dataset
Major
Minor
Total
FMAKv2
2,777
2,657
5,434
GTZAN-Key
357
473
830
GiantSteps
93
511
604
Appendix
Table 5 : Major–minor distribution of the key-detection datasets.
Genre
Tracks
Genre
Tracks
Blues
98
Jazz
79
Country
96
Metal
93
Disco
97
Pop
94
Hip-hop
79
Reggae
97
Rock
97
Total
830
Appendix
Table 6 : Genre distribution of GTZAN tracks.
Figure 5 : Activations of the selected silence and chord-boundary features. Both features are identified from a few music samples; in batched experiments, we approximate the same manual procedure using a small set of labeled anchors.
Figure 6 : The recovery of three core orbits across different topK settings and different-layer representations under different τ .
Figure 7 : Directed orbit graphs recovered from MuQ at K=32 with ne=10 and varying τ . Fixed points and components with fewer than three nodes are omitted, and “forth” refers to the subdominant.
Figure 8 : Directed orbit graphs recovered from MusicFM at K=32 with ne=10 and varying τ . Fixed points and components with fewer than three nodes are omitted.
Figure 9 : t-SNE projections of normalized SAE decoder vectors across different MuQ layers at K=32 . The columns correspond to different layers; the top row shows o-SAE, and the bottom row shows s-SAE. Black, red, and green points denote features in the subdominant, major-chord, and minor-chord orbits, respectively, while blue points denote the remaining active features.
Figure 10 : t-SNE visualization of the subdominant features across MuQ layers at K=32 . The top row shows o-SAE and the bottom row shows s-SAE. Directed edges follow the orbit order, and labels denote pitch classes.
Figure 11 : Frame-level activations of recovered MuQ (2) orbits on controlled synthetic excerpts in C major and A minor, rendered with piano, strings, brass, and woodwinds, using K=32 and τ=0.5 . The comparison isolates the effects of key and mode across different instrument families.
Figure 12 : Frame-level activations of all recovered MuQ (2) orbits with more than 10 features for three POP909 excerpts, using K=32 and τ=0.5 . The horizontal axis shows time with chord annotations, while feature rows are grouped by recovered orbit. All excerpts are real multi-instrument vocal recordings with complex arrangements rather than synthetic audio.
Figure 13 : Chord recognition results.
Figure 14 : Layer-wise key detection results derived from major- and minor-chord features.
Figure 15 : Effect of Top- K sparsity on layer-wise chord-recognition performance with Oracle Bd . Results are shown for K∈{32,48,64} , with the linear probe included as a reference.
Figure 16 : Effect of Top- K sparsity on layer-wise key-detection performance derived from chord features. Results are shown for K∈{32,48,64} , with the linear probe included as a reference.
Figure 17 : Effect of Top- K sparsity on layer-wise key-signature detection from subdominant features in MuQ. Results are shown for K∈{32,48,64} , with the linear probe included as a reference.
Table 7: Performance of the model trained on POP909
Figure 18 : Downstream performance across random seeds for MuQ (2) o-SAE at K=48 .
Figure 19 : Directed orbit graphs recovered from MuQ across random seeds at K=48 , ne=10 , and τ=0.5 . Fixed points and components with fewer than three nodes are omitted. All seeds are sampled randomly one-shot.
Figure 20 : Genre-wise key-detection results on GTZAN-Key.
Figure 21 : Activations of songs across different genres.