cs.SDOct 6, 2026

Geometric Representations for Transformed Pattern Matching in Music

Authors: David Meredith

Organizations: Aalborg University, Aalborg, Denmark

Abstract

We review the notion of representing music using point sets and argue that such representations are better adapted than sequential representations for matching patterns in unvoiced, polyphonic music, such as keyboard music. Musical transformations such as transposition, inversion, diminution, augmentation and retrograde can be modelled by geometric transformations in pitch-time representations that combine translation with scaling parallel to and reflection in the time axis. We identify eight types of geometric pitch-time representation that use chromatic pitch, morphetic pitch, morph or chroma to represent pitch and either onset time or midtime to represent time. We illustrate how these types of representation allow us to characterise different types of musical transformation. For example, by using midtime instead of onset time, we can precisely characterise certain retrograde relationships; and by using morph and chroma pitch representations we can characterise transformations involving octave displacements and duplications. We present the concept of a transformation class and consider the three specific classes, F2STRF_{\mathrm{2STR}}, F2STRMod7F_{\mathrm{2STRMod7}} and F2STRMod12F_{\mathrm{2STRMod12}}. We introduce the notion of an inter-pattern transformation graph for a set of patterns, SS, and a transformation class, FF. Each vertex in such a graph represents a pattern in SS and there is an edge in the graph from pattern P1P_1 to pattern P2P_2 if and only if P1P_1 can be mapped onto P2P_2 by a transformation in FF. We show, with the aid of such graphs, that the musical relationships between the occurrences of the HAYDN theme in Ravel's Menuet sur le nom d'Haydn can be precisely described in terms of transformations in F2STRF_{\mathrm{2STR}}, F2STRMod7F_{\mathrm{2STRMod7}} and F2STRMod12F_{\mathrm{2STRMod12}} within the pitch-time representations considered.

Figures & tables

Explore similar work

Aug 4, 2026cs.SD

Equivariant Music Transformer

Humans recognize a musical passage even when it is shifted in time or transposed in pitch, indicating a notion of equivariance in the representation space. Our analysis, however, shows that standard music transformers map such time-shifted or pitch-transposed inputs onto uncorrelated representations: these models become progressively less equivariant as they scale in size or train longer. This suggests that in standard music transformers, additional model capacity is allocated to memorizing absolute patterns rather than capturing shared musical structures. In this paper, we propose the Equivariant Music Transformer (EMT), which enforces equivariance through self-distillation by jointly optimizing a next-token-prediction and an auxiliary equivariance regularization loss. We find that the additional equivariance loss acts as a beneficial regularizer, simultaneously improving next-token prediction and producing equivariant latent representations. Through both objective and subjective evaluations, EMT demonstrates superior equivariance and generative capability compared to data augmentation, feature engineering, and state-of-the-art (SOTA) baselines. More broadly, our findings reveal that standard language modeling methods alone do not capture music's translational symmetries, and dedicated inductive biases are required to produce better music representations. The code, weights and demos are available online.
Oct 1, 2026cs.SD

From Isolated Feature to Orbits: Discovering Music Concepts via Multi-SAE Alignment

How can we understand what a music foundation model has learned \textit{internally}? Most interpretability approaches, such as probing and Sparse Autoencoders (SAEs), focus on identifying individual features with minimal structural assumptions. We argue that many concepts are better understood as \textit{structured relations} rather than isolated features. This is especially prominent in music, where tonal structures are organized in the space of pitch and time. For example, concepts such as chords or keys are naturally expressed as structured sets (e.g., the 12 transpositions of a chord or the diatonic system within a key), rather than isolated features. In this study, \textbf{we shift from feature identification to structure-based analysis}, asking whether the learned inner representations of music foundation model emerge as organized structures over features. To this end, we introduce a framework that uses pitch transposition as an inductive bias to induce ordered orbits via multi-view SAE alignment. Concretely, we generate pitch-shifted input pairs and align their SAE representations to discover structured groups of pitch-related features. Experimental results show that this approach recovers orbit structures corresponding to chords, keys, and melodic patterns across two state-of-the-art music foundation models, while requiring only minimal grounding (e.g., a few anchor examples) to interpret entire concept families.
May 22, 2026cs.SD

Rubato: Transcribing Piano Music with Timestamps

We consider the conversion of musical recordings into human-readable sheet music annotated with timestamps. Such output lets a listener clearly visualize rubato (temporally expressive playing), a learner diagnose ensemble precision and timing choices against the written music, and a musicology scholar compare performance styles across recordings of the same work. We introduce (1) a prompt-conditioned encoder-decoder model, named Rubato, trained to output (2) a new textual representation for polyphonic music, named InterMo, which we designed for compatibility with sequence-to-sequence training. Our experiments demonstrate that Rubato produces timestamped piano sheet music from audio with higher notational accuracy than the best existing approaches, which are based on cascades. We find that even if the cascade is given ground-truth MIDI instead of audio, Rubato performs better, suggesting that the ceiling of existing approaches is primarily representational, not acoustic. Further, because Rubato is trained on several related tasks (with prompts), it competes with or outperforms the best single-task systems on related but simpler tasks like MIDI note grounding and beat/downbeat detection. A demo is available at https://nctamer.github.io/rubato-transcription .