cs.SDMay 5, 2026

PHALAR: Phasors for Learned Musical Audio Representations

Authors: Davide MarincioneMichele MancusiGiorgio StranoLuca CerovazDonato CrisostomiRoberto RibuoliEmanuele Rodolà

Organizations: Department of Computer Science, Sapienza University of Rome, Italy · Moises Systems, Inc. · Paradigma, Inc.

Abstract

Stem retrieval, the task of matching missing stems to a given audio submix, is a key challenge currently limited by models that discard temporal information. We introduce PHALAR, a contrastive framework achieving a relative accuracy increase of up to 70%\approx 70\% over the state-of-the-art while requiring <50%<50\% of the parameters and a 7×\times training speedup. By utilizing a Learned Spectral Pooling layer and a complex-valued head, PHALAR enforces pitch-equivariant and phase-equivariant biases. PHALAR establishes new retrieval state-of-the-art across MoisesDB, Slakh, and ChocoChorales, correlating significantly higher with human coherence judgment than semantic baselines. Finally, zero-shot beat tracking and linear chord probing confirm that PHALAR captures robust musical structures beyond the retrieval task.

Explore similar work

CardsList
  1. FIGMA: Towards FIne-Grained Music retrievAl

    Jun 4, 2026Nishit Anand, Ashish Seth, Sreyan Ghosh +2Music Understanding