Morphing has recently gained renewed interest with the emergence of generative models, particularly in audio and image generation. In musical sound synthesis, morphing can generate intermediate sounds between two targets, helping musicians and sound engineers explore new sounds with interesting perceptual properties. As morphing is inherently defined in perceptual terms, evaluating this task is challenging. In this work, we introduce Sobolev Distances to Ideal Morphing (SDIM), a novel objective metric to quantify the regularity of audio morphing trajectories in perceptually relevant audio embedding spaces. Leveraging a physics-based sound synthesizer, we evaluate the discriminative power of SDIM on controlled morphing trajectories with varying degrees of regularity and compare it with that of existing audio morphing metrics. Results show that, contrary to state-of-the-art metrics, the proposed metric reliably discriminates desirable trajectories from adversarial ones.
Figures & tables
Figure 1: Morphing trajectories from two sounds s and t in a perceptually correlated embedding space. Ideal linear regular trajectory (solid green), good quality morphing trajectory (solid blue), and random Null trajectory (dotted red).
Figure 2: Adversarial morphing trajectories in a perceptually correlated embedding space. ( NUC , solid orange) Nonuniform Configuration where the angles are preserved but the spacing regularity is not; ( EQC , solid violet) Equidistant Configuration where the distance to the source is the same as in the ideal trajectory but the angle between smi and st is randomly sampled.
Table 1: Normalized SDIM metrics for different embeddings. Normalized SDIM metrics means (standard deviations) are estimated from 1,000LUM trajectories and 1,000Null trajectories.
Metric
LUM
NUC
EQC
Null
Corresp.SM↓
.19 (.02)
.38 (.01)
.14 (.03)
.03 (0.02)
Smooth.MF↑
.96 (.01)
.75 (.01)
.97 (.03)
.11 (0.17)
Interm.SM↓
.23 (.04)
.21 (.04)
1.15 (.24)
5.59 (0.73)
Smooth.SM↓
.02 (.00)
.02 (.00)
.09 (.02)
.52 (0.07)
10−2×S0,2↓
.69 (.09)
1.14 (.20)
1.72 (.32)
8.41 (1.27)
10−3×S1,2↓
.35 (.05)
.69 (.12)
2.59 (.51)
.013 (1.54)
Table 2: Average metrics on the four configurations of trajectories in the MERT embeddings space. Means (standard deviations) estimated from 1000 examples. Only the SDIM metrics (last two rows) are robust to all of the adversarial configurations. Underlined red font indicates failure in measuring the true regularity.
Figure 3: Angle dispersion for the LUM , EQC , and NUC configuration trajectories. Angles are measured through scalar products between the segment connecting the source and the intermediate point and the source and the target (ideal linear trajectory). Histograms are computed over 11 intermediate points across 1000 trajectories.
Sound morphing, generating intermediate sounds that transition from one sonic identity to another, can be a powerful tool for musical sound design. Existing diffusion-based morphing approaches entangle temporal structure and timbral identity, offering no mechanism to hold one fixed while transforming the other. We present smorph, a training-free guidance framework that preserves how a sound behaves over time while transforming what the sound is, allowing users to morph, for instance from brass to strings at a fixed pitch. We demonstrate across three morphing modes: prompt-to-prompt, audio-to-prompt, and audio-to-audio. Evaluations across diverse datasets show that smorph effectively produces smooth morph trajectories while substantially improving temporal-structure and source preservation over baselines, albeit with more conservative target-ward transformation in some settings. In an exploratory case study, musicians found smorph trajectories to be expressive and playable, suggesting structural anchoring can serve as a productive constraint for instrumental interaction.
Annie Chu, Hugo Flores García, Johannes Imort +5
Northwestern University, Evanston, IL, USA · Adobe Research, San Francisco, CA, USA · Wizdom Music, New York, NY, USA
Neural audio codecs were originally developed for high-fidelity compression; however, their latent token representations and expressive decoders also constitute a powerful substrate for controllable audio transformation. This work introduces Neural Morphing, a training-free token-domain audio effect that selects residual-vector-quantized (RVQ) token grains from a user palette and decodes the edited stream through a pretrained codec. The method combines an RVQ-group transfer policy that separates coarse, middle, and fine codebook groups with a continuity-constrained sequence matcher that replaces independent greedy selection with bounded beam search. The intended output is a controlled hybrid: the source preserves rhythmic organization while the palette contributes timbral color and residual detail. We focus on the implementation and realtime behavior of a deployable VST3/AU system, including chunked rendering, palette-size scaling, and backend health checks.
Emmanouil Karystinaios
Institute of Computational Perception Johannes Kepler University Linz, Austria
We introduce Spot, Separate, and Enhance (SSE), the first multimodal, user-guided generative model for audio remixing and enhancement. SSE enhances video content by rebalancing the audio, removing unwanted audio sources, and reducing reverberation, guided by both video and textual descriptions. To support its training and evaluation, we propose DegradedMix, a new dataset built on the audio remixing benchmark MuddyMix. We also adopt evaluation metrics from generative modeling, which better capture the creative nature of remixing than standard reconstruction-based metrics. SSE outperforms existing baselines in both controllability and remixing quality, as shown by extensive experiments. Project page: https://sse-ai.notion.site