Sound morphing, generating intermediate sounds that transition from one sonic identity to another, can be a powerful tool for musical sound design. Existing diffusion-based morphing approaches entangle temporal structure and timbral identity, offering no mechanism to hold one fixed while transforming the other. We present smorph, a training-free guidance framework that preserves how a sound behaves over time while transforming what the sound is, allowing users to morph, for instance from brass to strings at a fixed pitch. We demonstrate across three morphing modes: prompt-to-prompt, audio-to-prompt, and audio-to-audio. Evaluations across diverse datasets show that smorph effectively produces smooth morph trajectories while substantially improving temporal-structure and source preservation over baselines, albeit with more conservative target-ward transformation in some settings. In an exploratory case study, musicians found smorph trajectories to be expressive and playable, suggesting structural anchoring can serve as a productive constraint for instrumental interaction.
Figures & tables
Figure 1 : SoundCanvas by smorph . Four text-or-audio prompts define the corners of a 2D sound morphing space; the (x,y) cursor selects a position in the precomputed N=25 sample grid, and real-time bilinear crossfading between neighboring samples provides continuous timbral control. A MIDI keyboard triggers the currently active sound.
Figure 2 : Our structure-guided compositional guidance (See Section 3.1 ). The morph parameter α∈[0,1] interpolates audio-/text- endpoint conditions ( c1 , …, cN ), while a shared structural reference cS anchors the temporal envelope or pitch contour across all morph intermediates, fixed across all T denoising steps.
Figure 3 : Example morph trajectory with explicit structural control (top) and without (bottom). Each cell shows a mel-spectrogram with input (gray) vs generated (orange) RMS curves. Explicit control better preserves the intended temporal envelope across the morph trajectory.
Struct.
Timbral Trajectory
Joint
Qual.
Dataset
Method
Sctrl↑
Smtgt↑
ΔFLAM↑
Sctrl×Smtgt↑
KAD ↓
NSynth-NSynth fixed-pitch notes CTRL=Pitch
Crossfade
0.048
0.324
0.155
0.016
13.602
CoA (no ctrl)
0.012
0.442
0.154
0.005
14.058
smorph i (ours)
0.140
0.523
0.188
0.073
11.452
smorph e (ours)
0.094
0.307
0.130
0.029
12.587
Clotho-Clotho dynamic events CTRL=RMS
Crossfade
0.415
0.653
0.408
0.271
4.974
Table 1 : RQ1: Prompt-to-Prompt morphing with structural control: smorph variants outperform CoA [ 11 ] , but effects differ based on dataset settings.
Struct. & Src. Pres
Timbral Trajectory
Qual.
Dataset
Method
Sctrl↑
mel-L1 ↓
Sm tgt ↑
ΔFLAM↑
KAD ↓
NSynth-NSynth CTRL=Pitch
SDEdit
0.171
3.070
0.501
0.227
10.452
smorph (ours)
0.225
2.623
0.500
0.198
10.279
NSynth-Clotho CTRL = Pitch
SDEdit
0.144
3.542
0.741
0.540
4.158
smorph (ours)
0.187
3.053
0.673
0.466
5.579
Clotho-Clotho CTRL=RMS
SDEdit
0.356
1.518
0.695
0.392
4.420
Table 2 : RQ2: Audio-to-Prompt morphing. Structural control improves source preservation over SDEdit [ 15 ] , with a tradeoff in target transformation and quality.
Struct. & Src. Pres
Timbral Trajectory
Qual.
Dataset
Method
Sctrl↑
mel-L1 ↓
Sm tgt ↑
ΔFLAM↑
KAD ↓
NSynth-NSynth CTRL=Pitch
FLAM → FLAM
0.120
3.766
0.112
0.056
11.095
SDEdit+FLAM
0.204
2.772
0.150
0.080
10.173
NSynth-Clotho CTRL=Pitch
FLAM → FLAM
0.057
4.722
0.623
0.287
4.520
SDEdit+FLAM
0.159
3.429
0.589
0.271
4.485
Clotho-Clotho CTRL=RMS
FLAM → FLAM
0.666
1.177
0.489
0.256
6.092
Table 3 : RQ3: Audio-to-Audio morphing. Partial denoising improves preservation over FLAM-only conditioning, while trajectory & structure metrics vary by setting.
Morphing has recently gained renewed interest with the emergence of generative models, particularly in audio and image generation. In musical sound synthesis, morphing can generate intermediate sounds between two targets, helping musicians and sound engineers explore new sounds with interesting perceptual properties. As morphing is inherently defined in perceptual terms, evaluating this task is challenging. In this work, we introduce Sobolev Distances to Ideal Morphing (SDIM), a novel objective metric to quantify the regularity of audio morphing trajectories in perceptually relevant audio embedding spaces. Leveraging a physics-based sound synthesizer, we evaluate the discriminative power of SDIM on controlled morphing trajectories with varying degrees of regularity and compare it with that of existing audio morphing metrics. Results show that, contrary to state-of-the-art metrics, the proposed metric reliably discriminates desirable trajectories from adversarial ones.
Théo Chasle Cauchy, Modan Tailleur, Barbara Pascal +2
Nantes Universit´e, ´Ecole Centrale Nantes, CNRS, LS2N, UMR6004, F-44000 Nantes, France. · Arturia, Montbonnot Saint-Martin, France
Neural audio codecs were originally developed for high-fidelity compression; however, their latent token representations and expressive decoders also constitute a powerful substrate for controllable audio transformation. This work introduces Neural Morphing, a training-free token-domain audio effect that selects residual-vector-quantized (RVQ) token grains from a user palette and decodes the edited stream through a pretrained codec. The method combines an RVQ-group transfer policy that separates coarse, middle, and fine codebook groups with a continuity-constrained sequence matcher that replaces independent greedy selection with bounded beam search. The intended output is a controlled hybrid: the source preserves rhythmic organization while the palette contributes timbral color and residual detail. We focus on the implementation and realtime behavior of a deployable VST3/AU system, including chunked rendering, palette-size scaling, and backend health checks.
Emmanouil Karystinaios
Institute of Computational Perception Johannes Kepler University Linz, Austria
Zero-shot text-guided editing of real-world music recordings requires balancing semantic modification with faithful preservation of the original musical structure. Although recent diffusion transformers trained with rectified flow have achieved remarkable success in text-to-music generation, extending them to edit existing recordings remains challenging because editing requires accurate deterministic inversion, reliable structural preservation, and numerically stable integration throughout the inversion and generation processes. We present FlowSonic, a zero-shot music editing framework built upon a pretrained diffusion transformer trained with rectified flow. FlowSonic first deterministically inverts a real-world recording into the latent space and preserves its musical structure during editing by reusing cross-attention representations extracted during inversion. To improve the numerical reliability of inversion-based editing, we introduce a high-order ODE solver and systematically investigate how different numerical integration schemes influence trajectory stability, structural preservation, and semantic controllability. Comprehensive experiments on timbre-transfer and genre-modification tasks demonstrate that FlowSonic consistently outperforms existing music editing methods across semantic alignment, harmonic preservation, structural consistency, and perceptual audio quality. We further provide geometric and empirical analyses showing how the proposed numerical integration strategy improves latent trajectory stability and leads to more reliable music editing.