Sound morphing, generating intermediate sounds that transition from one sonic identity to another, can be a powerful tool for musical sound design. Existing diffusion-based morphing approaches entangle temporal structure and timbral identity, offering no mechanism to hold one fixed while transforming the other. We present smorph, a training-free guidance framework that preserves how a sound behaves over time while transforming what the sound is, allowing users to morph, for instance from brass to strings at a fixed pitch. We demonstrate across three morphing modes: prompt-to-prompt, audio-to-prompt, and audio-to-audio. Evaluations across diverse datasets show that smorph effectively produces smooth morph trajectories while substantially improving temporal-structure and source preservation over baselines, albeit with more conservative target-ward transformation in some settings. In an exploratory case study, musicians found smorph trajectories to be expressive and playable, suggesting structural anchoring can serve as a productive constraint for instrumental interaction.
Figures & tables
Figure 1 : SoundCanvas by smorph . Four text-or-audio prompts define the corners of a 2D sound morphing space; the (x,y) cursor selects a position in the precomputed N=25 sample grid, and real-time bilinear crossfading between neighboring samples provides continuous timbral control. A MIDI keyboard triggers the currently active sound.
Figure 2 : Our structure-guided compositional guidance (See Section 3.1 ). The morph parameter α∈[0,1] interpolates audio-/text- endpoint conditions ( c1 , …, cN ), while a shared structural reference cS anchors the temporal envelope or pitch contour across all morph intermediates, fixed across all T denoising steps.
Figure 3 : Example morph trajectory with explicit structural control (top) and without (bottom). Each cell shows a mel-spectrogram with input (gray) vs generated (orange) RMS curves. Explicit control better preserves the intended temporal envelope across the morph trajectory.
Struct.
Timbral Trajectory
Joint
Qual.
Dataset
Method
Sctrl↑
Smtgt↑
ΔFLAM↑
Sctrl×Smtgt↑
KAD ↓
NSynth-NSynth fixed-pitch notes CTRL=Pitch
Crossfade
0.048
0.324
0.155
0.016
13.602
CoA (no ctrl)
0.012
0.442
0.154
0.005
14.058
smorph i (ours)
0.140
0.523
0.188
0.073
11.452
smorph e (ours)
0.094
0.307
0.130
0.029
12.587
Clotho-Clotho dynamic events CTRL=RMS
Crossfade
0.415
0.653
0.408
0.271
4.974
Table 1 : RQ1: Prompt-to-Prompt morphing with structural control: smorph variants outperform CoA [ 11 ] , but effects differ based on dataset settings.
Struct. & Src. Pres
Timbral Trajectory
Qual.
Dataset
Method
Sctrl↑
mel-L1 ↓
Sm tgt ↑
ΔFLAM↑
KAD ↓
NSynth-NSynth CTRL=Pitch
SDEdit
0.171
3.070
0.501
0.227
10.452
smorph (ours)
0.225
2.623
0.500
0.198
10.279
NSynth-Clotho CTRL = Pitch
SDEdit
0.144
3.542
0.741
0.540
4.158
smorph (ours)
0.187
3.053
0.673
0.466
5.579
Clotho-Clotho CTRL=RMS
SDEdit
0.356
1.518
0.695
0.392
4.420
Table 2 : RQ2: Audio-to-Prompt morphing. Structural control improves source preservation over SDEdit [ 15 ] , with a tradeoff in target transformation and quality.
Struct. & Src. Pres
Timbral Trajectory
Qual.
Dataset
Method
Sctrl↑
mel-L1 ↓
Sm tgt ↑
ΔFLAM↑
KAD ↓
NSynth-NSynth CTRL=Pitch
FLAM → FLAM
0.120
3.766
0.112
0.056
11.095
SDEdit+FLAM
0.204
2.772
0.150
0.080
10.173
NSynth-Clotho CTRL=Pitch
FLAM → FLAM
0.057
4.722
0.623
0.287
4.520
SDEdit+FLAM
0.159
3.429
0.589
0.271
4.485
Clotho-Clotho CTRL=RMS
FLAM → FLAM
0.666
1.177
0.489
0.256
6.092
Table 3 : RQ3: Audio-to-Audio morphing. Partial denoising improves preservation over FLAM-only conditioning, while trajectory & structure metrics vary by setting.