cs.SDSep 24, 2026

Off-manifold robustness in synthesizer inversion with joint distribution flow matching

Authors: Ben Hayes

Organizations: Sony Computer Science Laboratories Paris

Abstract

Recent work on synthesizer inversion shows that generative models outperform deterministic approaches by explicitly modeling the ambiguity in mapping audio to parameters. Training such models, however, requires audio-parameter pairs, which are typically obtained by rendering sampled or preset parameters through the synthesizer itself. This creates a train-test mismatch that can degrade performance on off-manifold real-world recordings, for which ground-truth parameter annotations do not exist. To circumvent this obstacle, we propose to model the joint distribution of audio and parameters with a multi-modal continuous normalizing flow using independent noise schedules for each modality. This formulation allows us to train joint and conditional densities with paired synthesizer data, while unpaired real recordings can train the audio marginal alone, exposing the model to off-manifold signals without requiring parameter labels. Further, because the model learns to map from audio to parameters at all noise levels, we find that partially noising the audio reference at inference improves real-audio reconstruction, consistent with reducing sensitivity to distribution-specific detail while preserving coarse structure. Evaluating on Surge XT and Dexed, we find that modelling the joint distribution substantially improves both inversion of real-world and in-domain audio.

Figures & tables

Explore similar work

CardsList
  1. DDSynth-RL: Audio Synthesizer Inversion via Discrete Diffusion with Reinforcement Learning

    Aug 4, 2026Tristan Wu, Daniel Chin, Junan Zhang +3SynthesizerModern Generative Audio Models

  2. InvFlowFD: Reference-Free and Background-Set-Free Perceptual Music Quality Metric with Flow Matching Inversion

    Aug 4, 2026Alon Ziv, Harel Pogoda, Yossi AdiPerceptual QualityImage Quality Assessment

  3. UNITE-AUDIO: Joint Learning of Continuous Tokenization and Latent Flow Matching for Text-to-Audio Generation

    Sep 23, 2026Runwu Shi, Kai Li, Yujin Wang +8Modern Generative Audio ModelsAudio Tokenizers