Abstract
Recent work on synthesizer inversion shows that generative models outperform deterministic approaches by explicitly modeling the ambiguity in mapping audio to parameters. Training such models, however, requires audio-parameter pairs, which are typically obtained by rendering sampled or preset parameters through the synthesizer itself. This creates a train-test mismatch that can degrade performance on off-manifold real-world recordings, for which ground-truth parameter annotations do not exist. To circumvent this obstacle, we propose to model the joint distribution of audio and parameters with a multi-modal continuous normalizing flow using independent noise schedules for each modality. This formulation allows us to train joint and conditional densities with paired synthesizer data, while unpaired real recordings can train the audio marginal alone, exposing the model to off-manifold signals without requiring parameter labels. Further, because the model learns to map from audio to parameters at all noise levels, we find that partially noising the audio reference at inference improves real-audio reconstruction, consistent with reducing sensitivity to distribution-specific detail while preserving coarse structure. Evaluating on Surge XT and Dexed, we find that modelling the joint distribution substantially improves both inversion of real-world and in-domain audio.
Explore similar work
Aug 4, 2026cs.SD
Synthesizer inversion is challenging for two main reasons: 1) Distinct parameter configurations can produce perceptually similar sounds. 2) Parameter-space losses often fail to reflect rendered audio similarity, while the synthesizer being a non-differentiable black box prevents simple audio-domain supervision. To address the one-to-many mapping induced by the first challenge, we formulate synthesizer inversion as conditional generation over discrete synthesizer parameters and use masked discrete diffusion as the generator. This treatment additionally avoids the fixed-order assumption of autoregressive models and the continuous-relaxation mismatch of flow matching when modeling categorical synthesizer controls. To address the second challenge, we further fine-tune the model with GRPO-style audio-domain rewards computed from rendered outputs. Experiments on Dexed show that, after supervised training, the discrete diffusion model is competitive with autoregressive and flow-matching baselines, and reward-based fine-tuning further improves out-of-domain audio matching performance. Code and demos are available at: https://github.com/DDSynth-RL/DDSynthRL.
Tristan Wu, Daniel Chin, Junan Zhang +3
Computational Media and Art, The Hong Kong University of Science and Technology (Guangzhou) · New York University Shanghai · The Chinese University of Hong Kong, Shenzhen +2
Aug 4, 2026cs.SD
Existing reference-free methods for evaluating music perceptual quality alleviate the need for paired noisy-clean data, but they still rely on a background set, which is used to compute aggregated statistics of clean audio samples. In this work, we propose a novel approach that eliminates this requirement, achieving background-set-free and reference-free quality estimation using only a pre-trained Flow Matching backbone. We demonstrate that unconditional Flow Matching inversion via simple Euler integration is sufficient to detect various artificial distortions and accurately rank music generation models against human perceptual judgments. We introduce InvFlowFD, which performs flow inversion and compares a group of inverted samples to the prior distribution. We evaluate our method against prior work, quantitatively and with a thorough human study. Results suggest that InvFlowFD is highly correlated with human perception of sound distortions, as well as generative models' quality, while being more flexible and less restrictive than existing metrics.
Alon Ziv, Harel Pogoda, Yossi Adi
School of Computer Science and Engineering The Hebrew University of Jerusalem, Israel
Sep 23, 2026eess.AS
Text-to-audio (TTA) generation aims to synthesize realistic audio that faithfully reflects natural-language descriptions. Most TTA systems adopt a two-stage latent paradigm: an audio tokenizer is optimized for reconstruction and then frozen, after which a generative model is trained in the resulting latent space. However, reconstruction-oriented representations may be suboptimal for generation, motivating joint representation and generative learning. To this end, we introduce Unite-Audio, to our knowledge, is the first to jointly learn continuous audio representations and latent flow matching for TTA. By coupling reconstruction with self-supervised generative prediction, Unite-Audio allows the generative objective to directly shape the latent space rather than treating it as a fixed intermediate representation. We further employ Flow-GRPO post-training to improve text-conditioned generation. Experiments show competitive TTA performance with a compact latent flow model, while ablation studies confirm the benefit of jointly learning the audio representation and generative model. Audio samples are available at https://runwushi.github.io/Unite-Audio.
Runwu Shi, Kai Li, Yujin Wang +8
Institute of Science Tokyo · Tsinghua University · Wuhan University +2