AutoSynth: Learning to Generate Editable Synthesizer Programs from Audio and Text
Authors: Tristan Wu, Daniel Chin, Liwei Lin, Junan Zhang, Gus Xia
Organizations: The Hong Kong University of Science and Technology (Guangzhou) · New York University Shanghai · The Chinese University of Hong Kong, Shenzhen · Mohamed bin Zayed University of Artificial Intelligence
Audio generation models can translate natural-language descriptions into sound, but their outputs are typically waveforms. Their audio quality is constrained by audio compression, and their outputs do not readily support direct edits to notes, timbral parameters, or modulation relationships. We present AutoSynth, which represents MIDI performance events, fixed synthesizer parameters, and variable-length modulation routes as a unified sequence for a synthesizer, and learns their dependencies with an audio-conditioned autoregressive model. A single model supports both tasks. Given reference audio, the model directly predicts a synthesizer program; given text, it uses a pretrained audio generation model and converts the generated audio into a program. Training consists of two stages: supervised learning on large-scale audio-program pairs automatically constructed from a small set of native presets, followed by group-relative policy optimization with a mixed reward combining semantic similarity, pitch-related features, acoustic similarity, and sound usefulness. The pipeline requires neither paired text-target-program annotations nor a differentiable synthesizer. Experiments show that AutoSynth produces complete, editable synthesizer programs and achieves competitive results in both synthesizer inversion and text-driven generation. Audio demos and source code are available at https://auto-synth.github.io/.
Figures & tables
Figure 1: Overview of AutoSynth.
Figure 2: Autoregressive program generation and decoder architecture.
System
Audio CLAP ↓
LanguageBind ↓
wMFCC ↓
SOT ↓
RMS ↑
MSS ↓
Gemini–Woosh
AutoSynth
0.37
0.397
10.22
0.15
0.82
6.94
Synth Permutations
0.54
0.584
22.75
0.17
0.80
9.99
DDSynth-RL
0.49
0.482
10.34
0.16
0.87
7.04
FSD50K
AutoSynth
0.67
0.649
14.95
0.16
0.632
13.00
Table 1: Synthesizer inversion on Gemini–Woosh and FSD50K. Bold denotes the best value for each metric within each test set.
System
Text CLAP ↓
LanguageBind ↓
CE ↑
CU ↑
PQ ↑
FAD ↓
KL ↓
Gemini
AutoSynth
0.550
0.638
3.759
7.23
7.433
—
—
Synth Permutations
0.642
0.718
3.409
6.88
7.372
—
—
DDSynth-RL
0.607
0.678
3.253
6.81
7.276
—
—
CTAG
0.519
0.645
3.020
6.58
6.888
—
—
Stable Audio 3 ∗
0.546
0.614
3.891
7.16
7.373
—
—
Table 2: Text-driven generation on Gemini and FSD50K descriptions, corresponding to the audio in Table 1 .
Figure 3: Subjective synthesizer-inversion evaluation on Gemini–Woosh (left) and FSD50K (right). Error bars indicate 95% bootstrap confidence intervals for the mean.
Waveform compression frontend, 6 backbone blocks of width 768
53.73M
Yes
SAME-L
Waveform compression frontend, 12 backbone blocks of width 1536
426.06M
Yes
Appendix
Table A1: Encoder sizes. Counts exclude the program decoder and the SAME waveform decoder.
Figure A1: Validation curves for three encoders across two training stages. Left: in-domain SFT loss, shown up to 500,000 updates. Right: Stable Audio 3 validation reward. Horizontal axes count optimizer updates within each stage.
Condition
PQ ↑
CLAP distance ↓
Uncorrupted Woosh
7.18
0.000
Uncorrupted input → AutoSynth
7.43
0.372
Noisy input
5.07
0.306
Noisy input → AutoSynth
7.40
0.451
Low-pass input
7.10
0.134
Low-pass input → AutoSynth
7.34
0.375
Appendix
Table B1: Corrupted-input inversion and resynthesis, averaged over all 128 clips. CLAP distance uses the uncorrupted Woosh clip as reference; lower is better. Higher PQ is better.
Test set
Audio CLAP ↓
LanguageBind ↓
wMFCC ↓
SOT ↓
RMS ↑
MSS ↓
Gemini–Woosh
0.47 → 0.37
0.535 → 0.397
17.05 → 10.22
0.33 → 0.15
0.65 → 0.82
19.41 → 6.94
FSD50K
0.73 → 0.67
0.746 → 0.649
22.05 → 14.95
0.31 → 0.16
0.472 → 0.632
24.33 → 13.00
Helm in-domain
0.17 → 0.13
0.221 → 0.174
8.02 → 6.93
0.057 → 0.048
0.91 → 0.92
7.36 → 3.77
Appendix
Table C1: SFT/GRPO comparison for synthesizer inversion.
Test set
Text CLAP ↓
LanguageBind ↓
CE ↑
CU ↑
PQ ↑
FAD ↓
KL ↓
Gemini
0.601 → 0.550
0.671 → 0.638
3.469 → 3.759
6.92 → 7.23
7.278 → 7.433
—
—
FSD50K
0.768 → 0.771
0.776 → 0.793
3.24 → 3.48
6.73 → 6.91
6.97 → 7.07
8.83 → 7.00
5.04 → 5.27
Appendix
Table C2: SFT/GRPO comparison for text-driven generation.
Figure D1: In-domain validation loss for different autoregressive orders over the first 100,000 steps.
Table D1: Autoregressive order of fixed parameters, read left to right within each row.
Item
Setting
Audio input
44.1 kHz, 5 seconds, stereo
Mel spectrograms
128 bands per channel, 25 ms Hamming window, 100 frames/s
Encoder / decoder layers
12 / 12
Hidden width / heads / FFN width
512 / 8 / 1024
Normalization and activation
Pre-LayerNorm, GELU, dropout 0.1
Value output layer
Shared Linear(512, 128), with field-specific legal categories
Appendix
Table E1: Main model and optimization settings.
Component
Direction
Coefficient
CLAP distance
Negative
8
Basic Pitch hidden-feature distance
Negative
12
AudioBox CE, normalized to [0,1]
Positive
1
AudioBox CU, normalized to [0,1]
Positive
1
wMFCC
Negative
0.07
SOT
Negative
3.58
Appendix
Table E2: Coefficients used in Equation ( 3 ). All coefficients are positive; each metric’s direction determines its sign in the reward.
Dimension
Question
Evaluation criteria
Alignment: synthesizer inversion
Does this sound faithfully reproduce the reference audio?
Consider timbre, pitch, articulation, and changes over time; 1 means not at all and 5 means very faithfully.
Alignment: text-driven generation
Does this sound faithfully follow the text description?
Consider the described timbre, pitch, articulation, and temporal changes; 1 means not at all and 5 means very faithfully.
Content quality: both tasks
How good is this sound as material for music production or sound design?
Consider audio quality, unintended artifacts, and creative usefulness; 1 means very poor and 5 means excellent. Intentional distortion is not a defect, and quality is rated independently of alignment.
Synthesizer inversion is challenging for two main reasons: 1) Distinct parameter configurations can produce perceptually similar sounds. 2) Parameter-space losses often fail to reflect rendered audio similarity, while the synthesizer being a non-differentiable black box prevents simple audio-domain supervision. To address the one-to-many mapping induced by the first challenge, we formulate synthesizer inversion as conditional generation over discrete synthesizer parameters and use masked discrete diffusion as the generator. This treatment additionally avoids the fixed-order assumption of autoregressive models and the continuous-relaxation mismatch of flow matching when modeling categorical synthesizer controls. To address the second challenge, we further fine-tune the model with GRPO-style audio-domain rewards computed from rendered outputs. Experiments on Dexed show that, after supervised training, the discrete diffusion model is competitive with autoregressive and flow-matching baselines, and reward-based fine-tuning further improves out-of-domain audio matching performance. Code and demos are available at: https://github.com/DDSynth-RL/DDSynthRL.
Tristan Wu, Daniel Chin, Junan Zhang +3
Computational Media and Art, The Hong Kong University of Science and Technology (Guangzhou) · New York University Shanghai · The Chinese University of Hong Kong, Shenzhen +2
Song generation and editing have mostly been treated as separate tasks. Existing editing methods often require noise injection and regeneration or curated paired training data. We propose a unified approach for song generation and editing based on reconstructive pretraining, in which a model is trained to reconstruct audio from varying numbers of interpretable conditions. With conditions such as text and lyrics, the model learns to generate diverse songs. With dense conditions specifying fine-grained music attributes, the model learns to reconstruct the target and enables editing by modifying any single attribute while keeping others fixed. This leads to SongCraft, a latent flow matching based model trained for both generation and fine-grained editing. To improve song generation quality, we further introduce word-level phoneme alignment that improves pronunciation learning and accelerates convergence, beat conditioning that improves general musicality, and representation alignment on VAE latent space that produces semantically meaningful latents for improved generation quality. Experiments show that SongCraft achieves the lowest word error rate among evaluated song generation baselines while maintaining competitive audio quality. We further show that a single model can support editing of lyrics, vocal melody, beats, and singer identity, and we also study the trade-off between reconstruction quality and editability.
Text and lyrics specify broad musical characteristics and sung content but offer limited control over musical timing, melody, and reference-based style. We introduce DiffSynth-Music (https://modelscope.cn/models/DiffSynth-Studio/DiffSynth-Music), a framework that adds composable audio conditioning to a music synthesis backbone through layer-wise key-value injection. The three template models, Control, Prosody, and Reference, are initialized from the backbone diffusion transformer and trained with conditional flow matching. They support five control types: beats, vocals, accompaniment, prosody, and reference audio. A shared variational autoencoder maps conditioning waveforms into a common latent space, enabling their attention memories to be combined. With the template timestep fixed at the clean-data endpoint and other inputs held constant, each control cache is computed once and reused throughout sampling. Training pairs are derived from music recordings using beat extraction, source separation, vocal resynthesis, and reference-excerpt selection. Single-control evaluations on Mandarin and English songs demonstrate improved adherence across all five control types and better lyric fidelity under vocal conditioning relative to the backbone. Automatic music-quality and instruction-following scores remain broadly comparable to those of the evaluated base models, with metric-specific trade-offs. We release the three template models to support research and creative applications in controllable music generation.