AutoSynth: Learning to Generate Editable Synthesizer Programs from Audio and Text
Authors: Tristan Wu, Daniel Chin, Liwei Lin, Junan Zhang, Gus Xia
Organizations: The Hong Kong University of Science and Technology (Guangzhou) · New York University Shanghai · The Chinese University of Hong Kong, Shenzhen · Mohamed bin Zayed University of Artificial Intelligence
Audio generation models can translate natural-language descriptions into sound, but their outputs are typically waveforms. Their audio quality is constrained by audio compression, and their outputs do not readily support direct edits to notes, timbral parameters, or modulation relationships. We present AutoSynth, which represents MIDI performance events, fixed synthesizer parameters, and variable-length modulation routes as a unified sequence for a synthesizer, and learns their dependencies with an audio-conditioned autoregressive model. A single model supports both tasks. Given reference audio, the model directly predicts a synthesizer program; given text, it uses a pretrained audio generation model and converts the generated audio into a program. Training consists of two stages: supervised learning on large-scale audio-program pairs automatically constructed from a small set of native presets, followed by group-relative policy optimization with a mixed reward combining semantic similarity, pitch-related features, acoustic similarity, and sound usefulness. The pipeline requires neither paired text-target-program annotations nor a differentiable synthesizer. Experiments show that AutoSynth produces complete, editable synthesizer programs and achieves competitive results in both synthesizer inversion and text-driven generation. Audio demos and source code are available at https://auto-synth.github.io/.
Figures & tables
Figure 1: Overview of AutoSynth.
Figure 2: Autoregressive program generation and decoder architecture.
System
Audio CLAP ↓
LanguageBind ↓
wMFCC ↓
SOT ↓
RMS ↑
MSS ↓
Gemini–Woosh
AutoSynth
0.37
0.397
10.22
0.15
0.82
6.94
Synth Permutations
0.54
0.584
22.75
0.17
0.80
9.99
DDSynth-RL
0.49
0.482
10.34
0.16
0.87
7.04
FSD50K
AutoSynth
0.67
0.649
14.95
0.16
0.632
13.00
Table 1: Synthesizer inversion on Gemini–Woosh and FSD50K. Bold denotes the best value for each metric within each test set.
System
Text CLAP ↓
LanguageBind ↓
CE ↑
CU ↑
PQ ↑
FAD ↓
KL ↓
Gemini
AutoSynth
0.550
0.638
3.759
7.23
7.433
—
—
Synth Permutations
0.642
0.718
3.409
6.88
7.372
—
—
DDSynth-RL
0.607
0.678
3.253
6.81
7.276
—
—
CTAG
0.519
0.645
3.020
6.58
6.888
—
—
Stable Audio 3 ∗
0.546
0.614
3.891
7.16
7.373
—
—
Table 2: Text-driven generation on Gemini and FSD50K descriptions, corresponding to the audio in Table 1 .
Figure 3: Subjective synthesizer-inversion evaluation on Gemini–Woosh (left) and FSD50K (right). Error bars indicate 95% bootstrap confidence intervals for the mean.
Waveform compression frontend, 6 backbone blocks of width 768
53.73M
Yes
SAME-L
Waveform compression frontend, 12 backbone blocks of width 1536
426.06M
Yes
Appendix
Table A1: Encoder sizes. Counts exclude the program decoder and the SAME waveform decoder.
Figure A1: Validation curves for three encoders across two training stages. Left: in-domain SFT loss, shown up to 500,000 updates. Right: Stable Audio 3 validation reward. Horizontal axes count optimizer updates within each stage.
Condition
PQ ↑
CLAP distance ↓
Uncorrupted Woosh
7.18
0.000
Uncorrupted input → AutoSynth
7.43
0.372
Noisy input
5.07
0.306
Noisy input → AutoSynth
7.40
0.451
Low-pass input
7.10
0.134
Low-pass input → AutoSynth
7.34
0.375
Appendix
Table B1: Corrupted-input inversion and resynthesis, averaged over all 128 clips. CLAP distance uses the uncorrupted Woosh clip as reference; lower is better. Higher PQ is better.
Test set
Audio CLAP ↓
LanguageBind ↓
wMFCC ↓
SOT ↓
RMS ↑
MSS ↓
Gemini–Woosh
0.47 → 0.37
0.535 → 0.397
17.05 → 10.22
0.33 → 0.15
0.65 → 0.82
19.41 → 6.94
FSD50K
0.73 → 0.67
0.746 → 0.649
22.05 → 14.95
0.31 → 0.16
0.472 → 0.632
24.33 → 13.00
Helm in-domain
0.17 → 0.13
0.221 → 0.174
8.02 → 6.93
0.057 → 0.048
0.91 → 0.92
7.36 → 3.77
Appendix
Table C1: SFT/GRPO comparison for synthesizer inversion.
Test set
Text CLAP ↓
LanguageBind ↓
CE ↑
CU ↑
PQ ↑
FAD ↓
KL ↓
Gemini
0.601 → 0.550
0.671 → 0.638
3.469 → 3.759
6.92 → 7.23
7.278 → 7.433
—
—
FSD50K
0.768 → 0.771
0.776 → 0.793
3.24 → 3.48
6.73 → 6.91
6.97 → 7.07
8.83 → 7.00
5.04 → 5.27
Appendix
Table C2: SFT/GRPO comparison for text-driven generation.
Figure D1: In-domain validation loss for different autoregressive orders over the first 100,000 steps.
Table D1: Autoregressive order of fixed parameters, read left to right within each row.
Item
Setting
Audio input
44.1 kHz, 5 seconds, stereo
Mel spectrograms
128 bands per channel, 25 ms Hamming window, 100 frames/s
Encoder / decoder layers
12 / 12
Hidden width / heads / FFN width
512 / 8 / 1024
Normalization and activation
Pre-LayerNorm, GELU, dropout 0.1
Value output layer
Shared Linear(512, 128), with field-specific legal categories
Appendix
Table E1: Main model and optimization settings.
Component
Direction
Coefficient
CLAP distance
Negative
8
Basic Pitch hidden-feature distance
Negative
12
AudioBox CE, normalized to [0,1]
Positive
1
AudioBox CU, normalized to [0,1]
Positive
1
wMFCC
Negative
0.07
SOT
Negative
3.58
Appendix
Table E2: Coefficients used in Equation ( 3 ). All coefficients are positive; each metric’s direction determines its sign in the reward.
Dimension
Question
Evaluation criteria
Alignment: synthesizer inversion
Does this sound faithfully reproduce the reference audio?
Consider timbre, pitch, articulation, and changes over time; 1 means not at all and 5 means very faithfully.
Alignment: text-driven generation
Does this sound faithfully follow the text description?
Consider the described timbre, pitch, articulation, and temporal changes; 1 means not at all and 5 means very faithfully.
Content quality: both tasks
How good is this sound as material for music production or sound design?
Consider audio quality, unintended artifacts, and creative usefulness; 1 means very poor and 5 means excellent. Intentional distortion is not a defect, and quality is rated independently of alignment.
Computational Media and Art, The Hong Kong University of Science and Technology (Guangzhou) · New York University Shanghai · The Chinese University of Hong Kong, Shenzhen +2