cs.SDOct 7, 2026

AutoSynth: Learning to Generate Editable Synthesizer Programs from Audio and Text

Authors: Tristan Wu, Daniel Chin, Liwei Lin, Junan Zhang, Gus Xia

Organizations: The Hong Kong University of Science and Technology (Guangzhou) · New York University Shanghai · The Chinese University of Hong Kong, Shenzhen · Mohamed bin Zayed University of Artificial Intelligence

Abstract

Audio generation models can translate natural-language descriptions into sound, but their outputs are typically waveforms. Their audio quality is constrained by audio compression, and their outputs do not readily support direct edits to notes, timbral parameters, or modulation relationships. We present AutoSynth, which represents MIDI performance events, fixed synthesizer parameters, and variable-length modulation routes as a unified sequence for a synthesizer, and learns their dependencies with an audio-conditioned autoregressive model. A single model supports both tasks. Given reference audio, the model directly predicts a synthesizer program; given text, it uses a pretrained audio generation model and converts the generated audio into a program. Training consists of two stages: supervised learning on large-scale audio-program pairs automatically constructed from a small set of native presets, followed by group-relative policy optimization with a mixed reward combining semantic similarity, pitch-related features, acoustic similarity, and sound usefulness. The pipeline requires neither paired text-target-program annotations nor a differentiable synthesizer. Experiments show that AutoSynth produces complete, editable synthesizer programs and achieves competitive results in both synthesizer inversion and text-driven generation. Audio demos and source code are available at https://auto-synth.github.io/.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. DDSynth-RL: Audio Synthesizer Inversion via Discrete Diffusion with Reinforcement Learning

    Aug 4, 2026Tristan Wu, Daniel Chin, Junan Zhang +3Diffusion-Based Inverse ProblemsDiffusion Models

  2. DiffSynth-Music: Audio-Conditioned KV-Cache Adapters for Controllable Music Generation

    Sep 14, 2026Zhongjie Duan, Shengchuan Gao, Hong Zhang +1Controllable Music GenerationAudio Diffusion Models