eess.ASOct 6, 2026

CTAG-FX: Reinterpreting Synthesizer Parameter Spaces for Expressive Tone-Shaping Audio FX Design

Authors: Geonung Jo, Jongeun Choi

Organizations: Independent Researcher, Seoul, Republic of Korea · School of Mechanical Engineering, Yonsei University, Seoul 03722, Republic of Korea

Abstract

Recent text-guided audio FX research has focused on translating natural-language descriptions into parameters or chain configurations within existing FX systems. However, comparatively little attention has been paid to how the internal control space of an individual FX processor can itself be constructed. To address this gap, we propose CTAG-FX, which functionally reinterprets the roles of 78 parameters in a text-conditioned synthesizer configuration as controls of a tone-shaping FX processor. Each synthesizer configuration defines a fixed tone-shaping processor that can be repeatedly applied to new audio inputs rather than producing only a one-off rendered output. The text-conditioned synthesizer configurations are generated using A&R-CTAG, a retrieval-enhanced extension of CTAG. We evaluate CTAG-FX through signal-level analysis, together with a scrambled-mapping ablation in which parameter-to-control assignments are randomly permuted to assess the contribution of the proposed role assignment. Role-based mapping produces prompt-distinct spectral and nonlinear behavior, whereas scrambling reduces this prompt-specific differentiation and yields a more prompt-insensitive nonlinear profile. Overall, these results suggest that synthesizer parameter spaces can serve not only as sound-generation spaces but also as design resources for constructing new audio-FX control spaces. Audio samples are available at https://taylor3527eg-hub.github.io/ctag-fx-platform-demo/.

Figures & tables

Explore similar work

Jun 20, 2026cs.SD

InstructFX2FX: A Multi-Turn Text-to-Effect System for Sequential Audio Effect Refinement

We present InstructFX2FX, a system for sequential audio effect refinement through multi-turn natural-language instructions. Existing text-to-effect systems are largely single-shot, mapping one textual descriptor to one preset. Real audio engineering is instead sequential: engineers refine an existing effect chain through successive instructions. This poses a stateful problem that single-shot systems do not address: given the current effect parameters state and a new instruction, update the sound while preserving what earlier instructions already achieved. InstructFX2FX addresses this with a hybrid architecture that divides labor between a language model and CLAP-guided optimization. The LLM serves as a high-level planner that selects effects and proposes the initial parameter state, motivated by recent evidence that LLMs can outperform CLAP-based optimization for single-turn text-to-effect mapping; CLAP-guided optimization then refines the existing parameter state, providing a more stable and robust refinement mechanism than LLM reprompting. In the demo, attendees drive a dry recording through successive natural-language instructions: after each turn, they choose how strongly the effect is applied, then issue the next instruction based on what still differs from the sound they intend. In a preliminary evaluation on SocialFX-derived descriptor pairs, CLAP-guided refinement achieves lower DSP-feature MMD than an LLM+LLM initialize-then-reprompt baseline on 9 of 10 pairs. Trajectory analysis further shows that, for differentiable effects, optimization tends to gradually move the audio toward the new target while retaining the effects of the previous instruction, highlighting the potential for gradual refinement.
Jun 6, 2026cs.SD

FXplorer: A Map-Based Interface for Exploratory Audio Effect Design

Audio effects (FX) shape sound in contemporary music practice. However, most interfaces present them as discrete modules and parameters that favor targeted adjustment over exploratory listening. This separation can make it difficult to build intuition about the broader space of possible transformations or to move fluidly between searching and refinement. We present FXplorer, an interface that organizes audio effects within a perceptually informed 2D space, allowing sound transformations to be browsed as a continuous landscape rather than as isolated presets. By combining established spatial interaction approaches and interpretable DAW-style controls with recent embedding-based machine learning methods for similarity and semantic search, the system brings exploration and parameter refinement into a single workspace. FXplorer supports composition, production, or performance by allowing users to edit and interpolate between effect presets interactively.
Jul 17, 2026cs.SD

StemFX: Learning Mixing Style Representations via Autoregressive FX Chain Prediction on Source-Separated Stems

Audio mixing style encompasses the artistic and technical decisions a mix engineer makes, including level balancing, spatialization, and the choice, ordering, and parameterization of audio effects (FX) on each stem. FX chains are a key determinant of this style, yet existing approaches to modeling them remain limited. Some operate on stereo mixtures without explicit per-stem FX chain modeling, others fix the number or type of effects per track, and many require differentiable effect implementations or scarce multitrack datasets. We present StemFX, a framework that learns mixing style representations by autoregressively predicting variable-length FX chains on source-separated stems. A Transformer decoder predicts tokenized FX chains autoregressively, while a band-split multi-band CNN encoder with FiLM conditioning captures per-stem spectral structure. To enable large-scale paired training, we extract pseudo-stems from about 105K songs via source separation and augment them using MultiAFx, a toolkit unifying 85 audio effects from 7 Python libraries. Evaluated on mixing style retrieval, StemFX outperforms all baseline models across all tested chain lengths. On paired mixing style transfer, StemFX achieves the best spectral fidelity and the highest listener preference, over 4000 times faster than iterative optimization.