Recent text-guided audio FX research has focused on translating natural-language descriptions into parameters or chain configurations within existing FX systems. However, comparatively little attention has been paid to how the internal control space of an individual FX processor can itself be constructed. To address this gap, we propose CTAG-FX, which functionally reinterprets the roles of 78 parameters in a text-conditioned synthesizer configuration as controls of a tone-shaping FX processor. Each synthesizer configuration defines a fixed tone-shaping processor that can be repeatedly applied to new audio inputs rather than producing only a one-off rendered output. The text-conditioned synthesizer configurations are generated using A&R-CTAG, a retrieval-enhanced extension of CTAG. We evaluate CTAG-FX through signal-level analysis, together with a scrambled-mapping ablation in which parameter-to-control assignments are randomly permuted to assess the contribution of the proposed role assignment. Role-based mapping produces prompt-distinct spectral and nonlinear behavior, whereas scrambling reduces this prompt-specific differentiation and yields a more prompt-insensitive nonlinear profile. Overall, these results suggest that synthesizer parameter spaces can serve not only as sound-generation spaces but also as design resources for constructing new audio-FX control spaces. Audio samples are available at https://taylor3527eg-hub.github.io/ctag-fx-platform-demo/.
Figures & tables
Figure 1: Representative prior text-guided FX workflow.
Dimension
Prior text-guided FX systems
CTAG-FX
Strategy
Text → predefined FX configuration
Text → synth representation → FX processor design
Source of expressivity
Available FX modules and chains
Synthesizer parameter space
Design space
Predefined FX parameter or chain space
Reinterpreted synth-to-FX control space
Main mechanism
Predict / optimize / search / refine
Reinterpret / map / design
Design object
FX parameter or FX chain
Self-contained tone-shaping FX processor
Contribution
Semantic control within predefined FX spaces
Synth-derived FX control-space design
Table 1: Conceptual comparison of prior text-guided FX systems and CTAG-FX.
Figure 2: Overall CTAG-FX workflow. A&R-CTAG generates a text-conditioned synthesizer configuration, and CTAG-FX maps the functional roles of its parameter groups to controls in a tone-shaping processor for external audio.
Figure 3: A&R-CTAG pipeline for generating the synth-side character representation used by CTAG-FX.
Figure 4: Temporal and spectral characteristics of a plucked E4 guitar note motivating the ADSR-to-EQ reinterpretation: (a) normalized RMS envelope with temporal markers, (b) spectrogram with harmonic guides.
Figure 5: Conceptual VCO-to-drive reinterpretation. (a) Sine and SquareSaw VCOs generate comparatively smooth and harmonically rich sources, respectively. (b) Drive 1 and Drive 2/Fuzz apply analogous smooth and aggressive nonlinear coloration to a common input. Colors indicate corresponding roles; waveforms are illustrative, not equivalent.
Figure 6: Simplified CTAG-FX signal flow with synth-derived control routing. Parentheses identify the corresponding synthesizer groups.
Synth Group
Functional Reinterpretation
FX Control
MIDI f0
Pitch reference → input-response reference
Input Trim
Duration
Event length → output-scale reference
Output Trim
ADSR1
Temporal envelope → spectral contour
Pre-EQ
ADSR2
Secondary envelope → complementary tone contour
Tone Stack
VCO1/2
Waveform source → nonlinear harmonic character
Drive 1/2
Mixer
Source balance → processing-path balance
Clean / Processed / Noise Mixer
Table 2: Summary of synthesizer groups, functional reinterpretations, and FX controls in CTAG-FX.
Figure 7: Spectral characterization and mapping ablation for four tone prompts. (a) Mean output-to-input spectral difference across ten configurations. (b) Role-based versus scrambled-mapping responses. Shading indicates 95% confidence intervals.
Figure 8: Nonlinear descriptors under role-based and scrambled mappings. Markers and error bars show configuration-level means and 95% confidence intervals ( N=10 per prompt).
Metric
CTAG
A&R-CTAG
Change
Initial score mean
0.1449
0.2981
+105.7%
Final score mean
0.6166
0.7545
+22.4%
Final score SD
0.1246
0.0595
−52.2%
Score change mean
0.4718
0.4564
−3.3%
Table 3: LAION-CLAP similarity results for CTAG and A&R-CTAG across ten prompts under a 300-iteration budget. Mean scores and final-score SD are computed across prompts; Change is relative to CTAG.
We present InstructFX2FX, a system for sequential audio effect refinement through multi-turn natural-language instructions. Existing text-to-effect systems are largely single-shot, mapping one textual descriptor to one preset. Real audio engineering is instead sequential: engineers refine an existing effect chain through successive instructions. This poses a stateful problem that single-shot systems do not address: given the current effect parameters state and a new instruction, update the sound while preserving what earlier instructions already achieved. InstructFX2FX addresses this with a hybrid architecture that divides labor between a language model and CLAP-guided optimization. The LLM serves as a high-level planner that selects effects and proposes the initial parameter state, motivated by recent evidence that LLMs can outperform CLAP-based optimization for single-turn text-to-effect mapping; CLAP-guided optimization then refines the existing parameter state, providing a more stable and robust refinement mechanism than LLM reprompting. In the demo, attendees drive a dry recording through successive natural-language instructions: after each turn, they choose how strongly the effect is applied, then issue the next instruction based on what still differs from the sound they intend. In a preliminary evaluation on SocialFX-derived descriptor pairs, CLAP-guided refinement achieves lower DSP-feature MMD than an LLM+LLM initialize-then-reprompt baseline on 9 of 10 pairs. Trajectory analysis further shows that, for differentiable effects, optimization tends to gradually move the audio toward the new target while retaining the effects of the previous instruction, highlighting the potential for gradual refinement.
Song-Ze Yu, Milan Liessens Dujardin, Yuxuan Cai +4
Center for New Music and Audio Technologies (CNMAT) University of California, Berkeley
Audio effects (FX) shape sound in contemporary music practice. However, most interfaces present them as discrete modules and parameters that favor targeted adjustment over exploratory listening. This separation can make it difficult to build intuition about the broader space of possible transformations or to move fluidly between searching and refinement. We present FXplorer, an interface that organizes audio effects within a perceptually informed 2D space, allowing sound transformations to be browsed as a continuous landscape rather than as isolated presets. By combining established spatial interaction approaches and interpretable DAW-style controls with recent embedding-based machine learning methods for similarity and semantic search, the system brings exploration and parameter refinement into a single workspace. FXplorer supports composition, production, or performance by allowing users to edit and interpolate between effect presets interactively.
Audio mixing style encompasses the artistic and technical decisions a mix engineer makes, including level balancing, spatialization, and the choice, ordering, and parameterization of audio effects (FX) on each stem. FX chains are a key determinant of this style, yet existing approaches to modeling them remain limited. Some operate on stereo mixtures without explicit per-stem FX chain modeling, others fix the number or type of effects per track, and many require differentiable effect implementations or scarce multitrack datasets. We present StemFX, a framework that learns mixing style representations by autoregressively predicting variable-length FX chains on source-separated stems. A Transformer decoder predicts tokenized FX chains autoregressively, while a band-split multi-band CNN encoder with FiLM conditioning captures per-stem spectral structure. To enable large-scale paired training, we extract pseudo-stems from about 105K songs via source separation and augment them using MultiAFx, a toolkit unifying 85 audio effects from 7 Python libraries. Evaluated on mixing style retrieval, StemFX outperforms all baseline models across all tested chain lengths. On paired mixing style transfer, StemFX achieves the best spectral fidelity and the highest listener preference, over 4000 times faster than iterative optimization.
Yuan-Chiao Cheng, Jui-Te Wu, Brian Chen +3
Independent Researcher · National Taiwan University, Taiwan · AI-CoRE, National Taiwan University, Taiwan