eess.ASJul 14, 2026

Efficient Text-to-Audio Generation via Pruning

Authors: Arshdeep SinghYi YuanYun ChenWenwu WangMark D. Plumbley

Organizations: 1King’s College London (KCL) · University of Surrey, UK

Abstract

Diffusion-based text-to-audio generative models such as AudioLDM achieve high perceptual quality and strong semantic consistency; however, their practical deployment is hindered by the substantial computational cost of the U-Net denoising backbone. In this work, we apply model pruning to improve the computational efficiency of AudioLDM, a U-Net-based text-conditioned audio latent diffusion model. We analyse parameter redundancy across U-Net convolutional blocks and evaluate a filter-pruning strategy. Pruning is guided by norm-based criteria and followed by lightweight finetuning to recover performance losses. Experimental results demonstrate that up to 83% of the parameters and 39% of the multiply-accumulate operations of U-Net have been reduced while maintaining, and in some cases improving, generation quality compared to the baseline unpruned network. We find that pruning affects AudioLDM's ability to generate certain sound events including safety-critical sounds such as gunshots, sirens, and explosions, as well as mechanical sounds such as drills and sewing machines, and other sounds such as sprays and tick-tocks, which are mostly recovered by lightweight finetuning of the pruned model.

Explore similar work

CardsList
  1. Stable Audio 3

    May 18, 2026Zach Evans, Julian D. Parker, Matthew Rice +4Audio FlowAutoregressive Diffusion