cs.SDJun 22, 2026

Scaling Audio Models Efficiently: Joint Optimization of Scale, Resolution, Adaptation, Precision, and Sparsity

Authors: Vyom AgarwalMokshda GangradeSiddharth PalJerry Wu

Organizations: Computer, Mathematical, and Natural Sciences, University of Maryland, College Park, MD, USA · RTX BBN Technologies, Columbia, MD, USA · Electrical and Computer Engineering, University of Maryland, College Park, MD, USA

Abstract

Large automatic speech recognition (ASR) models such as Whisper must be deployed across hardware with widely varying memory and inference-speed constraints. We present a compression framework that jointly parametrizes Whisper deployment along \emph{six} dimensions: model size, temporal resolution, encoder token stride, low-rank adaptation capacity, weight precision and sparsity pattern. All axes are jointly optimized using NSGA-III with respect to three deployment objectives: word error rate (WER), inference FLOPs, and memory footprint. Across 50 of the 1,680 candidate configurations evaluated, we characterize the conditional effect of each axis and identify compression combinations that dominate naive single-axis scaling, while finding that 1:4 structured sparsity fails to recover acceptable accuracy under the tested recovery budgets. We report measured WER and resident memory, use analytical EffFLOPs as the search-time compute surrogate, and separately validate representative inference configurations using measured real-time factor (RTF).

Explore similar work

CardsList