eess.SPSep 24, 2026

Exact Factorisation and Fast Computation of Invertible Constant-Q Transforms

Authors: Facundo Franchino, Eloi Moliner, Vesa Välimäki

Organizations: Massachusetts Institute of Technology, Cambridge, MA, USA · Acoustics Lab, Dept. of Information and Communications Eng., Aalto University, Espoo, Finland

Abstract

The constant-Q transform (CQT) represents audio on a logarithmic frequency axis. Its nonstationary Gabor formulation is exactly invertible, but the unequal numbers of time coefficients in its bands complicate GPU computation. An exact factorisation combines spectral selection, conjugation, windowing, and reordering into a fixed map between one packed Fourier transform and the shorter band inverse transforms. The factors give waveform reconstruction, real adjoints for backpropagation, and bounds on arithmetic depth and block width; overlapping slices permit streaming with bounded memory. Tests on two GPU models show that Flash-CQT reduces analysis-synthesis round-trip time by factors of two to eight relative to a baseline computing the same CQT. The proposed implementation also uses over 30% less peak temporary workspace and reaches a negligible reconstruction error, with a signal-to-noise ratio of about 130 dB, in single-precision floating-point arithmetic. These advances make Flash-CQT a practical, computationally efficient front end for spectral analysis and modern audio machine-learning systems.

Figures & tables

Explore similar work

Sep 8, 2026cs.LG

KBBQ: A Predictive Noise Law and the Limits of Spectrum Flattening in FP4 Quantization

We develop a second-order theory of quantization noise in matrix multiplication in which the quantization format is characterized by the variance it assigns to each element. The constant variance profile of integer quantization recovers existing integer-noise theory, while the multiplicative profile of floating-point rounding reduces the data dependence to a scalar, the participation factor κκ, yielding a closed-form signal-to-noise-ratio law. The resulting functional also admits a closed-form upper bound κ∗κ^{*} that no function-preserving linear transform can exceed and that is attained by a recent state-of-the-art method. Building on this analysis, we introduce KBBQ (\textbf{K}appa-\textbf{B}raked \textbf{B}lockwise \textbf{Q}uantization), which parameterizes the extent to which a transform approaches this ceiling. At W4A4, across four base models and two FP4 formats, KBBQ outperforms the prior state of the art without additional deployment-time computation.
Jul 18, 2026cs.SD

HARP: Harmonic-Aware Residual Partitioning for Neural Audio Codecs

Neural audio codecs with residual vector quantization (RVQ) normally treat all frequencies uniformly, so their codebooks become spectrally entangled. Truncating stages then removes an unpredictable mix of frequencies. Parallel band decomposition addresses this by splitting audio into independent bands, but fragments the latent space and loses cross-frequency coherence. We introduce HARP (Harmonic-Aware Residual Partitioning), a training strategy that partitions RVQ stages into frequency-ordered groups where each group refines its target band while the decoder retains access to all lower frequencies. Overtones are reconstructed in the context of their fundamentals, preserving coherence that parallel methods lose. HARP requires no architectural changes; it only modifies the training loss, leaving inference identical to standard RVQ. On speech, music, and general audio, HARP outperforms both standard RVQ and parallel decomposition. MUSHRA listening tests also show perceptual improvements.
Jul 9, 2026cs.SD

Structural Bottlenecks on Frequency Representation in End-to-End Audio Models

End-to-end neural audio models achieve high-fidelity compression and generation. We might read that performance as evidence they directly represent interpretable features such as pitch and timbre, but a model can produce plausible outputs without doing so. A model may encode these features in any reachable basis, but regardless of which, the features are well described as compositions of time-frequency-localized primitives. Whether state-of-the-art encoders preserve access to these primitives, and thus to compositions of them, remains unclear. Through theoretical analysis and controlled experiments, we show that several state-of-the-art strided convolutional encoders impose two structural bottlenecks, both predictable from architecture and signal structure, on access to these primitives: (1) they collapse primitives into alias equivalence classes, establishing a bound on representational capacity, and (2) they limit the frequency resolution available to learned filters, restricting separability. For well structured data, we find collapse rates of 31-35% and filter bandwidths 10-35x above the theoretical resolution bound, confirming that both bottlenecks arise under realistic signal conditions. We then introduce Gabor Latent Refactorization (GLRF), a lightweight post-hoc intervention that re-expresses encoder latents in a frequency-localized basis, reducing filter bandwidths from 10-35x to 1.5-3x of the theoretical resolution bound while preserving reconstruction fidelity and improving control over attributes like pitch. These results show that the encoders in question predictably degrade access to frequency-localized primitives, entangling the features that depend on them, and that a lightweight, retraining-free intervention can recover much of that access, improving steerability and interpretability.