Exact Factorisation and Fast Computation of Invertible Constant-Q Transforms
Authors: Facundo Franchino, Eloi Moliner, Vesa Välimäki
Organizations: Massachusetts Institute of Technology, Cambridge, MA, USA · Acoustics Lab, Dept. of Information and Communications Eng., Aalto University, Espoo, Finland
The constant-Q transform (CQT) represents audio on a logarithmic frequency axis. Its nonstationary Gabor formulation is exactly invertible, but the unequal numbers of time coefficients in its bands complicate GPU computation. An exact factorisation combines spectral selection, conjugation, windowing, and reordering into a fixed map between one packed Fourier transform and the shorter band inverse transforms. The factors give waveform reconstruction, real adjoints for backpropagation, and bounds on arithmetic depth and block width; overlapping slices permit streaming with bounded memory. Tests on two GPU models show that Flash-CQT reduces analysis-synthesis round-trip time by factors of two to eight relative to a baseline computing the same CQT. The proposed implementation also uses over 30% less peak temporary workspace and reaches a negligible reconstruction error, with a signal-to-noise ratio of about 130 dB, in single-precision floating-point arithmetic. These advances make Flash-CQT a practical, computationally efficient front end for spectral analysis and modern audio machine-learning systems.
Figures & tables
Octave
Bands
Centres (Hz)
Mλ
Radix plan
1
8
43–81
32
32
2
8
88–165
64
8×8
3
8
180–337
128
16×8
4
16
345–666
128
16×8
5
16
695–1342
256
16×16
6
16
1402–2708
512
32×16
Table 1: Frequency band layout at N=65,536 and 44.1 kHz. DC and Nyquist sidebands have lengths 128 and 2048.
Multi-
SNR
Workspace
Round trip
Step
Method
res.
(dB)
(MB)
(ms)
(ms)
STFT (reference)
✗
133.1
4.3
0.051
0.838
CQT baseline
✓
126.8
4.2
0.601
10.003
Flash-CQT (proposed)
✓
129.8
2.6
0.116
1.969
Table 2: Reconstruction SNR, workspace per signal and B=1 timings on an NVIDIA A100. Round trips use graph replay; steps include eager backpropagation. Best values among the CQT implementations are bold.
We develop a second-order theory of quantization noise in matrix multiplication in which the quantization format is characterized by the variance it assigns to each element. The constant variance profile of integer quantization recovers existing integer-noise theory, while the multiplicative profile of floating-point rounding reduces the data dependence to a scalar, the participation factor κ, yielding a closed-form signal-to-noise-ratio law. The resulting functional also admits a closed-form upper bound κ∗ that no function-preserving linear transform can exceed and that is attained by a recent state-of-the-art method. Building on this analysis, we introduce KBBQ (\textbf{K}appa-\textbf{B}raked \textbf{B}lockwise \textbf{Q}uantization), which parameterizes the extent to which a transform approaches this ceiling. At W4A4, across four base models and two FP4 formats, KBBQ outperforms the prior state of the art without additional deployment-time computation.
Neural audio codecs with residual vector quantization (RVQ) normally treat all frequencies uniformly, so their codebooks become spectrally entangled. Truncating stages then removes an unpredictable mix of frequencies. Parallel band decomposition addresses this by splitting audio into independent bands, but fragments the latent space and loses cross-frequency coherence. We introduce HARP (Harmonic-Aware Residual Partitioning), a training strategy that partitions RVQ stages into frequency-ordered groups where each group refines its target band while the decoder retains access to all lower frequencies. Overtones are reconstructed in the context of their fundamentals, preserving coherence that parallel methods lose. HARP requires no architectural changes; it only modifies the training loss, leaving inference identical to standard RVQ. On speech, music, and general audio, HARP outperforms both standard RVQ and parallel decomposition. MUSHRA listening tests also show perceptual improvements.
Qiaoyu Yang, Lixing He, Binyue Deng +1
Georgia Institute of Technology, Atlanta, United States · The Chinese University of Hong Kong, Hong Kong, China · Tencent Music Entertainment, Shenzhen, China
End-to-end neural audio models achieve high-fidelity compression and generation. We might read that performance as evidence they directly represent interpretable features such as pitch and timbre, but a model can produce plausible outputs without doing so. A model may encode these features in any reachable basis, but regardless of which, the features are well described as compositions of time-frequency-localized primitives. Whether state-of-the-art encoders preserve access to these primitives, and thus to compositions of them, remains unclear. Through theoretical analysis and controlled experiments, we show that several state-of-the-art strided convolutional encoders impose two structural bottlenecks, both predictable from architecture and signal structure, on access to these primitives: (1) they collapse primitives into alias equivalence classes, establishing a bound on representational capacity, and (2) they limit the frequency resolution available to learned filters, restricting separability. For well structured data, we find collapse rates of 31-35% and filter bandwidths 10-35x above the theoretical resolution bound, confirming that both bottlenecks arise under realistic signal conditions. We then introduce Gabor Latent Refactorization (GLRF), a lightweight post-hoc intervention that re-expresses encoder latents in a frequency-localized basis, reducing filter bandwidths from 10-35x to 1.5-3x of the theoretical resolution bound while preserving reconstruction fidelity and improving control over attributes like pitch. These results show that the encoders in question predictably degrade access to frequency-localized primitives, entangling the features that depend on them, and that a lightweight, retraining-free intervention can recover much of that access, improving steerability and interpretability.