Neural audio codecs (NACs) are widely used in speech generation and audio-language modeling, yet how they encode speaker-trait information remains poorly understood. Prior work applied sparse autoencoders (SAEs) to investigate accent information in NACs through task-level analysis. Here, we extend this analysis to the waveform level and to age, gender, and accent, using SAE steering to probe trait-related information in sparse activations. We identify trait-associated dimensions, modify their activations, and evaluate the resulting reconstructed speech. Across five NACs, steering the selected dimensions induces target-directed shifts in speaker-trait predictions. A random-dimension baseline on Mimi produces smaller shifts, supporting the relevance of the selected dimensions. However, responses vary across codecs, traits, and steering directions, and increasing steering strength does not consistently amplify the intended shifts. Steering also generally increases word error rates and lowers predicted perceptual quality. These findings suggest that SAEs capture speaker-trait information in steerable activations, while the accompanying quality degradation highlights the need to better separate trait-related information from other information.
Figures & tables
Figure 1: Overview of the SAE steering framework. (A) Illustration of NAC encoding speech utterance into quantized frame-level representations. (B) A SAE decomposes each frame into sparse activations and reconstructs it. (C) Trait-associated activation dimensions are identified and modified to steer reconstructed speech from a source trait g1 toward a target trait g2 .
Figure 2: Steering results across five NACs for (a) gender, (b) age, (c) US–UK accent, and (d) US–SA accent. US, UK, and SA denote North American, British Isles, and Southern Asian accents, respectively. Diamonds indicate unsteered SAE reconstruction baselines; circles indicate steering results at steering strength α . Mimi Random uses randomly selected SAE dimensions from Mimi’s representation.
Waveform
Gender
Age
US–UK
US–SA
Female / Male
Young / Senior
US / UK
US / SA
Original
95.9 / 7.6
27.8 / 63.1
89.6 / 3.2
57.3 / 1.7
EnCodec 6 kbps
94.9 / 7.8
36.7 / 65.0
89.3 / 3.8
67.4 / 2.4
DAC
96.0 / 7.2
28.7 / 63.7
88.2 / 5.0
73.7 / 4.4
Mimi
95.3 / 8.0
31.5 / 62.0
88.9 / 3.6
60.6 / 1.8
SpeechTokenizer
93.9 / 7.9
33.1 / 62.7
86.8 / 4.8
57.8 / 3.4
Table 1: Speaker-trait predictions before steering. Gender and accent scores are mean female and pairwise US probabilities (%), respectively; age is reported in years.
Figure 3: Speech quality after steering across NACs. The top row reports SHEET quality scores ( ↑ ), and the bottom row reports WER ( ↓ ).
Neural Audio Codecs (NACs) are widely adopted in modern speech systems, yet how they encode linguistic and paralinguistic information remains unclear. Improving the interpretability of NAC representations is critical for understanding and deploying them in sensitive applications. Hence, we employ Sparse Autoencoders (SAEs) to decompose dense NAC representations into sparse, interpretable activations. In this work, we focus on a challenging paralinguistic attribute-accent-and propose a framework to quantify NAC interpretability. We evaluate four NAC models under 16 SAE configurations using a relative performance index. Our results show that DAC and SpeechTokenizer achieve the highest interpretability. We further reveal that acoustic-oriented NACs encode accent information primarily in activation magnitudes of sparse representations, whereas phonetic-oriented NACs rely more on activation positions, and that low-bitrate EnCodec variants show higher interpretability.
Neural audio codecs (NACs) convert speech into discrete token sequences, and prior work has reported that these sequences follow language-like statistical laws. This paper analyzes the token statistics of 13 NACs spanning multi-codebook residual vector quantization (RVQ), single-codebook VQ, and non-VQ designs, evaluated on three corpora under clean, white-noise, and real-world DEMAND-noise conditions. Zipf and Heaps parameters, unigram entropy, codebook occupancy, and Jensen-Shannon divergence (JSD) are estimated from matched token samples with explicit fit-validity safeguards and family-conditional n-gram orders. Corpus identity explains little variance in any metric, whereas acoustic condition and quantizer meta-category dominate in a metric-dependent way, and unigram entropy is the metric most strongly associated with meta-category. Clean-to-noise JSD computed at a common unigram order is associated with mel-cepstral distortion most clearly under DEMAND noise. The collapse and explosion degradation signatures previously reported for RVQ codecs concentrate in RVQ cells under white and DEMAND noise, respectively; explosion also occurs in non-VQ codecs, and single-codebook VQ codecs shift in occupancy and distribution shape without either signature. These results provide architecture-conditioned conventions for applying language-statistical analysis to NAC tokens.
Joonyong Park, Shinnosuke Takamichi, David M. Chan +3
Grad. School of Information Science and Technology, The University of Tokyo, Tokyo, Japan · Faculty of Science and Technology, Keio University, Tokyo, Japan · Berkeley Artificial Intelligence Research Lab (BAIR), University of California, Berkeley, CA, USA
Neural audio codecs (NACs) have become popular for obtaining speech representations as discrete tokens. Beyond compression, discrete tokens can be used to train self-supervised learning (SSL) models. Such models, referred to as codec-based SSL models, reduce data storage and computational cost, enabling scalable SSL pre-training. However, their language sensitivity remains unclear. When the language changes, codec-based SSL models may require retraining, which undermines their efficiency. In this paper, we present a systematic analysis of language sensitivity by varying either the NAC training language or the SSL pre-training language while keeping the other fixed. Experimental results show that downstream performance is insensitive to the NAC training language but strongly dependent on the SSL pre-training language. These findings suggest that a single NAC can be reused across languages, while aligning the SSL pre-training language with the target language is crucial.