Compositionality, the ability to represent complex acoustic scenes as combinations of simpler sound sources, is central to auditory perception and classical additive signal models. Still, it remains unclear whether modern pre-trained audio representations internalize additive structure without compositional supervision. Existing evaluation frameworks of audio compositional reasoning largely focus on cross-modal audio-text alignment, leaving open whether audio representations themselves exhibit additive compositional structure independent of text grounding, analogous to vector arithmetic in word representations. To investigate this, we adopt a two-step diagnostic for frozen audio representations. First, we quantify linear alignment between representations and sound source labels using canonical correlation analysis. Second, we test additive compositional generalization via leave-one-combination-out reconstruction, grouping clips by exact source-label set, averaging their representations, and predicting held-out means from per-source contributions fitted only on training combinations. With larger combination holdouts, CLAP outperforms the permuted and label-overlap baselines on FSD50K, while the speech models do not outperform the label-overlap baseline. We examine representations generated by Wav2Vec2, HuBERT, and CLAP on FSD50K and CHiME-Home datasets. All three models show consistently higher linear correlation and more accurate leave-one-combination-out reconstructions than the permuted baselines. However, only CLAP shows large cosine similarity gains, which could be associated with its training on many kinds of audio and text. Finally, we note that all three models exhibit reconstruction residuals, revealing limits of additive compositionality such as nonlinear or non-compositional audio structure.
Figures & tables
Fig. 1: Two-step evaluation framework for additive compositionality in audio representations. Step 1 (top): CCA quantifies linear alignment between the source labels and the representations. Step 2 (bottom): Leave-one-combination-out reconstruction tests whether a held-out combination can be predicted from source contributions learned on other combinations.
Fig. 2: Canonical correlations for CHiME-Home (7 components) and FSD50K (first 15 of 200). Solid: real; dashed: mean of 100 permutations. Legends give mean Pearson correlations (PCCs), real / permuted.
Dataset
Model
Cos. Sim. ( ↑ )
Rel. Frob. ( ↓ )
CKA ( ↑ )
KL ( ↓ )
Hits@5 ( ↑ )
(a) Leave-one-combination-out
FSD50K
CLAP
0.686 / 0.270
0.761 / 1.015
0.579 / 0.032
0.0007 / 0.0012
0.43 / 0.01
HuBERT
0.968 / 0.957
0.410 / 0.436
0.042 / 0.006
0.0093 / 0.0105
0.06 / 0.01
Wav2Vec2
0.942 / 0.930
0.470 / 0.501
0.055 / 0.005
0.0065 / 0.0075
0.04 / 0.01
CHiME-Home
CLAP
0.899 / 0.820
0.455 / 0.586
0.714 / 0.196
0.0002 / 0.0004
0.58 / 0.15
HuBERT
0.992 / 0.988
0.134 / 0.159
0.415 / 0.269
0.0009 / 0.0013
0.15 / 0.13
TABLE I: Held-out combination reconstruction. (a) LOO: additive / permutation. (b) Repeated holdout: additive mean ± SD / permutation / label-overlap over 20 splits. Bold in (b) indicates additive means better than both baselines before rounding.
Large audio-language models (LALMs) perform strongly on individual audio tasks, but whether these capabilities can be reliably composed remains underexplored. We conduct a controlled diagnostic study of capability composition in LALMs, requiring models to integrate audio-attribute recognition, cue-conditioned segment selection, and downstream ASR or question answering. We construct two-utterance inputs with distinct acoustic cues to evaluate composition across environmental sound, gender, and emotion cues, with ASR, Math QA, and Factual QA as downstream tasks. Across four open-source LALMs, compositional QA accuracy decreases in 39 of 40 model-task-cue settings, by an average of 26.7 percentage points. ASR exhibits a similarly consistent degradation, with WER increasing in 39 of 40 settings by an average of 28.5 percentage points, while the magnitude of degradation varies across models, cue types, and cue salience. We further probe these failures through output format, positional preference, and chain-of-thought (CoT) analyses. Our study reveals a systematic gap between possessing individual audio capabilities and reliably composing them.
Chien-Feng Liu, Chih-Kai Yang, Bo-Han Feng +3
National Taiwan University · ASUS Open Cloud Infrastructure Software Center · NTU Artificial Intelligence Center of Research Excellence (NTU AI-CoRE)
Audio-language pretraining (ALP) holds promise for learning general-purpose audio representation, yet remains underexplored. Crucially, there is no consensus on whether audio-language models can build effective general-purpose audio encoders, nor a systematic understanding of how pretraining objectives behave across diverse tasks and scales. We identify three key barriers: limited scale of audio-text corpora, limited coverage of audio attributes in existing caption corpora, and lack of systematic exploration and evaluation. To fill this gap, we present the first principled empirical study of ALP. We first introduce CaptionStew, a 10.7M caption dataset aggregating open-source audio-text corpora across multiple domains and captioning focuses. We then conduct the first comprehensive evaluation comparing contrastive and captioning objectives for learning audio representation across speech, music, and environmental sound tasks. Our results not only demonstrate that ALP yields competitive, transferable representations, but reveal critical trade-offs: contrastive learning offers superior data efficiency, while captioning exhibits better scalability. Furthermore, we find that the benefits of supervised initialization often diminish at larger scales, challenging common practices. By grounding these claims in empirical evidence, we establish a viable pathway toward general-purpose audio representation learning, guiding future research.
Wei-Cheng Tseng, Xuanru Zhou, Mingyue Huo +3
University of Texas at Austin · Tencent AI Lab Seattle · Zhejiang University +1
Recent Large Audio Language Models have demonstrated impressive capabilities in audio understanding. However, they often suffer from perceptual errors, while reliable audio reasoning is unattainable without first grounding the model's perception in structured auditory scenes. Inspired by Auditory Scene Analysis, we first introduce a Perception-Aware Question Answering (PAQA) dataset. PAQA implements a hierarchical decoupling strategy that separates speech from environmental sound and distinguishes multiple speakers, providing explicit perceptual reasoning for training. Building on this, we propose HyPeR, a two-stage Hybrid Perception-Reasoning framework. In Stage I, we finetune the model on PAQA to perceive acoustic attributes in complex audio. In Stage II, we leverage GRPO to refine the model's internal deliberation. We also introduce PAUSE tokens to facilitate latent computation during acoustically ambiguous phases and design perceptual consistency reward to align reasoning rationales with raw audio. Experiments across benchmarks demonstrate that HyPeR achieves absolute improvements over the base model, with performance comparable to large-scale models, stressing the effectiveness of hybrid perception-grounded reasoning for robust and multi-speaker audio understanding.
Jieyi Wang, Yazhe Niu, Dexuan Xu +1
Shanghai AI Laboratory · Peking University · CUHK MMLab +1