cs.LGMay 14, 2025

Disassociating performance from compositional feature learning

Authors: George DimitriadisSpyridon Samothrakis

Organizations: Sainsbury Wellcome Centre Gatsby Computational Neuroscience Unit University College London London · School of Computer Science and Electronic Engineering University of Essex Colchester

Abstract

Out-of-distribution (OOD) generalisation through composition requires a system to discover invariant properties from input-output associations and transfer them to novel inputs and unseen tasks. We argue that confirming compositional learning requires more than OOD evaluation alone: one must also verify that the learned features are genuinely compositional and that the system encodes their compositional rules. We demonstrate this through two tasks with clearly defined OOD metrics, generated via composable high-level abstractions, on which three standard architectures (MLP, CNN, Transformer) and an object-centric, slot-based architecture fail to generalise OOD. We pair these tasks with two novel attention-based architectures featuring an interpretable final hidden layer designed to expose whether compositional representations emerge. One architecture carries an engineered inductive bias that enables near-perfect OOD performance on one task. Our results show that even with appropriate biases and near-perfect OOD accuracy, a model can fail to learn the compositional feature structures necessary for systematic generalisation. The interpretable layer reveals that successful OOD performance is driven by task-specific biases rather than the discovery of reusable compositional primitives. These findings indicate that OOD benchmarks alone are insufficient for evaluating compositionality in neural networks.

Explore similar work

May 8, 2026cs.LG

Does Your Neural Network Extrapolate? Feature Engineering as Identifiability Bias for OOD Generalization

Successful deep neural networks discover salient features of data. We show when and why they fail to learn out-of-distribution (OOD)-relevant representations from an in-distribution (ID) training window. This requires decoupling feature learning from data-generating-process (DGP) identifiability. From a single training window, OOD extrapolation is non-identifiable: infinitely many DGPs are ε\varepsilon-observationally equivalent on the training data but diverge arbitrarily outside it, and no in-distribution criterion alone reliably breaks the tie. A structural commitment, the feature map, label map, and model class (φ,ψ,M)(\varphi, ψ, \mathcal{M}), dictates the assumed DGP and governs OOD generalization while leaving ID performance essentially unchanged. When architecture, pretraining, augmentation, input formats, or domain knowledge implicitly inject the missing commitment, the model succeeds. When it cannot infer OOD-relevant structure from ID evidence, it fails. Changing only the representation can make the same architecture, at the same in-distribution loss, differ by 520×{\sim}520\times out of distribution. When the commitment is correct and identifiable, OOD error vanishes. For example, Fourier coordinates turn periodic extrapolation into interpolation on S1\mathbb{S}^1. The same mechanism predicts outcomes in three natural-science settings (mass-action chemistry; Kepler's-third-law exoplanet prediction, n=2,362n=2{,}362; and cross-species coding-DNA detection) and in a 264-run positional-encoding study across Transformer, Mamba, and S4D. Finally, a controlled study shows: correct features are necessary but not sufficient. The model class must express the target, and the transformed training data must cover the relevant representation space.
Leonel Aguilar, Jan Nagler, Christoph Hoelscher +1
May 5, 2025cs.LG

A Theoretical Analysis of Provable Compositional Generalization in Neural Networks: A Necessary and Sufficient Condition

Compositional generalization\unicodex2013\unicode{x2013}the ability to systematically process novel combinations of known components\unicodex2013\unicode{x2013}is a hallmark of human intelligence; however, its theoretical foundation in neural networks is not yet well understood. This paper establishes a necessary and sufficient condition for provable compositional generalization, precisely characterizing its boundary. Conceptually, the condition consists of two principles: (i) structural alignment, where a model's computational graph aligns with a task's true compositional hierarchy, and (ii) unambiguous minimized representations, where each component encodes adequate but not redundant information on the training data. The result is fully proved and machine-verified in Lean 4 and holds even in few-shot and one-shot regimes. The necessity direction establishes that provable compositional generalization cannot circumvent these requirements, while the sufficiency direction yields a unified inductive bias that jointly governs architectural design, training data properties, and regularization strategies. Building on this condition, we develop an example algorithmic approach, illustrate it through a controlled minimal example, and further demonstrate the condition on the SCAN jump task. All conclusions are derived mathematically without reliance on empirical validation. Our work provides a theoretical characterization of provable compositional generalization.
Yuanpeng Li
Apr 23, 2026cs.CV

Component-Based Out-of-Distribution Detection

Out-of-Distribution (OOD) detection requires sensitivity to subtle shifts without overreacting to natural In-Distribution (ID) diversity. However, from the viewpoint of detection granularity, global representation inevitably suppress local OOD cues, while patch-based methods are unstable due to entangled spurious-correlation and noise. And neither them is effective in detecting compositional OODs composed of valid ID components. Inspired by recognition-by-components theory, we present a training-free Component-Based OOD Detection (CoOD) framework that addresses the existing limitations by decomposing inputs into functional components. To instantiate CoOD, we derive Component Shift Score (CSS) to detect local appearance shifts, and Compositional Consistency Score (CCS) to identify cross-component compositional inconsistencies. Empirically, CoOD achieves consistent improvements on both coarse- and fine-grained OOD detection.
Wenrui Liu, Hong Chang, Ruibing Hou +2