This paper studies the problem of learning disentangled representations of objects and their attributes from raw, unstructured image data. Slot-based methods have shown considerable success in unsupervised learning of object representations from images. Block-slot attention-based methods extend this framework to attribute representations by assuming a uniform factorization of object representations into attributes, which may be suboptimal and consequently limit the quality of the learned representations. We therefore investigate a framework for jointly discovering object and attribute representations. Our key contribution is leveraging the Linear Representation Hypothesis (LRH), which postulates that composable concepts can be represented as linearly additive subspaces in slot representations. Based on this insight, we propose a probabilistic model connecting images, slots (objects), and blocks (attributes). We present an architecture that leverages block attention to connect attribute representations to slots and incorporates LRH in both object and attribute representation spaces. This architecture effectively optimizes the Evidence Lower Bound (ELBO) of the proposed graphical model. Our experiments demonstrate (i) effective discovery of disentangled object and attribute representations, (ii) empirical evidence for LRH in slot space, and (iii) the ability to perform image editing owing to the disentangled and interpretable nature of the learned representations. Our experiments on multiple datasets demonstrate improvements in DCI scores over state-of-the-art methods.
Figures & tables
Figure 1: Overview of the proposed representation and generative model.
Figure 3: Overview of the inference: The image is encoded into spatial features E , from which slots S are inferred via slot attention. For each slot, projected features v are used to iteratively refine M attribute blocks through competitive attention, with each block constrained to a linear combination of shared concept vectors. The resulting blocks serve as disentangled attribute representations.
Figure 4: Object Manipulation via delta directions: Delta direction is estimated by subtracting slot of object indicated from first image from slot of the object indicated in second image. This delta direction is then added to the object indicated in the third image. The fourth image is decoded by slots manipulated slots from third image. First three images are reconstruction.
CLEVR Easy
CLEVR Hard
CLEVR-Tex
MOVi-C
Method
Slot Attn.
SLATE+
Slot Attn.
SLATE+
SLATE+
Slot Diff.
CODA
Same δ
0.6610
0.7465
0.6024
0.6511
0.5618
0.6356
0.2230
Different δ
0.2232
0.1647
0.1898
0.1493
0.1498
0.1719
0.1020
Table 1: Cosine similarity between two estimations of the same delta direction.
Figure 5: Quantitative Results: Comparison of DCI and FG-ARI scores for our model and baselines across all datasets.
Figure 6: Qualitative Results: Comparison of attribute swapping results between our model and SysBinder. Blue arrows indicate the objects participating in the swaps. Recon. denotes the image reconstructed by the model.
Table 7
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Module
Hyperparameter
Dataset
CLEVR
CLEVR
CLEVR
Easy
Hard
Tex
General
Batch Size
64
64
64
Image Size
128
128
128
Training Steps
200K
260K
315K
Encoder Learning Rate
0.0001
0.0001
0.0001
Appendix
Table 4: Hyperparameters of training for CLEVR-Easy, CLEVR-Hard, and CLEVR-Tex datasets.
Figure 7: Training phases of the model.
Method
FG-ARI
SLATE
0.8166
SLATE+
0.9369
Appendix
Table 5: SLATE vs SLATE+ Results on CLEVR-Easy
Dataset (model)
Factor
#Values
Stable rank
Part. ratio
CLEVR-Easy (Slot Attention, dslot=64 )
Position
cont.
3.21±0.38
5.58±0.53
Color
8
2.93±0.62
4.91±1.46
Shape
3
2.06±0.16
3.70±0.48
CLEVR-Hard (SLATE+, dslot=192 )
Position
cont.
3.89±0.05
8.08±0.04
Color
137
3.82±0.11
7.24±0.34
Shape
3
2.79±0.19
5.37±0.52
Appendix
Table 6: Stable rank and participation ratio of the delta-direction subspace of each factor (mean ± std over seeds).
Figure 8: Steering slots along principal axes of the delta space: We steer slots by adding a scaled multiple α of a principal axis computed from the delta directions associated with a particular factor. The slots corresponding to the objects highlighted by red boxes in the first column are steered, with α increasing from left to right. The resulting images exhibit meaningful and smooth changes along the corresponding factor, consistently across different objects, datasets, and methods, including Slot Attention, SLATE+, and Slot Diffusion.
School of Artificial Intelligence, Guilin University of Electronic Technology · Department of Computer Science, Aalto University · Center for Machine Vision and Signal Analysis, University of Oulu +1