Semantic Purification for Conditional Representation Learning
Authors: Jiaquan Wang, Yan Lyu, Chen Li, Yuheng Jia
Organizations: School of Computer Science and Engineering, Southeast University, China · Biomedicine Discovery Institute, Department of Biochemistry and Molecular Biology, Monash University, Australia
Conditional representation learning aims to extract criterion-specific features for customized tasks. Recent methods construct conditional subspaces spanned by criterion-specific text bases in the embedding space of vision-language models (VLMs). Image embeddings are then projected onto these subspaces to obtain conditional representations. However, since VLMs are not explicitly trained to disentangle semantics associated with different criteria, the corresponding conditional subspaces remain coupled. This coupling induces semantic leakage during projection, thereby degrading the semantic purity of conditional representations. To suppress semantic leakage, we propose Semantic Purification for Conditional Representation Learning (SP-CRL). Specifically, SP-CRL first decomposes the original text basis and performs curvature-based adaptive truncation on the resulting basis vectors to construct a purer conditional subspace. It then identifies an appropriate noise subspace and projects image embeddings onto its null space to remove irrelevant semantic components. Extensive experiments across customized clustering, customized few-shot classification, and customized retrieval tasks demonstrate that SP-CRL achieves state-of-the-art performance with superior generalization.
Figures & tables
Figure 1: (a) Inter-subspace coupling across criteria. We use normalized subspace affinity ( Soltanolkotabi et al., 2014 ) to quantify coupling between conditional subspaces. The red dashed line marks the expected affinity of randomly oriented subspaces with matched dimensionality. All pairwise affinities exceed this baseline, demonstrating that coupling is inherent and pervasive among conditional subspaces. (b) A geometric illustration of semantic leakage. Under coupled subspaces, the non-target semantic component possesses a non-zero projection onto the target subspace, introducing undesired semantics into the final conditional representation. (c) Semantic selectivity of basis vectors. As singular values decrease, basis vectors exhibit decreasing semantic discriminability for the target criterion (i.e., Color) but increasing semantic discriminability for the non-target criterion (i.e., Shape). The dashed line denotes the curvature-based truncation boundary. Further experimental details are provided in Appendix A . (d) Qualitative comparison on the customized retrieval task based on representation similarity. For color-based retrieval, CLIP ( Radford et al., 2021 ) is dominated by shape semantics and retrieves only gear-shaped samples with colors substantially different from the query. Orthogonal projection reduces the influence of shape semantics but still retrieves several gear-shaped, color-dissimilar samples. In contrast, SP-CRL effectively suppresses semantic leakage, retrieving shape-diverse samples with colors more similar to the query.
Figure 2: The framework of SP-CRL. In CSBR, the original text basis is decomposed via SVD into U , Σ , and V⊤ , where the top k∗ right singular vectors are retained as the refined orthogonal basis T∗ . The truncation boundary k∗ is determined by the maximum curvature of the cumulative energy curve. In NSDP, the candidate noise subspace with the highest denoising priority is identified as the appropriate noise subspace Sn . The image embeddings I are first projected onto the null space of the identified noise subspace to obtain the denoised embeddings I~ (➀), which are then projected onto the target subspace to extract the conditional representations Rt (➁).
Clevr4-10k
Method
Texture
Shape
Color
Count
Mean
NMI
ACC
ARI
NMI
ACC
ARI
NMI
ACC
ARI
NMI
ACC
ARI
SCAN
0.41
11.97
0.86
90.99
89.10
84.03
0.20
11.51
0.01
3.42
14.29
1.23
25.67
TAC
9.61
22.89
5.94
79.07
87.13
75.23
76.63
78.70
67.49
22.05
24.44
11.05
46.69
PCL
9.27
20.03
4.30
70.25
74.71
61.51
74.36
66.93
58.71
27.19
29.06
14.12
42.54
Multi-Map
3.77
17.25
1.81
67.48
66.01
57.40
56.83
56.46
45.73
11.38
20.13
7.67
34.33
Table 1: Performance on the task of customized clustering.
Clevr4-10k
Cards
Flowers
Method
Shape
Color
Number
Color
Mean
1
5
10
1
5
10
1
5
10
1
5
10
CLIP
58.16
83.17
89.47
26.85
57.33
70.00
20.63
33.73
41.84
42.48
71.73
81.84
56.44
CRL
58.69
86.63
92.28
65.71
89.13
93.21
17.65
44.51
51.08
54.55
80.40
87.62
68.46
SP-CRL
68.81
90.26
94.83
82.43
91.43
93.42
33.50
49.55
54.28
72.23
86.34
90.56
75.64
Table 2: Performance on the task of customized few-shot classification.
Method
Texture
Fabric
Shape
Part
Style
Mean
Triplet
13.26
6.28
9.49
4.43
3.33
7.36
CSN
14.09
6.39
11.07
5.13
3.49
8.01
ASEN
15.13
7.11
12.39
5.51
3.56
8.74
ASEN++
15.60
7.67
14.31
6.60
4.07
9.64
RPF
15.62
8.30
15.02
7.38
4.77
10.22
CLIP
9.14
4.68
7.86
4.26
4.48
6.08
Table 3: Performance on the task of customized retrieval. Symbol ‡ signifies that training is conducted.
Components
Clevr4-10k
Cards
CSBR
NSDP
Shape
Color
Number
Suit
NMI
ACC
ARI
NMI
ACC
ARI
NMI
ACC
ARI
NMI
ACC
ARI
81.72
82.12
75.00
73.91
70.78
63.99
36.86
39.38
23.27
35.99
55.41
28.44
✓
82.62
85.65
77.85
87.87
85.06
80.35
38.10
40.76
24.08
39.65
70.34
39.00
✓
82.48
82.16
76.14
78.95
75.68
68.91
37.33
39.97
25.78
39.39
58.89
33.18
✓
✓
86.67
90.69
84.11
89.99
89.88
85.09
38.78
44.45
28.40
41.61
73.22
42.20
Table 4: Ablation study on components.
Clevr4-10k
Cards
Retention Ratio
Shape
Color
Number
Suit
NMI
ACC
ARI
NMI
ACC
ARI
NMI
ACC
ARI
NMI
ACC
ARI
20%
81.38
85.38
76.40
86.86
84.63
79.28
22.57
26.95
11.85
38.70
67.90
37.43
40%
81.99
84.70
77.05
85.10
82.08
76.55
31.74
34.72
18.30
35.82
58.23
31.16
60%
81.89
82.57
75.68
80.85
76.96
70.88
38.04
40.58
23.53
37.09
57.86
30.77
80%
82.78
84.08
77.19
76.63
72.55
66.07
37.85
40.66
23.86
36.05
57.51
30.08
Table 5: Ablation study on basis truncation.
Clevr4-10k (Target Subspace)
Noise Subspace
Texture
Shape
Color
Count
NMI
ACC
ARI
NMI
ACC
ARI
NMI
ACC
ARI
NMI
ACC
ARI
None
13.80
27.86
8.60
82.62
85.65
77.85
87.87
85.06
80.35
25.74
28.56
14.22
Texture
-
-
-
82.55
84.16
76.92
87.97
88.28
82.04
26.38
29.33
14.88
Shape
10.45
24.75
6.05
-
-
-
89.99 ⋆
89.88 ⋆
85.09 ⋆
29.00
29.16
15.01
Color
14.27 ⋆
28.27 ⋆
9.53 ⋆
86.67 ⋆
90.69 ⋆
84.11 ⋆
-
-
-
26.61
27.49
13.28
Table 6: Ablation study on noise subspace identification.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Semantic discriminability of SVD-derived basis vectors across different criteria. Basis vectors are ranked by descending singular values. Solid curves represent the mean η2 for each bin of 10 basis vectors, while shaded regions denote the corresponding standard deviation. Dashed vertical lines indicate the curvature-based truncation boundary k∗ determined by CSBR.
Dataset
Target / Non-target
CLIP
Orthogonal Projection
SP-CRL
Target ↑
Non-target ↓
Target ↑
Non-target ↓
Target ↑
Non-target ↓
Cards
Number / Suit
77.00
72.82
64.31
60.85
62.61
53.96
Suit / Number
88.99
56.19
84.16
42.02
82.89
33.02
Clevr4-10k
Color / Shape
98.84
98.98
98.68
97.85
97.52
66.38
Shape / Color
99.07
98.68
98.21
97.77
98.45
89.76
Mean
90.98
81.67
86.34
74.62
85.37
60.78
Appendix
Table 7: Linear probing prediction accuracy under target and non-target criteria.
Contrastive Language-Image Pretraining learns a shared representation space through large-scale contrastive learning. However, existing methods that enforce global consistency regularization overlook a key challenge: the inherent information asymmetry between images and text: captions typically describe only one specific aspect of an image, thus images with similar visual content can be paired with completely divergent textual content and semantic information. Consequently, global regularizers inadvertently impose constraints between visually similar images whose captions describe divergent aspects, introducing semantic distortion into the representation space. We propose AspectCLIP, a framework that reformulates consistency regularization to respect this one-to-many structure. AspectCLIP first partitions training samples into attribute clusters based on textual similarity to identify aspect-coherent groups, then applies full cyclic consistency within each cluster while restricting cross-cluster regularization to prototype-level comparisons. This aspect-guided regularization enforces strict geometric alignment only when images and texts describe a consistent facet, while allowing flexibility across divergent aspects. Extensive experiments on downstream tasks demonstrate that AspectCLIP consistently outperforms traditional methods and achieves a more structured representation space.
Yiyang Yao, Shanglin Liu, Jianming Lv +4
School of Computer Science and Technology, South China University of Technology, Guangzhou, China
Human perception of visual similarity is inherently adaptive and subjective, depending on the users' interests and focus. However, most image retrieval systems fail to reflect this flexibility, relying on a fixed, monolithic metric that cannot incorporate multiple conditions simultaneously. To address this, we propose CLAY, an adaptive similarity computation method that reframes the embedding space of pretrained Vision-Language Models (VLMs) as a text-conditional similarity space without additional training. This design separates the textual conditioning process and visual feature extraction, allowing highly efficient and multi-conditioned retrieval with fixed visual embeddings. We also construct a synthetic evaluation dataset CLAY-EVAL, for comprehensive assessment under diverse conditioned retrieval settings. Experiments on standard datasets and our proposed dataset show that CLAY achieves high retrieval accuracy and notable computational efficiency compared to previous works.
Image-based Joint-Embedding Predictive Architecture (I-JEPA) offers a promising approach to visual self-supervised learning through masked feature prediction. However with the inherent visual uncertainty at masked positions, feature prediction remains challenging and may fail to learn semantic representations. In this work, we propose Text-Conditional JEPA (TC-JEPA) that uses image captions to reduce the prediction uncertainty. Specifically, we modulate the predicted patch features using a fine-grained text conditioner that computes sparse cross-attention over input text tokens. With such conditioning, patch features become predictable as a function of text, thus are more semantically meaningful. We show TC-JEPA improves downstream performance and training stability, with promising scaling properties. TC-JEPA also offers a new vision-language pretraining paradigm based on feature prediction only, outperforming contrastive methods on diverse tasks, especially those requiring fine-grained visual understanding and reasoning.