Semantic Purification for Conditional Representation Learning
Authors: Jiaquan Wang, Yan Lyu, Chen Li, Yuheng Jia
Organizations: School of Computer Science and Engineering, Southeast University, China · Biomedicine Discovery Institute, Department of Biochemistry and Molecular Biology, Monash University, Australia
Conditional representation learning aims to extract criterion-specific features for customized tasks. Recent methods construct conditional subspaces spanned by criterion-specific text bases in the embedding space of vision-language models (VLMs). Image embeddings are then projected onto these subspaces to obtain conditional representations. However, since VLMs are not explicitly trained to disentangle semantics associated with different criteria, the corresponding conditional subspaces remain coupled. This coupling induces semantic leakage during projection, thereby degrading the semantic purity of conditional representations. To suppress semantic leakage, we propose Semantic Purification for Conditional Representation Learning (SP-CRL). Specifically, SP-CRL first decomposes the original text basis and performs curvature-based adaptive truncation on the resulting basis vectors to construct a purer conditional subspace. It then identifies an appropriate noise subspace and projects image embeddings onto its null space to remove irrelevant semantic components. Extensive experiments across customized clustering, customized few-shot classification, and customized retrieval tasks demonstrate that SP-CRL achieves state-of-the-art performance with superior generalization.
Figures & tables
Figure 1: (a) Inter-subspace coupling across criteria. We use normalized subspace affinity ( Soltanolkotabi et al., 2014 ) to quantify coupling between conditional subspaces. The red dashed line marks the expected affinity of randomly oriented subspaces with matched dimensionality. All pairwise affinities exceed this baseline, demonstrating that coupling is inherent and pervasive among conditional subspaces. (b) A geometric illustration of semantic leakage. Under coupled subspaces, the non-target semantic component possesses a non-zero projection onto the target subspace, introducing undesired semantics into the final conditional representation. (c) Semantic selectivity of basis vectors. As singular values decrease, basis vectors exhibit decreasing semantic discriminability for the target criterion (i.e., Color) but increasing semantic discriminability for the non-target criterion (i.e., Shape). The dashed line denotes the curvature-based truncation boundary. Further experimental details are provided in Appendix A . (d) Qualitative comparison on the customized retrieval task based on representation similarity. For color-based retrieval, CLIP ( Radford et al., 2021 ) is dominated by shape semantics and retrieves only gear-shaped samples with colors substantially different from the query. Orthogonal projection reduces the influence of shape semantics but still retrieves several gear-shaped, color-dissimilar samples. In contrast, SP-CRL effectively suppresses semantic leakage, retrieving shape-diverse samples with colors more similar to the query.
Figure 2: The framework of SP-CRL. In CSBR, the original text basis is decomposed via SVD into U , Σ , and V⊤ , where the top k∗ right singular vectors are retained as the refined orthogonal basis T∗ . The truncation boundary k∗ is determined by the maximum curvature of the cumulative energy curve. In NSDP, the candidate noise subspace with the highest denoising priority is identified as the appropriate noise subspace Sn . The image embeddings I are first projected onto the null space of the identified noise subspace to obtain the denoised embeddings I~ (➀), which are then projected onto the target subspace to extract the conditional representations Rt (➁).
Clevr4-10k
Method
Texture
Shape
Color
Count
Mean
NMI
ACC
ARI
NMI
ACC
ARI
NMI
ACC
ARI
NMI
ACC
ARI
SCAN
0.41
11.97
0.86
90.99
89.10
84.03
0.20
11.51
0.01
3.42
14.29
1.23
25.67
TAC
9.61
22.89
5.94
79.07
87.13
75.23
76.63
78.70
67.49
22.05
24.44
11.05
46.69
PCL
9.27
20.03
4.30
70.25
74.71
61.51
74.36
66.93
58.71
27.19
29.06
14.12
42.54
Multi-Map
3.77
17.25
1.81
67.48
66.01
57.40
56.83
56.46
45.73
11.38
20.13
7.67
34.33
Table 1: Performance on the task of customized clustering.
Clevr4-10k
Cards
Flowers
Method
Shape
Color
Number
Color
Mean
1
5
10
1
5
10
1
5
10
1
5
10
CLIP
58.16
83.17
89.47
26.85
57.33
70.00
20.63
33.73
41.84
42.48
71.73
81.84
56.44
CRL
58.69
86.63
92.28
65.71
89.13
93.21
17.65
44.51
51.08
54.55
80.40
87.62
68.46
SP-CRL
68.81
90.26
94.83
82.43
91.43
93.42
33.50
49.55
54.28
72.23
86.34
90.56
75.64
Table 2: Performance on the task of customized few-shot classification.
Method
Texture
Fabric
Shape
Part
Style
Mean
Triplet
13.26
6.28
9.49
4.43
3.33
7.36
CSN
14.09
6.39
11.07
5.13
3.49
8.01
ASEN
15.13
7.11
12.39
5.51
3.56
8.74
ASEN++
15.60
7.67
14.31
6.60
4.07
9.64
RPF
15.62
8.30
15.02
7.38
4.77
10.22
CLIP
9.14
4.68
7.86
4.26
4.48
6.08
Table 3: Performance on the task of customized retrieval. Symbol ‡ signifies that training is conducted.
Components
Clevr4-10k
Cards
CSBR
NSDP
Shape
Color
Number
Suit
NMI
ACC
ARI
NMI
ACC
ARI
NMI
ACC
ARI
NMI
ACC
ARI
81.72
82.12
75.00
73.91
70.78
63.99
36.86
39.38
23.27
35.99
55.41
28.44
✓
82.62
85.65
77.85
87.87
85.06
80.35
38.10
40.76
24.08
39.65
70.34
39.00
✓
82.48
82.16
76.14
78.95
75.68
68.91
37.33
39.97
25.78
39.39
58.89
33.18
✓
✓
86.67
90.69
84.11
89.99
89.88
85.09
38.78
44.45
28.40
41.61
73.22
42.20
Table 4: Ablation study on components.
Clevr4-10k
Cards
Retention Ratio
Shape
Color
Number
Suit
NMI
ACC
ARI
NMI
ACC
ARI
NMI
ACC
ARI
NMI
ACC
ARI
20%
81.38
85.38
76.40
86.86
84.63
79.28
22.57
26.95
11.85
38.70
67.90
37.43
40%
81.99
84.70
77.05
85.10
82.08
76.55
31.74
34.72
18.30
35.82
58.23
31.16
60%
81.89
82.57
75.68
80.85
76.96
70.88
38.04
40.58
23.53
37.09
57.86
30.77
80%
82.78
84.08
77.19
76.63
72.55
66.07
37.85
40.66
23.86
36.05
57.51
30.08
Table 5: Ablation study on basis truncation.
Clevr4-10k (Target Subspace)
Noise Subspace
Texture
Shape
Color
Count
NMI
ACC
ARI
NMI
ACC
ARI
NMI
ACC
ARI
NMI
ACC
ARI
None
13.80
27.86
8.60
82.62
85.65
77.85
87.87
85.06
80.35
25.74
28.56
14.22
Texture
-
-
-
82.55
84.16
76.92
87.97
88.28
82.04
26.38
29.33
14.88
Shape
10.45
24.75
6.05
-
-
-
89.99 ⋆
89.88 ⋆
85.09 ⋆
29.00
29.16
15.01
Color
14.27 ⋆
28.27 ⋆
9.53 ⋆
86.67 ⋆
90.69 ⋆
84.11 ⋆
-
-
-
26.61
27.49
13.28
Table 6: Ablation study on noise subspace identification.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Semantic discriminability of SVD-derived basis vectors across different criteria. Basis vectors are ranked by descending singular values. Solid curves represent the mean η2 for each bin of 10 basis vectors, while shaded regions denote the corresponding standard deviation. Dashed vertical lines indicate the curvature-based truncation boundary k∗ determined by CSBR.
Dataset
Target / Non-target
CLIP
Orthogonal Projection
SP-CRL
Target ↑
Non-target ↓
Target ↑
Non-target ↓
Target ↑
Non-target ↓
Cards
Number / Suit
77.00
72.82
64.31
60.85
62.61
53.96
Suit / Number
88.99
56.19
84.16
42.02
82.89
33.02
Clevr4-10k
Color / Shape
98.84
98.98
98.68
97.85
97.52
66.38
Shape / Color
99.07
98.68
98.21
97.77
98.45
89.76
Mean
90.98
81.67
86.34
74.62
85.37
60.78
Appendix
Table 7: Linear probing prediction accuracy under target and non-target criteria.