Given a target dataset, such as faces with eyeglasses, and a background dataset, such as faces without, contrastive analysis separates \textit{salient} factors specific to the target from \textit{common} content shared by both. We aim for salient representations that capture target-specific detail in each image, such as the shape, color, and position of the glasses, so that they reveal subtypes without subtype labels and guide the generation of new examples of a discovered subtype, even one with no name or text description. We introduce SAGE, which learns both factors directly in the high-dimensional spatial latent of a frozen representation autoencoder and conditions a diffusion transformer on the learned salient representation of a reference image. On Digits-ImageNet and FFHQ eyeglasses, SAGE combines high-fidelity \textit{reconstruction} (rFID below 2) with unsupervised \textit{subtype discovery}, recovering the digits better than baselines (probe accuracy 0.950 vs.\ at most 0.281) and revealing eyewear types, finer sunglasses styles, and mislabeled images; salient-conditioned \textit{generation} raises Digits-ImageNet subtype accuracy over the unfactorized latent (90.5% vs.\ 27.7%) and diversity on both datasets. On retinal OCT, SAGE's salient space separates three diseases using only normal/disease labels.
Figures & tables
Figure 1: Overview of SAGE. Given background (BG) and target (TG) datasets, SAGE uses only dataset labels to decompose frozen self-supervised visual representations into common and salient factors, capturing shared content and target-specific variation, respectively. The decomposition supports high-fidelity reconstruction, subtype discovery in the salient space, and salient-conditioned generation that preserves target attributes while varying common content.
Figure 2: Stage 1 of SAGE for salient factor discovery. Trainable common and salient encoders split latents of the frozen RAEv2 encoder (DINOv3); the frozen RAEv2 decoder maps factors back to images. Top: reconstruction ( Lrec ) and swap adversarial training ( LswapG ). Bottom: swapped images are re-encoded for Lcyc and LcNCE , while Llsc and the sparsity priors act directly on the factors. Snowflakes and flames mark frozen and trainable modules.
Figure 3: Stage 2 of SAGE for salient-conditioned generation. The frozen RAEv2 and salient encoders extract a reference salient representation, max-pooled to cs . An RAEv2-initialized diffusion transformer is trained with conditional flow matching in this latent. Top right: generated samples; each row keeps the eyewear type or digit of the reference at left while faces and scenes vary.
Digits-ImageNet
FFHQ
Reconstruction
Salient representation
Reconstruction
Salient representation
Method
rFID ↓
PSNR ↑
SSIM ↑
LPIPS ↓
BG/TG Sil. ↑
LP Acc. ↑
ARI/NMI ↑
rFID ↓
PSNR ↑
SSIM ↑
LPIPS ↓
BG/TG Sil. ↑
LP Acc. ↑
ARI/NMI ↑
cVAE ( Abid and Zou, 2019 )
151.0
18.33
0.478
0.673
0.006
0.135
0.001/0.003
188.2
18.42
0.567
0.590
0.027
0.898
0.001/0.001
SepVAE ( Louiset et al., 2023 )
162.5
17.50
0.461
0.696
0.006
0.145
0.009/0.019
197.4
17.67
0.552
0.598
0.206
0.897
-0.001/0.000
Double-InfoGAN ( Carton et al., 2024 )
171.4
15.97
0.375
0.668
0.013
0.281
0.001/0.003
122.5
16.73
0.424
0.458
0.050
0.964
0.298/0.299
SepCLR ( Louiset et al., 2024 )
N/A
N/A
N/A
N/A
0.546
0.180
0.000/0.001
N/A
N/A
N/A
N/A
0.752
0.970
0.822/0.696
Table 1: Benchmark comparison on Digits-ImageNet and FFHQ. rFID: reconstruction Fréchet Inception Distance; N/A: no decoder; Double-InfoGAN uses its native 128×128 resolution; all other methods use 256×256 . BG/TG Sil.: silhouette score of background versus target salient representations (Appendix Table 5 ). LP Acc.: linear-probe subtype accuracy, five-fold on Digits-ImageNet and balanced on FFHQ (Appendix B.4 ). ARI/NMI: k -means on the ℓ2 -normalized, unprojected salient representations, with k set to the subtype count (Appendix B.5 ). FFHQ labels are audit-corrected.
Figure 4: Reconstruction, common-only, and salient-only decoding. Columns show the input, the full reconstruction D(zc,zs) , common-only decoding D(zc,0) , and salient-only decoding D(0,zs) on Digits-ImageNet (left) and FFHQ (right); rows compare cVAE, SepVAE, Double-InfoGAN, and SAGE. A clean salient factor renders only the digit or eyeglasses.
Figure 5: Subtype discovery in the salient space. Top: t-SNE of SAGE’s salient representations separates background and target images (a, d) and organizes digit and eyewear subtypes (b, e), whereas raw DINOv3 features mix them (c, f). Bottom: examples from the k -means clusters marked in (b, e) show coherent digit and eyewear groups (g, h); FFHQ cluster C contains images mislabeled as wearing glasses, and red points are audited no-glasses images.
Figure 6: Salient swapping and erasure.
Figure 7: Salient interpolation. For two held-out target images xA and xB , salient-only rows decode D(zcA,zsα) with zsα=(1−α)zsA+αzsB ; full-latent rows interpolate both factors.
Figure 8: OCT disease subtypes and salient edits. Left: salient t-SNE of disease scans, colored by disease subtype labels unused in training. BG → TG adds a disease scan’s salient factor to a normal scan; TG → BG removes a disease scan’s salient factor.
Dataset
Conditioning
Uncond. gFID ↓
Cond. gFID ↓
Subtype Acc. ↑
Vendi ↑
Cos-to-ref ↓
Digits-ImageNet
Real images
–
–
99.3
15.79
–
Raw DINOv3 GMP
8.86
2.40
27.70
2.65
0.794
SAGE Salient GMP
4.76
4.09
90.48
12.50
0.116
FFHQ
Real images
–
–
99.1
10.09
–
Raw DINOv3 GMP
16.07
11.17
98.5
2.32
0.779
SAGE Salient GMP
15.80
16.03
96.6
6.96
0.470
Table 2: Salient-conditioned generation. Raw DINOv3 GMP conditions on the unfactorized latent. Subtype Acc.: share of samples assigned the reference subtype by a DINOv3-L classifier trained on real targets. Vendi (diversity) and Cos-to-ref (reference similarity) use 16 samples per reference and DINOv2-L features. Real images : classifier accuracy and within-subtype Vendi (Appendix D.2 ).
Full Reconstruction
Common-only
Salient Linear Probe
Clustering
Configuration
rFID ↓
PSNR ↑
SSIM ↑
LPIPS ↓
rFID ↓
SSIM ↑
Digit ↑
ImageNet ↓
ARI/NMI ↑
SAGE (all objectives)
1.78
20.56
0.571
0.216
1.89
0.569
0.950
0.045
0.337/0.472
w/o swap adversarial ( LswapG )
1.01
20.67
0.566
0.210
4.88
0.544
0.123
0.169
0.000/0.001
w/o target sparsity ( LTG-sp )
1.15
21.06
0.598
0.193
1.36
0.585
0.923
0.302
0.274/0.409
w/o Cycle NCE ( LcNCE )
1.29
20.78
0.576
0.205
1.58
0.573
0.950
0.100
0.348/0.479
w/o image-space cycle consistency ( Lcyc )
2.84
19.88
0.562
0.236
3.46
0.559
0.951
0.086
0.360/0.496
Table 3: Loss ablations on Digits-ImageNet. Common-only: rFID and SSIM of D(zc,0) for target images against their digit-free ImageNet backgrounds. Salient probes measure digit identity (target content, higher is better) and ImageNet category (shared content, lower is better). ARI/NMI are computed as in Table 1 ; their standard deviations over three k -means seeds are at most 0.0007 .
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9: Additional SAGE factor-decoding examples. Randomly selected images. For Digits-ImageNet (a–d) and FFHQ (e–h), complete decoding reconstructs the target image, common-only decoding removes the digit or eyeglasses while preserving the scene or identity, and salient-only decoding retains the digit’s shape and location or the specific eyewear style with little common content.
Digits-ImageNet ( k=10 , n=24,739 )
FFHQ ( k=2 , n=3,398 )
FFHQ ( k=3 , n=3,437 )
Method
Dim.
PCA-2D
PCA-50
Raw
PCA-2D
PCA-50
Raw
PCA-2D
PCA-50
Raw
ARI ↑
cVAE
64
0.000
0.001
0.001
0.001
0.000
0.001
0.036
0.017
0.016
SepVAE
64
0.006
0.009
0.009
-0.001
-0.001
-0.001
0.012
0.019
0.017
Double-InfoGAN
64
0.005
0.001
0.001
0.324
0.292
0.298
0.247
0.303
0.303
SepCLR
32
0.000
0.000
0.000
0.808
0.822
0.822
0.441
0.455
0.455
Appendix
Table 4: Clustering readout ablation on ℓ2 -normalized salient representations (ARI/NMI). All vectors are normalized before optional projection. Raw denotes no dimensionality reduction; PCA-50 retains min(50,Dim.) components. Scores are means over k -means seeds {0,1,2} . FFHQ uses manually verified annotations, with 39 no-glasses images excluded for k=2 and included for k=3 .
Figure 10: t-SNE visualizations of salient representations across contrastive analysis methods. Columns show cVAE, SepVAE, Double-InfoGAN, SepCLR, DINOv3 + SepCLR, and SAGE. Rows show Digits-ImageNet background-versus-target and digit-colored target representations, then their FFHQ counterparts with audited eyewear labels. Blue/orange denote background/target; green/purple/red denote reading glasses/sunglasses/audited no-glasses images.
Method
Digits t-SNE
Digits original
FFHQ t-SNE
FFHQ original
cVAE
0.022
0.006
0.004
0.027
SepVAE
0.003
0.006
0.394
0.206
Double-InfoGAN
0.248
0.013
0.188
0.050
SepCLR
0.381
0.546
0.449
0.752
DINOv3 + SepCLR
0.365
0.642
0.387
0.551
SAGE
0.417
0.820
0.462
0.873
Appendix
Table 5: Silhouette scores for BG-vs.-TG salient separation in t-SNE and original spaces. Scores treat background (BG) and target (TG) salient representations as the two groups. The silhouette coefficient ( Rousseeuw, 1987 ) measures within-group cohesion relative to separation; higher is better.
Figure 11: Fine-grained sunglasses structure in SAGE’s salient representation. Left: FFHQ target embeddings colored by audited eyewear labels, with C1–C3 marking subgroups within sunglasses. Right: example images from each subgroup: Wayfarer-like (C1, n=168 ), mixed styles (C2, n=390 ), and Aviator-like (C3, n=175 ). Group names describe the displayed examples; no style labels are used in training.
Figure 12: Azure eyewear annotation errors identified by manual inspection of FFHQ. Left: t-SNE of SAGE’s target salient representations, colored by manually verified labels; outlined points mark the 125 Azure label disagreements. Right: all 125 corrected images, grouped by verified label: sunglasses (65), no glasses (39), and reading glasses (21). Tile titles give the original Azure category and image ID.
Figure 13: Remaining SAGE errors among manually verified no-glasses FFHQ images. (a) Target salient-space t-SNE, colored by audited labels: reading glasses (green, 2,636), sunglasses (purple, 762), and no glasses (red, 39). Black-outlined red points identify five no-glasses cases associated with glasses groups. (b) Each numbered row shows one query and its six nearest neighbors; border colors indicate human labels and values report cosine similarity. Arrows mark the group each query is associated with. (c) The other 34 no-glasses images form a separate group.
Figure 14: Additional salient swapping and erasure examples. Left (BG → TG): D(zcb,zst) combines a background image’s common factor with a target image’s salient factor. Right (TG → BG): D(zct,0) decodes the target common factor alone. Rows show Digits-ImageNet and FFHQ examples; red borders mark decoded outputs.
Method
rFID ↓
PSNR ↑
SSIM ↑
LPIPS ↓
LP Acc. ↑
ARI/NMI ↑
SAGE (Ours)
27.21
24.55
0.448
0.273
0.960
0.415/0.415
Appendix
Table 6: Stage-1 metrics on OCT-Kermany. Evaluation uses the official test set ( Kermany et al., 2018 ) , with 250 normal images and 250 images from each of CNV, DME, and DRUSEN, from patients not included in training or validation. LP Acc. reports mean stratified five-fold linear-probe accuracy on the 750 target images; ARI/NMI use k -means directly on SAGE’s ℓ2 -normalized, 1,024-dimensional salient representations obtained by global max pooling (GMP), without dimensionality reduction. With only 1,000 evaluation images, rFID is sensitive to finite-sample bias and is not directly comparable to scores computed on larger evaluation sets.
Figure 15: Generation conditioned on a reference image. For FFHQ (left) and Digits-ImageNet (right), each column shows a reference (top) and one sample conditioned on Raw DINOv3 GMP (middle) or on SAGE’s salient representation (bottom).
Full reconstruction
Common-only
Digit LP
ImageNet LP
Clustering
Configuration
rFID ↓
PSNR ↑
SSIM ↑
LPIPS ↓
rFID ↓
SSIM ↑
zs ↑
zc ↓
zc ↑
zs ↓
ARI/NMI ↑
Raw z
0.39
22.18
0.613
0.157
N/A
N/A
0.328
N/A
0.765
N/A
0.000/0.001
SAGE (all)
1.78
20.56
0.571
0.216
1.89
0.569
0.950
0.336
0.765
0.045
0.337/0.472
w/o LswapG
1.01
20.67
0.566
0.210
4.88
0.544
0.123
0.356
0.763
0.169
0.000/0.001
w/o LTG-sp
1.15
21.06
0.598
0.193
1.36
0.585
0.923
0.656
0.758
0.302
0.274/0.409
w/o LcNCE
1.29
20.78
0.576
0.205
1.58
0.573
0.950
0.335
0.767
0.100
0.348/0.479
Appendix
Table 7: Full loss ablation of SAGE on Digits-ImageNet. ARI/NMI use Euclidean k -means on ℓ2 -normalized original salient representations, without dimensionality reduction (Appendix B.5 ). ARI/NMI are means over k -means seeds {0,1,2} ; standard deviations are omitted because they are at most 0.0007 . Digit/ImageNet LP: linear-probe accuracy for digit or ImageNet labels from the indicated representation; Raw z : the unfactorized latent, whose probes use z itself. Objectives are named as in Table 3 .
Split
NORMAL
CNV
DME
DRUSEN
Total
Patients
Train
46,026
33,485
10,213
7,754
97,478
4,772
Val.
5,114
3,720
1,135
862
10,831
3,329
Test
250
250
250
250
1,000
635
Appendix
Table 8: OCT-Kermany split statistics after preprocessing. CNV, DME, and DRUSEN are the target diseases; normal scans are background.
Component
DINOv3 + SepCLR
SAGE
Frozen visual backbone
DINOv3-L (MLS)
DINOv3-L (MLS)
Latent normalization
RAEv2
RAEv2
Common/salient Transformers
Shared architecture
Shared architecture
Evaluation pooling
GAP
GMP
Evaluation representation
32 dimensions
1,024 dimensions
Training objectives
SepCLR objectives
SAGE objectives
Appendix
Table 9: Components of DINOv3 + SepCLR and SAGE.
Figure 16: Common and salient encoder architecture. Input and output convolutions use kernel size 3. The Transformer block applies layer normalization, self-attention, and a feed-forward network (FFN), with residual connections around the attention and FFN sublayers. The two encoders produce the common and salient representations, respectively. In the salient encoder, both the weights and biases of Conv_out are initialized to zero.
Figure 17: Shared multi-expert discriminator architecture. A frozen DINOv1 ViT-S/8 backbone provides features from layers 2, 5, 8, 11, and the final layer to separate background (BG) and target (TG) experts. Each expert contains trainable heads for the corresponding feature levels, enabling group-specific real/fake discrimination. Snowflake and flame symbols indicate frozen and trainable components, respectively.