cs.CVSep 30, 2026

SAGE: Salient Factor Discovery and Generation with Visual Foundation Representations

Authors: Shuang Liang, Lejun Liao, Shiyuan Zhang, Max C. Zhang, Xiaolong Luo, Han Wang, Stefano Anzellotti, Yuan Yuan

Organizations: HKU · Boston College · University of Virginia · Harvard University

Abstract

Given a target dataset, such as faces with eyeglasses, and a background dataset, such as faces without, contrastive analysis separates \textit{salient} factors specific to the target from \textit{common} content shared by both. We aim for salient representations that capture target-specific detail in each image, such as the shape, color, and position of the glasses, so that they reveal subtypes without subtype labels and guide the generation of new examples of a discovered subtype, even one with no name or text description. We introduce SAGE, which learns both factors directly in the high-dimensional spatial latent of a frozen representation autoencoder and conditions a diffusion transformer on the learned salient representation of a reference image. On Digits-ImageNet and FFHQ eyeglasses, SAGE combines high-fidelity \textit{reconstruction} (rFID below 22) with unsupervised \textit{subtype discovery}, recovering the digits better than baselines (probe accuracy 0.9500.950 vs.\ at most 0.2810.281) and revealing eyewear types, finer sunglasses styles, and mislabeled images; salient-conditioned \textit{generation} raises Digits-ImageNet subtype accuracy over the unfactorized latent (90.5%90.5\% vs.\ 27.7%27.7\%) and diversity on both datasets. On retinal OCT, SAGE's salient space separates three diseases using only normal/disease labels.

Figures & tables

Appendix figures & tables16 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder

    Jun 9, 2026Yitong Chen, Zijie Diao, Junke Wang +5Autoregressive Image GenerationAutoencoder Architectures

  2. Self-Supervised Representation Learning via Hyperspherical Density Shaping

    Apr 27, 2026Esteban Rodríguez-Betancourt, Edgar Casasola-MurilloSelf-Supervised RepresentationsDiscriminative Congruence Transform

  3. HiRAE: Hierarchical Representation Autoencoding with Residual Budgets

    Sep 29, 2026Xuanyu Zhu, Yan Bai, Yang Shi +5Autoregressive Image GenerationAutoencoder Architectures