cs.LGOct 7, 2026

Shared Gaussianization: What Gaussian Regularizers Certify About Contrastive Learning, and What They Miss

Authors: Ruoyu Zhao, Yuting Chen, Jinheng Zhang, Zhehao Zou, Tong Che

Organizations: City University of Hong Kong · Georgia Institute of Technology · University of Pennsylvania · The Chinese University of Hong Kong · NVIDIA Research

Abstract

What can a distribution-matching regularizer such as SIGReg in LeJEPA certify about contrastive learning? We study shared Gaussianization (SG), a characteristic-function Gaussianity test on the average of two normalized views, scaled by an independent χdχ_d radius. Because disagreeing views shorten the average, one test detects both misalignment and non-uniformity. SG vanishes exactly at the aligned, uniform minimizers of population InfoNCE, and under equal marginals it bounds the InfoNCE excess by 4⋅33/4β4\cdot 3^{3/4}β times the square root of the SG loss, plus a term linear in the loss. The square-root rate and this dimension-free constant are sharp, and no squared mean-embedding distance on view pairs achieves a faster rate. With an explicit alignment term, a rotation-invariant uniformity test gives a linear bound if and only if its spectrum dominates that of InfoNCE's kernel eβu⊤ve^{βu^\top v}; SG's own test does, Gaussian kernels e−γ∥u−v∥2e^{-γ\|u-v\|^2} qualify exactly when γ≥β/2γ\ge β/2, and moment matching never does. Away from the optimum, the objectives differ. Along an isotropic nuisance channel, pure SG lowers its loss by adding per-view nuisance whenever the shared code is non-uniform. An alignment weight above the channel's gain makes the nuisance-free solution a strict local minimizer; for LeJEPA, the same rule gives a critical SIGReg weight that decreases with the batch size. At finite batch size, an off-diagonal U-statistic removes a plug-in bias toward misalignment. In controlled latent-variable models, pure SG retains per-view style, an alignment weight above the measured gain removes it, and for LeJEPA at three batch sizes the measured gain separates the encoders that retain style from those that do not. InfoNCE training also reaches a lower SG0.2_{0.2} loss than SG0.2_{0.2} training from scratch, which points to an optimization gap.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Similarity search generalisation in contrastive learning with InfoNCE loss

    Jul 10, 2026Nick WhiteleyContrastive LearningText Embeddings

  2. The Geometric Mechanics of Contrastive Representation Learning: Alignment Potentials, Entropic Dispersion, and Cross-modal Divergence

    Jan 27, 2026Yichao Cai, Zhen Zhang, Yuhang Liu +1Multimodal Contrastive LearningContrastive Learning