Many machine learning systems try to explain complex data - like images or financial time series - in terms of hidden, independent factors that generated them. Recovering the true underlying factors, rather than some scrambled version of them, is the central challenge of nonlinear Independent Component Analysis (nICA). We prove identifiability (exact recovery) up to trivial ambiguities for real analytic generating functions when source probability density functions have a finite number of discontinuities in the first derivative. The Laplace distribution is the most prominent example satisfying this assumption. Our proof relies on the contrast between kinks in the source distribution and the smoothness of real analytic functions. Real analytic functions comprise a broad class of generating mechanisms, and can be approximated with Normalizing Flows or Variational Autoencoders with standard activation functions (e.g., tanh, softplus, GELU), so our result applies with minimal changes to existing training pipelines. We perform experiments on real and synthetic data with both Normalizing Flows and Variational Auto-Encoders demonstrating their identifiability properties. In experiments on CelebA data we recover several interpretable latent factors controlling unique attributes across the dataset.
Figures & tables
Figure 1: Common monotonic activation functions
k=5
k=10
Activation Function
Mixture A
Mixture B
Mixture A
Mixture B
Tanh
0.944 ± 0.017
0.945 ± 0.007
0.930 ± 0.023
0.920 ± 0.037
Sigmoid
0.895 ± 0.030
0.925 ± 0.031
0.842 ± 0.085
0.895 ± 0.042
Softplus
0.833 ± 0.075
0.916 ± 0.014
0.765 ± 0.110
0.859 ± 0.032
LeakyHardTanh
0.431 ± 0.198
0.577 ± 0.081
0.448 ± 0.057
0.502 ± 0.030
LeakyReLU
0.143 ± 0.159
0.067 ± 0.057
0.257 ± 0.107
0.192 ± 0.188
Table 1: Average MCC ± standard deviation for Laplace source model with different activation functions in the Normalizing Flow. Real analytic activation functions achieved substantially higher MCCs than non-real analytic activation functions.
k=15
k=25
Distribution
Mixture A
Mixture B
Mixture A
Mixture B
Laplace
0.965 ± 0.004
0.925 ± 0.015
0.953 ± 0.006
0.916 ± 0.037
LMM
0.937 ± 0.021
0.919 ± 0.027
0.939 ± 0.013
0.904 ± 0.025
Gaussian
0.514 ± 0.011
0.503 ± 0.034
0.417 ± 0.011
0.423 ± 0.018
Gamma
0.511 ± 0.029
0.442 ± 0.015
0.432 ± 0.021
0.399 ± 0.020
Triangular
0.453 ± 0.010
0.461 ± 0.013
0.385 ± 0.005
0.395 ± 0.007
Table 2: Average MCC ± standard deviation for different source distributions. The Laplace source distribution results in high MCCs even for higher dimensions.
Figure 2: Correlation plot between learned and ground truth sources for Laplace source (left) and Gamma source (right). The identifiable Laplace model is able to recover the true sources, while the Gamma model is not.
Source Distribution
Avg MCC ± SE
Laplace
0.871 ± 0.012
Gaussian
0.516 ± 0.003
Exponential
0.447 ± 0.003
Gamma
0.492 ± 0.003
GMM
0.480 ± 0.044
Table 3: Average ± standard error (SE) of MCC between model pairs trained on the Yahoo stock returns dataset for different source distributions. The Laplace distribution is the only one that satisfies Assumption 2 and thus has the highest average MCC.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Example Log Density
k=5
k=10
Activation Function
Mixture A
Mixture B
Mixture A
Mixture B
Tanh
0.931 ± 0.015
0.993 ± 0.003
0.928 ± 0.008
0.983 ± 0.006
Sigmoid
0.926 ± 0.016
0.987 ± 0.002
0.860 ± 0.037
0.987 ± 0.006
Softplus
0.877 ± 0.063
0.991 ± 0.002
0.829 ± 0.042
0.980 ± 0.007
LeakyHardTanh
0.656 ± 0.080
0.660 ± 0.154
0.561 ± 0.092
0.540 ± 0.039
LeakyReLU
0.136 ± 0.116
0.066 ± 0.041
0.095 ± 0.105
0.038 ± 0.011
Appendix
Table 4: Average MCC ± standard deviation for Laplace source model with different activation functions in the Normalizing Flow. Real analytic activation functions achieved substantially higher MCCs than non-real analytic activation functions.
AAPL
MSFT
NVDA
AMZN
Apple Inc.
Microsoft Corp.
NVIDIA Corp.
Amazon Inc.
ALL
JPM
PFE
JNJ
The Allstate Corp.
JPMorgan Chase & Co.
Pfizer Inc.
Johnson & Johnson
CVX
WMT
CAT
DIS
Chevron Corp.
Walmart Inc.
Caterpillar Inc.
The Walt Disney Co.
XOM
AMD
BRK-A
Appendix
Table 5: Company names of stock data in the Yahoo Stock experiment. For easier replication we provide the Yahoo Finance Symbol over the company names.
Figure 4: Auto-correlation and trace plots of stock returns data
We present the mathematical foundations of linear independent component analysis (ICA) models based on standard literature in a self-contained note. It is aimed at readers with a background in measure-theoretic probability theory. We first develop the theory of the characteristic functions of probability measures on Rd, including their analyticity and the way in which they determine and characterise the distributions. We then focus on several identifiability results of ICA models with successively strengthened assumptions on the sources: from merely non-constant, to non-Gaussian, to Gaussian-free independent sources. Under the strictest assumptions, we show that the independent sources are identifiable up to translation, permutation, scales and signs, and this even in the presence of additive Gaussian noise. Furthermore, we present the online equivariant gradient descent ICA algorithm for recovering the independent sources from data, in the standard complete noiseless non-Gaussian ICA setting.
Patrick Forré
AI4Science Lab Korteweg-de Vries Institute for Mathematics University of Amsterdam
In this work, we establish the sufficient conditions under which nonlinear Canonical Correlation Analysis (CCA) recovers ground-truth latent factors up to an affine transformation. By transporting the analysis from the observation space to the source space, we extend classical statistical results on orthogonal polynomial expansions of bivariate distributions to representation learning, proving affine identifiability under specific distributional priors. We formally demonstrate that whitening is strictly necessary to ensure the boundedness and well-conditioning of the learned mappings. Furthermore, we bridge the gap between theory and practice by proving that ridge-regularized empirical CCA converges to its population counterpart in the finite-sample regime. Finally, our findings provide a rigorous theoretical foundation explaining the empirical success of recent correlation-based non-contrastive learning methods. Experiments on synthetic and rendered image datasets, alongside systematic ablations, validate the predicted recovery behavior and illustrate the failure modes that arise when the assumptions are violated.
Zhiwei Han, Stefan Matthes, Hao Shen
fortiss GmbH, Munich, Germany · Technical university of Munich
This paper explores unsupervised disentangled representation learning from a functional perspective. We define latent concepts as factors that influence observations through locally orthogonal directions, formalized as an orthogonality constraint on the Jacobian of the generative mapping. We prove that this condition yields identifiability of general nonlinear generative models, without requiring statistical independence or causal assumptions, provided the latent domain admits all combinations of factor values. Experiments with orthogonality-regularized normalizing flows empirically confirm the theory, demonstrate reliable recovery of ground-truth factors, and shed light on the success of VAEs. These findings challenge the prevailing impossibility claims for unsupervised disentanglement and provide a principled alternative foundation.
Mathieu Cyrille Simon, Pascal Frossard, Christophe De Vleeschouwer