The presence of noise that depends on the latent variables poses a fundamental challenge to identifiability. Existing results rely on conditional independence among the observations given the latent variables. We study a more general \emph{misspecified structure}, where this conditional factorization does not hold, and establish both precise and approximate identifiability guarantees. We characterize structural misspecification as a perturbed factor analysis problem. For precise identifiability, we establish subspace identifiability under spectral separation and controlled perturbation, followed by component-wise identifiability under structural sparsity. When the precise condition is not guaranteed, we derive an approximate subspace-identifiability theorem. Based on these results, we develop an unsupervised variational estimator for recovering latent variables. Experiments demonstrate the effectiveness of the proposed framework.
Figures & tables
Figure 1: Visualization of the data generations of Eq. 1 under misspecified structure.
Figure 2: The overall framework of our proposed approach consists of: (1) two encoders qψ and qϕ that map observations xt to z^ and ϵ^ , respectively; (2) a decoder that reconstructs observations x^ from z^ and ϵ^ ; and (3) two prior estimation modules pδ and pγ that models the prior of z and ϵ , respectively. We train the framework by LRecon along with LKLD .
Figure 3: (a) Visualization of the correlations between each component of true latent variables ( zi ) and their corresponding component of estimated latent variables ( z^i ) using our approach. The green bounding boxes highlight the components that are identified. (b) Mean Correlation Coefficient (MCC) scores comparing our framework with state-of-the-art approaches, including IndVAE, MCRL, BetaVAE, iVAE, and SlowVAE, as well as the ablation baselines W/O e and W/O s .
Figure 4: R2 scores for comparison methods.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Table of notions
Variables
x∈RK
Observations
x^∈RK
Reconstructions
κ
∣∣Lxa∣z∣∣op∣∣Lxa∣z−1∣∣op
ε
Identifiability error
z∈RN
Latent variables
z^∈RN
Latent variable estimations
ϵ
True dependent noise term
ϵ^
Estimation of ϵ
η
Auxiliary variable
η^
Estimation of η
Appendix
Table 5
Configuration
Description
Output dimensions
Encoder qψ
Input: x
BS×K
Dense
128 neurons, LeakyReLU
BS×128
Dense
128 neurons, LeakyReLU
BS×128
Dense
Output embeddings
BS×2N
Bottleneck
Compute mean and variance of posterior
μz , σz
Appendix
Table 1: The details of our network architectures for the experiment on MSTM-17 dataset, where BS means batch size, N=32 , M=32 and K=1280 .
Figure 5: Visualization of the data generating process for Person Classification task
Methods
Acc
AGW ( Ye et al., 2021 )
85.5 ± 1.2
TransReID ( He et al., 2021 )
87.8 ± 0.5
CLIPReID ( Li et al., 2023 )
90.1 ± 0.3
GTL Yang et al. (2025)
91.5 ± 1.2
MCRL ( Sun et al., 2025 )
92.6 ± 0.9
IndVAE ( Hu and Schennach, 2008 )
93.1 ± 0.5
Appendix
Table 2: Comparison of Top-1 Accuracy on MSMT17 dataset
N
IndVAE
Ours
8
0.64±0.06
0.80±0.02
12
0.51±0.03
0.68±0.05
18
0.47±0.04
0.61±0.05
Appendix
Table 3: MCC on the synthetic dataset for increasing latent dimensionality N .
Hyperparameter
Value
MCC
λ
1
0.87±0.04
0.1
0.82±0.03
0.01
0.70±0.07
10
0.68±0.05
β1
0.02
0.87±0.04
1
0.73±0.02
Appendix
Table 4: Sensitivity of MCC to regularization hyperparameters.
Methods
Acc
AGW ( Ye et al., 2021 )
87.6±0.8
TransReID ( He et al., 2021 )
90.2±0.9
CLIPReID ( Li et al., 2023 )
91.7±1.1
GTL Yang et al. (2025)
93.9±0.5
MCRL ( Sun et al., 2025 )
94.5±1.0
IndVAE ( Hu and Schennach, 2008 )
95.8±0.8
Appendix
Table 5: Comparison of Top-1 Accuracy on the RobotPKU dataset.
Methods
Acc
LDP-net Zhou et al. (2023)
91.7±1.1
Style Fu et al. (2023)
92.8±0.8
CLIPReID ( Li et al., 2023 )
94.1±1.0
GTL Yang et al. (2025)
95.7±0.4
MCRL ( Sun et al., 2025 )
96.4±0.8
IndVAE ( Hu and Schennach, 2008 )
96.8±0.5
Appendix
Table 6: Comparison of Top-1 Accuracy on the SYSU-MM01 dataset.
Figure 6: Visualization of the data generations of Eq. 52 .
Given only observational data X=g(Z), where both the latent variables Z and the generating process g are unknown, recovering Z is ill-posed without additional assumptions. Existing methods often assume linearity or rely on auxiliary supervision and functional constraints. However, such assumptions are rarely verifiable in practice, and most theoretical guarantees break down under even mild violations, leaving uncertainty about how to reliably understand the hidden world. To make identifiability actionable in the real-world scenarios, we take a complementary view: in the general settings where full identifiability is unattainable, what can still be recovered with guarantees, and what biases could be universally adopted? We introduce the problem of diverse dictionary learning to formalize this view. Specifically, we show that intersections, complements, and symmetric differences of latent variables linked to arbitrary observations, along with the latent-to-observed dependency structure, are still identifiable up to appropriate indeterminacies even without strong assumptions. These set-theoretic results can be composed using set algebra to construct structured and essential views of the hidden world, such as genus-differentia definitions. When sufficient structural diversity is present, they further imply full identifiability of all latent variables. Notably, all identifiability benefits follow from a simple inductive bias during estimation that can be readily integrated into most models. We validate the theory and demonstrate the benefits of the bias on both synthetic and real-world data.
Given a generalist model, learning a task-relevant specialist representation is fundamental for downstream applications. Identifiability, the asymptotic guarantee of recovering the ground-truth representation, is critical because it sets the ultimate limit of any model, even with infinite data and computation. We study this problem in a completely nonparametric setting, without relying on interventions, parametric forms, or structural constraints. We first prove that the structure between time steps and tasks is identifiable in a fully unsupervised manner, even when sequences lack strict temporal dependence and may exhibit disconnections, and task assignments can follow arbitrarily complex and interleaving structures. We then prove that, within each time step, the task-relevant latent representation can be disentangled from the irrelevant part under a simple sparsity regularization, without any additional information or parametric constraints. Together, these results establish a hierarchical foundation: task structure is identifiable across time steps, and task-relevant latent representations are identifiable within each step. To our knowledge, each result provides a first general nonparametric identifiability guarantee, and together they mark a step toward provably moving from generalist to specialist models.
This paper explores unsupervised disentangled representation learning from a functional perspective. We define latent concepts as factors that influence observations through locally orthogonal directions, formalized as an orthogonality constraint on the Jacobian of the generative mapping. We prove that this condition yields identifiability of general nonlinear generative models, without requiring statistical independence or causal assumptions, provided the latent domain admits all combinations of factor values. Experiments with orthogonality-regularized normalizing flows empirically confirm the theory, demonstrate reliable recovery of ground-truth factors, and shed light on the success of VAEs. These findings challenge the prevailing impossibility claims for unsupervised disentanglement and provide a principled alternative foundation.
Mathieu Cyrille Simon, Pascal Frossard, Christophe De Vleeschouwer