Causal representation learning (CRL) is the process of recovering causally-related latent variables from high-dimensional observations. As a label-free inference method, CRL is particularly attractive for applications where data labels are unavailable or impractical to obtain. While there has been significant progress in understanding the identifiability guarantees of CRL, such guarantees often hold under highly stylized assumptions, which temper the direct application to real-world problems. This paper has a two-fold objective for interventional CRL. First, it establishes identifiability guarantees for substantially weaker interventional assumptions, resulting in block disentanglement of the causal variables, where the block structure depends on the realistically available intervention mechanisms. Secondly, the block disentanglement framework is used for embodied visual state estimation, in which the objective is to recover the latent physical variables of a robotic system directly from visual data (images and videos) without labeled data. These two components are critically complementary. The block disentanglement theory delineates identifiability guarantees under weakened assumptions, and the application demonstrates that the resulting objective remains effective in a controlled embodied setting despite further assumption violations, providing a theory-to-practice bridge needed to translate the promise of label-free CRL into practical problems.
Figures & tables
Figure 1 : Training pipeline. A shared Autoencoder 1 (AE1) compresses each 128×128×1 image X into an 8×8×1 feature vector Y . Features from both cameras are concatenated along the channel axis, forming an 8×8×2 input to the LDR and AE2. AE2 then outputs the final db×1 latent estimate Z^ .
Model
Training Labels
Joint 1
Joint 2
Joint 3
Joint 4
Joint 5
Joint 6
Ours
none during training
0.083
0.015
0.217
0.035
0.198
0.080
RoboPEPP
1% of dataset
0.136
0.039
0.237
0.053
0.253
0.200
5% of dataset
0.075
0.017
0.134
0.024
0.153
0.081
10% of dataset
0.030
0.010
0.072
0.022
0.091
0.063
100% of dataset
0.003
0.001
0.007
0.003
0.010
0.011
Table 1 : Comparison of MSE ( rad2 ) for RoboPEPP and our method in assumption-matched regime.
λ
db
Active R2(ZI∣Z^I^)
All-latent R2(ZI∣Z^)
Feature R2(ZI∣Y)
0
16
0.33
0.72
0.77 (AE1 frozen)
50
16
0.58
0.63
100
16
0.63
0.66
0
32
0.23
0.77
50
32
0.61
0.67
100
32
0.55
0.66
Table 2 : Embodied setting results across score regularization weight λ and bottleneck dimension db .
Appendix figures & tables33 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 2 : Training dynamics across bottleneck dimensions db ( λ=50 ).
Figure 3 : Training dynamics across score-regularization weights λ ( db=8 ).
Figure 4 : Training dynamics across score-regularization weights λ ( db=16 ).
Figure 5 : Training dynamics across score-regularization weights λ ( db=32 ).
Figure 6 : R2 under multi-dominant-joint data ( db=16 )
Figure 7 : The six primary joint angles of the Franka Emika Panda arm targeted for intervention, together with their respective axes of rotation. Note that the final (7th) joint is kept fixed.
Joint
Scenario
Distribution
2
Observational
TN[−1.5,1.5](0,1)
Intervention 1
TN[−1.5,1.5](−0.75,0.5)
Intervention 2
TN[−1.5,1.5](0.75,0.5)
4
Observational
TN[−1.5,1.5](0,1)
Intervention 1
TN[−1.5,1.5](−0.75,0.5)
Intervention 2
TN[−1.5,1.5](0.75,0.5)
Appendix
Table 3: Sampling distributions for observational and interventional settings for single camera setup.
Joint
Scenario
Distribution
1
Observational
TN[0,3](1.2,0.4)
Intervention 1
TN[0,3](2.0,0.4)
Intervention 2
TN[0,3](0.6,0.4)
2
Observational
TN[−1.5,1.5](0,0.4)
Intervention 1
TN[−1.5,1.5](0.7,0.4)
Intervention 2
TN[−1.5,1.5](−0.7,0.4)
Appendix
Table 4: Sampling distributions for the in-distribution (ID) dataset. These truncated normal distributions define the observational and interventional data used for the two-camera Independent and Occlusion experiments.
Joint
OOD Observational Distribution
1
TN[0,3](1.2,0.4)
2
TN[−1.5,1.5](0.0,0.4)
3
TN[−1.5,1.5](0.0,0.4)
4
TN[−1.5,1.5](0.8,0.4)
5
TN[−1.5,1.5](0.0,0.4)
6
TN[0,3](0.5,0.4)
Appendix
Table 5 : Observational sampling distributions for the out-of-distribution (OOD) dataset. These parameters define the OOD test sets for the Two-Camera Independent, Causal, and Occlusion experiments. The interventional distributions for the OOD dataset remain identical to those of the in-distribution dataset, as defined in Table 4 .
Figure 8 : The causal model of the robot joints used to generate the dataset. In this graph, J i represents the angle of joint i .
Joint
Scenario
Distribution
1
Observational
TN[0,3](1.2,0.4)
Intervention 1
TN[0,3](2.0,0.4)
Intervention 2
TN[0,3](0.6,0.4)
2
Observational
TN[−1.5,1.5](0,0.4)
Intervention 1
TN[−1.5,1.5](0.7,0.4)
Intervention 2
TN[−1.5,1.5](−0.7,0.4)
Appendix
Table 6: Sampling distributions for observational and interventional settings for causal dataset corresponding to the causal graph Figure 8
Joint 1
Joint 2
Joint 3
Joint 4
Joint 5
Joint 6
Model Setup
MCC
MSE
MCC
MSE
MCC
MSE
MCC
MSE
MCC
MSE
MCC
MSE
1C, indep.
–
–
0.949
0.053
–
–
0.975
0.029
–
–
0.957
0.049
2C, indep.
0.874
0.083
0.979
0.015
0.634
0.217
0.950
0.035
0.679
0.198
0.884
0.080
2C, causal
0.921
0.058
0.966
0.020
0.788
0.106
0.976
0.019
0.742
0.051
0.756
0.070
2C, indep., occl.
0.844
0.101
0.964
0.025
0.568
0.245
0.884
0.082
0.617
0.225
0.768
0.145
Appendix
Table 7 : MCC and MSE in assumption-matched regime. MSE is reported in radians squared.
Figure 9 : Single Camera: Scatter-plots of ground-truth vs. estimated angles for joints 2, 4, and 6
Figure 10 : Scatter plots evaluating the two-camera Independent model on the in-distribution (ID) test set ( Table 4 ) . Each plot visualizes the relationship between a learned latent variable and its corresponding ground-truth joint angle. The displayed Correlation and MSE values correspond to the single best trial out of 15 runs, while the results presented in Tables 7 and 9 correspond to the mean statistics.
Figure 11 : Scatter plots evaluating the two-camera causal model on the causally-generated test set. Each plot visualizes the relationship between a learned latent variable and its corresponding ground-truth joint angle. The displayed Correlation and MSE values correspond to the single best trial out of 15 runs, while the results presented in Tables 7 and 9 correspond to the mean statistics.
Experiment
Distribution
Patch Size
Joint 1
Joint 2
Joint 3
Joint 4
Joint 5
Joint 6
Independent
ID
–
0.083
0.015
0.217
0.035
0.198
0.080
Causal
ID
–
0.058
0.020
0.106
0.019
0.051
0.070
Occlusion
ID
16
0.089
0.018
0.225
0.044
0.200
0.082
Occlusion
ID
32
0.101
0.025
0.245
0.082
0.225
0.145
Occlusion
ID
64
0.186
0.077
0.322
0.258
0.298
0.273
Independent
OOD
–
0.084
0.017
0.239
0.048
0.199
0.120
Appendix
Table 8 : Mean Squared Error (MSE) in radians squared ( rad2 ) for each joint under various experimental conditions for two camera angles. The table compares performance on in-distribution (ID) test set ( Table 4 ) and out-of-distribution (OOD) test set ( Table 5 ) for independent, causal, and occluded inference models.
2C, indep.
2C, causal
Joint Angle
MCC
MSE
MCC
MSE
Joint 1
0.874±0.004
0.083±0.006
0.921±0.003
0.058±0.003
Joint 2
0.979±0.001
0.015±0.001
0.966±0.002
0.020±0.001
Joint 3
0.634±0.010
0.217±0.011
0.788±0.003
0.106±0.005
Joint 4
0.950±0.002
0.035±0.002
0.976±0.001
0.019±0.001
Joint 5
0.679±0.010
0.198±0.011
0.742±0.006
0.051±0.003
Appendix
Table 9 : Comparison of MCC and MSE ( rad2 ) for each joint across two model settings with error bars. Mean and Std Dev are calculated across the 15 runs. Values are reported as Mean ± Std Dev.
Dataset
Distribution
Epochs
Patch Size
Joint 1
Joint 2
Joint 3
Joint 4
Joint 5
Joint 6
2.6k
ID
100
-
0.136
0.039
0.237
0.053
0.253
0.200
2.6k
ID
100
16
0.166
0.030
0.287
0.072
0.309
0.218
2.6k
ID
100
32
0.186
0.060
0.305
0.125
0.331
0.263
2.6k
ID
100
64
0.412
0.358
0.462
1.328
0.515
0.573
2.6k
OOD
100
-
0.202
0.054
0.374
0.101
0.378
0.454
2.6k
OOD
100
16
0.220
0.058
0.413
0.117
0.412
0.432
Appendix
Table 10 : Per-joint Mean Squared Error (MSE) for the RoboPEPP model in radians squared ( rad2 ). The table presents an ablation study on the effect of varying training data labels, evaluated on both in-distribution (ID) test set ( Table 4 ) and out-of-distribution (OOD) test set ( Table 5 ).
Model
Experiment
Patch Size
Joint 1
Joint 2
Joint 3
Joint 4
Joint 5
Joint 6
Our method
Independent
–
0.083
0.015
0.217
0.035
0.198
0.080
Causal
–
0.058
0.020
0.106
0.019
0.051
0.070
Occlusion
16
0.089
0.018
0.225
0.044
0.200
0.082
Occlusion
32
0.101
0.025
0.245
0.082
0.225
0.145
Occlusion
64
0.186
0.077
0.322
0.258
0.298
0.273
RoboPEPP
2.6k Dataset
–
0.136
0.039
0.237
0.053
0.253
0.200
Appendix
Table 11 : Comparison of In-Distribution (ID) Mean Squared Error (MSE) in radians squared ( rad2 ) for the our score-based method and RoboPEPP models. Results are shown per joint under various experimental conditions.
Model
Experiment
Patch Size
Joint 1
Joint 2
Joint 3
Joint 4
Joint 5
Joint 6
Our method
Independent
–
0.084
0.017
0.239
0.048
0.199
0.120
Causal
–
0.108
0.044
0.241
0.085
0.219
0.116
Occlusion
16
0.092
0.019
0.249
0.054
0.205
0.109
Occlusion
32
0.103
0.024
0.288
0.079
0.226
0.140
Occlusion
64
0.201
0.087
0.356
0.227
0.331
0.216
RoboPEPP
2.6k Dataset
–
0.202
0.054
0.374
0.101
0.378
0.454
Appendix
Table 12 : Comparison of Out-of-Distribution (OOD) Mean Squared Error (MSE) in radians squared ( rad2 ) for our method and RoboPEPP. Results are shown per joint under various experimental conditions.
Figure 12 : A step-by-step visualization of the reconstruction process for an occluded input. The final reconstruction from autoencoder-2, generated by passing its output through the decoder of autoencoder-1, successfully inpaints the occluded region.
Figure 13 : Evolution of MCC scores during training for AE2 using a 2-camera setup on the independent dataset. The plot compares performance across various learning rates. Training is performed with a batch size of 128, and sparsity loss weight of λ=3 .
Figure 14 : Evolution of MCC scores during training for AE2 using a 2-camera setup on the independent dataset. The plot compares performance across various batch sizes. Training is performed with a learning rate of 5e−5 and sparsity loss weight of λ=3 .
Figure 15 : Evolution of MCC scores during training for AE2 using a 2-camera setup on the independent dataset. The plot compares performance across various combinations of reconstruction loss weights and sparsity loss weights. Training is performed with a batch size of 128 and a fixed learning rate of 5e−5 .
Figure 16 : Ablation showing the advantage of having residual connections for better reconstruction loss and MCC score
Joint 1
Joint 2
Joint 3
Joint 4
Joint 5
Joint 6
Calibration Samples
MCC
MSE
MCC
MSE
MCC
MSE
MCC
MSE
MCC
MSE
MCC
MSE
1000 samples
0.874
0.083
0.979
0.015
0.634
0.217
0.950
0.035
0.679
0.198
0.884
0.080
500 samples
0.863
0.094
0.979
0.015
0.618
0.234
0.945
0.037
0.677
0.191
0.895
0.075
100 samples
0.881
0.079
0.975
0.019
0.612
0.227
0.942
0.039
0.661
0.199
0.879
0.084
Appendix
Table 13 : MCC and MSE of our method across different Calibration settings. MSE is reported in radians squared.
Joint 1
Joint 2
Joint 3
Joint 4
Joint 5
Joint 6
Lighting Condition
MCC
MSE
MCC
MSE
MCC
MSE
MCC
MSE
MCC
MSE
MCC
MSE
original
0.874
0.083
0.979
0.015
0.634
0.217
0.950
0.035
0.679
0.198
0.884
0.080
bright
0.874
0.072
0.978
0.015
0.640
0.207
0.950
0.033
0.677
0.180
0.874
0.094
dark
0.811
0.107
0.940
0.046
0.520
0.211
0.860
0.093
0.199
0.303
0.526
0.265
Appendix
Table 14 : MCC and MSE of our method across different lighting condition. MSE is reported in radians squared.
Figure 31
Figure 19 : Visual comparison of the reconstruction quality at each stage of our pipeline for the single camera setup. (a) The original input image. (b) The reconstruction from the first autoencoder (AE1). (c) The final reconstruction from the second autoencoder (AE2)
Figure 20 : Visual comparison of the reconstruction quality at each stage of our pipeline using the two camera independent model. (a) The original input image. (b) The reconstruction from the first autoencoder (AE1). (c) The final reconstruction from the second autoencoder (AE2)
Causal representation learning (CRL) seeks to uncover meaningful latent variables and their corresponding causal structure from high-dimensional observational data. Although its significance, CRL identifiability remains a crucial property, as it ensures the recovery of the mechanisms behind the data generation process, and hence the interpretability and robustness of the representation. Proving identifiability in CRL is intrinsically difficult, and we address in this work an even more challenging setting: multimodality. We consider multimodal observed data with a latent partially shared structure. Each modality is generated, through non linear mixing functions, from a specific subset of causal latent variables. Under flexible assumptions and without imposing any parametric distribution on the latent variables, we establish component-wise identifiability guarantees for the causal latent representation. Our identifiability results, furthermore, apply to the undercomplete scenario where we have, for each modality, more observed than latent variables. To instantiate our theoretical analysis, we introduce a Wasserstein-based module to recover the partially shared latent structure. Due to its differentiability, the latter can be easily integrated into all types of architecture, only requiring minimal changes. Extensive experiments on synthetic and realistic datasets validate the superiority of our approach over SOTA methods.
Manal Benhamza, Marianne Clausel, Myriam Tami
Paris-Saclay University, CentraleSupélec, MICS Lab, France · Lorraine University, CRAN, France
Causal representation learning (CRL) and traditional representation learning have largely developed along different trajectories. Traditional representation learning has been driven mainly by applications and empirical objectives, whereas CRL has focused more on theoretical questions, particularly identifiability. This difference in emphasis has created a gap between the two fields in terminology, problem formulation, and evaluation, limiting communication and sometimes leading to disconnected or redundant efforts. In this paper, we argue that these two fields should be brought into dialogue rather than treated as separate paradigms. To this end, we introduce a unified formulation in which the representation learning is characterized by two components: a task component, which specifies what information the learned representation is required to preserve, and a constraint component, which specifies what structure is imposed on the latent space. Under this formulation, the benefits run in both directions. CRL provides theoretical tools for understanding when structured latent constraints are useful or necessary, while traditional representation learning offers practical insights on task design and objective choice that can improve the development of CRL methods. To illustrate this interaction, we experimentally study how different task components affect the behavior of CRL methods under different structured constraints. Results on CausalVerse show that the effectiveness of causal constraints depends strongly on the tasks with which they are paired.
Yan Li, Yuewen Sun, Shaoan Xie +4
1Mohamed bin Zayed University of Artificial Intelligence · 2Carnegie Mellon University
Causal representation learning aims to discover robust features by exploiting the causal structure underlying data generation. Existing methods require specifying the causal structure a priori, yet different structures demand fundamentally incompatible invariance constraints, and misspecification leads to representations that discard predictive information. We introduce SaCRL, a framework that jointly identifies the causal structure and learns the corresponding invariant representation without prior structural knowledge. Our approach formulates structure selection as a soft optimization over candidate invariances using HSIC-based violation metrics, with adaptive weights that automatically concentrate on the achievable structure. We provide theoretical guarantees for structure identification, including under random-feature approximation, invariance satisfaction, and out-of-distribution generalization. Empirically, SaCRL recovers the true structure on synthetic and semi-synthetic Bayesian-network benchmarks, outperforms fixed-invariance baselines on Colored MNIST, achieves state-of-the-art accuracy on three DomainBed benchmarks (PACS, VLCS, OfficeHome), and degrades gracefully under structural misspecification and limited environment diversity. Code is available at: https://github.com/ArmanBehnam/sacrl.
Arman Behnam, Binghui Wang
Department of Computer Science Illinois Institute of Technology Chicago, IL, USA