ARO: Aligned Representation learning for multi-Omics data
Authors: Amogh Singh, Yash Shah, Chiara D'Ercoli, Arash Mehrjou, Patrick Schwab, Timothy Jones, Pietro Liò
Organizations: Department of Computer Science, University of Cambridge, Cambridge, UK · DIAG, University of Rome “Sapienza”, Rome, Italy · Max Planck Institute for Intelligent Systems
The high cost of functional molecular assays, and prevalence of missing modalities and unmatched samples in computational biology, create significant barriers to comprehensive multi-omic profiling, essential for capturing and reasoning over molecules, cells, tissues, and organisms. This work proposes a model that learns meaningful representations from multi-omics cancer data supporting the reconstruction of missing and unpaired modalities. Contrary to increasingly complex, larger models, e.g. Foundation Models (FMs), ARO prioritizes practical applicability in limited or incomplete data settings. ARO optimally reconstructs missing modalities (MSE of 0.15 on the validation and test data in the Unmasked settings), with its learned latent embeddings enabling a downstream cancer classification task. Our findings indicate that analyzing diverse molecular layers as a single integrated system offers a reliable and cost-efficient approach, reducing dependence on large-scale experimental testing, while still supporting multi-omic exploration in limited data settings.
Figures & tables
Figure 1 : Overview of ARO. Multi-omics cancer datasets are first harmonized and aggregated. Next, pre-processing involving median imputation, Winsorization, normalization, and log transformation is performed. Features are then optionally masked (shown by dotted outlines) and passed to the model. The loss is computed, and learned latent representations can be used for downstream tasks.
Figure 2 : ARO performance in discriminating each omics-modality within the latent space, visualized using PCA (Figure 2(a) ) as well as across downstream tasks including binary cancer classification (Figure 2(b) ) and cancer laterality prediction (Figure 2(c) ).
Train
Validation
Test
ARO Masked
0.046
0.14
0.14
ARO Unmasked
0.064
0.15
0.15
PCA
3.1e−30
0.14
0.14
Table 1 : Reconstruction performance using MSE Loss averaged over 5 runs
Train
Validation
Test
ARO Masked
0.022
0.070
0.068
ARO Unmasked
0.032
0.071
0.07
PCA
1.5e−30
0.068
0.065
Table 2 : Reconstruction performance using Huber Loss averaged over 5 runs
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Layer
Shape
Parameters
inp_lin.weight
(2048, 71094)
145,600,512
inp_lin.bias
(2048, )
2048
encoder_layers.0.weight
(2048, 2048)
4,194,304
encoder_layers.0.bias
(2048, )
2048
encoder_layers.1.weight
(2048, 2048)
4,194,304
encoder_layers.1.bias
(2048, )
2048
Appendix
Table 3 : Parameter numbers of each block of the autoencoder.
Figure 3 : Iterations for hyperparameter tuning performed over validation loss, optimizing for dropout configuration, hidden dimensions, learning rate, modality-specific masking, number of hidden layers, batch normalization usage and weight decay. The configuration with the lowest validation loss is chosen from each setup.
Train
Validation
Test
ARO Masked
0.046±56.0e−04
0.14±31.0e−05
0.14±21.0e−05
ARO Unmasked
0.064±31.0e−05
0.15±21.0e−05
0.15±24.0e−05
PCA
3.1e−30
0.14
0.14
Appendix
Table 4 : Reconstruction performance using MSE Loss averaged over 5 runs
Train
Validation
Test
ARO Masked
0.022±28.0e−04
0.070±5.6e−05
0.068±20.0e−05
ARO Unmasked
0.032±16.0e−05
0.071±10.0e−05
0.07±6.1e−05
PCA
1.5e−30
0.068
0.065
Appendix
Table 5 : Reconstruction performance using Huber Loss averaged over 5 runs
Figure 4 : t-SNE cluster of the learned embeddings from ARO
Figure 5 : Confusion matrices of downstream tasks using PCA projections.
Data set
Samples
Feature Count
( Li et al., 2024 ) - Transcriptomics
108
45886
( Li et al., 2024 ) - Proteomics
109
10998
( Li et al., 2024 ) - Metabolomics
87
67
( Wang et al., 2021 ) - Transcriptomics
240
26941
( Wang et al., 2021 ) - Proteomics
333
12309
( Wang et al., 2021 ) - Metabolomics
36
216
Appendix
Table 6 : Details of the Multi-omics datasets considered and analysed in the study.
The Alan Turing Institute, London, United Kingdom · University of Manchester, Manchester, United Kingdom · The Institute of Cancer Research, London, United Kingdom +1
Institute of Medical Robotics, Shanghai Jiao Tong University · Institute of Medical Robotics, Shanghai Jiao Tong University, Shanghai 200240, China · Institute of Data Science, The University of Hong Kong +1