ARO: Aligned Representation learning for multi-Omics data
Authors: Amogh Singh, Yash Shah, Chiara D'Ercoli, Arash Mehrjou, Patrick Schwab, Timothy Jones, Pietro Liò
Organizations: Department of Computer Science, University of Cambridge, Cambridge, UK · DIAG, University of Rome “Sapienza”, Rome, Italy · Max Planck Institute for Intelligent Systems
The high cost of functional molecular assays, and prevalence of missing modalities and unmatched samples in computational biology, create significant barriers to comprehensive multi-omic profiling, essential for capturing and reasoning over molecules, cells, tissues, and organisms. This work proposes a model that learns meaningful representations from multi-omics cancer data supporting the reconstruction of missing and unpaired modalities. Contrary to increasingly complex, larger models, e.g. Foundation Models (FMs), ARO prioritizes practical applicability in limited or incomplete data settings. ARO optimally reconstructs missing modalities (MSE of 0.15 on the validation and test data in the Unmasked settings), with its learned latent embeddings enabling a downstream cancer classification task. Our findings indicate that analyzing diverse molecular layers as a single integrated system offers a reliable and cost-efficient approach, reducing dependence on large-scale experimental testing, while still supporting multi-omic exploration in limited data settings.
Figures & tables
Figure 1 : Overview of ARO. Multi-omics cancer datasets are first harmonized and aggregated. Next, pre-processing involving median imputation, Winsorization, normalization, and log transformation is performed. Features are then optionally masked (shown by dotted outlines) and passed to the model. The loss is computed, and learned latent representations can be used for downstream tasks.
Figure 2 : ARO performance in discriminating each omics-modality within the latent space, visualized using PCA (Figure 2(a) ) as well as across downstream tasks including binary cancer classification (Figure 2(b) ) and cancer laterality prediction (Figure 2(c) ).
Train
Validation
Test
ARO Masked
0.046
0.14
0.14
ARO Unmasked
0.064
0.15
0.15
PCA
3.1e−30
0.14
0.14
Table 1 : Reconstruction performance using MSE Loss averaged over 5 runs
Train
Validation
Test
ARO Masked
0.022
0.070
0.068
ARO Unmasked
0.032
0.071
0.07
PCA
1.5e−30
0.068
0.065
Table 2 : Reconstruction performance using Huber Loss averaged over 5 runs
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Layer
Shape
Parameters
inp_lin.weight
(2048, 71094)
145,600,512
inp_lin.bias
(2048, )
2048
encoder_layers.0.weight
(2048, 2048)
4,194,304
encoder_layers.0.bias
(2048, )
2048
encoder_layers.1.weight
(2048, 2048)
4,194,304
encoder_layers.1.bias
(2048, )
2048
Appendix
Table 3 : Parameter numbers of each block of the autoencoder.
Figure 3 : Iterations for hyperparameter tuning performed over validation loss, optimizing for dropout configuration, hidden dimensions, learning rate, modality-specific masking, number of hidden layers, batch normalization usage and weight decay. The configuration with the lowest validation loss is chosen from each setup.
Train
Validation
Test
ARO Masked
0.046±56.0e−04
0.14±31.0e−05
0.14±21.0e−05
ARO Unmasked
0.064±31.0e−05
0.15±21.0e−05
0.15±24.0e−05
PCA
3.1e−30
0.14
0.14
Appendix
Table 4 : Reconstruction performance using MSE Loss averaged over 5 runs
Train
Validation
Test
ARO Masked
0.022±28.0e−04
0.070±5.6e−05
0.068±20.0e−05
ARO Unmasked
0.032±16.0e−05
0.071±10.0e−05
0.07±6.1e−05
PCA
1.5e−30
0.068
0.065
Appendix
Table 5 : Reconstruction performance using Huber Loss averaged over 5 runs
Figure 4 : t-SNE cluster of the learned embeddings from ARO
Figure 5 : Confusion matrices of downstream tasks using PCA projections.
Data set
Samples
Feature Count
( Li et al., 2024 ) - Transcriptomics
108
45886
( Li et al., 2024 ) - Proteomics
109
10998
( Li et al., 2024 ) - Metabolomics
87
67
( Wang et al., 2021 ) - Transcriptomics
240
26941
( Wang et al., 2021 ) - Proteomics
333
12309
( Wang et al., 2021 ) - Metabolomics
36
216
Appendix
Table 6 : Details of the Multi-omics datasets considered and analysed in the study.
Foundation models (FMs) have emerged as powerful representation extractors for medical data, yet their generalizability to datasets under distribution shift remains underexplored. This work systematically evaluates FM-based representations on a suite of computational pathology tasks across two real-world commercial cohorts, IH-BC and IH-NSCLC, drawn from the licensed in-house (IH) oncology dataset. The analysis focuses on two modalities, whole-slide images and transcriptomic profiles, drawn from the IH multimodal data. We first benchmark unimodal probing performance across five FMs on eight downstream classification tasks, and find that image and omics representations carry complementary predictive signals. Then we investigate whether multimodal fusion can yield additional gains over unimodal baselines by comparing three image-omics fusion strategies built on paired representations. The trustworthiness of selected unimodal and multimodal pipelines is further assessed through conformal prediction. Our results show that FM representations achieve competitive performance on out-of-distribution data and that multimodal fusion helps mainly when no single modality dominates the signal. Conformal prediction reveals that in the majority of cases where a point prediction fails, the true diagnosis remains recoverable within the prediction set, reinforcing the value of uncertainty-aware inference for clinical support.
Jingyu Hu, Giuseppe Tripodi, Reed Naidoo +2
The Alan Turing Institute, London, United Kingdom · University of Manchester, Manchester, United Kingdom · The Institute of Cancer Research, London, United Kingdom +1
We study multimodal learning under missing modalities, with particular motivation from bioscience applications in which heterogeneous modalities are often only partially available when decisions need to be made. We propose Latent World Recovery (LWR), a framework built on two key ideas: (i) modality-specific embeddings from different modalities are aligned in a shared latent space, and (ii) a unified representation is constructed by fusing only the embeddings of the modalities that are actually available at both training and inference time. Rather than imputing missing modalities or requiring a fixed modality set, LWR treats each modality as a partial perception of an underlying latent state and performs availability-aware representation learning directly from the observed modalities. This combination of neighbor-based latent alignment and availability-aware modality fusion enables robust multimodal prediction under partial observation, while avoiding error propagation from explicit reconstruction of missing modalities. We evaluate the proposed framework on real-world incomplete multi-omics benchmarks and demonstrate that it provides an effective approach to downstream tasks such as cancer phenotype classification and survival prediction.
Hui Wang, Tianyu Ren, Joseph Butler +3
Queen’s University Belfast, Belfast, United Kingdom
Attention mechanisms have been widely utilized in modern deep learning, and many existing multi-omics models inherit their conventional use to allow unrestricted bidirectional interactions. However, the fundamental logic of life is directional. Existing designs often overlook the directionality suggested by the central dogma, potentially limiting transfer across heterogeneous cancers, downstream tasks, and incomplete modality settings. In this work, we present DoGMA, a central-dogma-guided foundation model for pan-cancer multi-omics analysis, arguing that robust transfer requires representations with domain-specific inductive bias. Concretely, we build it on a Transformer-MoE architecture where directed attention biases inter-omics communication toward central-dogma information flow. We further pretrain our model with masked hierarchical omics reconstruction to guide it toward learning central-dogma-consistent interactions. Across diverse downstream tasks, including cancer representation learning, survival prediction, and metastasis prediction, DoGMA consistently demonstrates strong predictive performance. Ablations and analyses further suggest that the performance gains arise from the synergy between central-dogma-guided directed attention and reconstruction-based pretraining, which together promote more biologically consistent cross-omics information exchange. Overall, DoGMA demonstrates that domain-specific inductive biases can improve the robustness and transferability of multi-omics foundation models, offering new insights into the design of attention mechanisms for multi-omics representation learning.
Junfei Ling, Bangzheng Pu, Bingsen Xue +3
Institute of Medical Robotics, Shanghai Jiao Tong University · Institute of Medical Robotics, Shanghai Jiao Tong University, Shanghai 200240, China · Institute of Data Science, The University of Hong Kong +1