q-bio.QMDec 8, 2024

Batch effects can impair federated learning in multi-center omics studies

Authors: Yuliya Burankova, Julian Klemm, Jens J. G. Lohmann, Anne Hartebrodt, Ahmad Taheri, Niklas Probul, Jan Baumbach, Olga Zolotareva

Organizations: Institute for Computational Systems Biomedicine, University of Hamburg, 22761 Hamburg, Germany · Chair of Proteomics and Bioanalytics, TUM School of Life Sciences, Technical University of Munich, 85354 Freising, Germany · Department of Mathematics and Computer Science, University of Southern Denmark, 5230 Odense, Denmark · Biomedical Network Science Lab, Department Artificial Intelligence in Biomedical Engineering, Friedrich-Alexander-Universität Erlangen-Nürnberg, 91052 Erlangen, Germany · Data Science in Systems Biology, TUM School of Life Sciences, Technical University of Munich, 85354 Freising, Germany · Institute of Clinical Molecular Biology (IKMB), Kiel University and University Medical Center Schleswig-Holstein, Rosalind‑Franklin‑Str. 12, 24105 Kiel, Germany

Abstract

Federated learning (FL) enables collaborative analysis of biomedical data without exchanging sensitive patient-level information, but its performance in multi-center studies may be compromised by batch effects which can obscure biological signals. Here, we systematically assess the impact of uncorrected batch effects on FL outcomes using four multi-center omics datasets, including transcriptomic, proteomic, and metabolomic data, and two representative algorithms: federated k-means clustering and federated random forest classification. Our results demonstrate that uncorrected batch effects undermine unsupervised FL and can substantially degrade supervised FL performance, indicating that privacy-aware batch-effect correction is essential for reliable FL. To enable privacy-preserving BEC in distributed bulk omics data, we introduce fedRBE ( https://featurecloud.ai/app/fedrbe ), a federated implementation of limma's removeBatchEffect() method enhanced by secure multi-party computation, suitable for datasets with missing values and non-identical feature sets across clients, including proteomics and metabolomics data.

Explore similar work

CardsList
  1. Fed-ReMasker: Federated Tabular Imputation under Feature-Level Missingness

    Sep 23, 2026Ioannis Papathanail, Rooholla Poursoleymani, Lubnaa Abdur Rahman +1Missing ModalitiesFeature Engineering