Making Cross-Continental Federated Learning Repeatable with FLIP: a Multi-Application Study
Authors: Rafael Garcia-Dias, Alexandre Triay Bagur, Chayanin Tangwiriyasakul, Virginia Fernandez, Parhom Esmaeili, Piyalitt Ittichaiwong, Yang Li, Lawrence Adams, +15 more
Organizations: School of Biomedical Engineering & Imaging Sciences, King’s College London, London, UK · Bangkok Dusit Medical Services, Bangkok, Thailand
Federated learning (FL) in healthcare remains challenging, as the overhead of rebuilding governance guarantees for every collaboration stops most projects at the proof-of-concept stage. Here we present FLIP (Federated Learning Interoperability Platform), an open-source, multi-application platform that makes FL training and evaluation repeatable. FLIP implements common FL workflows as a set of composable services: cohort queries against per-site structured databases, on-demand DICOM retrieval from institutional PACS, per-site project approval, and reusable FL job types. To demonstrate FLIP, we ran two distinct use cases, federated fine-tuning and federated evaluation, on synthetic chest X-ray cohorts across two client nodes based in the United Kingdom (UK) and Thailand. In FLIP, each institution independently approves its participation in each project and operates its own node under local IT security processes. This study makes an operational rather than an algorithmic claim. It does not compare federated with centralised training; for that question, we refer the reader to existing systematic reviews and meta-analyses. The central result is evidence that such platforms enable international FL collaboration and improve repeatability, auditability, and site-specific governance. We also present a comprehensive comparison of existing platforms to help researchers and operators choose the right platform for their use case.
Figures & tables
Figure 1: FLIP architecture. The Central Hub (left, cloud-hosted) handles coordination and the FL server, while each Trust Node (middle) hosts an OMOP database, XNAT, and an FL client attached to GPU compute. Data leaving the TN are limited to model weight updates, aggregate cohort counts, per-round scalar metrics, and execution logs; no row-level records or image data are transferred.
Figure 2: Screenshot of the cohort query for the training set, with the aggregated per-site record counts on the right. A WHERE clause separates the training and holdout sets. The platform accepts SQL SELECT statements only; the full query is available in the FLIP repository.
Platform
Sites (countries)
Data
Workloads shown
FLIP (ours)
2 (2: UK, TH)
synthetic CXR
fine-tuning, evaluation
FeTS [ 17 ]
71 (6 continents)
real MRI, 6314 cases
segmentation
RACOON [ 3 ]
6 of 38 (1: DE)
real CT, 582 scans
segmentation
MedPerf [ 43 ]
9 (13)
real, multi-site
evaluation only
CODA [ 39 ]
8 of 9 (1: CA)
1.09M patients
analytics; FL simulated
JIP [ 41 ]
11 (1: DE)
real, multi-centre
imaging
Table 3: Deployment evidence reported by each platform. Scale and data realism are the weakest aspects of this study, and we report them explicitly rather than leave them to be inferred. Sites are those reported as deployed . Where a platform’s published federated training ran on a subset of them, or on a simulated partition of one dataset, this is stated, since the two are routinely conflated.
Class
UK Synthetic (n=478)
Thai Synthetic (n=468)
Pretr.
F-tuned
p
Pretr.
F-tuned
p
Effusion
0.998
1.000
0.12 n.s.
1.000
1.000
1.00 n.s.
Consolidation
0.989
0.997
0.027
0.998
1.000
0.31 n.s.
Infiltration
0.863
0.998
1.8×10−14
0.878
1.000
1.2×10−15
Nodule or Mass
0.970
1.000
5.9×10−6
0.955
1.000
7.2×10−7
Pneumothorax
0.971
1.000
0.0065
0.997
1.000
0.19 n.s.
Table 4: Per-class AUROC of pretrained and FedAvg fine-tuned Ark+ on UK and Thai holdout sets. p -values are from two-sided DeLong tests, Benjamini–Hochberg (BH) corrected within each site. The testing family is the 5 classes for the single model pair at that site ( m=5 ), so the BH threshold for the k -th smallest p -value is 0.05k/5 , ranging from 0.01 to 0.05 . Four comparisons survive correction at the UK site and two at the Thai site, whereas effusion (both sites) and Thai consolidation and pneumothorax do not (n.s.). Bold values mark the best model per class and site at Q=0.05 .
Figure 3: FLIP interface showing federated training results, UK–Thailand deployment.
Phase
Depl. UK–TH
Depl. same-net
Sim. UK
Sim. TH
Rounds 1–49 ( Tround )
745.8±0.7
181.7±0.8
275.4±0.5
1,605.2±2.3
Round 0 (init.)
3,830.0
334.6
294.7
1,696.2
Aggregation ( Tagg )
0.19±0.06
0.18±0.04
0.22±0.05
0.39±0.09
Checkpoint persist. ( Tpersist )
5.52±0.04
—
—
—
Inter-round gap ( Tgap )
0.19±0.02
0.19±0.02
0.21±0.02
0.29±0.03
Model transfer ( Tdown )
<0.2
—
—
—
Table 5: Round-loop timing over 50 rounds (mean ± std, seconds; gap n=49 ). Each simulator serialises both clients on one GPU, so its round time is approximately the sum of both sites’ Tcompute . Client turnaround runs from task dispatch to acceptance of the contribution, and therefore spans Tdown+Tcompute+Tup . Dashes mark phases not yet re-derived from the corresponding job export, as only the UK–TH deployment has been decomposed at this granularity.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Platform
Layer
PubMed
GitHub
arXiv
medRxiv
APPFL
deployed
yes
yes
no
no
CODA
deployed
yes
–
no
no
FLA 3
deployed
–
yes
yes
no
FeatureCloud
deployed
–
yes
yes
no
FeTS
deployed
yes
yes
yes
no
FL4Health
deployed
–
yes
no
no
Appendix
Table 6: Retrieval of each comparator by the four searches, measured against the saved candidate pools. yes the pool contained an artefact for that platform; no it did not; – the platform has no artefact of that kind to retrieve (no indexed paper, or no public repository). The PubMed and arXiv columns are assessed over the 1,485-record and 1,190-preprint pools respectively, the GitHub column over the 2,330-repository union, and the medRxiv column over the 167-preprint match pool. A platform paper here means a record describing the platform, not a study merely citing it.
1STADIUS, Department of Electrical Engineering (ESAT), KU Leuven, Leuven, Belgium · 2Biomedical Research Institute (BIOMED), Hasselt University, Hasselt, Belgium · 3Data Science Institute (DSI), Hasselt University, Hasselt, Belgium +1
Dept. of Applied Mathematics and Theoretical Physics University of Cambridge Cambridge, UK · Translational AI Laboratory, Dept. of Laboratory Medicine Amsterdam UMC Amsterdam, The Netherlands · Precision Health University Research Institute Queen Mary Univ. of London London, UK +3