ReLaG: A Scalable Framework Generalizing Random Splits to Data with Latent Relations
Organizations: Department of Computer Science Université Laval Québec, QC, Canada · Department of Biochemistry, Microbiology and Bioinformatics Université Laval Québec, QC, Canada · Department of Molecular Medicine Université Laval Québec, QC, Canada
Abstract
Random splitting can yield non-independent train--test subsets when a dataset contains related samples, as is common in certain applications such as biochemical studies. This leads to overly optimistic generalization estimates. Here, we introduce ReLaG, a modality-agnostic framework that models sample relatedness through a hierarchical latent-variable process and infers groups of related samples using proximity graphs and community detection to produce independent train--test subsets. Across molecular and protein datasets, ReLaG matches existing relation-aware methods while scaling substantially better, enabling splits at previously impractical dataset sizes. We further introduce a label-free procedure that adapts the splitting resolution to production data, aligning evaluation with the intended deployment setting. ReLaG's inferred groups provide a cheap estimate of effective dataset size, enabling diversity-aware dataset scaling. ReLaG is open source and can be installed with pip install relag.
Figures & tables
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Modality | Task | Metric | Encoder | |
|---|---|---|---|---|---|
| cyp2c19_veith | molecule | 12,665 | classification | AUROC | ChemBERTa |
| ames | molecule | 7,278 | classification | AUROC | ChemBERTa |
| sr_are | molecule | 5,832 | classification | AUROC | ChemBERTa |
| lipophilicity | molecule | 4,200 | regression | PCC | ChemBERTa |
| pgp_broccatelli | molecule | 1,218 | classification | AUROC | ChemBERTa |
| caco2_wang | molecule | 910 | regression | PCC | ChemBERTa |
| Percentile | train train | train prod | prod train |
|---|---|---|---|
| 1 | 0.127 | 0.000 | |
| 5 | 0.500 | 0.516 | |
| 10 | 0.600 | 0.600 | |
| 25 | 0.650 | 0.650 | |
| 50 | 0.700 | 0.700 | |
| 75 | 0.750 | 0.750 |
| Dataset | Metric | Singleton | Community | |||
|---|---|---|---|---|---|---|
| caco2_wang | PCC | 0.6 | ||||
| lipophilicity | PCC | 0.4 | ||||
| sr_are | AUROC | 0.4 | ||||
| cyp2c19_veith | AUROC | 0.5 | ||||
| dbaasp | PCC | 0.3 | ||||
| ames | AUROC | 0.7 |