cs.LGSep 29, 2026

ReLaG: A Scalable Framework Generalizing Random Splits to Data with Latent Relations

Authors: Anthony Lavertu, Jacob Cote, Sophie Gobeil, Jacques Corbeil, Isabeau Premont-Schwarz, Pascal Germain

Organizations: Department of Computer Science Université Laval Québec, QC, Canada · Department of Biochemistry, Microbiology and Bioinformatics Université Laval Québec, QC, Canada · Department of Molecular Medicine Université Laval Québec, QC, Canada

Abstract

Random splitting can yield non-independent train--test subsets when a dataset contains related samples, as is common in certain applications such as biochemical studies. This leads to overly optimistic generalization estimates. Here, we introduce ReLaG, a modality-agnostic framework that models sample relatedness through a hierarchical latent-variable process and infers groups of related samples using proximity graphs and community detection to produce independent train--test subsets. Across molecular and protein datasets, ReLaG matches existing relation-aware methods while scaling substantially better, enabling splits at previously impractical dataset sizes. We further introduce a label-free procedure that adapts the splitting resolution to production data, aligning evaluation with the intended deployment setting. ReLaG's inferred groups provide a cheap estimate of effective dataset size, enabling diversity-aware dataset scaling. ReLaG is open source and can be installed with pip install relag.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Less Data, Faster Training: repeating smaller datasets speeds up learning via sampling biases

    May 19, 2026Jingwen Liu, Ezra Edelman, Surbhi Goel +1Inductive BiasBiases

  2. Enhancing Automated Machine Learning via Homogeneous Train-Test Splitting Methods

    Jul 29, 2026Yearn Tan Yin Tze, Charles GrelloisModel EvaluationBenchmark Datasets

  3. idSCD: Identifying Training Datasets through Semantic Correlation Descriptors

    May 28, 2026Andrada Gobeaja, Ionut Hodoroaga, Elena Burceanu +1Membership InferenceCross-Dataset Benchmark