ReLaG: A Scalable Framework Generalizing Random Splits to Data with Latent Relations
Authors: Anthony Lavertu, Jacob Cote, Sophie Gobeil, Jacques Corbeil, Isabeau Premont-Schwarz, Pascal Germain
Organizations: Department of Computer Science Université Laval Québec, QC, Canada · Department of Biochemistry, Microbiology and Bioinformatics Université Laval Québec, QC, Canada · Department of Molecular Medicine Université Laval Québec, QC, Canada
Random splitting can yield non-independent train--test subsets when a dataset contains related samples, as is common in certain applications such as biochemical studies. This leads to overly optimistic generalization estimates. Here, we introduce ReLaG, a modality-agnostic framework that models sample relatedness through a hierarchical latent-variable process and infers groups of related samples using proximity graphs and community detection to produce independent train--test subsets. Across molecular and protein datasets, ReLaG matches existing relation-aware methods while scaling substantially better, enabling splits at previously impractical dataset sizes. We further introduce a label-free procedure that adapts the splitting resolution to production data, aligning evaluation with the intended deployment setting. ReLaG's inferred groups provide a cheap estimate of effective dataset size, enabling diversity-aware dataset scaling. ReLaG is open source and can be installed with pip install relag.
Figures & tables
Figure 1: Illustration of two generative processes that produce datasets containing related samples.
Figure 2: Two-level latent-variable model. Observed samples sij are generated conditionally on group-specific latent variables gi , which are drawn i.i.d. from a fixed population-level distribution G . The objective is to infer this group structure from the observed samples.
Figure 3: ReLaG algorithm overview. (A) The algorithm requires a distance function that correlates with the probability that two samples are unrelated. (B) A proximity graph is built using this distance function. (C) Communities are identified, then graph edges that are unlikely to represent relationships are removed. Each community is an estimate of a group, i.e., a set of samples conditioned on the same latent variable gi . (D) The dataset is partitioned along the community boundaries.
Figure 4: Threshold selection intuition. A wrong threshold makes train communities closer or further away from each other than they are from production communities. The right threshold makes production communities integrate nicely in the community distribution.
Figure 5: Relation-aware splits yield evaluations that match production performance. In each of 100 repeats, we sample independently two synthetic peptide datasets (15,000 sequences, 1,000 families each) under the latent-variable model: an annotated dataset to split and a production dataset.
Figure 6: Performances on real data. Downstream MLP test performance of dataset-splitting strategies across eight datasets (six molecular and two protein datasets). Results are averaged over 10 split seeds, with error bars denoting standard deviations. We report AUROC for classification tasks and Pearson correlation coefficient for regression tasks.
Figure 7: Threshold selection characterizes the relatedness regime. Top: single-linkage distances from pure-training and pure-production communities to training communities (15th–85th percentile shaded). Bottom: KS statistic.
Figure 8: Scaling on PeptideAtlas subsets. Runtime as a function of dataset size for ReLaG and Hestia, the two fastest compared methods. Runs were limited to one day; values above the dashed line are extrapolated. Log–log fits indicate approximately linear scaling for ReLaG and quadratic scaling for Hestia ( R2=0.998 for both).
Figure 9: Community structure captures an aspect of training-set diversity that is not reflected by the raw sample count and is associated with generalization performance.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Modality
N
Task
Metric
Encoder
cyp2c19_veith
molecule
12,665
classification
AUROC
ChemBERTa
ames
molecule
7,278
classification
AUROC
ChemBERTa
sr_are
molecule
5,832
classification
AUROC
ChemBERTa
lipophilicity
molecule
4,200
regression
PCC
ChemBERTa
pgp_broccatelli
molecule
1,218
classification
AUROC
ChemBERTa
caco2_wang
molecule
910
regression
PCC
ChemBERTa
Appendix
Table 1: Datasets used in the experiments. N is the number of samples after filtering.
Figure 10: Scaling of ReLaG on molecular modality. Both axes are logarithmic. Exponents in the legend are fitted on the log–log space.
Figure 11: Threshold selection on large-scale and synthetic data. Upper panels: single-linkage distances from pure-training communities (train ↔ train) and from pure-production communities (train ↔ production) to training-containing communities, with shading denoting the 15th–85th percentile range. Lower panels: the corresponding Kolmogorov–Smirnov statistic. Dashed lines mark the selected threshold τ⋆ .
Percentile
train ↔ train
train ↔ prod
prod − train
1
0.127
0.000
−0.127
5
0.500
0.516
+0.016
10
0.600
0.600
0.000
25
0.650
0.650
0.000
50
0.700
0.700
0.000
75
0.750
0.750
0.000
Appendix
Table 2: Nearest-train-neighbor distance on UniRef50: percentiles of the distance from a training sequence to its closest other training sequence (train ↔ train), and from a production benchmark sequence to its closest training sequence (train ↔ prod).
Figure 12: ProtSpaM distance against global alignment distance, on 39,380 pairs built from real UniRef50 seeds mutated at a sweep of rates (BLOSUM62 substitutions and indels) so the alignment axis is covered end to end. The two are monotonically related (Pearson r=0.954 , Spearman r=0.911 ).
Dataset
Metric
τ
Singleton
Community
Δ
p
caco2_wang
PCC
0.6
0.473±0.122
0.417±0.103
+0.056
1.1×10−2
lipophilicity
PCC
0.4
0.540±0.037
0.500±0.054
+0.040
2.6×10−5
sr_are
AUROC
0.4
0.710±0.030
0.682±0.030
+0.029
1.5×10−6
cyp2c19_veith
AUROC
0.5
0.779±0.045
0.796±0.020
−0.017
3.8×10−2
dbaasp
PCC
0.3
0.483±0.061
0.520±0.041
−0.037
1.3×10−5
ames
AUROC
0.7
0.530±0.075
0.640±0.064
−0.109
8.8×10−5
Appendix
Table 3: Singleton information content. Comparison of models trained on singleton samples and equally sized samples from non-singleton communities, evaluated on the same validation set. Δ is the performance difference (singleton − community), with Δ>0 indicating greater information content in singletons. Mean over 5 splits × 6 repeats. p -values are from two-sided t -tests. Rows are grouped by higher, lower, or indistinguishable singleton performance.
Figure 13: Family leakage on the synthetic dataset. Fraction of training samples whose family also appears in the test subset. Random: 93.8%, ReLaG: 0.62%, DataSAIL: 21.01%, Hestia: 0.00%, Oracle: 0.00%
Figure 14: Train–test distance across data-splitting methods. ECDFs of the minimum distance from each test sample to its nearest training sample, pooled within modality: (a) six small-molecule datasets and (b) two protein-sequence datasets. Distance is 1− similarity, using Tanimoto similarity for molecules and global sequence identity ( Needleman & Wunsch, 1970 ) for proteins. Dashed lines indicate the splitting thresholds ( 0.60 and 0.50 , respectively); the legend reports the fraction of test samples closer to any train samples than the corresponding threshold.
Figure 15: Effect on test performance of dropping random samples or full communities on different datasets. Neff corresponds to the number of communities remaining in the dataset.
This work investigates the ``small-vs-large gap'', where repeating on fewer samples can lead to compute saving during training compared to using a larger dataset. This is observed across algorithmic tasks, architectures and optimizers and cannot be explained using prior theory. We argue that the speedup comes from appropriate layer-wise growth enabled by sampling biases, which is more pronounced when the dataset size is smaller. We provide both theoretical analysis and empirical evidence from various interventions. Our results suggest that using a smaller dataset with more repetitions is not just a fallback strategy under data scarcity, but can be proactively leveraged as a favorable inductive biases for optimization, particularly in reasoning tasks.
Jingwen Liu, Ezra Edelman, Surbhi Goel +1
Columbia University · University of Pennsylvania · Kempner Institute, Harvard University
Accurate model evaluation in machine learning depends critically on how datasets are split into training and testing subsets. Standard random splitting assumes that both partitions share the same underlying distribution, an assumption often violated in datasets with class imbalance, natural clustering, or spatial autocorrelation. This paper investigates the role of statistical similarity in train-test splitting and its consequences for AutoML model evaluation. Five established strategies are compared across fifteen UCI benchmark datasets: random splitting, stratified sampling, Kennard-Stone, Duplex, and SPXY. Similarity is assessed using chi-square, Kolmogorov-Smirnov, and Maximum Mean Discrepancy (MMD) tests. Geometry-based methods consistently produce near-zero MMD scores, introducing instability into downstream performance estimates. The proposed Optimised-Distribution method treats similarity as an explicit optimisation objective and achieves the highest mean MMD similarity, 89.0%, across all strategies evaluated.
Yearn Tan Yin Tze, Charles Grellois
School of Computer Science, University of Sheffield, U.K.
Can a dataset be recognized from the spurious correlations it induces during training? We argue that datasets leave dataset-specific traces in a model's learned semantic correlation structure: incidental regularities that are predictive within a dataset, but not causal for the underlying task, can be internalized during training. We use this insight to study dataset-level membership inference, moving beyond existing methods that rely on behavioral or distributional evidence such as confidence scores, losses, margins, generated samples, or query responses. We introduce a white-box semantic fingerprinting approach based on semantic correlation descriptors (SCDs), which capture the semantic correlation structure learned by a model and make it comparable across dataset mixtures. In a controlled leave-one-dataset-out diagnostic, SCDs recover dataset-specific changes and perfectly separate matching from non-matching dataset pairs. We then propose a practical SCD-based membership score that tests whether a target dataset is part of a model's training mixture using only the model's SCD and the target dataset's standalone SCD, without requiring leave-one-dataset-out models. Across three diverse experimental settings, with dataset groups for natural language inference, emotion classification, and medical text classification, we test both the advantages and limitations of SCD-based membership inference with different degrees of semantic separation and keyword support between dataset splits. On average, the classifier based on this score achieves the highest performance and the lowest std, outperforming black-box baselines RMIA, Attack-P, and LiRA, as well as the white-box SIF baseline. These results show that dataset membership can be traced through internal semantic correlations, with the largest relative gain exceeding 60% in ROC-AUC when dataset groups expose distinct semantic particularities.
Andrada Gobeaja, Ionut Hodoroaga, Elena Burceanu +1
POLITEHNICA University of Bucharest · Bitdefender, Romania · Institute of Mathematics of the Romanian Academy