Beyond Marginal Monitoring: Distributed Joint-Distribution Testing for Data Concept Drift in Large Scale E-Commerce Operations
Authors: Cagdas Pullu, Mahmut Emir Arslan, Bugra Balkac, Aylin Ondersev Balta, Cihangir Celal Palaci, Fikri Cem Yilmaz, Altan Cakir
Organizations: Data & Analytics, Trendyol Group, Istanbul, 34485, Istanbul, Türkiye. · Department of Data Science and Analytics, Istanbul Technical University, Maslak, Istanbul, 34469, Istanbul, Türkiye.
Concept drift threatens production machine learning, yet the empirical behavior of multivariate two-sample drift detectors at scale remains under-characterized. Existing benchmarks rarely address the hundreds of millions of rows and high-cardinality features typical of industrial-operational datasets. We evaluate five multi-column two-sample tests (marginal, projection-based, and kernel embedding methods) across three complementary environments: the Harvard Dataverse, a validated Failing Loudly reproduction (mean absolute error between 0.030 and 0.053), and a novel synthetic-injection benchmark on the 137.5-million-row Trendyol collection-ranking feature table. Testing four drift types across two severity-scope regimes, we demonstrate that distributed Maximum Mean Discrepancy with Random Fourier Features on Apache Spark scales robustly. Averaged over the four drift types in the strong regime and under a calibrated threshold, it achieves a Pearson correlation of r = 0.940 with expected drift magnitude, an 80.4% true positive rate, and a 3.2% false positive rate. Conversely, the per-dimension Kolmogorov-Smirnov test failed due to statistic saturation from ID-like columns under asymmetric sampling, establishing a critical constraint for large-scale sampling design. At weak configurations (realized-flip fractions of at most 0.57%), detectors struggled to reliably discriminate, highlighting the need for future intensity-grid power analyses to distinguish fundamental sensitivity bounds from scalable threshold shifts.
Figures & tables
Figure 1: Temporal concept-drift profiles, following the taxonomy of Gama et al. (2014) : (a) sudden drift replaces the old concept at once; (b) gradual drift alternates between the old and new concepts until the new one prevails; (c) incremental drift moves through intermediate states; (d) recurring drift makes previously seen concepts reappear. The dashed line marks the drift onset. Terminology varies across surveys; for example, Hinder et al. (2024) describe incremental drift as sampling from both concepts with varying probabilities.
Test
Family
Vector
Decision Rule / Calibration
Complexity
Univariate marginal ×6 1 1 1 Chi-square, Jensen–Shannon, Kullback–Leibler, PSI, KS, and univariate Wasserstein-1, applied per feature.
Marginal per-feature
Xk
Fixed threshold or asymptotic p
O(n) per column
Per-dim KS
Marginal aggregation
(X,Y) stack
Bonferroni: minkpk<α0/d
O(ndlogn)
Projected KS (SRP+KS)
Projection-based
(X,Y)S , S∈Rd×32
Bonferroni on projected dims
O(nd+nlogn)
MMD-RBF
Multivariate kernel
(X,Y) stack
Permutation ( B=1,000 )
O(n2d) ; kernel matrix
RFF-MMD
Multivariate kernel
(X,Y) stack
Permutation ( B=200 )
O(nDd) ; D=500
Sliced Wasserstein Distance
Projection-based
(X,Y) projections
Fixed τSW=0.05
O(nPlogn) ; P=100
Table 1: Overview of two-sample tests evaluated in this study. Family classification follows Section 3.1 (marginal per-feature; marginal per-feature aggregation on the stacked vector; projection-based; multivariate embedding/kernel). Vector column indicates the input to which the test is applied. Complexity is per-position at reference size nR and current-window size nC . Parameter values are those of the Failing Loudly reproduction; the Trendyol scan settings are given in Section 6.3.1 .
Test Name
Sub-Family
Core Statistic / Mechanism
Threshold / Calibration
Execution Paradigm
Univariate (6 variants)
Marginal
Marginal PR(Xk) vs PC(Xk) per feature
Fixed α0=0.05 or lit. τ
Hybrid (Driver aggregated)
Per-dim KS
Marginal Agg.
maxkDk across d raw dimensions
Bonferroni ( α0/d )
Driver-side (Subsampled)
Projected KS (SRP)
Projection
max KS across K=32 sparse random projections
Bonferroni ( α0/K )
Driver-side (Subsampled)
Sliced Wasserstein Distance
Projection
Avg. Wasserstein-1 over 100 random unit directions
Fixed distance ( τ=0.05 )
Spark Distributed
MMD-RBF
Kernel Embed.
MMD2 via Gaussian RBF ( O(n2d) )
Permutation ( B=1000 )
Driver-side (Subsampled)
RFF-MMD
Kernel Embed.
∥hˉ(XR)−hˉ(XC)∥22 mapping ( D=500 )
Permutation ( B=200 )
Spark Distributed
Table 2: Summary of Evaluated Two-Sample Tests and Execution Architecture
Layer
Source / Family
Configs
Dim
Ground Truth
Harvard Dataverse
SEA, SINE, STAGGER, MIXED, Random Tree
20
2–9
3 temporal changepoints
Failing Loudly
MNIST, CIFAR-10 (10 shift types)
60
784 / 3072
Image perturbation labels
Trendyol benchmark
Collection-ranking feature table
8
11
Injected row-level drift (Eqs. 6 – 8 )
Table 3: Dataset inventory and feature dimensionality. Feature dimensionality is reported after one-hot encoding of categorical covariates.
Figure 2: Evaluation protocols of this study within the taxonomy of unsupervised drift detectors of Gemaque et al. (2020) (adapted from their Fig. 2, CC BY 4.0). Batch detectors test an accumulated window as a whole; sliding-window scans advance the detection window along the stream and differ in whether the reference window stays fixed or slides with it. The survey’s online detectors check every arriving instance, whereas our scans evaluate at a stride Δ (Harvard Dataverse) or per position bucket (Trendyol benchmark). Partial-batch detection, which tests only a selected subset of each batch, is not evaluated here.
N=50 position buckets. Fixed reference table (first 100,000 rows; 10,000 for driver-side tests). 8 scenarios.
MMD-RBF, Per-dim KS, SRP+KS, RFF-MMD, SWD
Pearson r vs. mˉb , TPR, FPR, Runtime
Failing Loudly
Static sizes N∈{10…10,000} . 5 random splits, 3 fractions.
MMD-RBF, Per-dim KS, SRP+KS, RFF-MMD, SWD
Acc(N) , MAE vs literature
Table 4: Summary of experimental protocols and evaluation metrics.
Drift Type (Regime)
Severity ( σ )
Scope ( ϕ )
Mean αˉs
FlipRatescenarioexpected
sudden_strong
0.75
0.60
0.400
10.3%
sudden_weak
0.25
0.10
0.400
0.57%
gradual_strong
0.75
0.60
0.325
8.4%
gradual_weak
0.25
0.10
0.325
0.47%
incremental_strong
0.75
0.60
0.325
8.4%
incremental_weak
0.25
0.10
0.325
0.47%
Table 5: Trendyol Benchmark Scenario Configurations. The strong regime measures clean detection capabilities, while the weak regime establishes an empirical low-signal failure point.
Test Category
SEA
SINE
STAGGER
MIXED
Random Tree (RT)
Multivariate (Kernel & Proj.)
100%
100%
100%
100%
100%
Per-dim KS (Bonferroni)
100%
0%
100%
0%
100%
Univariate Marginal
0%
0%
0%
0%
100% 1 1 1 Recall artifact resulting from a generator-induced class-marginal rescaling at changepoints.
Table 6: Aggregated Batch Recall on Harvard Dataverse Streams. The bifurcation in recall demonstrates the structural failure of marginal tests under stationary P(X) environments, contrasted against the robustness of joint multivariate embeddings.
Figure 3: Sliding-scan p -value trajectories on Harvard Dataverse streams. Each panel aggregates streams within one family and plots the median p -value at each scan position for MMD-RBF, projected KS, and per-dim KS. Shaded bands are interquartile ranges over streams. Red dashed lines mark ground-truth changepoints; the horizontal line is α0=0.05 on a symmetric-log axis. STAGGER is omitted because its whole-vector detection uses distance-based statistics without p -values after one-hot encoding of the categorical features.
Figure 4: Detection accuracy by test and shift family at N=1000 on the Failing Loudly benchmark, disaggregated by dataset. Rows are the three reproduced baselines (per-dimension KS on raw pixels, projected KS with SRP, MMD-RBF). Columns are the five shift families of Failing Loudly. Each cell aggregates all comparisons of one family over the three perturbation fractions and five splits: 45 for the Gaussian and image families (three intensities each), 30 for the combined family (two variants), and 15 for the adversarial and class-knockout families.
Figure 5: Reproduced p -value curves under the Gaussian-noise shift class of Rabanser et al. (2019) on MNIST and CIFAR-10. Panels correspond to affected-fraction {0.10,0.50,1.00} . Lines are median p -values across 30 shift instances per sample size (3 intensities × 5 splits × 2 datasets). Shaded bands are the interquartile range. The horizontal reference line is α0=0.05 . This figure is the reproduction analog of Failing Loudly Figure 3.
Figure 6: Failing Loudly reproduction: image-transformation shifts (rotation, translation, zoom at three intensities). Panels correspond to {10%,50%,100%} perturbed test data.
Figure 7: Failing Loudly reproduction: adversarial shifts generated with the fast gradient sign method. Panels correspond to {10%,50%,100%} perturbed test data. Adversarial shifts are the hardest single-class shift family at low perturbation fractions.
Figure 8: Failing Loudly reproduction: class knock-out (label) shifts. Panels correspond to {10%,50%,100%} perturbed test data. Because knock-out shifts change P(Y) rather than P(X) , MMD-based tests detect the shift through its indirect effect on the observed covariate distribution.
Figure 9: Failing Loudly reproduction: combined image transformation + class knock-out shifts. Panels correspond to {10%,50%,100%} perturbed test data. Combined shifts show the sharpest median- p descent because covariate and label distributions shift jointly.
Test
TPR (%)
FPR (%)
Ratio
r
Time (s)
Per-dim KS
100.0
100.0
0.91
−0.085
0.07
MMD-RBF
100.0
98.4
1.99
−0.022
36.20
Projected KS
100.0
100.0
1.25
+0.109
0.63
Spark RFF-MMD
40.2
3.2
8.41
+0.427
2.68
Spark SWD
48.2
10.0
1.84
+0.422
5.50
Table 7: Sliding-scan detection metrics on the Trendyol collection-ranking feature table averaged over all 8 test scenarios (4 drift types × 2 severity regimes). Grand means include the weak regime, where both RFF-MMD and SWD fail to discriminate; see Table 8 for the per-regime breakdown. 1 1 footnotetext: TPR = rejection rate over injection-active positions ( P1={b:mˉb>0} ); FPR = injection-relative rejection rate over injection-null positions ( P0={b:mˉb=0} ; this partition rather than a chronological pre/post-drift split is used because it correctly handles recurring drift). The injection-relative FPR is not the classical statistical Type-I error, since injection-null positions may still carry natural production-side covariate heterogeneity as documented in Section 6.3.5 ; Ratio = injection-active mean statistic divided by injection-null mean statistic; r = Pearson correlation between test statistic and per-bucket expected drift magnitude mˉb (Eq. 7 ); Time = median compute time per position on the Spark configuration described in Section 4.2 . Spark RFF-MMD uses the calibrated threshold 3×10−3 (Section 6.3.1 ); the other tests use their implementation defaults. The three tests reporting near-100% FPR and TPR simultaneously do not perform meaningful binary discrimination at their default thresholds; their failure mechanisms are discussed in Section 6.3.6 .
Sudden
Gradual
Incremental
Recurring
Test
Metric
Strong
Weak
Strong
Weak
Strong
Weak
Strong
Weak
RFF-MMD
r
+0.948
−0.094
+0.944
−0.081
+0.942
−0.083
+0.926
−0.084
TPR
95.2
0.0
71.4
0.0
71.4
0.0
83.3
0.0
FPR
3.4
3.4
3.4
3.4
3.4
3.4
2.6
2.6
SWD
r
+0.844
+0.033
+0.844
+0.095
+0.820
+0.041
+0.767
−0.069
TPR
95.2
9.5
76.2
14.3
76.2
14.3
83.3
16.7
Table 8: Strong- vs weak-regime detection performance by drift type on the Trendyol collection-ranking feature table. RFF-MMD and SWD are the two tests capable of producing meaningful discrimination; the remaining three (per-dim KS, MMD-RBF, projected KS) exhibit systematic failures at all regimes (see Section 6.3.6 ). 1 1 footnotetext: Metric definitions ( r , TPR, FPR and the P1 / P0 partition) as in Table 7 . Each cell is a single scenario; TPR and FPR in %. RFF-MMD uses the calibrated threshold 3×10−3 (Section 6.3.1 ); SWD uses the fixed implementation threshold τSW=0.05 . With 21 active and 29 null positions per scenario (12 and 38 for recurring drift), Wilson 95% intervals are wide, e.g., 20/21=95.2% [77.3, 99.2], 15/21=71.4% [50.0, 86.2], 1/29=3.4% [0.6, 17.2], and 3/29=10.3% [3.6, 26.4].
Figure 10: Sliding-scan trajectories of Spark RFF-MMD (red) and Sliced Wasserstein Distance (blue) across 8 scenarios. Rows correspond to the four drift types; the left column is the strong regime and the right column the weak regime. Solid lines are the test statistic on a shared log y-axis; dashed horizontal lines are the drift threshold of each test in the matching color (RFF-MMD: 3×10−3 , calibrated per Section 6.3.1 ; SWD: 5×10−2 , implementation default). The dotted vertical line marks the initial drift onset (for recurring drift, subsequent returns between concepts are additional true transition points not marked). Per-panel Pearson r is shown in each subplot title.
Figure 11: Per-position runtime aggregated across all 8 scenarios ( n=400 measurements per test: 50 positions × 8 scenarios). Bars are medians; whiskers span the interquartile range. Each test is assigned a distinct color that is reused consistently in Figure 12 . MMD-RBF dominates the range at ∼ 36 seconds per position, roughly 6.5 × slower than the next test (Sliced Wasserstein Distance) and roughly 480 × slower than per-dim KS.
Figure 12: Per-position runtime distribution disaggregated by scenario. Each box aggregates n=50 per-position measurements from one scenario. Colors match Figure 11 . The y-axis is broken to accommodate MMD-RBF’s range ( ∼31 –53 s, top panel) while preserving resolution on the four smaller-scale tests (0–10.5 s, bottom panel). Distributions are similar across scenarios, confirming that runtime is a function of test complexity and execution mode, independent of drift structure or severity regime.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning
Defined in
Two-sample testing
DR,DC
Reference and current samples
Sec. 3.1
PR,PC
Distributions of the reference and current samples
Eq. ( 1 )
z ; X,Y
Tested vector ( (X,Y) or Xk ); features and label
Sec. 3.1
d
Feature dimensionality
Table 1
s ; p
Test statistic; p -value
Sec. 3.1
Appendix
Table 9: Notation used throughout the paper.
Name in Text
Symbol
Role
Original Column
Source-table columns
Seed-content identifier
–
Metadata (vector)
seed
Owner identifier
–
Metadata (vector)
ownerid
Channel identifier
–
Metadata (vector)
channelId
Collection size
–
Metadata (vector)
content_cnt
Image cosine similarity
simimg
Similarity (vector); Concept A
image_cosine_similarity
Appendix
Table 10: Feature and injection-variable glossary for the Trendyol benchmark. Descriptive names and symbols used in the text are mapped to the original column names of the source table.
Data stream mining is fundamentally challenged by concept drift, where distributional changes can degrade model performance. Despite the proliferation of drift detection methods, progress in the field is hindered by inconsistent evaluation practices: studies rely on oversimplified synthetic data generators, adopt incompatible metrics, and lack transparency in hyperparameter selection, making fair comparisons difficult. We address this gap with a novel benchmarking framework comprising three contributions: (1) a drift simulation method that injects controlled distributional changes into real-world datasets via Monte Carlo trials, enabling supervised evaluation while preserving real-world data complexity; (2) an evaluation protocol for drift detection with timing-aware criteria, including the derivation of new metrics (e.g., F1 detection score, normalized detection time) that are comparable across streams; and (3) we advocate for a leave-one-dataset-out hyperparameter optimization protocol for drift detection methods that promotes configuration robustness across heterogeneous stream dynamics. We benchmark 14 widely used drift detection methods on 7 realworld datasets across 4 drift types (class prior, label swap, feature permutation, feature filtering), each under both abrupt and gradual transitions. Our experimental results provide insights into the strengths and weaknesses of current drift detection approaches while establishing baseline performance metrics for future research in this area. All code and experiments are publicly available.
Vitor Cerqueira, Heitor Murilo Gomes, Marco Heyden +2
University of Coimbra · Victoria University of Wellington · Commerzbank +2
Machine learning algorithms deployed for evolving streaming environments must handle the non-stationary data distributions, commonly referred to as concept drift. The presence of concept drift poses a major challenge for many real-world applications because it can severely degrade their predictive performance, hindering their ability to support robust decision-making. Consequently, the timely and efficient detection of drift events is critical for sustaining high accuracy over time. This study examines theoretically the concept drift characteristics and numerous drift detection algorithms across several categories. Furthermore, we evaluate their performance on both synthetic and real-world datasets exhibiting diverse streaming scenarios and drift characteristics, such as abrupt and gradual changes. This study aims to enhance understanding of the complex notion of concept drift characteristics and behavior of drift detectors, along with their applicability to diverse contexts.
Md Moman Ul Haque Khan, Samira Sadaoui
Department of Computer Science, University of Regina, Canada
Drift detection is a core component of production machine learning monitoring systems, where detectors are used to compare incoming data with a reference distribution and trigger alerts when changes occur. However, these detectors are often evaluated in research settings that emphasize detection accuracy under synthetic shifts, while overlooking false alarms under continuous monitoring. In production environments, models are monitored repeatedly over time and across many features, and even small false positive rates can accumulate into frequent alerts, leading to alarm fatigue. We empirically analyze false positive behavior across five commonly used drift detectors: PSI, KS, MMD, LSDD, and adversarial validation. Consistent with existing literature, PSI exhibits strong sensitivity to batch size, producing frequent false alarms at small sample sizes; however, we further observe that its behavior stabilizes and improves substantially once batch sizes exceed approximately 200 samples. In contrast, KS, MMD, and LSDD display persistent fluctuations across batch sizes, while remaining comparatively more reliable than PSI in low-data regimes. Applying a Bonferroni correction reduces false positive rates, but often at the cost of reduced true positive sensitivity, reinforcing the well-known stability - sensitivity trade-off in drift detection. This work provides a systematic comparison of false positive behavior across multiple drift detectors under continuous monitoring conditions. We identify tradeoffs across detector families and provide practical guidelines for selecting and calibrating drift detectors in production ML systems.
Raj Shekhar Singh
Indian Institute of Technology, Roorkee Roorkee, Uttarakhand, India