cs.DCOct 6, 2026

Beyond Marginal Monitoring: Distributed Joint-Distribution Testing for Data Concept Drift in Large Scale E-Commerce Operations

Authors: Cagdas Pullu, Mahmut Emir Arslan, Bugra Balkac, Aylin Ondersev Balta, Cihangir Celal Palaci, Fikri Cem Yilmaz, Altan Cakir

Organizations: Data & Analytics, Trendyol Group, Istanbul, 34485, Istanbul, Türkiye. · Department of Data Science and Analytics, Istanbul Technical University, Maslak, Istanbul, 34469, Istanbul, Türkiye.

Abstract

Concept drift threatens production machine learning, yet the empirical behavior of multivariate two-sample drift detectors at scale remains under-characterized. Existing benchmarks rarely address the hundreds of millions of rows and high-cardinality features typical of industrial-operational datasets. We evaluate five multi-column two-sample tests (marginal, projection-based, and kernel embedding methods) across three complementary environments: the Harvard Dataverse, a validated Failing Loudly reproduction (mean absolute error between 0.030 and 0.053), and a novel synthetic-injection benchmark on the 137.5-million-row Trendyol collection-ranking feature table. Testing four drift types across two severity-scope regimes, we demonstrate that distributed Maximum Mean Discrepancy with Random Fourier Features on Apache Spark scales robustly. Averaged over the four drift types in the strong regime and under a calibrated threshold, it achieves a Pearson correlation of r = 0.940 with expected drift magnitude, an 80.4% true positive rate, and a 3.2% false positive rate. Conversely, the per-dimension Kolmogorov-Smirnov test failed due to statistic saturation from ID-like columns under asymmetric sampling, establishing a critical constraint for large-scale sampling design. At weak configurations (realized-flip fractions of at most 0.57%), detectors struggled to reliably discriminate, highlighting the need for future intensity-grid power analyses to distinguish fundamental sensitivity bounds from scalable threshold shifts.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. A Framework for Evaluating and Benchmarking Concept Drift Detection Methods

    Jun 5, 2026Vitor Cerqueira, Heitor Murilo Gomes, Marco Heyden +2Synthetic Data

  2. Learner-based Concept Drift Detection: Analysis and Evaluation

    Jun 18, 2026Md Moman Ul Haque Khan, Samira SadaouiNon-Stationarity

  3. When Drift Detectors cry Wolf: False Alarm Rates in continuous ML Monitoring

    Jul 19, 2026Raj Shekhar SinghFalse-Negative RateAlarm Fatigue