stat.MLJul 18, 2025

Conformal Data Contamination Tests for In-distribution Data Acquisition

Authors: Martin V. Vejling, Shashi Raj Pandey, Christophe A. N. Biscio, Petar Popovski

Organizations: Department of Mathematical Sciences and Department of Electronic Systems Aalborg University Aalborg, Denmark · Department of Electronic Systems Aalborg University Aalborg, Denmark · Department of Mathematical Sciences Aalborg University Aalborg, Denmark

Abstract

The amount of quality data in many machine learning tasks is limited to what is available locally to data owners. The set of quality data can be expanded through trading or sharing with external data agents. However, external data may be contaminated or introduce undesirable sample diversity which can degrade performance of personalized machine learning tasks, as in diagnosis of a rare disease or recommendation systems. Therefore, data buyers need quality guarantees prior to data acquisition. Previous works primarily rely on distributional assumptions about data from different agents, relegating quality checks to post-hoc steps involving costly data valuation procedures. We propose a distribution-free, contamination-aware data acquisition framework that, by inspecting only a small volume of data, identifies external data agents whose data is most valuable for model personalization. To achieve this, we introduce novel two-sample testing procedures, preceding full data acquisition, grounded in rigorous theoretical foundations for conformal outlier detection, to determine whether an agent's data exceeds a contamination threshold. The proposed tests, termed conformal data contamination tests, remain valid under arbitrary contamination levels and the novel Storey-type test provably enables finite-sample false discovery rate control via the Benjamini-Hochberg procedure. Empirical evaluations across diverse collaborative learning scenarios demonstrate the robustness and effectiveness of our approach. Overall, the conformal data contamination test distinguishes itself as a generic procedure for aggregating data with statistically rigorous quality guarantees.

Explore similar work

CardsList
  1. Multi-Agent Conformal Prediction with Personalized Statistical Validity

    May 30, 2026Martin V. Vejling, Christophe A. N. Biscio, Adrien Mazoyer +2Weighted Conformal PredictionConformal Prediction

  2. Multi-Distribution Robust Conformal Prediction

    Jan 6, 2026Yuqi Yang, Ying JinDistribution Shift RobustnessConformal Prediction

  3. Provable Joint Decontamination for Benchmarking Multiple Large Language Models

    May 20, 2026Zhenlong Liu, Hao Zeng, Hongxin WeiLLM EvaluationBenchmark Contamination